Training method, method and apparatus for predicting regional rainwater runoff pollution load.
Patent Information
- Application Number
- JP2026507242
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-27
- Filing Date
- 2024-08-27
- Publication Date
- 2026-09-01
Smart Images

Figure 2026529586000001_ABST
Abstract
Description
Technical Field
[0001] The present application claims the priority right of the Chinese patent application filed with the China National Intellectual Property Administration on March 27, 2024, with the application number 202410354280.X and the application title "Training method, method and apparatus for predicting regional rainwater runoff pollution load", the entire content of which is incorporated into the present application by reference.
[0002] The present application relates to the technical field of water environment management, and in particular to a training method, a method and an apparatus for predicting regional rainwater runoff pollution load. Background Art
[0003] Calculation of rainwater runoff pollution load is a fundamental basis for water environment management. Conventional calculation of rainwater runoff pollution load is usually realized by performing statistical analysis and calculation based on a large amount of actually measured data (at least dozens of times) over a long period of time (more than one year). However, this method requires on-site sampling and monitoring of rainwater runoff pollution load in real time, so it has problems such as long cycle and high cost. In addition, rainfall has seasonality and randomness every year, and the error of results obtained in the actual operation process is large, which cannot meet the requirements of timeliness and accuracy of water environmental management engineering design.
[0004] In conventional methods, it is further proposed to reduce the sampling cycle and cost by predicting the rainwater runoff pollution load through a model. However, existing models still rely on a large amount of monitoring data of the region to be predicted. Due to the uncertainty of rainfall monitoring, existing monitoring data often contain a large number of abnormalities and invalid data, and existing models still lack screening training for abnormal and invalid data, which ultimately affects the prediction accuracy of the model, and the processing of abnormal and invalid data leads to a decrease in the processing efficiency of the model. Summary of the Invention Problem to be Solved by the Invention
[0005] This application provides a training method, a method and apparatus for predicting regional rainwater runoff pollution load, to solve the problem of large data processing volume and low accuracy in prediction models. [Means for solving the problem]
[0006] In a first aspect, the present invention provides a predictive model training method comprising the steps of: acquiring historical stormwater runoff pollution load data and historical feature data for the region, wherein the historical feature data includes feature data representing attribute data categories and feature data representing meteorological data categories; screening the historical feature data to obtain training feature data based on a correlation value between the historical stormwater runoff pollution load data and the historical feature data; and training an initial predictive model based on the training feature data and the historical stormwater runoff pollution load data to obtain a predictive model for predicting the stormwater runoff pollution load for the region.
[0007] In one feasible form, the historical feature data includes the surface data of the plurality of regions, and the step of obtaining historical stormwater runoff pollution load data for the regions includes the step of obtaining single stormwater runoff pollution load data for the surface corresponding to the surface data, and the step of obtaining historical stormwater runoff pollution load data representing the stormwater runoff pollution load for the region based on the surface area in the surface data and the single stormwater runoff pollution load data.
[0008] In one feasible form, the number of regions is multiple, and the step of screening from the historical feature data to obtain training feature data based on the correlation value between the historical stormwater runoff pollution load data and the historical feature data includes the steps of performing clustering based on the historical feature data of each region to obtain a regional category corresponding to each region, and screening from the historical feature data in each regional category based on the correlation value between the historical stormwater runoff pollution load data and the historical feature data in each regional category to obtain training feature data corresponding to each regional category, and training based on the training feature data of each regional category and the historical stormwater runoff pollution load data to obtain a predictive model for the corresponding regional category.
[0009] In one feasible form, the historical feature data includes feature data corresponding to a plurality of feature categories, and the step of screening the historical feature data to obtain training feature data based on a correlation value between the historical stormwater runoff pollution load data and the historical feature data includes the steps of obtaining a correlation value between the historical stormwater runoff pollution load data and each feature data, and screening the feature data corresponding to the plurality of feature categories to obtain training feature data based on the magnitude of the correlation value corresponding to each feature data.
[0010] In a second aspect, the present invention provides a method for predicting rainwater runoff pollution load in a region, comprising the steps of: acquiring target feature data for a target region, wherein the target feature data includes feature data representing attribute data categories and feature data representing meteorological data categories; acquiring pre-configured feature categories based on the target feature data; screening the target region feature data based on the pre-configured feature categories; inputting the screened target feature data into a pre-trained prediction model; and obtaining rainwater runoff pollution load data for the target region output by the prediction model, wherein the pre-configured feature categories are obtained based on correlation values between historical rainwater runoff pollution load data and feature data corresponding to each feature category in historical feature data.
[0011] In one feasible form, the prediction model comprises multiple prediction models, each corresponding to a different regional category, each corresponding to a different pre-defined feature category, and the steps include: obtaining the pre-defined feature category based on the target feature data; screening the target regional feature data based on the pre-defined feature category; inputting the screened target feature data into a pre-trained prediction model; and obtaining the target region's storm runoff pollution load data output by the prediction model. The steps include: classifying the target region based on the target feature data and obtaining the target region category for the target region; screening the target region data for feature data corresponding to the feature category based on the pre-defined feature category in the target region category; obtaining the screened target region feature data; and inputting the screened target region feature data into a trained model corresponding to the target region category; and obtaining the target region's storm runoff pollution load data.
[0012] In a third aspect, the present invention provides a prediction model training device comprising: a historical data acquisition module configured to acquire historical stormwater runoff pollution load data and historical feature data of a region, wherein the historical feature data includes feature data representing attribute data categories and feature data representing meteorological data categories; a data processing module configured to obtain training target feature data by screening the historical feature data based on a correlation value between the historical stormwater runoff pollution load data and the historical feature data; and a model training module configured to train an initial prediction model based on the training target feature data and the historical stormwater runoff pollution load data to obtain a prediction model for predicting the stormwater runoff pollution load of a region.
[0013] In a fourth aspect, the present invention provides a regional rainwater runoff pollution load prediction device, comprising: a target data acquisition module configured to acquire target feature data for a target region, wherein the target data includes feature data representing attribute data categories and feature data representing meteorological data categories; and a prediction module configured to acquire pre-set feature categories based on the target feature data, screen the target region feature data based on the pre-set feature categories, input the screened target region feature data into a pre-trained prediction model, and obtain rainwater runoff pollution load data for the target region output by the prediction model.
[0014] In a fifth aspect, the present invention provides an electronic device comprising a processor and a memory, the memory being used to store instructions, and the processor executing the instructions in the memory to cause the electronic device to perform the predictive model training method described in the first aspect or the regional rainwater runoff pollution load prediction method described in the second aspect.
[0015] In the sixth aspect, the present invention provides a computer-readable storage medium on which computer execution instructions are stored, and which is used to implement the predictive model training method described in the first aspect or the regional rainwater runoff pollution load prediction method described in the second aspect when the computer execution instructions are executed by a processor.
[0016] In the seventh aspect, the present application provides a computer program product that, when executed by a processor, realizes the predictive model training method described in the first aspect or the regional rainwater runoff pollution load prediction method described in the second aspect. [Effects of the Invention]
[0017] The training method, regional rainwater runoff pollution load prediction method, and apparatus according to the present application acquire historical feature data representing regional attributes and weather, acquire historical rainwater runoff pollution load data, calculate a correlation value between the historical feature data and the historical rainwater runoff pollution load data, screen for training target feature data with a high correlation to the historical pollution load data, thereby removing invalid data in the historical feature data, improving the accuracy of the training data, improving the accuracy of the trained prediction model, reducing the volume of training target feature data for training the prediction model, reducing the workload of model training, and improving the efficiency of model training.
[0018] The drawings herein are incorporated into the specification and constitute part of this specification, illustrating embodiments conforming to the present application and are used together with the specification to interpret the principles of the present application. [Brief explanation of the drawing]
[0019] [Figure 1] This is a structural diagram of the implementation environment shown in one exemplary embodiment of the present application. [Figure 2] This is a flowchart of a predictive model training method shown in one exemplary embodiment of the present invention. [Figure 3]It is a schematic diagram of clustering analysis results shown in one exemplary embodiment of the present application. [Figure 4] It is a schematic diagram of the fitting result between predicted values and observed values shown in one exemplary embodiment of the present application. [Figure 5] It is a flowchart of a method for predicting regional rainwater runoff pollution load shown in one exemplary embodiment of the present application. [Figure 6] It is a result diagram of principal coordinate analysis shown in one exemplary embodiment of the present application. [Figure 7] It is a comparison result diagram between actually measured values and predicted values shown in one exemplary embodiment of the present application. [Figure 8] It is a structural schematic diagram of a prediction model training apparatus shown in one exemplary embodiment of the present application. [Figure 9] It is a structural schematic diagram of an apparatus for predicting regional rainwater runoff pollution load shown in one exemplary embodiment of the present application. [Figure 10] It is a block diagram of an electronic device shown in one exemplary embodiment of the present application. DETAILED DESCRIPTION OF EMBODIMENTS
[0020] The above drawings show clear embodiments of the present application, which will be described in more detail hereinafter. These drawings and textual descriptions are not intended to limit the scope of the inventive concept of the present application in any way, but rather to explain the concept of the present application to those skilled in the art with reference to specific embodiments.
[0021] Herein, exemplary embodiments are described in detail, examples of which are shown in the drawings. When reference is made to the drawings in the following description, unless otherwise stated, the same numerals in different drawings refer to the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the present application as detailed in the appended claims.
[0022] First, terms related to the present application are defined.
[0023] EMC refers to the average concentration per rainfall and is used to evaluate and calculate the large amount of pollutants contained in rainwater runoff on the ground surface. These pollutants may include various contaminants such as COD (organic pollutants), SS (suspended solids), NH3-N (ammonia nitrogen), TN (total nitrogen), and TP (total phosphorus).
[0024] In this application, rainwater runoff pollution load data generally refers to data related to EMC, and specifically, data from a single rainfall is not very meaningful and is inaccurate because it is random. Historical rainwater runoff pollution load data may be average EMC data over a specific time period (e.g., a year, a quarter, a three-year period), that is, obtained by averaging several EMC data over that time period, specifically by weighting the average concentration per rainfall over that time period, and if the time period is one year, the corresponding historical rainwater runoff is obtained by weighting the EMC data of several single rainfalls over that year.
[0025] As a result, the rainwater runoff pollution load data for a target area output by the prediction model in this application is also the average EMC data for that time unit. Of course, in some embodiments, the EMC data reflects the concentration of pollutants. However, when managing rainwater pollution, pollution data can generally show the wastewater situation in the area more intuitively, thereby specifying the corresponding management method. The pollution data is generally obtained by multiplying the EMC data by the runoff flow rate. Specifically, in the case of a single rainfall, the pollution data for that rainfall is obtained by multiplying the runoff flow rate / rainfall amount of the rainwater corresponding to the EMC data of that rainwater. On the other hand, in the case of pollution data for a specific time unit, the pollution data for that time unit is the sum of the products of the runoff flow rates of the number of rainfalls corresponding to the EMC data of each rainfall in that time unit. The pollution data for each time unit is the sum of the pollution data for each rainfall event in that time unit, and in some embodiments, when using the prediction model, a set of data (the runoff flow rate for the target area in that time unit, where the runoff flow rate refers to the sum of the runoff flow rates for all rainfall events in that time unit) can be added, and when the prediction model processes, it first obtains the average EMC data for that time unit, and then multiplies the average EMC data for that time unit by the runoff flow rate for that time unit to obtain the pollution data for the target area in that time unit, and the prediction model takes the pollution data for the target area in that time unit as input, and in some special embodiments, the rainwater runoff pollution load data for the target area output by the prediction model is the pollution data for the target area in that time unit.
[0026] The problem of rainwater runoff pollution caused by rainfall is becoming increasingly serious. Calculating regional rainwater runoff pollution loads is a fundamental basis for water environment management. However, because rainwater runoff pollution loads have complex pollution sources and many influencing factors, calculating and predicting them is difficult. Generally, monitoring is complex, inaccurate, and has a short shelf life.
[0027] Currently, rainwater runoff pollution load predictions typically employ real-time sampling and monitoring methods, obtaining the Event Mean Concentration (EMC) per rainfall event after calculation. To obtain relatively accurate weighted average concentrations of pollutants for each type of ground surface, it is necessary to monitor typical and representative rainfall runoff data from at least 15-20 events throughout the year in real time. This method requires repeated data measurements for different ground surfaces and different regions, making it time-consuming, resulting in long monitoring cycles, difficulty in controlling the normative nature of sampling, and large errors during the actual operation.
[0028] Rainwater runoff pollution load prediction may employ modeling methods, or modify models based on real-time monitoring data. For example, models such as SWMM (Dynamic Precipitation-Runoff Simulation Model), SWAT (Watershed-Scale Model), and HSPF (Hydrological Simulation Model) undergo certain calibrations and modifications in combination with basic data and monitoring data, taking into account local conditions. However, this method still places a high demand on basic data, requiring large amounts of data on meteorology, hydrology, land use, pollutant accumulation-erosion, and drainage network construction. This basic data is generally obtained from historical data monitored over the past several years, or by obtaining relevant basic data through additional monitoring as needed for the model, thereby meeting the needs for model calculations and parameter calibration and ensuring the accuracy of model calculations.
[0029] Consequently, conventional regional stormwater runoff pollution load prediction technologies, whether model-based or experimental, rely on large amounts of monitoring data from target areas. This necessitates repeated monitoring of large amounts of data for each target area, and it is not possible to mine common and characteristic elements, or common and characteristic data, across different areas. Furthermore, it is not possible to reduce the workload of repeated monitoring by performing monitoring data mining, screening, and application as disclosed in previous research. From the above, it can be seen that the shortcomings of regional stormwater runoff pollution load prediction technologies include three aspects. Firstly, whether relying on large amounts of real-time monitoring data or historical monitoring data, the data is generally incomplete and needs to be supplemented with real-time monitoring data. The amount of monitoring data that needs supplementation is large, and the efficiency is somewhat low. In addition, due to the seasonality and randomness of rainfall, the timeframe for acquiring large amounts of actual monitoring data is long (e.g., more than one year), resulting in a high workload, being time-consuming, and failing to meet the timeframe needs of engineering management. Secondly, accuracy in field monitoring data such as hydrology and water quality is a major challenge due to the influence of monitoring environmental conditions and the randomness of sampling operations. Regardless of whether it is real-time or historical monitoring data, a large amount of anomaly and invalid data exists, and conventional methods generally lack data screening and scientific removal of invalid or interfering data, which affects the accuracy of the results. Thirdly, conventional models are all built for specific target areas, rely on large amounts of data from those target areas, and cannot be used for other target areas (e.g., different cities). Since it is necessary to retrain a single model based on large amounts of data from other target areas, the methods have low versatility and accuracy, and the repetitive workload of monitoring and calculation is high.
[0030] Based on the above problems, this embodiment proposes a predictive model training method and a method for predicting regional rainwater runoff pollution load. The predictive model training method makes full use of historical data from multiple regions as well as the target region, significantly increasing the volume of monitoring data, and enables the scientific use of historical data from different regions through specific clustering. Furthermore, the predictive model training method can screen high-quality training data and reduce the volume of training data, thereby improving the accuracy of the predictive model. As a result, the rainwater runoff pollution load prediction method built on this predictive model can efficiently obtain highly accurate regional rainwater runoff pollution load data.
[0031] The following will describe in detail, with reference to specific embodiments, the technical solution of the present application and how it solves the above technical problem. Several of the following specific embodiments can be combined with each other, and in some embodiments, the same or similar concepts or processes may not be explained redundantly. The embodiments of the present application will be described below with reference to the drawings.
[0032] It should be noted that the stormwater runoff pollution in this embodiment generally refers to pollution caused by initial rainwater. In actual operations, pollution caused by initial rainwater is severe and worth studying and monitoring. However, as rainfall progresses, the obtained stormwater runoff pollution load decreases and is not worth studying. The historical stormwater runoff pollution load data and the stormwater runoff pollution load data for the target area in this embodiment generally refer to pollution loads in specific time units, such as pollution loads over a year, a quarter, or two quarters. The pollution load per rainfall is somewhat random and does not have much guiding significance for water environment management engineering.
[0033] The predictive model training method and the regional rainwater runoff pollution load prediction method according to this application can both be applied to the implementation environment shown in Figure 1. As shown in Figure 1, the implementation environment includes terminal 1 and server-side 2, and terminal 1 is connected to server-side 2 in communication. Of course, the number of terminals 1 and server-side 2 shown in Figure 1 is illustrative, and in other embodiments, the implementation environment may include a different number of terminals 1 and server-side 2, and is not specifically limited here.
[0034] In this embodiment, terminal 1 can be understood as various electronic devices such as wired terminals or wireless terminals, and includes, but is not specifically limited to, devices such as smartphones, computers, smart home appliances, and in-vehicle terminals. Server-side 2 may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server, but is not specifically limited to, the same.
[0035] When the prediction model training method in this embodiment is applied to the implementation environment shown in Figure 1, terminal 1 can acquire regional historical stormwater runoff pollution load data and regional historical feature data. The historical feature data includes feature data representing attribute data categories and feature data representing meteorological data categories. Terminal 1 transmits the regional historical stormwater runoff pollution load data and regional historical feature data to server 2. After acquiring the regional historical stormwater runoff pollution load data and regional historical feature data, server 2 screens the historical feature data to obtain training target feature data based on the correlation value between the historical stormwater runoff pollution load data and the historical feature data. Based on the training target feature data and the historical stormwater runoff pollution load data, it trains an initial prediction model to obtain a prediction model for predicting regional stormwater runoff pollution load. Correspondingly, the prediction model training device is installed within server 2 and completes the training of the prediction model.
[0036] When the regional rainwater runoff pollution load prediction method in this embodiment is applied to the implementation environment shown in Figure 1, terminal 1 can acquire target feature data for the target region. The target feature data includes feature data representing attribute data categories and feature data representing meteorological data categories. Terminal 1 transmits the target feature data for the target region to server 2. After acquiring the target feature data for the target region, server 2 acquires pre-configured feature categories based on the target feature data, screens the target region feature data based on the pre-configured feature categories, inputs the screened target feature data into a pre-trained prediction model, and obtains rainwater runoff pollution load data for the target region output by the prediction model. The pre-configured feature categories are obtained based on the correlation values between historical rainwater runoff pollution load data and the feature data corresponding to each feature category in the historical feature data. Accordingly, the regional rainwater runoff pollution load prediction device is installed within server 2 and completes the rainwater runoff pollution load prediction for the target region.
[0037] Of course, the implementation environment shown in Figure 1 for this embodiment is merely illustrative, and in other embodiments, the predictive model training method and the regional rainwater runoff pollution load prediction method can be applied to other implementation environments, and the implementation environments for the predictive model training method and the regional rainwater runoff pollution load prediction method may be the same or different.
[0038] Figure 2 shows a predictive model training method according to one embodiment of the present invention. As shown in Figure 2, the predictive model training method may specifically include steps S210 to S250, and a detailed explanation is as follows.
[0039] Step S210: Obtain historical stormwater runoff pollution load data and historical characteristic data for the region.
[0040] In this embodiment, the historical feature data includes feature data representing attribute data categories and feature data representing meteorological data categories, and the historical stormwater runoff pollution load data represents the average EMC data for a specific time unit in a region, i.e., it is obtained by weighting the stormwater runoff pollution load data. For example, if one year is considered as one time unit, the historical stormwater runoff pollution load data represents the annual average EMC data, i.e., it is obtained by weighting the stormwater runoff pollution load data for that year.
[0041] In this embodiment, when calculating the historical rainwater runoff pollution load data, instead of using all single rainwater runoff pollution load data for that time unit, a screening is first performed to obtain representative single rainwater runoff pollution load data, and then the historical rainwater runoff pollution load data for that time unit is calculated.
[0042] The characteristic data in the attribute data category includes characteristic data such as regional topography and economic level. Regional topography may include regional ground surface data, average gradient data, etc. Ground surface data may be further subdivided into data such as ground surface type, ground surface impermeability, ground surface gradient, and ground surface area. The characteristic data for economic level may include characteristic data that may affect EMC, such as regional per capita income.
[0043] The characteristic data in the meteorological data category includes characteristic data that may affect EMC, such as annual precipitation, annual average temperature, rainfall amount per rain, duration of rainfall per rain, average rainfall intensity per rain, and period of clear weather preceding rainfall. Rainfall meteorological data can correspond to ground surface data; for example, for each surface, there is characteristic data such as rainfall amount per rain and duration of rainfall per rain for that ground surface scene.
[0044] Specifically, historical feature data includes multiple data categories such as attribute data categories and weather data categories. Each data category contains multiple feature data, and the feature data is distinguished by the feature category. This feature category can be considered the smallest category unit in historical feature data. For example, the impermeability of the ground surface and the average temperature mentioned above can both be considered as two feature categories. In other words, historical feature data can be considered as a division of multiple feature categories under each data category, with feature data associated with each feature category, and multiple feature data constitute the historical feature data for a region.
[0045] In this embodiment, the historical feature data for a region is historical monitoring data for that region, which can be obtained by querying literature databases, statistical yearbooks, or statistical publications. In this embodiment, the historical feature data and historical stormwater runoff pollution load data obtained should cover long periods, such as 5, 6, or 10 years, to ensure that the obtained historical feature data and historical stormwater runoff pollution load data have sufficient data volume and temporal representativeness for training the prediction model.
[0046] This period can be considered the time for acquiring historical feature data and historical stormwater runoff pollution load data. If the data acquisition period described above is 5 years, historical feature data and historical stormwater runoff pollution load data for the past 5 years are acquired. Then, using a time unit as a representative unit, the average value of the average historical feature data and the average value of the historical stormwater runoff pollution load data for a specific time unit are calculated, and these average values are used as the historical feature data and historical stormwater runoff pollution load data corresponding to the region.
[0047] The time unit is a pre-set value, thereby reducing the randomness of pollution load data per rainfall. Generally, the time unit is long, for example, a year or a quarter. If the time unit is a year, and the period is 5 years, the corresponding historical feature data and historical stormwater runoff pollution load data obtained are the annual average historical feature data and historical stormwater runoff pollution load data. Thus, the annual average data in the historical stormwater runoff pollution load data and the historical feature data of the region (for example, data such as annual precipitation and annual average temperature are considered to be the annual average data over 5 years), and the corresponding historical stormwater runoff pollution load data is data obtained by averaging the stormwater runoff pollution load data over 5 years. Of course, in some embodiments, representative data may be obtained on a quarterly basis. In this case, the obtained data such as annual precipitation and annual average temperature are modified to correspond to quarterly precipitation and quarterly average temperature, and the data acquisition method is not limited here.
[0048] In one specific embodiment, a method for calculating historical stormwater runoff pollution load data is proposed, which includes steps S10 to S12. Step S10: Obtain single stormwater runoff pollution load data for the ground surface corresponding to the ground surface data. Step S12: Obtain historical stormwater runoff pollution load data representing the regional stormwater runoff pollution load based on the ground surface area and single stormwater runoff pollution load data in the ground surface data.
[0049] Specifically, for each surface in a region, rainwater runoff pollution load data for the corresponding ground surface can be acquired. This rainwater runoff pollution load data may be obtained simultaneously with the acquisition of historical feature data, or it may be acquired after acquiring the historical feature data, based on the ground surface data. In other words, in this case, rainwater runoff pollution load data for each surface can be obtained, along with corresponding data such as rainfall amount, rainfall duration, average rainfall intensity, and clear weather period preceding the rain.
[0050] In this embodiment, historical rainwater runoff pollution load data representing the regional rainwater runoff pollution load can be obtained based on the ground surface area and single-rainfall rainwater runoff pollution load data in the ground surface data, and historical rainwater runoff pollution load data can be calculated by taking a weighted average of the single-rainfall rainwater runoff pollution load data.
[0051] Specifically, based on the ground surface data and the rainwater runoff pollution load data, the average rainwater runoff pollution load data for the corresponding ground surface can be calculated. In the case of historical rainwater runoff pollution load data with a year as the unit of time, multiple rainwater runoff pollution load data for a specific ground surface in a specific year can be obtained, and the average value can be calculated to obtain historical rainwater runoff pollution load data for that ground surface in that year. If the period is 5 years, the rainwater runoff pollution load data for that ground surface in each year of the 5 years can be calculated.
[0052] Subsequently, weights are obtained for different ground surfaces in the region based on the size of the ground surface area in that region. Larger ground surface areas have higher weights, and smaller areas have lower weights. Historical stormwater runoff pollution load data is calculated using these weights and historical stormwater runoff pollution load data for each surface. The obtained historical stormwater runoff pollution load data is historical stormwater runoff pollution load data for the corresponding time unit. The predictions made by a predictive model trained using this data are also historical stormwater runoff pollution load data for the target region for the same time unit, for example, data that predicts the stormwater runoff pollution load for a specific year in the target region.
[0053] Of course, in order to ensure the accuracy of the training data, after obtaining the historical stormwater runoff pollution load data on an hourly basis, the average value may be calculated based on the period, thereby obtaining the average value of the historical stormwater runoff pollution load data on an hourly basis and when subsequently training the model, for each set of data input to the model, the historical stormwater runoff pollution load data on an hourly basis corresponds to the feature data to be trained, and the average value of the historical stormwater runoff pollution load data on an hourly basis corresponds to the average value of the feature data to be trained on an hourly basis. In one embodiment, if the period is 5 years and the hourly basis is years, the historical stormwater runoff pollution load data for each year and the feature data to be trained on the corresponding year may be input to the model as a set of training data for training, the average value of the historical stormwater runoff pollution load data obtained by averaging the historical stormwater runoff pollution load data for 5 years and the average value of the data to be trained on 5 years may be input to the model as a set of training data for training, or the average value of the historical stormwater runoff pollution load data for any 2 years or more over 5 years and the average value of the data to be trained on 5 years may be input to the model as a set of training data for training.
[0054] In this embodiment, in the case of a single region, there may be multiple different categories of ground surfaces within that region, and there may be multiple rainwater runoff pollution load data for each surface. Especially when the period is long, the number of rainwater runoff pollution load data for each surface that need to be processed accordingly becomes even larger, resulting in a large amount of data to process. In this case, first, the rainwater runoff pollution load data for each surface can be screened, representative rainwater runoff pollution load data for each surface can be selected, and then historical rainwater runoff pollution load data can be calculated according to the data after screening. This screening method may be obtained by manually screening according to descriptions in past literature, or it may be possible to input the rainwater runoff pollution load data for each different ground surface into a machine model, analyze the characteristics of each rainwater runoff pollution load data using the machine model, and screen representative rainwater runoff pollution load data. Screening single-rainwater runoff pollution load data using a machine model may involve analyzing the correlation between each single-rainwater runoff pollution load data based on the machine model and screening highly correlated single-rainwater runoff pollution load data as representative single-rainwater runoff pollution load data, or, after learning ground surface data with a machine model, scoring each single-rainwater runoff pollution load data and screening the highest-scoring single-rainwater runoff pollution load data as representative single-rainwater runoff pollution load data, and this is not specifically limited here. Of course, it may also be done by randomly selecting from ground surface single-rainwater runoff pollution load data, and this is not specifically limited here.
[0055] Step S230: Based on the correlation value between historical rainwater runoff pollution load data and historical feature data, the feature data to be used for training is screened from the historical feature data.
[0056] In this embodiment, the historical feature data has a wide range and a long period, resulting in a large data volume and some invalid feature data unrelated to historical stormwater runoff pollution load. This invalid feature data slows down model training and affects the model training results. Therefore, it is possible to screen the feature data to be trained based on the correlation between each feature data in the historical feature data and the historical stormwater runoff pollution load data, and to exclude feature data that is uncorrelated or has a low correlation, thereby improving the efficiency of model training. The correlation value represents the degree of correlation between the historical feature data and the historical stormwater runoff pollution data, with a higher value indicating a higher correlation.
[0057] This allows us to obtain correlation values between historical stormwater runoff pollution load data and each feature data, and based on the magnitude of the correlation value corresponding to each feature data, we screen the feature data to be trained from the feature data corresponding to multiple feature categories. The feature data to be trained is the feature data that has a high correlation with the historical stormwater runoff pollution load data for the corresponding region. This allows us to remove invalid data in the historical stormwater runoff pollution load data, reduce the data volume, and improve the reliability of the training data.
[0058] Of course, in this embodiment, when training the prediction model, historical stormwater runoff pollution load data and historical feature data corresponding to multiple regions can be acquired, and multiple regions can be classified. As a result, a prediction model for a corresponding regional category can be obtained based on data training for that regional category. Subsequently, when predicting stormwater runoff pollution load data for a target region, the prediction module for the target region's regional category can be used directly. This prediction module is highly versatile and can reduce the workload of stormwater runoff pollution load prediction.
[0059] Specifically, for at least one region of the same g-region type, the historical rainwater load data and historical feature data for all regions within that region type are used as the historical rainwater load data and historical feature data for that region type. Then, correlation data is calculated between each historical rainwater load data in that region type and each feature data in the historical feature data of the corresponding region. The degree of correlation between the feature data of each region within the region type and the historical rainwater load data of the corresponding region can be obtained, thereby allowing for screening and obtaining the feature data to be used for training in the corresponding region type.
[0060] It can be understood that the classification method in this embodiment may be clustering analysis, similarity matching, regression, etc., and is classified based on historical feature data and historical rainwater runoff pollution load data for each region, and is not specifically limited thereto.
[0061] Step S250: Train an initial prediction model based on the training target feature data and historical storm runoff pollution load data to obtain a prediction model for predicting the storm runoff pollution load in the region.
[0062] In this embodiment, an initial prediction model is first constructed, and then the initial prediction model is trained based on the training target feature data and historical stormwater runoff pollution load data to obtain a prediction model for predicting the stormwater runoff pollution load in a region.
[0063] In this embodiment, an initial prediction model can be constructed for one region, and of course, from the standpoint of versatility, an initial prediction model can be constructed for a corresponding region category for one region category, and then trained to obtain a prediction model for the corresponding region category. Subsequently, when making predictions for a target region, the region category of the target region can be detected, and the prediction model for the corresponding region category can be used to predict the rainwater runoff pollution load of the target region.
[0064] Specifically, the initial prediction model may be a random forest model, a linear regression model, a neural network model, or the like, and is not limited to any particular model here.
[0065] In this embodiment, the feature data to be trained is used as the independent variable, and the historical rainwater runoff pollution load data is used as the dependent variable. The feature data to be trained and the historical rainwater runoff pollution load data corresponding to each initial prediction model are divided into a training set and a test set (data volume ratio of 4:1). The initial prediction model is trained using the training set, and the trained prediction model is used to predict samples in the test set.
[0066] In one specific example, the initial prediction model is a random forest model. A random forest regression model is constructed using software or a language such as RStudio, SPSS, or Python. The number of decision trees in the model, ntree, is set to 500. When the decision trees branch, the number of candidate variables, mtry, is tried until an ideal value is found according to the actual situation, and the remaining parameters can be set to the model's default values. In evaluating and optimizing the model, the accuracy and generalization ability of the model, such as mean squared error, mean absolute error, and coefficient of determination, are used, and the parameters of the random forest can be adjusted according to the evaluation results to optimize the model's performance.
[0067] In this embodiment, a predictive model training method is proposed. By making full use of historical stormwater runoff pollution load data and historical feature data as basic database data, and by constructing a correlation between historical stormwater runoff pollution load data and historical feature data, non-major influencing elements in the historical feature data to the historical stormwater load data are removed, and feature data with a high correlation to the historical stormwater load data is screened. This significantly simplifies the conditions for subsequent model training, reduces data volume, lowers the data processing time cost for training, meets the timeliness needs of engineering management, and improves the reliability and accuracy of the training data, thereby improving the accuracy of model training.
[0068] Furthermore, this invention further classifies multiple regions based on historical stormwater runoff pollution load data and historical feature data using clustering analysis, trains an initial prediction model for the corresponding region type using the feature data to be trained in the classified region type, obtains a prediction model for the corresponding region category, and then obtains a corresponding prediction model based on the region category of the target region. This improves the versatility of stormwater runoff pollution load in regions, eliminates the need to repeat steps such as acquiring monitoring data and training models for each region, significantly reduces the monitoring workload and redundant work, and solves the problem of not being able to train a highly accurate prediction model in regions with little or no monitoring data.
[0069] In one embodiment, a method is further proposed for obtaining training target feature data by calculating correlation values in regional categories. Specifically, there are multiple regions, and in this case, step S230 may include steps S21 to S23. A detailed explanation is as follows.
[0070] Step S21: Clustering is performed based on the historical feature data of each region to obtain the regional category corresponding to each region.
[0071] In this embodiment, historical feature data and historical rainwater runoff pollution load data for each region are used, and a hierarchical clustering analysis method is employed to classify all target regions based on similarity.
[0072] The number of clusters is determined using a "dendritic method" (creating a dendritic structure diagram and cutting off parts of the dendritic structure to find the optimal number of clusters), the distance metric is determined according to the final clustering effect, and various distances such as Euclidean distance, Manhattan distance, and Canberra distance can be selected, and data analysis can be performed using data analysis software such as Spss, Past, and RStudio.
[0073] Based on the clustering analysis results, it is possible to analyze the typical characteristics of various types of regions (for example, from the perspectives of topography, climate, economy, and pollution) and organize the reasons for the classification.
[0074] Step S23: Based on the correlation values between historical stormwater runoff pollution load data and historical feature data in each regional category, screen the historical feature data in each regional category to obtain training feature data corresponding to each regional category, and train a predictive model for the corresponding regional category based on the training feature data and historical stormwater runoff pollution load data for each regional category.
[0075] In this embodiment, for a single regional category, a correlation value is calculated between the historical rainwater runoff pollution load data for each region in that regional category and the characteristic data for the corresponding characteristic category. The characteristic data with the highest correlation is then used as the training data for each region in the regional category. The training data for each region in a single regional category is then used as the training data for that regional category, and the training data for a regional category corresponds to the historical rainwater runoff pollution load data for a specific region in that regional category.
[0076] In one embodiment, the present invention further proposes a method for screening feature data to be trained, wherein the historical feature data includes feature data corresponding to multiple feature categories, and includes steps S31 to S32, the specific details of which are as follows.
[0077] Step S31: Obtain the correlation values between the historical rainwater runoff pollution load data and each characteristic data.
[0078] In this embodiment, correlation analysis and partial correlation analysis methods can be employed for one region or one region type to screen for feature categories that have a high impact on historical stormwater runoff pollution load data. Specifically, a correlation value is calculated between the historical stormwater runoff pollution load data and the feature data of each feature category in the corresponding region, and this correlation value may be the Pearson correlation coefficient.
[0079] In one embodiment, the historical stormwater runoff pollution load data may include historical stormwater runoff pollution load data for pollution factors such as COD, SS, NH3-N, TN, and TP. Therefore, calculating the correlation value between the stormwater runoff pollution load data and each feature data may be equivalent to calculating the correlation value between the historical stormwater runoff pollution load data for each pollution factor and each feature data.
[0080] Step S32: Based on the magnitude of the correlation value corresponding to each feature data, the feature data to be used for training is screened from the feature data corresponding to multiple feature categories.
[0081] In this embodiment, one threshold can be set. If the correlation value is greater than the threshold, the correlation between the feature data of the corresponding feature category and the historical rainwater runoff pollution load data is not significant, and the feature data can be removed. If the correlation value is less than or equal to the threshold, the correlation between the feature data of the corresponding feature category and the historical rainwater runoff pollution load data is significant, and the feature data can be used as training feature data.
[0082] In this example, we propose screening the feature data to be trained based on correlation values. This reduces the amount of data required for model training, improves model training efficiency, and removes uncorrelated data, thereby improving the reliability of the feature data to be trained and further enhancing the accuracy of the predictive model.
[0083] Based on the above-described prediction model training method, this embodiment discloses a specific example, obtaining relevant literature on stormwater runoff pollution for D1, D2, D3, D4, D5, D6, D7, D8, D9, D10, D11, D12, and D13 by querying a literature database, obtaining 584 sets of stormwater runoff pollution load data (including COD, SS, TN, and TP pollution factors) for different ground surfaces for each rainfall event, calculating historical stormwater runoff pollution load data for all rainfall events in each region by weighted averaging, obtaining corresponding historical feature data (feature data such as precipitation, precipitation duration, average rainfall intensity and preceding clear weather period, ground surface type, impermeability and gradient, average gradient, annual precipitation, annual average temperature, and economic level) from databases and yearbooks, the economic level may be shown as data such as GDP (Gross Domestic Product, Gross Domestic Product of the corresponding region), and the data period is 15 years, ensuring a sufficient amount of data and yearly representativeness.
[0084] Specifically, the data shown in Table 1 below can be obtained. Of course, Table 1 shows only a portion of the historical stormwater runoff pollution load data and historical characteristic data. In Table 1, COD, SS, TN, and TP are the historical stormwater runoff pollution load data for the corresponding pollutants.
[0085] [Table 1]
[0086] The 13 regions listed above were classified using hierarchical clustering, and combined with the data in Table 1, clustering analysis was performed using RStudio software to obtain the clustering analysis results shown in Figure 3. The number of clusters was determined using the "dendritic method," and Figure 3 clearly shows that the 13 regions can be classified into three types of regional categories. Depending on the clustering analysis results, typical characteristics of each type of region (e.g., from the perspectives of topography, climate, economy, and pollution) can be analyzed, and characteristic data for several typical characteristic categories of the corresponding regional categories can be obtained, as shown in Table 2. Of course, the data in Table 2 is also illustrative and does not represent all characteristic data for all characteristic categories of each regional category.
[0087] SS in Table 2 ave TN ave , TP ave and COD ave These correspond to historical rainwater runoff pollution load data for different pollutants.
[0088] [Table 2]
[0089] Correlation analysis was applied to each of the three regional categories to analyze the correlation between historical stormwater runoff pollution load data and the characteristic data of the corresponding characteristic category. In this embodiment, in the three regional categories, annual precipitation and rainfall per month all showed a negative correlation with historical stormwater runoff pollution load data, while average rainfall intensity per month, preceding clear weather period, and ground surface impermeability all showed a positive correlation with historical stormwater runoff pollution load data. Per capita GDP (Gross Domestic Product) showed a positive correlation only with the historical stormwater runoff pollution load data for COD. As a result, the characteristic data of the above characteristic categories were screened as important characteristic data that affects stormwater surface runoff pollution load and were selected as training characteristic data. Average gradient, average temperature, precipitation duration, and gradient did not correlate with historical stormwater runoff pollution load data and were therefore removed in subsequent model training and construction.
[0090] After obtaining the feature data to be trained for each feature category, a Random Forest model is used to model and predict rainwater runoff pollution load prediction data for single rainfall events in different regional categories. The Random Forest model is built using the RandomForest package of the RStudio software platform, with the feature data to be trained, screened by correlation analysis, as the independent variable and the historical rainwater runoff pollution load data as the dependent variable. Each set of data is split into a training set and a test set (4:1) to construct a Random Forest regression model. Debugging revealed that the model performs well when the number of decision trees (ntree) is set to 500 and the number of candidate variables (mtry) is set to 3 when the decision trees branch. The remaining parameters are all set to the model's default values, and predictive models corresponding to three regional categories are obtained.
[0091] Furthermore, we evaluated prediction models corresponding to three regional categories and obtained the trend of change in predicted and observed values of pollution factors (COD, SS, TN, and TP) for each of the three regional categories. We obtained a schematic diagram of the fitting results between predicted and observed values, as shown in Figure 4. In Figure 4, the horizontal coordinates are the observed values of rainwater runoff pollution load data actually observed in the region of the corresponding regional category, the vertical coordinates are the predicted values calculated according to the prediction model of the corresponding regional category, each point is a fitting point corresponding to the observed and predicted values, and the straight line is the fitting line between the observed and predicted values. It was found that the fitting between the predicted and observed values of these prediction models corresponding to the three regional categories is good, and the degree of fitting (R) between predicted and observed values is good. 2 The minimum fitting degree was 0.68, and most fitting degrees were 0.75 or higher, indicating that the overall predictive model for the three regional categories is relatively reasonable.
[0092] Figure 5 shows a method for predicting the pollution load of rainwater runoff in a region, as shown in Figure 5. Specifically, this method for predicting the pollution load of rainwater runoff in a region may include steps S510 to S530, and a detailed explanation is as follows.
[0093] Step S510: Obtain target feature data for the target region.
[0094] In this embodiment, when it is necessary to predict the rainwater runoff pollution load for a target area, target feature data for the target area is obtained. This target feature data includes feature data representing attribute data categories and feature data representing meteorological data categories, and specifically, historical feature data can be referenced.
[0095] Of course, in some embodiments, if the target region is the region shown in Figure 2, the historical feature data of the corresponding region can be directly selected and used for prediction.
[0096] In some embodiments, the target feature data can refer to historical feature data, and in some other embodiments, since the data predicts the stormwater runoff pollution load of the target area, unlike historical feature data, it is not necessary to obtain features and ground surface data that correlate with precipitation indicators such as precipitation amount, precipitation duration, average rainfall intensity, and preceding clear weather period of the target area. In other words, the target feature data may include only data such as average gradient, multi-year precipitation, multi-year average temperature, and economic level.
[0097] Step S530: Obtain pre-configured feature categories based on target feature data, screen the target regional feature data based on the pre-configured feature categories, input the screened target feature data into a pre-trained predictive model, and obtain rainwater runoff pollution load data for the target region output by the predictive model.
[0098] In this embodiment, the pre-configured feature categories are obtained based on the correlation values between the historical rainwater runoff pollution load data and the feature data corresponding to each feature category in the historical feature data.
[0099] Specifically, the prediction model may include multiple models, each corresponding to a different regional category, and each regional category corresponding to a different pre-configured feature category. This pre-configured feature category may be recorded when the corresponding regional category screens the feature data to be trained. In other words, when acquiring the feature data to be trained using correlation values, the feature category corresponding to the feature data to be trained is set to the pre-configured feature category corresponding to the regional category.
[0100] In one specific embodiment, step S530 may include steps S41 to S43, specifically as follows:
[0101] Step S41: Classify the target region based on the target feature data and obtain the target region category for the target region.
[0102] In this embodiment, classifying the target region may involve determining the target region category of the target region based on the regional categories corresponding to the existing prediction model. Specifically, in Figure 2, a principal coordinate analysis is performed based on the historical feature data and target feature data of the region where the prediction model is trained, and the target category to which the target region belongs is determined from among the existing regional categories.
[0103] The principal coordinate analysis may be performed using software such as RStudio, the distance metric may be selected according to the clustering effect, and the influence of differences in numerical scales between variables may be mitigated by a single numerical standardization. Of course, the principal coordinate analysis is illustrative, and other classification methods may be used in other embodiments, and are not specifically limited here.
[0104] After classification, analysis may be performed based on historical feature data within the same target region category as the target feature data to determine whether the classification of the target region is reasonable.
[0105] In some other embodiments, the classification of the target region does not have to be limited to existing regional categories. If the distance between the target region and an existing regional category is too large, other regions other than those shown in Figure 2 may be acquired. Based on the steps in Figure 2, a predictive model for the other regional categories is trained to obtain the target regional category for the target region and the predictive model corresponding to the target regional category.
[0106] Of course, the classification of target regions here refers to regions that are not used when training the predictive model. If the target region is used when training the predictive model, then it is sufficient to base the classification directly on the regional category and data of the target region during the predictive model's training.
[0107] Step S42: Based on the pre-defined feature categories in the target region category, feature data corresponding to the feature categories is screened from the target region data to obtain the screened target region feature data.
[0108] In this embodiment, the predictive feature categories are the feature categories of the training feature data in the corresponding regional category. For each regional category, there is a high correlation between the feature data of the corresponding predictive feature category and the rainwater runoff pollution load data in the region of the corresponding regional category. This allows target feature data to be screened using pre-configured feature categories, reducing the amount of data processing required for predicting rainwater runoff pollution load in the target region and improving prediction efficiency.
[0109] Of course, in some embodiments, the target feature data may not contain feature data for predefined feature categories. In this case, it is sufficient to screen the target feature data for feature data that does contain the predicted feature categories.
[0110] Step S43: Input the target regional characteristic data after screening into the trained model corresponding to the target regional category, and obtain rainwater runoff pollution load data for the target region.
[0111] Predictions of storm runoff pollution loads in target areas are similarly based on training models for the target area category, thereby improving the accuracy of predictions for storm runoff pollution loads in target areas.
[0112] In this embodiment, the rainwater runoff pollution load data output by the prediction model is hourly rainwater runoff pollution load data, which is the time unit used when training the prediction model. If the time unit is a year, the target feature data is relevant data for a specific year in which prediction of the target area is required. The corresponding prediction model outputs rainwater runoff pollution load data for the target area in that year. If it is necessary to calculate pollution amount data for the target area in that year, this can be obtained by multiplying the rainwater runoff pollution load data by the rainfall or runoff flow rate of the target area in that year.
[0113] In this embodiment, we propose a specific embodiment based on the three prediction models described above (i.e., prediction models corresponding to the three regional categories A, B, and C, which correspond to the schematic diagram of the fitting results between predicted and observed values shown in Figure 4).
[0114] In this embodiment, the target region is D9, and since the target region is the region used in the predictive model training process corresponding to three regional categories A, B, and C, the predictive model for the corresponding regional category is directly selected.
[0115] Of course, this embodiment may further analyze whether the division of D9 into three regional categories, A, B, and C, is rational by principal coordinate analysis. Euclidean distance is selected as the distance metric, and the influence of differences in numerical scales between variables is reduced by a single numerical standardization, and the principal coordinate analysis results shown in Figure 6 are obtained. The areas enclosed by solid lines are regions of regional category C, the areas enclosed by dashed lines are regions of regional category A, and the areas enclosed by dotted lines are regions of regional category B. The variance contribution rate of the first principal coordinate (i.e., the horizontal coordinate in Figure 6) is 51.12%, the variance contribution rate of the second principal coordinate is 31.99% (i.e., the vertical coordinate in Figure 6), and the cumulative variance contribution rate is 83.11%. This indicates that the dimensionality reduction effect of the information is relatively good and that the regional categories obtained by the classification of the regions are rational.
[0116] Subsequently, D9 was selected and sampling was carried out on-site, establishing a total of three sampling points: an asphalt road surface, a concrete road surface, and a green space. Storm drains collected initial rainwater, and target feature data corresponding to the three sampling points were obtained. For example, the rainfall was 20.4 mm, moderate rain, with an average rainfall intensity of 0.028 mm / min and a preceding period of clear weather of 15 days. Predictions were made using a prediction model to obtain predicted values. Water samples were collected at the openings of the storm drains within 0, 3, 6, 9, 12, 15, 20, 25, 30, and 60 minutes after runoff began to form on the road surface. Stormwater runoff pollution load data for related pollutants such as COD, SS, TN, and TP were detected. Then, the stormwater runoff pollution load data for each rainfall was calculated using the following formula to obtain measured values.
[0117]
number
[0118] The comparison between measured and predicted values is shown in Figure 7. As can be seen from Figure 7, the predicted values of the rainwater runoff pollution load data for D9 obtained using the prediction model according to this embodiment all showed an error of 15% or less compared to the measured values for three different ground surfaces, demonstrating the high accuracy of this prediction model.
[0119] Figure 8 shows a predictive model training device according to one embodiment of the present invention. As shown in Figure 8, the predictive model training device 800 specifically includes a historical data acquisition module 810 configured to acquire regional historical stormwater runoff pollution load data and regional historical feature data, wherein the historical feature data includes feature data representing attribute data categories and feature data representing meteorological data categories; a data processing module 830 configured to obtain training target feature data by screening from the historical feature data based on correlation values between historical stormwater runoff pollution load data and historical feature data; and a model training module 850 configured to train an initial predictive model based on the training target feature data and historical stormwater runoff pollution load data to obtain a predictive model for predicting regional stormwater runoff pollution load.
[0120] In one feasible form, the historical data acquisition module 810 includes a ground surface data acquisition unit configured to acquire single-rainwater runoff pollution load data for the ground surface corresponding to ground surface data, and a historical rainwater runoff pollution load data acquisition unit configured to acquire historical rainwater runoff pollution load data representing the regional rainwater runoff pollution load based on the ground surface area and single-rainwater runoff pollution load data in the ground surface data.
[0121] In one feasible configuration, there are multiple regions, and the data processing module 830 includes a clustering unit configured to perform clustering based on historical feature data for each region to obtain a regional category corresponding to each region, and a first data processing unit configured to screen the historical feature data for each regional category based on the correlation value between historical stormwater runoff pollution load data and historical feature data for each regional category to obtain training target feature data corresponding to each regional category, and to train a predictive model for the corresponding regional category based on the training target feature data and historical stormwater runoff pollution load data for each regional category.
[0122] In one feasible form, the historical feature data includes feature data corresponding to multiple feature categories, and the data processing module 830 includes a correlation processing unit configured to obtain correlation values between historical stormwater runoff pollution load data and each feature data, and a second data processing unit configured to obtain training feature data by screening the feature data corresponding to multiple feature categories based on the magnitude of the correlation value corresponding to each feature data.
[0123] The predictive model training device according to this embodiment can be used to execute the predictive model training method described above, and its implementation principle and technical effects are the same; therefore, in this embodiment, redundant explanations will not be provided here.
[0124] Figure 9 shows a regional rainwater runoff pollution load prediction device as an embodiment of the present invention. Specifically, the regional rainwater runoff pollution load prediction device 900 includes a target data acquisition module 910 configured to acquire target feature data for a target region, wherein the target data includes feature data representing attribute data categories and feature data representing meteorological data categories. The prediction module 930 is configured to acquire pre-set feature categories based on the target feature data, screen the target region feature data based on the pre-set feature categories, input the screened target region feature data into a pre-trained prediction model, and obtain target region rainwater runoff pollution load data output by the prediction model.
[0125] In one feasible form, the prediction model includes multiple prediction models, each corresponding to a different regional category, and each regional category corresponding to a different pre-configured feature category. The prediction module 930 includes a target regional category acquisition unit configured to classify target regions based on target feature data and obtain target regional categories for the target regions; a target regional feature data acquisition unit configured to screen feature data corresponding to feature categories from target regional data based on pre-configured feature categories in the target regional categories and obtain screened target regional feature data; and a prediction unit configured to input the screened target regional feature data into a trained model corresponding to the target regional category and obtain rainwater runoff pollution load data for the target regions.
[0126] The regional rainwater runoff pollution load prediction device according to this embodiment can be used to implement the above-mentioned regional rainwater runoff pollution load prediction method, and its implementation principle and technical effects are the same; therefore, in this embodiment, a redundant explanation will not be given here.
[0127] Figure 10 is a block diagram of an electronic device shown in one exemplary embodiment, and as shown in Figure 10, the electronic device 1000 may include a processor 101 and a memory 102, the processor 101 and the memory 102 being communicative, and exemplary, the processor 101 and the memory 102 being communicative via a communication bus 103, the memory 102 being used to store instructions, and the processor 101 being used to call instructions in the memory to execute the predictive model training method or the regional rainwater runoff pollution load prediction method shown in any one of the above method embodiments.
[0128] The processor may be a Central Processing Unit (CPU), another general-purpose processor, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), etc. The general-purpose processor may be a microprocessor, or the processor may be any ordinary processor, etc. The steps of the method disclosed herein may be performed by a hardware processor or directly implemented by a combination of processor hardware and software modules.
[0129] This application provides a computer-readable storage medium in which computer execution instructions are stored, and when executed by a processor, the computer execution instructions are used to realize a predictive model training method or a regional rainwater runoff pollution load prediction method in any one of the above embodiments.
[0130] Embodiments of the present invention provide a computer program product which includes instructions, and when these instructions are executed, causes the computer to execute the above-described predictive model training method or regional rainwater runoff pollution load prediction method.
[0131] A person skilled in the art will readily conceive of other means of carrying out the application disclosed herein by considering the specification. The application is intended to cover any variations, uses, or adaptive changes of the application, which may include well-known or commonly used means in the art that are not disclosed herein, in accordance with the general principles of the application. The specification and examples are to be considered merely illustrative, and the actual scope and intent of the application are set forth in the following claims.
[0132] It should be understood that this application is not limited to the exact structure described above and shown in the drawings, and that various modifications and changes can be made without exceeding its scope. The scope of this application is limited to the attached claims.
Claims
1. A predictive model training method, A step of obtaining historical stormwater runoff pollution load data and historical characteristic data of the region, wherein the historical characteristic data includes characteristic data representing attribute data categories and characteristic data representing meteorological data categories, Based on the correlation value between the historical rainwater runoff pollution load data and the historical feature data, the step of screening the historical feature data to obtain feature data to be used for training, A method for training a predictive model, comprising the steps of: training an initial predictive model based on the characteristic data of the training target and the historical rainwater runoff pollution load data to obtain a predictive model for predicting the rainwater runoff pollution load of the region.
2. The aforementioned historical feature data includes the ground surface data of the region, and the step of obtaining the aforementioned historical rainwater runoff pollution load data of the region is, The steps include obtaining single-cycle rainwater runoff pollution load data for the ground surface corresponding to the aforementioned ground surface data, The method according to claim 1, comprising the step of obtaining historical rainwater runoff pollution load data representing the rainwater runoff pollution load of the region based on the area of the ground surface in the ground surface data and the rainwater runoff pollution load data for a single instance.
3. The number of the aforementioned regions is multiple, and the step of obtaining the training target feature data by screening from the historical feature data based on the correlation value between the historical rainwater runoff pollution load data and the historical feature data is, The steps include: performing clustering based on historical feature data for each region to obtain a regional category corresponding to each region; The method according to 1 or 2, comprising the steps of: screening from the historical feature data in each regional category based on the correlation value between the historical stormwater runoff pollution load data and the historical feature data in each regional category to obtain the feature data to be trained corresponding to each regional category; and training to obtain the predictive model for the corresponding regional category based on the feature data to be trained for each regional category and the historical stormwater runoff pollution load data.
4. The aforementioned historical feature data includes feature data corresponding to multiple feature categories, and the step of obtaining the feature data to be trained by screening the historical feature data based on the correlation value between the historical rainwater runoff pollution load data and the historical feature data is as follows: The steps include obtaining correlation values between the historical rainwater runoff pollution load data and each characteristic data, The method according to any one of claims 1 to 3, comprising the step of obtaining the training target feature data by screening from the feature data corresponding to the plurality of feature categories based on the magnitude of the correlation value corresponding to each feature data.
5. A method for predicting the pollution load of rainwater runoff in a region, A step of obtaining target feature data for a target region, wherein the target feature data includes feature data representing attribute data categories and feature data representing weather data categories. A method for predicting rainwater runoff pollution load in a region, comprising the steps of: obtaining pre-configured feature categories based on the target feature data; screening target region feature data based on the pre-configured feature categories; inputting the screened target feature data into a pre-trained prediction model; and obtaining rainwater runoff pollution load data for the target region output by the prediction model, wherein the pre-configured feature categories are obtained based on correlation values between historical rainwater runoff pollution load data and feature data corresponding to each feature category in the historical feature data.
6. The aforementioned prediction model includes multiple models, each corresponding to a different regional category, and each regional category corresponding to a different pre-configured feature category. The steps include obtaining the pre-configured feature category based on the target feature data, screening the target regional feature data based on the pre-configured feature category, inputting the screened target feature data into the pre-trained prediction model, and obtaining the rainwater runoff pollution load data for the target region output by the prediction model. The steps include: classifying the target region based on the target characteristic data and obtaining the target region category for the target region; The steps include: screening the target region data for feature data corresponding to the pre-defined feature categories in the target region category, and obtaining the target region feature data after screening; The method according to claim 5, further comprising the step of inputting the target area characteristic data after screening into a trained model corresponding to the target area category, thereby obtaining rainwater runoff pollution load data for the target area.
7. A predictive model training device, A historical data acquisition module configured to acquire historical stormwater runoff pollution load data and historical characteristic data of a region, wherein the historical characteristic data includes characteristic data representing attribute data categories and characteristic data representing meteorological data categories, A data processing module configured to obtain training target feature data by screening the historical feature data based on the correlation value between the historical rainwater runoff pollution load data and the historical feature data, A predictive model training device comprising a model training module configured to train an initial predictive model based on the characteristic data of the training target and the historical rainwater runoff pollution load data, thereby obtaining a predictive model for predicting the rainwater runoff pollution load of the region.
8. A regional rainwater runoff pollution load prediction device, A target data acquisition module configured to acquire target feature data for a target region, wherein the target data includes feature data representing attribute data categories and feature data representing weather data categories, A regional rainwater runoff pollution load prediction device, comprising: a prediction module configured to acquire pre-configured feature categories based on the target feature data, screen the target regional feature data based on the pre-configured feature categories, input the screened target regional feature data into a pre-trained prediction model, and obtain rainwater runoff pollution load data for the target region output by the prediction model.
9. An electronic device comprising a processor and a memory, wherein the memory is used for storing instructions, and the processor is used to cause the electronic device to execute the predictive model training method according to any one of claims 1 to 4 or the regional rainwater runoff pollution load prediction method according to any one of claims 5 to 6 by executing the instructions in the memory.
10. A computer-readable storage medium that stores computer execution instructions, and is used to implement the predictive model training method described in any one of claims 1 to 4 or the rainwater runoff pollution load prediction method for the region described in any one of claims 5 to 6 when the computer execution instructions are executed by a processor.
11. A computer program product comprising a computer program, wherein when the computer program is executed by a processor, it realizes the predictive model training method described in any one of claims 1 to 4 or the rainwater runoff pollution load prediction method for the region described in any one of claims 5 to 6.