A Crowdsourcing Rainfall Data Quality Control Method
By extracting multiple statistical features of rainfall observation points and machine learning algorithm training models, the subjectivity and large-scale data processing problems of crowdsourcing rainfall data quality judgment are solved, and the data quality control capabilities of hydrological simulation and smart water conservancy are improved.
Patent Information
- Application Number
- CN202411461603.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-10-18
AI Technical Summary
In the prior art, the quality judgment of crowdsourcing rainfall data is subjective and cannot automatically process large-scale data, which affects the application of hydrological simulation and smart water conservancy.
By extracting multiple statistical features of rainfall observation points, a crowdsourcing rainfall data set with mathematical attributes is established, and a machine learning algorithm is used to train a data quality control model, combining SHAP and additional tree algorithms to sort performance impact factors, and calibrate the model to improve data quality.
It realizes automated control of crowdsourcing rainfall data quality, improves the accuracy of hydrological simulation and the data governance capabilities of smart water conservancy, and reduces the subjectivity of data processing.
Smart Images

Figure CN119443918B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data analysis, and particularly relates to a crowdsourcing rainfall data quality control method. Background Art
[0002] Rainfall is the most active link in the water cycle process and is also a key input to the basin hydrological system. Its spatio-temporal accuracy directly affects the accuracy of hydrological simulation results. The abnormal spatio-temporal distribution of precipitation is also an important factor leading to natural disasters such as floods and landslides. Therefore, effective and accurate observation of the spatio-temporal distribution characteristics of rainfall plays an extremely important role in early warning and decision-making in aspects such as hydrology, meteorology, and agriculture. The main ways to obtain traditional rainfall data include station, radar, and satellite observations, etc. Their spatial and temporal resolutions have their own characteristics, but most of them are not suitable for research on small and medium-scale hydrological simulations that require high spatio-temporal rainfall field data. The acquisition of a high spatio-temporal rainfall field can effectively improve the accuracy and predictability of flood process simulations, increase the power generation of hydropower stations, and generate significant economic benefits. For urban construction management, the acquisition of a high spatio-temporal rainfall field can meet the needs of meteorological fine services in urbanized areas and improve the modernization level of urban disaster prevention, mitigation, and emergency management.
[0003] In the big data era, with the rapid development of the Internet of Things, non-traditional and diverse data acquisition technologies such as data crowdsourcing have promoted the generation of massive data. Different from professional data collection methods, data crowdsourcing mainly involves the public's participation in the data collection process. Currently, data crowdsourcing has received more and more research worldwide, and crowdsourcing rainfall data has the potential to make up for this shortcoming of traditional monitoring methods. Therefore, how to use widely available crowdsourcing data in the big data era to promote the improvement of the spatio-temporal resolution of rainfall observations is the most important and key issue. At the same time, the quality control of crowdsourcing data has become the key to determining the further application of crowdsourcing science. Summary of the Invention
[0004] Based on this, it is necessary to provide a crowdsourcing rainfall data quality control method to solve the subjectivity of previous crowdsourcing data quality judgment and the problem of being unable to automatically process large-scale data, and provide key technical support for the collection and application of hydrological crowdsourcing data, data governance of smart water conservancy, etc.
[0005] A crowdsourcing rainfall data quality control method includes the following steps:
[0006] S1. Extract multiple rainfall data statistical features of rainfall observation points to establish a crowdsourcing rainfall data set with mathematical attributes;
[0007] S2. Based on a machine learning algorithm, train and preprocess the rainfall data statistical features in the crowdsourcing rainfall data set to train a crowdsourcing rainfall data quality control model;
[0008] S3. Rank the importance of the performance impact factors of the crowdsourced rainfall data quality control model to calibrate the crowdsourced rainfall data quality control model.
[0009] In one embodiment, the step S1 of extracting multiple statistical characteristics of rainfall data includes:
[0010] S11. Based on the areal rainfall fields with different coverage ranges centered on the same preset rainfall observation point, with the preset rainfall observation point as the center, extract the rainfall mathematical attributes of different-sized windows, where the rainfall mathematical attributes include at least one of variance, maximum value, minimum value, range, average value, standard deviation, and absolute deviation;
[0011] S12. Based on the areal rainfall fields within the coverage ranges at different radial distances centered on the same preset rainfall observation point, extract the rainfall mathematical correlation characteristics between the preset rainfall observation point and other rainfall observation points, where the rainfall mathematical correlation characteristics include at least one of difference, change range, and gradient;
[0012] S13. Based on the rainfall fields formed at different times centered on the same preset rainfall observation point, extract the correlation coefficients of adjacent two areal rainfall fields.
[0013] In one embodiment, the machine learning algorithm includes supervised machine learning algorithms and unsupervised machine learning algorithms;
[0014] The step S2 includes:
[0015] S21. Based on a supervised machine learning algorithm, in the way of using calibration error, randomly add noise values to the rainfall observation points by introducing noise values and calibrate them, and train the multiple statistical characteristics of rainfall data of the learning error rainfall observation points;
[0016] S22. Based on an unsupervised machine learning algorithm, introduce the noise data into the multiple statistical characteristics of rainfall data for automatic training and preprocessing.
[0017] In one embodiment, the step S3 includes:
[0018] S31. Use the SHAP interpretable algorithm to quantitatively interpret and rank the performance impact factors of the crowdsourced rainfall data quality control model;
[0019] S32. Use the extra trees algorithm to identify the best combination of performance impact factors of the crowdsourced rainfall data quality control model.
[0020] In one embodiment, in step S31, the calculation of the SHAP value is based on the following formula:
[0021]
[0022] Among them, f(x) is the prediction of the model for the input x; x i is the i-th feature of the input x; M is the number of features; φ0 is a constant representing the feature base value; φ i is the SHAP value of the i-th feature, representing the contribution of this feature to the prediction f(x);
[0023]
[0024] Among them, S is the feature set, representing the subset of features considered during model prediction; |S| is the number of features in the set S; f x (S) is the prediction of the model when only considering the features in the set S; f x (S ∪ {i}) is the prediction of the model when considering the features in the set S and the feature i.
[0025] In one embodiment, in step S1, a part of the original rainfall data of the multiple rainfall data statistical features is directly collected by a fixed sensor, and another part of the original rainfall data of the multiple rainfall data statistical features is obtained by preset processing of the data directly collected by a movable sensor.
[0026] In one embodiment, the steps of preset processing the data directly collected by the movable sensor include:
[0027] Couple the mobile path model associated with the movable sensor into the data quality control architecture to simulate the dynamic movement path of the participant, and achieve the deep coupling of the mobile path model and the crowdsourcing rainfall quality control model. The control equations and parameters involved are as follows:
[0028]
[0029]
[0030] Among them, S i,t represents the position of participant i at time step t; L i,t represents the moving distance of participant i at time step t; θ i,t represents the moving direction of participant i at time step t; x and y are the coordinates of the rainfall observation point in space respectively; α and φ represent the moving distance L i,t and the moving direction θ i,t respectively; ω i,t represents the boolean moving state factor of participant i at time step t; the moving distance L i,t is assumed to follow L i,t~U(0,L max ), U() represents a uniform distribution, and L max represents the maximum distance that the participant can move within the time step; the movement direction θ i,t is assumed to follow θ i,t ~U(0, 2π), represents the movement probability of the rainfall observation point; and represent the maximum movement probability and the minimum movement probability of the rainfall observation point respectively; cx and cy represent the positions of the center of the rainfall observation point in the x and y directions respectively; r represents the radius of the center of the rainfall observation point; ζp represents the observation density gradient from the center of the rainfall observation point to the edge of the study area.
[0031] In one of the embodiments, it further includes:
[0032] S4. Use a preset evaluation index to evaluate the prediction result of the crowdsourced rainfall data quality control model.
[0033] In one of the embodiments, the evaluation index includes WCFG and WCPJ. WCFG is used to describe the ability to capture rainfall variability on the hourly and spatial scales, and WCPJ is used to describe the relative deviation of the total rainfall amount of the rainfall event captured by the site within the duration and spatial coverage of the rainfall;
[0034] The calculation formulas of WCFG and WCPJ are as follows respectively:
[0035]
[0036] Among them, is the rainfall intensity estimated based on the crowdsourced observation at time t; y i is the relevant ground truth rainfall intensity; i represents the serial number of the rainfall observation point.
[0037] In one of the embodiments, it further includes:
[0038] S5. Use the ΔWCFG and ΔWCPJ evaluation indexes to evaluate the performance improvement of the crowdsourced rainfall data quality control model after noise removal;
[0039] The calculation formulas of ΔWCFG and ΔWCPJ are as follows respectively:
[0040]
[0041] Among them, WCFG before represents the WCFG value of the crowdsourced rainfall intensity before noise removal; WCFG after represents the WCFG value of the crowdsourced rainfall intensity after noise removal; WCPJ beforeIndicates the WCPJ value of crowdsourced rainfall intensity before noise rejection; WCPJ after Indicates the WCPJ value of crowdsourced rainfall intensity after noise rejection.
[0042] The crowdsourced rainfall data quality control method provided by this application first extracts multiple statistical characteristics of rainfall data at rainfall observation points to establish a crowdsourced rainfall dataset with mathematical attributes. Then, based on machine learning algorithms, it trains and preprocesses the statistical characteristics of rainfall data in the crowdsourced rainfall dataset to obtain a crowdsourced rainfall data quality control model. Finally, by ranking the importance of the performance impact factors of the crowdsourced rainfall data quality control model, the crowdsourced rainfall data quality control model is calibrated. The crowdsourced rainfall data quality control model obtained through the above training and calibration solves the subjectivity of previous crowdsourced data quality judgment and the problem of being unable to automatically process large-scale data, providing key technical support for the collection and application of hydrological crowdsourced data, data governance of smart water conservancy, etc. At the same time, it is of great significance for the popularization and application of big data, artificial intelligence, and Internet of Things technologies in the construction of smart water services. Brief Description of the Drawings
[0043] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0044] Figure 1 Is a flowchart of a crowdsourced rainfall data quality control method in an embodiment;
[0045] Figure 2 Is an interpolated surface rainfall field diagram in an embodiment;
[0046] Figure 3 Is a surface rainfall amount field diagram of identified noise observations in an embodiment;
[0047] Figure 4 Is a distribution diagram of the importance values of the top 20 performance impact factors of the crowdsourced rainfall data quality model in an embodiment;
[0048] Figure 5 Is an interaction influence diagram of the change ratio of WCFG, the change ratio of WCPJ, and the noise level and noise amount of the crowdsourced rainfall data quality model in an embodiment;
[0049] Figure 6 Is a distribution diagram of the change ratio of WCFG, the change ratio of WCPJ, and the AUC value of the crowdsourced rainfall data quality control model trained under different scenario conditions in an embodiment;
[0050] Figure 7 The distribution diagrams of the change ratio of WCFG, the change ratio of WCPJ, and the AUC value of another crowdsourced rainfall data quality control model trained under different scenario conditions in an embodiment;
[0051] Figure 8 The distribution diagrams of ΔWCFG and ΔWCPJ values of a crowdsourced rainfall data quality control model trained based on different machine learning algorithms under the same benchmark scenario in an embodiment. Detailed implementation manners
[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0053] It should be noted that all directional indications (such as up, down, left, right, front, back...) in the embodiments of the present invention are only used to explain the relative position relationship and movement conditions between components in a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly.
[0054] In addition, the descriptions involving "first", "second", etc. in the present invention are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, "and / or" throughout the text includes three solutions. Taking A and / or B as an example, it includes the technical solution of A, the technical solution of B, and the technical solution that both A and B are satisfied at the same time. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0055] As Figure 1 shown, the present application provides a method for controlling the quality of crowdsourced rainfall data, including the following steps:
[0056] S1. Extract multiple statistical characteristics of rainfall data at rainfall observation points to establish a crowdsourced rainfall data set with mathematical attributes;
[0057] S2. Based on a preset machine learning algorithm, train and preprocess the statistical characteristics of rainfall data in the crowdsourced rainfall data set to train a crowdsourced rainfall data quality control model;
[0058] S3. Rank the importance of the performance impact factors of the crowdsourced rainfall data quality control model to calibrate the crowdsourced rainfall data quality control model.
[0059] The crowdsourced rainfall data quality control method provided in this application first extracts multiple rainfall data statistical features of rainfall observation points to establish a crowdsourced rainfall data set with mathematical attributes, and then trains and preprocesses the rainfall data statistical features in the crowdsourced rainfall data set based on machine learning algorithms to obtain a crowdsourced rainfall data quality control model. Finally, by ranking the importance of the performance impact factors of the crowdsourced rainfall data quality control model, the crowdsourced rainfall data quality control model is calibrated. The crowdsourced rainfall data quality control model obtained through the above training and calibration solves the subjectivity of previous crowdsourced data quality judgment and the problem of being unable to automatically process large-scale data, providing key technical support for the collection and application of hydrological crowdsourced data, data governance of smart water conservancy, etc. At the same time, it is of great significance for the popularization and application of big data, artificial intelligence, and Internet of Things technologies in the construction of smart water services.
[0060] Specifically, establishing the above-mentioned crowdsourced rainfall data set with mathematical attributes provides a data basis for subsequent training of the crowdsourced rainfall data quality control model. Optionally, step S1 of extracting multiple rainfall data statistical features includes:
[0061] S11. Based on the areal rainfall fields with different coverage ranges centered on the same preset rainfall observation point, taking the preset rainfall observation point as the center of the circle, extract the rainfall mathematical attributes of different-sized windows. The rainfall mathematical attributes include at least one of variance, maximum value, minimum value, range, average value, standard deviation, and absolute deviation;
[0062] S12. Based on the areal rainfall fields within the coverage ranges at different radial distances centered on the same preset rainfall observation point, extract the rainfall mathematical correlation features between the preset rainfall observation point and other rainfall observation points. The rainfall mathematical correlation features include at least one of difference, change range, and gradient;
[0063] S13. Based on the rainfall fields formed at different times centered on the same preset rainfall observation point, extract the correlation coefficients of adjacent two areal rainfall fields.
[0064] Furthermore, in step S1, a part of the original rainfall data of the multiple rainfall data statistical features is directly collected by fixed sensors (such as fixed cameras), and another part of the original rainfall data of the multiple rainfall data statistical features is obtained by preset processing of the data directly collected by mobile sensors (such as umbrella sensors).
[0065] The steps for performing preset processing on the data directly collected by the mobile sensor include:
[0066] Couple the movement path model associated with the mobile sensor into the data quality control architecture to simulate the dynamic movement path of the participant, realizing the deep coupling of the movement path model and the crowdsourcing rainfall quality control model. The control equations and parameters involved are as follows:
[0067]
[0068] Among them, S i,t represents the position of participant i at time step t; L i,t represents the movement distance of participant i at time step t; θ i,t represents the movement direction of participant i at time step t; x and y are the coordinates of the rainfall observation point in space respectively; α and φ represent the adjustment factors of the movement distance L i,t and the movement direction θ i,t respectively. The initial values of α and φ are defaulted to 1 and 0 respectively. When S i,t+1 is outside the study area, the values of α and φ will be adjusted so as to move only within the boundary of the study scope; ω i,t represents the Boolean movement state factor of participant i at time step t; the movement distance L i,t is assumed to follow L i,t ~U(0, L max ), where U() represents following a uniform distribution, and L max represents the maximum distance that the participant can move within the time step; the movement direction θ i,t is assumed to follow θ i,t ~U(0, 2π), represents the movement probability of the rainfall observation point; and represent the maximum movement probability and the minimum movement probability of the rainfall observation point respectively; cx and cy represent the positions of the center of the rainfall observation point in the x and y directions respectively; r represents the radius of the center of the rainfall observation point; ζp represents the observation density gradient from the center of the rainfall observation point to the edge of the study area.
[0069] Specifically, as described above, the rainfall data may be directly collected by fixed devices (such as fixed cameras, which are immovable and have a constant position), that is, fixed sensors are arranged at fixed rainfall observation points for rainfall data collection;
[0070] Rainfall data may also be obtained through preset processing of data collected by devices equipped with movable sensors such as umbrellas or cars (the devices are movable). In this case, the rainfall observation points are variable, and the devices equipped with movable sensors are also controlled by participants. For example, a participant holds an umbrella equipped with an umbrella sensor and moves it to collect rainfall data during the movement of the umbrella sensor. Another example is that a participant drives a car equipped with a car wiper sensor to move forward to collect rainfall data during the movement of the car wiper sensor.
[0071] Furthermore, rainfall observations are usually point observations. Therefore, the calculation of areal rainfall requires interpolation processing. Specifically, for the method of rainfall spatial interpolation, whether it is global interpolation (such as boundary interpolation, trend surface analysis) or local interpolation (such as Kriging method, Thiessen polygon), it is based on the geographical location relationship between the interpolation points and the rainfall observation points, and uses the interpolation method to interpolate the rainfall data.
[0072] Furthermore, the purpose of the machine learning algorithm is to label each crowdsourced rainfall observation result as a noise observation result or a normal observation result. A noise observation result means that the observed data shows anomalies. For example, if a tipping bucket rain gauge is broken, it will report a zero value. In the case where there is data in the surrounding rainfall observations, this will be identified as an incorrect observation, that is, a noise observation.
[0073] In step S2, the statistical characteristics of each rainfall data in the crowdsourced rainfall dataset are trained through a machine learning algorithm. It can be understood that the statistical characteristics of each rainfall data are used as independent variables, and each crowdsourced rainfall observation result is used as a dependent variable, and the machine learning algorithm is used to identify the complex relationship between the statistical characteristics of rainfall data and the crowdsourced rainfall observation results.
[0074] Figure 2 It is the areal rainfall field map of interpolation in an embodiment. Figure 3 It is the areal rainfall field map of identifying noise observations in an embodiment. In Figure 3 , the noise observation points are represented by black triangles. The machine learning algorithm includes supervised machine learning algorithms and unsupervised machine learning algorithms. Specifically, the supervised machine learning algorithms include multi-level neural networks and K-nearest neighbor algorithms, and the unsupervised machine learning algorithms include isolation forest and K-means clustering algorithms. Step S2 includes:
[0075] S21. Based on the supervised machine learning algorithm, in the way of using the calibration error, randomly add noise values to the rainfall observation points by introducing noise values and calibrate them, and train the statistical characteristics of multiple rainfall data of the learning error rainfall observation points;
[0076] Specifically, step S21 refers to introducing outliers artificially into normal observations based on a supervised machine learning algorithm, and then calibrating them as abnormal observations for supervised training to learn the characteristics of such abnormalities; that is, based on the known feature mapping relationship of noise points (i.e., the mapping relationship between rainfall data statistical characteristics and crowdsourced rainfall observation results), training is carried out through a supervised machine learning algorithm to identify abnormal characteristics (such as sudden 0 values (possibly due to rain gauge blockage), or extremely abnormal high values), and then using these abnormal characteristics to identify noise points and then removing them.
[0077] S22. Based on an unsupervised machine learning algorithm, introduce noise data into multiple rainfall data statistical characteristics for automatic training and preprocessing.
[0078] Specifically, step S22 refers to introducing outliers artificially into normal observations based on an unsupervised machine learning algorithm, without the need for calibration, and automatically identifying noise data; that is, unsupervised training means not knowing the feature mapping relationship of noise points in advance, ranking the importance of various rainfall data statistical characteristics, and then selecting the characteristics with higher importance for model application to remove and identify noise points. Here, the model refers to a crowdsourced rainfall data quality control model trained with historical data, and this model can be directly used to predict noise observations of subsequent rainfall.
[0079] Specifically, a random data splitting method can be used to train the supervised learning algorithm and the unsupervised learning algorithm. 70% of the collected data is used for training, and 30% is used for testing. Among them, the training set is used to train the model, and the test set is used to verify the prediction accuracy of the model. The 5-Folds cross-validation method is used for training to avoid overfitting. Before training, the input feature values are normalized to 0-1 through the min-max scaling method. The step of feature normalization enables the trained crowdsourced rainfall data quality control model to exclude the influence of the data ranges of different rainfall elements, making it transferable and comparable between different regions and scenarios, even if their original feature values are in different ranges.
[0080] Specifically, the formula involved in the min-max standard normalization method is as follows:
[0081]
[0082] Among them, X represents the value of the current data point; X norm represents the current data value after normalization; X min represents the minimum value of the current data set; X max represents the maximum value of the current data set.
[0083] In addition, the zero-mean normalization method can be used for normalization processing, and the specific formula involved is as follows:
[0084]
[0085] Among them, μ and σ respectively represent the mean and variance of the original training data set.
[0086] Optionally, data augmentation is performed on the data in the training data set. For example, data augmentation can be performed by means of cropping, adding noise, etc. Using data augmentation techniques mainly adds minor perturbations or changes to the training data. On the one hand, it can increase the training data, thereby improving the generalization ability of the model. On the other hand, it can increase the noise data, thereby enhancing the robustness of the model.
[0087] In one embodiment, step S3 includes:
[0088] S31. Using the SHAP explainable algorithm to quantitatively interpret and rank the performance impact factors of the crowdsourcing rainfall data quality control model;
[0089] Specifically, the contribution degree of each performance impact factor of the crowdsourcing rainfall data quality control model to the prediction of the crowdsourcing rainfall data quality control model is calculated through the SHAP explainable algorithm, so that it can be understood how each performance impact factor affects the prediction of the crowdsourcing rainfall data quality control model.
[0090] In step S31, the calculation of the SHAP value is based on the following formula:
[0091]
[0092] Among them, f(x) is the prediction of the model for the input x; x i is the i-th feature of the input x; M is the number of features; φ0 is a constant, representing the feature base value; φ i is the SHAP value of the i-th feature, representing the contribution of this feature to the prediction f(x);
[0093]
[0094] Among them, S is the feature set, representing the subset of features considered when the model makes a prediction; |S| is the number of features in the set S; f x (S) is the prediction of the model when only considering the features in the set S; f x (SU{i}) is the prediction of the model when considering the features in the set S and the feature i. Specifically, x in the SHAP value calculation formula i refers to the rainfall data statistical features in the present invention, and the calculation result of this formula is the importance ranking of each rainfall data statistical feature in Figure 4
[0095] S32. Use the Extra Trees algorithm to identify the best combination of performance impact factors for the crowdsourced rainfall data quality control model.
[0096] The Extra Trees algorithm is an ensemble learning method. By constructing decision trees and selecting feature subsets for node splitting, its principle can be summarized as randomly selecting m features from the feature set F, where m = |F|. For the selected feature f, a random split point split is chosen. f , and according to the split point, the data set is divided into two subsets D left and D right , where:
[0097] D left = {x ∈ D|x f ≤ split f}
[0098] D right = {x ∈ D|x f > split f}
[0099] For the noise judgment problem involved in this case, the Extra Trees algorithm is obtained through the majority voting of the prediction results of all trees.
[0100]
[0101] where I is the indicator function, and if then it is 1, otherwise it is 0. Specifically, in the present invention, the feature set F represents multiple rainfall data statistical features. The feature selection step iteratively deletes the least important performance impact factors (measured by the feature importance in the Extra Trees) until the prediction performance of the crowdsourced rainfall data quality model decreases, and the best combination of performance impact factors is selected through the accuracy rate.
[0102] By using the prediction accuracy value of the crowdsourced rainfall data quality model as the measurement feature of the performance of the crowdsourced rainfall data quality model, according to the feature selection results, the number of performance impact factors selected by the supervised learning algorithm is 20 (the 20 features in Figure 4 ), and the number of performance impact factors selected by the unsupervised learning algorithm is 5 (the first 5 features in Figure 4 ), Figure 4It is a distribution diagram of the importance values of the top 20 performance impact factors of the crowdsourced rainfall data quality model in an embodiment. The importance values of the top 20 performance impact factors of the crowdsourced rainfall data quality control model are calculated through the extra trees algorithm, and the importance ranking of the top 20 performance impact factors of the crowdsourced rainfall data quality control model is carried out according to the calculation results of the importance values. Among all the candidate performance impact factors, Iad, C, and Iad,w_9 are the top two contributing performance impact factors. These two performance impact factors are both related to the difference in the average rainfall conditions of the target observation value within a certain spatial area, and this difference can fully represent the mutation brought by the noisy observation.
[0103] Table 1. Definitions of each performance impact factor of the crowdsourced rainfall data quality model
[0104]
[0105]
[0106] Referring to Table 1, Table 1 is the definition of each performance impact factor of the crowdsourced rainfall data quality model. The maximum and minimum values of the window domain can be defined as the maximum rainfall intensity and the minimum rainfall intensity within different window ranges respectively;
[0107] The calculation formula for the standard deviation value of the window domain is as follows:
[0108]
[0109] Among them, σ represents the standard deviation, N represents the total number of values in the data set, xi represents each value in the data set, and μ represents the average value of the data set.
[0110] The calculation formula for the average value of the window domain is as follows:
[0111] average = sum(Rn) / n
[0112] Among them, Rn represents the rainfall intensity of each rainfall observation point, and n represents the number of observation points.
[0113] It should be noted that I av,C (absolute value) refers to: the absolute value of the difference between the value range within the preset coverage and the sampling point. Among them, the value range is the range interval of the maximum rainfall intensity and the minimum rainfall intensity of each rainfall observation point within the preset coverage, and the sampling point difference is the absolute value of the difference between the value of this monitoring point and the value range. The relevant formula is as follows:
[0114]
[0115] For example, within a certain range, the rainfall intensity values at each rainfall observation point are (1, 3, 7, 5, 8, 10) respectively, and the value range is 10 - 1 = 9; for the point where the observed value at the sampling point is 3, I av,c = |3 - 9| = 6, indicating whether the sample point value is close to the interval extreme value.
[0116] It should be noted that the last digit in the mathematical symbol of the performance impact factor represents the size of the window range of the window-based feature. For example, for the performance impact factor I ad,w_9 , the number 9 represents that the size of the window-based range is 9 * 9 spatial grids.
[0117] Optionally, the crowdsourced rainfall data quality control method further includes:
[0118] S4. Use a preset evaluation index to evaluate the prediction result of the crowdsourced rainfall data quality control model.
[0119] Specifically, the evaluation indexes include WCFG (root mean square error) and WCPJ (relative mean error). WCFG is used to describe the ability to capture rainfall variability on the hourly and spatial scales, and WCPJ is used to describe the relative deviation of the total rainfall amount of the rainfall event captured by the site within the duration and spatial coverage of the rainfall.
[0120] The calculation formulas of WCFG and WCPJ are as follows respectively:
[0121]
[0122]
[0123] Among them, is the rainfall intensity estimated according to the crowdsourced observation at time t; y i is the relevant ground truth rainfall intensity; i represents the serial number of the rainfall observation point.
[0124] Furthermore, the crowdsourced rainfall data quality control method further includes:
[0125] S5. Use the evaluation indexes of ΔWCFG (change ratio of WCFG) and ΔWCPJ (change ratio of WCPJ) to evaluate the performance improvement of the crowdsourced rainfall data quality control model after noise removal;
[0126] The calculation formulas of ΔWCFG and ΔWCPJ are as follows respectively:
[0127]
[0128] Among them, WCFG before represents the WCFG value of the crowdsourced rainfall intensity before noise removal; WCFG afterThe WCFG value of the crowdsourced rainfall intensity after noise rejection; WCPJ before The WCPJ value of the crowdsourced rainfall intensity before noise rejection; WCPJ after The WCPJ value of the crowdsourced rainfall intensity after noise rejection.
[0129] In the analysis of the interactive effects of noise level and the number of noise observations, as Figure 5 shown, the change ratio (specifically, the reduction ratio) of WCFG and the change ratio (specifically, the reduction ratio) of WCPJ increase with the increase of noise level and noise amount. Specifically, the reduction ratios of WCFG and WCPJ are as high as 55.49% and 69.25% respectively, and are achieved at relatively high noise levels and noise amounts. At the same time, there is no clear interaction between the noise level and the application impact of the noise amount on the crowdsourced rainfall data quality control method of the present invention. This observation can be partly explained by the relatively stable ability of the crowdsourced rainfall data quality control method of the present invention in identifying noisy crowdsourced observations.
[0130] The crowdsourced rainfall data quality control method of the present invention has portability, which means directly applying the crowdsourced rainfall data quality control model trained under the benchmark scenario to other regions or other noise level scenarios. The transfer application of the crowdsourced rainfall data quality control model helps to solve the problem of low data quality or data scarcity in some regions where retraining is not possible, while saving computational costs. In the transfer application of the crowdsourced rainfall data quality control model, the crowdsourced rainfall data quality control model trained in one region is transferred to a region with significantly different rainfall climate characteristics. The results of the transfer application show that the crowdsourced rainfall data quality control method still plays a role in reducing the error of the rainfall field and identifying noise observations.
[0131] Figure 6 Distribution diagrams of the change ratio of WCFG, the change ratio of WCPJ, and the AUC value of the crowdsourced rainfall data quality control model trained under different scenario conditions in an embodiment; It can be seen from the figure that the S1 benchmark scenario shows good performance in both the error reduction ratio and AUC. Among them, the average reduction of WCFG is close to 46.2%, the average reduction of WCPJ is close to 48.6%, and the AUC judgment value is close to 0.7, demonstrating good recognition effects. In addition, the transfer effect in S2 - S4 is slightly worse than the local training application effect, and the variance of the transfer application effect for different regions is also slightly larger, but it can still reduce errors to a certain extent and ensure a certain recognition effect; Figure 7Distribution of the change ratio of WCFG, the change ratio of WCPJ, and the AUC value for another crowdsourced rainfall data quality control model trained under different scenario conditions in an embodiment; where S1 represents the crowdsourced rainfall data quality control model trained under the baseline scenario (abbreviated as the baseline crowdsourced rainfall data quality control model); S2 represents the baseline crowdsourced rainfall data quality control model migrated and applied to Region 1; S3 represents the baseline crowdsourced rainfall data quality control model migrated and applied to Region 2; S4 represents the baseline crowdsourced rainfall data quality control model migrated and applied to Region 3; it should be noted that Figure 6 and Figure 7 The machine learning algorithms relied on by the crowdsourced rainfall data quality control models obtained through training are different. Specifically, Figure 6 and Figure 7 The crowdsourced rainfall data quality control models shown respectively rely on the kNN and MLPs algorithms for training.
[0132] In machine learning, AUC is an important feature for evaluating the performance of classification models. Especially in binary classification problems, the AUC value is the non-parametric estimate of the area under the ROC curve. The ROC curve is constructed by plotting the True Positive Rate (TPR) against the False Positive Rate (FPR) at different thresholds. The calculation method of AUC can be summarized as:
[0133]
[0134] where TPR is the True Positive Rate, specifically FPR is the False Positive Rate, specifically TP is the true positive, FN is the false negative, FP is the false positive, and TN is the true negative. Specifically, TP, FN, FP, and TN are the confusion matrix in the machine learning algorithm; in the present invention, TP (true positive) is the normal rainfall observation itself and the recognition and judgment result is also a normal observation; TN (true negative) is the noise observation itself and the recognition and judgment result is also a noise observation; FN (false negative) is the normal observation itself, but the recognition and judgment result is a noise observation; FP (false positive) is the noise observation itself, but the recognition and judgment result is a normal observation.
[0135] Considering the randomness of the simulation of the crowdsourcing rainfall data quality control model, the performance of the crowdsourcing rainfall data quality control model may be affected by the uncertainty caused by the random selection of crowdsourcing observations. By calculating the ΔWCFG and ΔWCPJ values of the crowdsourcing rainfall data quality control models trained based on different machine learning algorithms in the same benchmark scenario, the impact of the uncertainty caused by the random selection of crowdsourcing observations is quantified. Specifically, the impact of this uncertainty is simulated and analyzed by repeating the benchmark scenario 1000 times. This application uses a total of 4 machine learning algorithms, namely kNN, MLPs, iForest, and K-Means algorithms. Among them, kNN refers to the K-Nearest Neighbor algorithm, MLPs refers to the multi-layer neural network algorithm, iForest refers to the Isolation Forest algorithm, and K-Means refers to the K-Means clustering algorithm. As Figure 8 shown, the ΔWCFG and ΔWCPJ values of the crowdsourcing rainfall data quality control models trained by each machine learning algorithm approximately form a normal distribution, and the standard deviation is relatively small, which also indicates that the performance ranking of these machine learning algorithms is not affected by uncertainty.
[0136] It should be noted that in addition to the 4 classical supervised / unsupervised machine learning algorithms (kNN, MLPs, iForest, K-Means) used in the present invention, other algorithms such as support vector machines, random forests, and Gaussian process algorithms can also be used to train the crowdsourcing rainfall data quality control model.
[0137] The above is only the preferred embodiment of the present invention, and does not limit the patent scope of the present invention. Any equivalent structural transformation made by using the description and drawings of the present invention under the inventive concept of the present invention, or direct / indirect application in other related technical fields is included in the patent protection scope of the present invention.
Claims
1. A crowdsourcing rainfall data quality control method, characterized in that, It includes the following steps: S1. Extract multiple rainfall data statistical features of rainfall observation points to establish a crowdsourced rainfall dataset with mathematical attributes; S2. Based on machine learning algorithms, train and preprocess various rainfall data statistical features in the crowdsourced rainfall dataset to train a crowdsourced rainfall data quality control model; S3. Rank the importance of performance impact factors of the crowdsourced rainfall data quality control model to calibrate the crowdsourced rainfall data quality control model; The step S3 includes: S31. Use the SHAP interpretable algorithm to quantitatively interpret and rank the performance impact factors of the crowdsourced rainfall data quality control model; S32. Use the extra tree algorithm to identify the best combination of performance impact factors of the crowdsourced rainfall data quality control model; In step S32, the Extra Trees algorithm is an ensemble learning method that constructs decision trees and selects feature subsets for node splitting. Its principle can be summarized as randomly selecting m features from the feature set F, where m = |F|. For the selected feature f, a split point split is randomly chosen f , and the data set is divided into two subsets D left and D right , where: D left = {x ∈ D | x f ≤ split f} D right = {x ∈ D | x f > split f}; The feature set F represents multiple rainfall data statistical features. The feature selection step iteratively deletes the least important performance impact factors until the prediction performance of the crowdsourced rainfall data quality model decreases, and selects the best combination of performance impact factors through accuracy; In step S1, a part of the original rainfall data of the multiple rainfall data statistical features is directly collected by fixed sensors, and another part of the original rainfall data of the multiple rainfall data statistical features is obtained by preset processing of the data directly collected by mobile sensors; The steps for preset processing of the data directly collected by the mobile sensors include: Couple the mobile path model associated with the mobile sensor into the data quality control architecture to simulate the dynamic movement path of participants and achieve deep coupling of the mobile path model and the crowdsourced rainfall quality control model. The control equations and parameters involved are as follows: Among them, S i,t represents the position of participant i at time step t; L i,t represents the moving distance of participant i at time step t; θ i,t represents the moving direction of participant i at time step t; x and y are the coordinates of the rainfall observation point in space respectively; α and φ represent the adjustment factors of the moving distance L i,t and the moving direction θ i,t respectively; ω i,t represents the Boolean moving state factor of participant i at time step t; the moving distance L i,t is assumed to follow L i,t ~U(0, L max ), where U() represents following a uniform distribution, and L max represents the maximum distance that the participant can move within the time step; the moving direction θ i,t is assumed to follow θ i,t ~U(0, 2π), represents the moving probability of the rainfall observation point; and represent the maximum moving probability and the minimum moving probability of the rainfall observation point respectively; cx and cy represent the positions of the center of the rainfall observation point in the x and y directions respectively; r represents the radius of the center of the rainfall observation point; ζp represents the observation density gradient from the center of the rainfall observation point to the edge of the study area; The crowdsourced rainfall data quality control method has transferability, so that the crowdsourced rainfall data quality control model trained under the benchmark scenario can be applied to other regions or other noise level scenarios.
2. The crowdsourcing rainfall data quality control method according to claim 1, wherein The step S1 of extracting multiple rainfall data statistical features includes: S11. Based on the areal rainfall fields with different coverage ranges centered on the same preset rainfall observation point, with the preset rainfall observation point as the center, extract the rainfall mathematical attributes of different-sized windows. The rainfall mathematical attributes include at least one of variance, maximum value, minimum value, range, average value, standard deviation, and absolute deviation; S12. Based on the areal rainfall fields within the coverage ranges at different radial distances centered on the same preset rainfall observation point, extract the rainfall mathematical correlation features between the preset rainfall observation point and other rainfall observation points. The rainfall mathematical correlation features include at least one of difference, change range, and gradient; S13. Based on the rainfall fields formed at different times centered on the same preset rainfall observation point, extract the correlation coefficients of adjacent areal rainfall fields.
3. The crowdsourcing rainfall data quality control method according to claim 1, characterized in that The machine learning algorithms include supervised machine learning algorithms and unsupervised machine learning algorithms; The step S2 includes: S21. Based on the supervised machine learning algorithm, in the way of using the calibration error, randomly add noise values to the rainfall observation points by introducing noise values and calibrate them, and train the statistical characteristics of multiple rainfall data of the learning error rainfall observation points; S22. Based on the unsupervised machine learning algorithm, introduce the noise data into the statistical characteristics of multiple rainfall data for automatic training and preprocessing.
4. The crowdsourcing rainfall data quality control method according to claim 1, wherein In step S31, the SHAP value is calculated based on the following formula: where \(f(x)\) is the prediction of the model for the input \(x\); \(x i _i_ is the \(i\)-th feature of the input \(x\); \(M\) is the number of features; \(\varphi_0\) is a constant representing the feature base value; \(\varphi i _i_ is the SHAP value of the \(i\)-th feature, representing the contribution of this feature to the prediction \(f(x)\); Among them, S is the feature set, representing the subset of features considered during model prediction; |S| is the number of features in set S; f x (S) is the prediction of the model when only considering the features in set S; f x (S ∪ {i}) is the prediction of the model when considering the features in set S and feature i.
5. The crowdsourcing rainfall data quality control method according to claim 1, wherein It also includes: S4. Use the preset evaluation index to evaluate the prediction result of the crowdsourced rainfall data quality control model.
6. The crowdsourcing rainfall data quality control method according to claim 5, wherein, The evaluation index includes WCFG and WCPJ. WCFG is used to describe the ability to capture rainfall variability on the hourly and spatial scales, and WCPJ is used to describe the relative deviation of the total rainfall amount of the rainfall event captured by the site within the duration and spatial coverage of the rainfall; The calculation formulas of WCFG and WCPJ are as follows respectively: Among them, is the rainfall intensity estimated at time t based on crowdsourced observations; y i is the relevant ground truth rainfall intensity; i represents the serial number of the rainfall observation point.
7. The crowdsourcing rainfall data quality control method according to claim 6, characterized in that It also includes: S5. Use the ΔWCFG and ΔWCPJ evaluation indexes to evaluate the performance improvement of the crowdsourced rainfall data quality control model after noise removal; The calculation formulas of ΔWCFG and ΔWCPJ are as follows respectively: Among them, WCFG before represents the WCFG value of crowdsourced rainfall intensity before noise rejection; WCFG after represents the WCFG value of crowdsourced rainfall intensity after noise rejection; WCPJ before represents the WCPJ value of crowdsourced rainfall intensity before noise rejection; WCPJ after represents the WCPJ value of crowdsourced rainfall intensity after noise rejection.