Processing method of precipitation data applied to flood disaster model and related device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA REINSURANCE (GROUP) CORPORATION
- Filing Date
- 2023-12-04
- Publication Date
- 2026-08-07
AI Technical Summary
[0003]但是,历史降水数据可能存在质量问题,导致历史降水数据不满足洪涝巨灾模型处理数据的要求,影响洪涝巨灾模型仿真的准确程度
[0044]本申请提供应用于洪涝巨灾模型的降水数据的处理方法及相关装置,该方法中,获取原始降水数据;对原始降水数据进行数据集成,得到包括空间维度和时间维度的矩阵数据;对矩阵数据进行数据清洗,得到待转换降水数据;将待转换降水数据转换为符合正态分布的降水数据。符合正态分布的降水数据用于被洪涝巨灾模型处理。如此处理后得到的降水数据的数据质量较高。符合正态分布的降水数据,能够减少降雨数据的非连续性和无雨日对洪涝巨灾模型的仿真的不良影响,便于洪涝巨灾模型对降水数据进行分析和处理,进而得到更为准确的洪涝巨灾模型的仿真结果,提高对洪涝灾害的模拟准确程度。
Smart Images

Figure CN117610301B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, specifically to a method and related apparatus for processing precipitation data applied to flood disaster models. Background Technology
[0002] The flood catastrophic model models the hazard of disaster-causing factors by simulating flood disaster events. Based on historical precipitation data, the flood catastrophic model simulates random precipitation events over tens of thousands of years to simulate conditions excluding extreme weather events. The simulated dataset of random precipitation events over tens of thousands of years is then input into a hydrological and hydraulic model. The hydrological and hydraulic model is used to calculate the flood process, such as inundation depth and extent.
[0003] However, historical precipitation data may have quality issues, causing it to fail to meet the data processing requirements of flood disaster models and affecting the accuracy of flood disaster model simulations. Summary of the Invention
[0004] In view of this, this application provides a method and related apparatus for processing precipitation data applied to flood disaster models, aiming to obtain precipitation data that meets the requirements of flood disaster models for processing data, thereby facilitating the processing of precipitation data by flood disaster models.
[0005] Based on this, the technical solution provided in this application is as follows:
[0006] In a first aspect, this application provides a method for processing precipitation data applied to a flood disaster model, the method comprising:
[0007] Obtain raw precipitation data;
[0008] The raw precipitation data is integrated to obtain matrix data, which includes spatial and temporal dimensions.
[0009] The matrix data is cleaned to obtain precipitation data to be converted;
[0010] The precipitation data to be converted is transformed into precipitation data that conforms to a normal distribution. The precipitation data that conforms to a normal distribution is used to generate a dataset of random precipitation events using a flood disaster model.
[0011] In one possible implementation, converting the precipitation data to be transformed into precipitation data conforming to a normal distribution includes:
[0012] The precipitation data to be converted is transformed into precipitation data that conforms to a normal distribution using a conversion formula. The independent variable of the conversion formula is the precipitation data to be converted, and the dependent variable is the precipitation data that conforms to a normal distribution.
[0013] In one possible implementation, the transformation formula is generated based on a Gaussian latent variable function, which is used to simulate actual precipitation data based on pseudo precipitation data conforming to a normal distribution and transformation parameters. The pseudo precipitation data conforming to a normal distribution is determined based on constructed data conforming to a normal distribution and the probability of rainless days in the sample actual precipitation data. The daily precipitation of rainless days in the sample actual precipitation data is less than or equal to a precipitation threshold.
[0014] In one possible implementation, the Gaussian hidden variable function is: P =
[0015]
[0016] Where P represents the simulated actual precipitation data, P m Q is the precipitation threshold, Q is the pseudo precipitation data that conforms to the standard normal distribution, Q0 is the preset threshold, and a, b and c are parameters to be determined. The difference between the simulated actual precipitation data and the sample actual precipitation data satisfies the difference condition.
[0017] In one possible implementation, the parameters to be determined in the Gaussian latent variable function are obtained by fitting a global optimization algorithm, with the goal of fitting such that the difference between the simulated actual precipitation data and the sampled actual precipitation data satisfies the difference condition.
[0018] In one possible implementation, the spatial dimension is the station that generated the original precipitation data, and the step of cleaning the matrix data to obtain the precipitation data to be transformed includes:
[0019] Based on the missing data included in the precipitation data of the stations, determine the annual missing rate and the duration of missing data;
[0020] For each station's precipitation data, perform the following operations to generate precipitation data to be converted:
[0021] If the annual missing rate is greater than the first threshold and the missing duration is greater than the second threshold, the site will be deleted.
[0022] If the duration of missing measurements is greater than the third threshold and less than or equal to the second threshold, the missing measurement data of the station is replaced by precipitation data from the backup station of the station. The distance between the backup station and the station is less than the distance threshold and the backup condition is met. The third threshold is less than the second threshold.
[0023] If the duration of missing data is less than or equal to the third threshold, the missing data of the site is supplemented using preset data.
[0024] In one possible implementation, the step of cleaning the matrix data to obtain precipitation data to be converted includes:
[0025] Outlier processing is performed on precipitation data where the daily precipitation exceeds a precipitation threshold to obtain precipitation data to be converted. The precipitation threshold is determined based on the average annual precipitation of the station.
[0026] Secondly, this application provides a processing apparatus for precipitation data applied to a flood disaster model, the apparatus comprising:
[0027] The acquisition unit is used to acquire raw precipitation data;
[0028] An integration unit is used to integrate the raw precipitation data to obtain matrix data, which includes spatial and temporal dimensions.
[0029] A cleaning unit is used to clean the matrix data to obtain precipitation data to be converted.
[0030] The conversion unit is used to convert the precipitation data to be converted into precipitation data that conforms to a normal distribution. The precipitation data that conforms to a normal distribution is used to generate a dataset of random precipitation events using a flood disaster model.
[0031] In one possible implementation, the conversion unit is used to convert the precipitation data to be converted into precipitation data conforming to a normal distribution, including:
[0032] The conversion unit is used to convert the precipitation data to be converted into precipitation data that conforms to a normal distribution using a conversion formula, wherein the independent variable of the conversion formula is the precipitation data to be converted, and the dependent variable is the precipitation data that conforms to a normal distribution.
[0033] In one possible implementation, the transformation formula is generated based on a Gaussian latent variable function, which is used to simulate actual precipitation data based on pseudo precipitation data conforming to a normal distribution and transformation parameters. The pseudo precipitation data conforming to a normal distribution is determined based on constructed data conforming to a normal distribution and the probability of rainless days in the sample actual precipitation data. The daily precipitation of rainless days in the sample actual precipitation data is less than or equal to a precipitation threshold.
[0034]
[0035] The simulated precipitation data is used, where Q0 is a preset threshold, and a, b, and c are parameters to be determined. The difference between the simulated actual precipitation data and the sampled actual precipitation data satisfies the difference condition.
[0036] In one possible implementation, the parameters to be determined in the Gaussian latent variable function are obtained by fitting a global optimization algorithm, with the goal of fitting such that the difference between the simulated actual precipitation data and the sampled actual precipitation data satisfies the difference condition.
[0037] In one possible implementation, the spatial dimension is the station that generated the original precipitation data. The cleaning unit is specifically used to determine the annual missing rate and missing duration based on the missing data included in the precipitation data of the station; for the precipitation data of each station, the following operations are performed to generate precipitation data to be converted: if the annual missing rate is greater than a first threshold and the missing duration is greater than a second threshold, the station is deleted; if the missing duration is greater than a third threshold and less than or equal to the second threshold, the missing data of the station is replaced with precipitation data from a backup station of the station, wherein the distance between the backup station and the station is less than a distance threshold and meets the backup condition, and the third threshold is less than the second threshold; if the missing duration is less than or equal to the third threshold, the missing data of the station is supplemented with preset data.
[0038] In one possible implementation, the cleaning unit is specifically used to perform outlier processing on precipitation data where the daily precipitation exceeds a precipitation threshold, to obtain precipitation data to be converted, wherein the precipitation threshold is determined based on the average annual precipitation of the station.
[0039] Thirdly, this application provides an apparatus, including: a processor, a memory, and a system bus;
[0040] The processor and the memory are connected via the system bus;
[0041] The memory is used to store one or more programs, the one or more programs including instructions that, when executed by the processor, cause the processor to perform the method described in any of the embodiments of the first aspect above.
[0042] Fourthly, this application provides a computer-readable storage medium storing instructions that, when executed on a terminal device, cause the terminal device to perform the method described in any of the embodiments of the first aspect.
[0043] Therefore, this application has the following beneficial effects:
[0044] This application provides a method and related apparatus for processing precipitation data for flood disaster models. The method involves: acquiring raw precipitation data; integrating the raw precipitation data to obtain matrix data including spatial and temporal dimensions; cleaning the matrix data to obtain precipitation data to be transformed; and converting the precipitation data to be transformed into precipitation data conforming to a normal distribution. The normally distributed precipitation data is then used for processing by the flood disaster model. The resulting precipitation data has higher quality. Normally distributed precipitation data reduces the adverse effects of discontinuous rainfall data and rainless days on the simulation of flood disaster models, facilitating the analysis and processing of precipitation data by the flood disaster model, thereby obtaining more accurate simulation results and improving the accuracy of flood disaster simulation. Attached Figure Description
[0045] Figure 1 A flowchart illustrating the method for processing precipitation data applied to a flood disaster model, as provided in an embodiment of this application;
[0046] Figure 2 This is a schematic diagram of a precipitation data processing device for a flood disaster model, provided in an embodiment of this application. Detailed Implementation
[0047] To facilitate understanding and explanation of the technical solutions provided in the embodiments of this application, the background technology of this application will be described first.
[0048] The disaster module of a flood catastrophic model mainly consists of two parts: a precipitation event set and a hydrological and hydraulic model. The precipitation event set, for example, is a 10,000-year precipitation event set. This 10,000-year precipitation event set needs to be constructed based on historical precipitation data through data analysis. However, historical precipitation data has some quality issues. For example, there may be missing or incorrect measurements, data from different sources, short timeframes that cannot reflect extreme precipitation scenarios, and different data formats.
[0049] The 10,000-year precipitation event set is constructed based on a spatiotemporal statistical model of historical precipitation elements. It treats variables as random functions in the spatiotemporal domain to construct a multi-element random event set. The elements need to possess both temporal and spatial characteristics. However, the periodicity and spatial non-stationarity of historical precipitation data cannot meet the requirements of flood catastrophic models.
[0050] The characteristics of daily precipitation sequences composed of precipitation data are often difficult to describe using simple and universal single-distribution probability models or mixed-distribution probability models. Precipitation data that include daily or hourly precipitation amounts often contain a large number of zero values indicating no rainfall and exhibit right-skewed effects. Precipitation data also exhibits temporal intermittency. Furthermore, spatial intermittency of precipitation is a problem in large-scale precipitation simulations.
[0051] Flood disaster models built using historical precipitation data with quality issues may suffer from inaccurate simulations.
[0052] Based on this, embodiments of this application provide a method and related apparatus for processing precipitation data applied to a flood disaster model. The method involves: acquiring raw precipitation data; integrating the raw precipitation data to obtain matrix data including spatial and temporal dimensions; cleaning the matrix data to obtain precipitation data to be transformed; and converting the precipitation data to be transformed into precipitation data conforming to a normal distribution. The normally distributed precipitation data is used for processing by the flood disaster model. The resulting precipitation data has higher quality. Normally distributed precipitation data reduces the discontinuity of rainfall data and the adverse effects of rainless days on the simulation of the flood disaster model, facilitating the analysis and processing of precipitation data by the flood disaster model, thereby obtaining more accurate simulation results and improving the accuracy of flood disaster simulation.
[0053] This application provides a method for processing precipitation data in a flood disaster model, applicable to electronic devices with data processing capabilities. These electronic devices can be, for example, servers or terminals. Terminals include, but are not limited to, smartphones, tablets, laptops, personal digital assistants (PDAs), or smart wearable devices. Servers can be cloud servers, such as central servers in a central computing cluster or edge servers in an edge computing cluster. Alternatively, servers can be located in a local data center. A local data center refers to a data center directly controlled by the user.
[0054] Electronic devices acquire raw precipitation data; the raw precipitation data is integrated to obtain matrix data including spatial and temporal dimensions; the matrix data is cleaned to obtain precipitation data to be transformed; the precipitation data to be transformed is converted into precipitation data that conforms to a normal distribution.
[0055] Those skilled in the art will understand that the above application scenarios are merely one example of how the embodiments of this application can be implemented. The scope of application of the embodiments of this application is not limited by any aspect of this framework.
[0056] To facilitate understanding of the technical solutions provided in the embodiments of this application, the processing method of precipitation data applied to the flood disaster model provided in the embodiments of this application will be described below with reference to the accompanying drawings.
[0057] See Figure 1 As shown in the figure, this figure is a flowchart illustrating a method for processing precipitation data applied to a flood disaster model according to an embodiment of this application, including S101-S104.
[0058] S101: Obtain raw precipitation data.
[0059] Raw precipitation data refers to historically generated precipitation data that requires processing. Raw precipitation data includes daily precipitation. This application does not limit the source of the raw precipitation data. In one possible implementation, the raw precipitation data is obtained from the observation station that generated the precipitation data. In another possible implementation, the raw precipitation data is obtained from an existing meteorological or precipitation dataset database.
[0060] S102: Data integration of raw precipitation data to obtain matrix data, which includes spatial and temporal dimensions.
[0061] The format of raw precipitation data may vary. Common formats include station-level or grid-level data. Integrating the raw precipitation data yields a matrix dataset with both spatial and temporal dimensions. The spatial dimension is represented by the stations or grid points that generated the raw precipitation data. The integrated matrix dataset might look like this: [M] station ×N time Furthermore, the data index uses station number indexing. This standardizes the format of precipitation data, facilitating further processing of the data.
[0062] S103: Perform data cleaning on the matrix data to obtain precipitation data to be converted.
[0063] Data cleaning is performed on the matrix data. This application does not limit the specific implementation of data cleaning in its embodiments. In one possible implementation, data cleaning includes one or more of two methods: handling missing values and handling outliers.
[0064] The precipitation data obtained after data cleaning is of higher quality, making it easier to process the precipitation data in subsequent steps.
[0065] As an example, embodiments of this application provide specific processes for handling missing values and outliers.
[0066] The first method: Data cleaning includes handling missing values.
[0067] This application provides a specific implementation method for cleaning matrix data to obtain precipitation data to be transformed, including:
[0068] A1: Determine the annual missing rate and duration of missing data based on the missing data included in the precipitation data of the stations.
[0069] The precipitation data for each station in the matrix is processed to determine the missing data for each station. For example, missing data in the precipitation data may be represented by the values 99999, NaN, or empty.
[0070] Determine the site's annual missing rate and missing duration. The annual missing rate is the ratio of the number of days with missing data to the total number of days. The missing duration is, for example, the number of abnormal years. An abnormal year is, for example, a year where the annual missing rate exceeds the missing threshold. The missing threshold can be a pre-set parameter. As an example, the missing threshold is the first threshold. The missing duration is, for example, the number of days with missing data.
[0071] A2: For each station's precipitation data, perform the following operations to generate precipitation data to be converted.
[0072] A3: If the annual missing rate is greater than the first threshold and the missing duration is greater than the second threshold, the site will be deleted.
[0073] If the annual missing rate is greater than the first threshold and the missing duration is greater than the second threshold, it indicates that the missing situation at the station is quite serious, and there may be an observational anomaly at the station. Therefore, the precipitation data of the station should be deleted.
[0074] The first and second thresholds are pre-set parameters. For example, the missing data duration is the number of years of missing data, the first threshold is 50%, and the second threshold is 5 years. That is, for any station with an annual missing data rate greater than 50% and more than 5 years of such missing data, the station will be deleted, and its precipitation data will no longer be processed.
[0075] A4: If the missing measurement duration is greater than the third threshold and less than or equal to the second threshold, the missing measurement data of the station shall be replaced by the precipitation data of the station's backup station. The distance between the backup station and the station shall be less than the distance threshold, and the annual missing measurement rate and missing measurement duration shall meet the backup conditions.
[0076] If the duration of missing data exceeds the third threshold but is less than or equal to the second threshold, it indicates that the station has some missing data. In this case, precipitation data from a backup station is used to supplement the missing data. The distance between the backup station and the original station is less than a distance threshold. This ensures that the precipitation data from the backup station is representative of the original station's precipitation data and reflects the precipitation situation in the same region. For example, the backup station could be the closest station in the same region. Furthermore, the backup station must meet backup criteria, such as an annual missing data rate less than the first threshold. This ensures that the precipitation data from the backup station is usable.
[0077] The third threshold is a parameter less than the second threshold. By limiting the duration of missing data, it is also necessary to meet the condition of being greater than the third threshold to avoid supplementing precipitation data with data from backup stations for too short a duration of missing data. The duration of missing data is the number of consecutive days of missing data; for example, the third threshold is 30 days.
[0078] A5: If the missing test duration is less than or equal to the third threshold, the missing test data of the site will be supplemented using preset data.
[0079] If the duration of missing data is less than or equal to the third threshold, it indicates that the missing data situation at that site is relatively minor, and the missing data at that site will be supplemented using preset data. The preset data is, for example, 0.
[0080] The second method: Data cleaning includes outlier handling.
[0081] Precipitation data with daily precipitation exceeding a precipitation threshold are considered outliers. The precipitation threshold is, for example, 50% of the average annual precipitation. Outliers are then processed. This application does not limit the method of outlier processing. In one possible implementation, different outlier processing methods are predefined, and processing is performed based on these methods. In another possible implementation, the outlier is displayed so the user can process it. The user-input outlier processing result is obtained, and the outlier is processed based on this result. This approach is applicable to the processing of precipitation data from extreme rainfall events, preventing such data from being treated as outliers. This facilitates the analysis of extreme rainfall events by flood disaster models based on precipitation data, improving the accuracy of flood disaster model simulations.
[0082] S104: Convert the precipitation data to be transformed into precipitation data that conforms to a normal distribution. The precipitation data that conforms to a normal distribution is used to generate a dataset of random precipitation events using a flood disaster model.
[0083] The precipitation data to be transformed may conform to a mixed distribution, which does not meet the requirements for simulation by the flood disaster model. The precipitation data to be transformed into precipitation data conforming to a normal distribution is then used as input data for the flood disaster model, enabling the model to process the precipitation data and simulate random precipitation event datasets.
[0084] In one possible implementation, the precipitation data to be converted at each station is processed to obtain precipitation data that conforms to a normal distribution for each station.
[0085] The embodiments of this application do not limit the method of converting precipitation data to conform to a normal distribution.
[0086] In one possible implementation, the distribution of the precipitation data to be transformed is first determined. Based on the transformation relationship between this distribution and the normal distribution, the precipitation data is transformed to obtain precipitation data conforming to a normal distribution. The transformation relationship can be represented by a pre-established numerical transformation table, which includes the transformation relationships between numerical values.
[0087] In another possible implementation, the precipitation data to be transformed is transformed using a pre-constructed transformation formula. This formula converts data that does not conform to a normal distribution into data that does. The independent variable of the transformation formula is the non-normally distributed precipitation data (the data to be transformed), and the dependent variable is the normally distributed precipitation data.
[0088] This application does not limit the construction method of the conversion formula. As an example, this application provides a method for generating conversion formulas based on Gaussian latent variable functions. Gaussian latent variable functions are used to simulate actual precipitation data based on pseudo-precipitation data conforming to a normal distribution and conversion parameters. The pseudo-precipitation data conforming to a normal distribution is determined based on constructed data conforming to a normal distribution and the probability of rainless days in the sample actual precipitation data. Using pseudo-precipitation data, normally distributed precipitation data can be constructed, and then conversion parameters can be obtained through fitting, establishing a conversion relationship between normally distributed and non-normally distributed precipitation data. The daily precipitation on rainless days in the sample actual precipitation data is less than or equal to a precipitation threshold. This filters out some data with low precipitation, facilitating the fitting of normally distributed precipitation data.
[0089] For example, the Gaussian hidden variable function is shown in formula (1).
[0090]
[0091] Where P represents the simulated actual precipitation data. The difference between P and the sample actual precipitation data used to construct formula (1) satisfies the difference condition. m Q is the precipitation threshold, Q is pseudo-precipitation data conforming to a standard normal distribution, Q0 is the preset threshold, and a, b, and c are parameters to be determined.
[0092] As an example, the sample actual precipitation data is historical precipitation data used to construct the Gaussian latent variable function. The sample actual precipitation data can be selected from historical precipitation data. For example, the sample actual precipitation data is precipitation data generated by a certain station and obtained through data integration and data cleaning.
[0093] p mThe value is, for example, 0.1 mm. If the daily precipitation in the actual precipitation data of the sample is less than or equal to 0.1 mm, then that day is considered a rainless day.
[0094] The ratio of rainless days to the total number of days in the actual precipitation data of the sample is calculated to obtain norain_P. Q0 is determined based on norain_P. As an example, Q0 is the inverse function value of the standard normal cumulative distribution function (CDF) with probability value norain_P.
[0095] Q is determined based on constructed data conforming to a normal distribution and the probability of rainless days in the actual precipitation data of the sample. As an example, the expression for Q is shown in formula (2):
[0096] Q=(1-norain_P)d+norain_P (2)
[0097] d represents the constructed data that conforms to a standard normal distribution. As an example, d is a generated set of data that follows a standard normal distribution (μ = 0, σ = 1). For instance, d might contain 1000 data points.
[0098] Based on the quantiles determined by (Q-Q0), the corresponding quantiles of the actual precipitation data in the sample are determined, forming a correspondence between the actual precipitation data of the sample and (Q-Q0). For example, the 10th percentile of (Q-Q0) is matched with the 10th percentile of the actual precipitation data in the sample.
[0099] Using the corresponding (Q-Q0), the actual precipitation data of the sample, and formula (1), the specific values of the parameters a, b, and c to be determined are determined.
[0100] In one possible implementation, the parameters to be determined are fitted using a global optimization algorithm (Shuffled Complex Evolution, SCE-UA). The fitting objective is to ensure that the difference between the simulated sample precipitation data and the actual sample precipitation data satisfies a difference condition. The simulated sample precipitation data is the value of P calculated by substituting (Q-Q0) as the dependent variable into formula (1). The actual sample precipitation data is the sample precipitation data corresponding to (Q-Q0). The difference condition is, for example, that the difference is less than a threshold. Alternatively, the difference condition is, for example, that the difference is minimized. Or, the difference condition is, for example, that the number of iterations is reached. As an example, the objective function is the sum of squares of the errors between the simulated sample precipitation data and the actual sample precipitation data. The difference condition is that the independent variable of the objective function takes the minimum value.
[0101] This application does not limit the method of using the SCE-UA algorithm to fit the parameters to be determined. As an example, the parameter range, number of complexes, number of points inside the complexes, sample size, upper and lower limits of parameters, convergence criteria parameters, etc., are preset. The SCE-UA algorithm is applied to randomly generate random numbers for a given number of complexes within the parameter range, and divide the complexes into several population partitions. The parameters of each partition are independent and each searches for its own optimum. However, the populations can pass the searched information to each other to update the searched partition parameters. After reaching the difference condition, the parameters to be determined simultaneously reach the overall optimum, and the fitted values of the parameters a, b, and c are obtained.
[0102] Based on the relevant content in S101-S104 above, it can be seen that through data integration, data cleaning, and distribution transformation, high-quality precipitation data can be obtained after processing. Precipitation data that conforms to a normal distribution can reduce the adverse effects of discontinuity in rainfall data and rainless days on the simulation of flood disaster models, making it easier for flood disaster models to analyze and process precipitation data, thereby obtaining more accurate simulation results and improving the accuracy of flood disaster simulation.
[0103] Based on the above-described method embodiment, which provides a method for processing precipitation data applied to a flood disaster model, this application embodiment also provides a device for processing precipitation data applied to a flood disaster model. The device for processing precipitation data applied to a flood disaster model will be described below with reference to the accompanying drawings.
[0104] See Figure 2 As shown in the figure, this is a schematic diagram of the structure of a precipitation data processing device applied to a flood disaster model, according to an embodiment of this application. Figure 2 As shown, the precipitation data processing device applied to the flood disaster model includes:
[0105] Acquisition unit 201 is used to acquire raw precipitation data;
[0106] The integration unit 202 is used to integrate the raw precipitation data to obtain matrix data, which includes spatial and temporal dimensions.
[0107] The cleaning unit 203 is used to clean the matrix data to obtain precipitation data to be converted;
[0108] The conversion unit 204 is used to convert the precipitation data to be converted into precipitation data that conforms to a normal distribution. The precipitation data that conforms to a normal distribution is used to generate a random precipitation event dataset by processing with a flood disaster model.
[0109] In one possible implementation, the conversion unit 204 is used to convert the precipitation data to be converted into precipitation data conforming to a normal distribution, including:
[0110] The conversion unit 204 is used to convert the precipitation data to be converted into precipitation data that conforms to a normal distribution using a conversion formula, wherein the independent variable of the conversion formula is the precipitation data to be converted, and the dependent variable is the precipitation data that conforms to a normal distribution.
[0111] In one possible implementation, the transformation formula is generated based on a Gaussian latent variable function, which is used to simulate actual precipitation data based on pseudo precipitation data conforming to a normal distribution and transformation parameters. The pseudo precipitation data conforming to a normal distribution is determined based on constructed data conforming to a normal distribution and the probability of rainless days in the sample actual precipitation data. The daily precipitation of rainless days in the sample actual precipitation data is less than or equal to a precipitation threshold.
[0112] In one possible implementation, the Gaussian hidden variable function is:
[0113] Where P represents the simulated actual precipitation data, P m Q is the precipitation threshold, Q is the pseudo precipitation data that conforms to the standard normal distribution, Q0 is the preset threshold, and a, b and c are parameters to be determined. The difference between the simulated actual precipitation data and the sample actual precipitation data satisfies the difference condition.
[0114] In one possible implementation, the parameters to be determined in the Gaussian latent variable function are obtained by fitting a global optimization algorithm, with the goal of fitting such that the difference between the simulated actual precipitation data and the sampled actual precipitation data satisfies the difference condition.
[0115] In one possible implementation, the spatial dimension is the station that generated the original precipitation data. The cleaning unit 203 is specifically used to determine the annual missing rate and missing duration based on the missing data included in the precipitation data of the station; for the precipitation data of each station, the following operations are performed to generate precipitation data to be converted: if the annual missing rate is greater than a first threshold and the missing duration is greater than a second threshold, the station is deleted; if the missing duration is greater than a third threshold and less than or equal to the second threshold, the missing data of the station is replaced with precipitation data from a backup station of the station, wherein the distance between the backup station and the station is less than a distance threshold and meets the backup condition, and the third threshold is less than the second threshold; if the missing duration is less than or equal to the third threshold, the missing data of the station is supplemented with preset data.
[0116] In one possible implementation, the cleaning unit 203 is specifically used to perform outlier processing on precipitation data where the daily precipitation is greater than a precipitation threshold, to obtain precipitation data to be converted, wherein the precipitation threshold is determined based on the average annual precipitation of the station.
[0117] Based on the above-described method embodiments, this application provides a method for processing precipitation data applied to a flood disaster model. The method includes a processor, a memory, and a system bus.
[0118] The processor and the memory are connected via the system bus;
[0119] The memory is used to store one or more programs, the one or more programs including instructions, which, when executed by the processor, cause the processor to perform the precipitation data processing method applied to the flood disaster model as described in any of the above embodiments.
[0120] Based on the above-described method embodiments, this application provides a computer-readable storage medium storing instructions. When the instructions are executed on a terminal device, the terminal device performs the above-described method for processing precipitation data for a flood disaster model.
[0121] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems or apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and relevant parts can be referred to the method section.
[0122] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0123] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0124] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0125] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for processing precipitation data applied to a flood disaster model, characterized in that, The method includes: Obtain raw precipitation data; The raw precipitation data is integrated to obtain matrix data, which includes spatial and temporal dimensions. The matrix data is cleaned to obtain precipitation data to be converted; The precipitation data to be transformed is converted into precipitation data conforming to a normal distribution using a transformation formula. The normally distributed precipitation data is used to generate a random precipitation event dataset using a flood disaster model. The independent variable of the transformation formula is the precipitation data to be transformed, and the dependent variable is the normally distributed precipitation data. The transformation formula is generated based on a Gaussian latent variable function. The Gaussian latent variable function is used to simulate actual precipitation data based on normally distributed pseudo-precipitation data and transformation parameters. The normally distributed pseudo-precipitation data is determined based on normally distributed constructed data and the probability of rainless days in the sample actual precipitation data. The daily precipitation of rainless days in the sample actual precipitation data is less than or equal to a precipitation threshold.
2. The method according to claim 1, characterized in that, The Gaussian hidden variable function is: ; in, The data is based on simulated actual precipitation. For precipitation threshold, To obtain pseudo-precipitation data that conforms to a standard normal distribution, For the preset threshold, b and c are parameters to be determined, and the difference between the simulated actual precipitation data and the sample actual precipitation data satisfies the difference condition.
3. The method according to claim 2, characterized in that, The parameters to be determined in the Gaussian latent variable function are obtained by fitting based on a global optimization algorithm. The goal of the fitting is to ensure that the difference between the simulated actual precipitation data and the sample actual precipitation data satisfies the difference condition.
4. The method according to claim 1, characterized in that, The spatial dimension refers to the stations that generated the original precipitation data. The step of cleaning the matrix data to obtain the precipitation data to be transformed includes: Based on the missing data included in the precipitation data of the stations, determine the annual missing rate and the duration of missing data; For each station's precipitation data, perform the following operations to generate precipitation data to be converted: If the annual missing rate is greater than the first threshold and the missing duration is greater than the second threshold, the site will be deleted. If the missing measurement duration is greater than the third threshold and less than or equal to the second threshold, the missing measurement data of the station is replaced by precipitation data from the backup station of the station. The distance between the backup station and the station is less than the distance threshold and the backup condition is met. The third threshold is less than the second threshold. If the duration of missing data is less than or equal to the third threshold, the missing data of the site is supplemented using preset data.
5. The method according to claim 1, characterized in that, The step of cleaning the matrix data to obtain precipitation data to be converted includes: Outlier processing is performed on precipitation data where the daily precipitation exceeds a precipitation threshold to obtain precipitation data to be converted. The precipitation threshold is determined based on the average annual precipitation of the station.
6. A device for processing precipitation data applied to a flood disaster model, characterized in that, The device includes: The acquisition unit is used to acquire raw precipitation data; An integration unit is used to integrate the raw precipitation data to obtain matrix data, which includes spatial and temporal dimensions. A cleaning unit is used to clean the matrix data to obtain precipitation data to be converted. A transformation unit is used to convert the precipitation data to be transformed into precipitation data conforming to a normal distribution using a transformation formula. The normally distributed precipitation data is used to generate a random precipitation event dataset using a flood disaster model. The independent variable of the transformation formula is the precipitation data to be transformed, and the dependent variable is the normally distributed precipitation data. The transformation formula is generated based on a Gaussian latent variable function. The Gaussian latent variable function is used to simulate actual precipitation data based on normally distributed pseudo-precipitation data and transformation parameters. The normally distributed pseudo-precipitation data is determined based on normally distributed constructed data and the probability of rainless days in the sample actual precipitation data. The daily precipitation of rainless days in the sample actual precipitation data is less than or equal to a precipitation threshold.
7. A device, characterized in that, include: Processor, memory, system bus; The processor and the memory are connected via the system bus; The memory is used to store one or more programs, the one or more programs including instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on a terminal device, cause the terminal device to perform the method described in any one of claims 1-5.
Citation Information
Patent Citations
New random generation method of rainfall events
CN107423496A