Method for determining the sample for verifying the proportion of houses damaged by rainstorms and floods
By constructing high-dimensional spatial optimization samples and combining them with the Block Kriging model, the problems of large sample dependency and inaccurate assessment in traditional methods are solved. This enables rapid and accurate assessment of the proportion of house damage with low sample size, improving the scientific nature and timeliness of disaster assessment.
Patent Information
- Application Number
- CN202211666034.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-23
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2042-12-23
AI Technical Summary
Existing disaster assessment methods for rainstorm and flood disasters suffer from large sample dependence, inaccurate assessment results, and timeliness issues. Traditional sampling methods fail to effectively consider the spatial correlation of building damage.
A high-dimensional spatial optimization sample is constructed. By considering the average rainfall, maximum precipitation, runoff, and the proportion of brick and wood structure houses, the estimated variance of the sample combination is calculated using the Block Kriging model to determine the optimal sample combination, thereby reducing the sample size while ensuring the accuracy of the assessment.
By reducing the sample size, the accuracy and timeliness of the assessment are improved, enabling a rapid and accurate assessment of the proportion of damaged houses, reducing survey costs, and enhancing the scientific rigor of disaster assessment.
Smart Images

Figure CN115935117B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of geospatial information technology, specifically to a method for determining the sample for verifying the proportion of house damage caused by rainstorms and floods. Background Technology
[0002] Rainstorms and floods are a major natural disaster that seriously threatens human life and property. When floods occur, the accuracy, timeliness, and scientific rigor of disaster assessment are urgent needs for society and industry sectors. Among existing methods, local reporting and sampling surveys are two commonly used assessment methods. However, local reporting may have certain uncertainties due to issues such as statistical standards and the accuracy of information, necessitating the combination of local reporting and sampling survey methods.
[0003] Traditional random sampling or sampling methods based on expert knowledge / experience do not adequately consider the spatial correlation of disaster loss indicators such as the proportion of damaged houses across different statistical units. Although more and more new technologies and methods have been applied in disaster assessment in recent major natural disasters, the timeliness and accuracy of disaster assessment remain challenges in disaster loss assessment. Moreover, current accuracy generally relies on a large sample size, which seriously affects timeliness. Further research is needed on a method that can reduce the reliance on sample size while still ensuring the accuracy of assessment results. Summary of the Invention
[0004] To address the shortcomings of the existing technology, the present invention aims to provide a method for determining the sample of houses damaged by rainstorms and floods. This method considers key influencing factors on house damage, such as average rainfall, maximum precipitation, runoff, and brick-and-wood structure houses, and constructs a high-dimensional spatially optimized sample to make the spatial correlation of the measurement sample more reasonable. This method can ensure the accuracy of loss assessment with the fewest possible samples and enable timely and accurate assessment after a disaster.
[0005] Specifically, the present invention provides a method for determining the sample of houses damaged by rainstorms and floods, which includes the following steps:
[0006] S1. Collect the average rainfall, maximum rainfall, runoff, proportion of brick-and-wood structure houses, and proportion of damaged houses for each administrative unit to be investigated. Record the number of administrative units to be investigated as n.
[0007] S2. Normalize the average rainfall, maximum rainfall, runoff and the proportion of brick and wood structure houses to the range of [0,1], and use the normalized data to form a four-dimensional attribute space matrix denoted as X, with the size of X being n×4. Perform logarithmic preprocessing on the house damage proportion data, and the processed house damage proportion matrix is denoted as Y, with the size of Y being n×1.
[0008] S3. Calculate and generate a semi-variant scatter plot of the house damage ratio matrix Y as a function of distance in the four-dimensional attribute space matrix X, and fit the theoretical semi-variant function. The horizontal axis of the semi-variant scatter plot is the distance between the administrative units to be investigated, and the vertical axis is half of the square of the difference in the house damage ratio of the corresponding administrative units to be investigated.
[0009] Among them, the distance d between the administrative units to be investigated ij The calculation formula is shown in equation (3) below:
[0010]
[0011] Where, d ij It is the distance between the i-th administrative unit and the j-th administrative unit. These are the coordinates of the i-th administrative unit and the k-th administrative unit, respectively, where k = 1, ..., 4;
[0012] The square of the difference in the proportion of damaged houses in the administrative units to be investigated is r. ij The calculation formula is shown in formula (4):
[0013]
[0014] In the formula, r ij y′ is the semivariance of the housing damage ratio between the i-th and j-th administrative units. i y j ′ are the logarithmic values of the housing damage ratio for the i-th and j-th administrative units, respectively;
[0015] S4. Determine the sampling sample, which includes the following sub-steps:
[0016] S41. Starting from a selected minimum sample size, let the sample size be m. How many combinations are there of selecting m samples from n administrative units to be investigated? Iterate through these combinations, and each time arbitrarily select a non-repeating combination from these combinations. Calculate the estimated variance of different sample combinations for estimating the overall average housing damage ratio based on the horizontal and vertical coordinates calculated in step S3. Select the sample combination with the smallest estimated variance as the optimal sample combination for this sample size.
[0017] S42. Repeat step S41 after gradually increasing the sample size, and record the optimal sample combination for each sample size and its estimated variance for the overall housing damage ratio.
[0018] S43. Based on the estimated variance corresponding to different sample sizes, plot a scatter plot of sample size versus estimated variance to obtain the fitting function σ(m) = a·m between sample size m and estimated variance σ(m). bWhere a and b are the fitting parameters;
[0019] S44. When the estimated variance corresponding to a certain sample size m satisfies the following formula (5), the optimal sample combination corresponding to the sample size m is taken as the final sampling sample.
[0020] (σ m -σ m+1 ) / σ m <θ (5)
[0021] Where, σ m Let σ be the estimated variance of the m-th sample size. m+1 Let θ be the estimated variance of the (m+1)th sample size, and θ be the threshold.
[0022] Preferably, the threshold θ is 0.01 or 0.05.
[0023] Preferably, the formula for data normalization in step S2 is shown in equation (1) below:
[0024]
[0025] Where x i 、x′ i Let x and y be the i-th original data and the i-th normalized data, respectively, and max(x) and min(x) be the maximum and minimum values of the explanatory variable x, respectively.
[0026] Preferably, the formula for logarithmic preprocessing of the house damage ratio data in step S2 is shown in equation (2) below:
[0027] y′ i =log 10 (y i +10 -6 (2)
[0028] In the formula, y i y′ i These are the raw data and logarithmic data of the housing damage ratio for the i-th survey administrative unit, respectively.
[0029] Preferably, the specific method for calculating the estimated variance of the overall average housing damage ratio under different sample combinations in step S41 is to use the Block Kriging model to calculate the estimated variance of the overall average housing damage ratio under different combinations.
[0030] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0031] (1) When constructing the sample, the present invention considers the key factors affecting the damage to the house, such as average rainfall, maximum precipitation, runoff and brick and wood structure houses, so as to construct a high-dimensional spatial optimization sample, ensure the rationality of the sample, and achieve the purpose of scientifically and rationally selecting the sample.
[0032] (2) Compared with traditional methods, the method of the present invention makes the measurement of spatial correlation of samples more reasonable and has stronger interpretability. It can reduce the sample size required in actual surveys, reduce dependence on a large number of samples, and meet the accuracy of assessment on the basis of low sample size. It will not affect the assessment results due to the small sample size. It helps to make a faster and more accurate assessment of losses after rainstorm and flood disasters, and ensures the correctness of the assessment results. It can assess the losses after the disaster in a timely and accurate manner. Attached Figure Description
[0033] Figure 1 This is a schematic diagram of the process of the present invention;
[0034] Figure 2 The following are detailed step diagrams of an embodiment of the present invention;
[0035] Figure 3 This is a scatter plot illustrating the sample size-estimated variance relationship in an embodiment of the present invention. Detailed Implementation
[0036] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings.
[0037] This invention provides a method for determining the sample for verifying the proportion of house damage caused by rainstorm and flood disasters, such as... Figure 1 As shown, it includes the following steps:
[0038] S1. Collect the average rainfall, maximum rainfall, runoff, proportion of brick-and-wood structure houses, and proportion of damaged houses for each administrative unit to be investigated after the rainstorm and flood disaster. Record the number of administrative units to be investigated as n.
[0039] S2. Normalize the average rainfall, maximum precipitation, runoff, and proportion of brick-and-wood houses to the range of [0,1]. Use the normalized data to form a four-dimensional attribute space matrix, denoted as X, with a size of n×4. Perform logarithmic preprocessing on the house damage proportion data, and denot the processed house damage proportion matrix as Y, with a size of n×1.
[0040] The formula for data normalization is shown in equation (1) below:
[0041]
[0042] Where x i 、x′i Let x and y be the i-th original data and the i-th normalized data, respectively, and max(x) and min(x) be the maximum and minimum values of the explanatory variable x, respectively.
[0043] The formula for preprocessing the logarithmic proportion data of houses is shown in the following equation (2):
[0044] y′ i =log 10 (y i +10 -6 (2)
[0045] In the formula, y i y′ i These are the raw data and logarithmic data of the housing damage ratio for the i-th survey administrative unit, respectively.
[0046] S3. Calculate and generate a semi-variant scatter plot of the house damage ratio matrix Y as a function of distance within the four-dimensional attribute space matrix X, and fit it with the theoretical semi-variant function. The horizontal axis of the semi-variant scatter plot is the distance between the administrative units to be investigated, and the vertical axis is half of the square of the difference in the house damage ratio of the corresponding administrative units to be investigated.
[0047] Among them, the distance d between the administrative units to be investigated ij The calculation formula is shown in equation (3) below:
[0048]
[0049] Where, d ij It is the distance between the i-th administrative unit and the j-th administrative unit. These are the coordinates of the i-th administrative unit and the k-th administrative unit, respectively, where k = 1, ..., 4.
[0050] The square of the difference in the proportion of damaged houses in the administrative units to be investigated is r. ij The calculation formula is shown in formula (4);
[0051]
[0052] In the formula, r ij y′ is the semivariance of the housing damage ratio between the i-th and j-th administrative units. i y′ j These are the logarithmic values of the building damage ratio for the i-th and j-th administrative units, respectively.
[0053] S4. Determine the sampling sample, which includes the following sub-steps:
[0054] S41. Starting from a selected minimum sample size, let the sample size be m. How many combinations are there of selecting m samples from n administrative units to be investigated? Iterate through these combinations, and each time arbitrarily select a unique combination, calculate the estimated variance of the overall average housing damage ratio for different sample combinations, and select the sample combination with the smallest estimated variance as the optimal sample combination for that sample size.
[0055] In practical applications, the specific method for calculating the estimated variance of the overall average housing damage ratio for different sample combinations is to use the Block Kriging model. The x-axis and y-axis are respectively taken as half the distance between the administrative units to be investigated, calculated in step S3, and half the square of the difference in housing damage ratios between the administrative units to be investigated. The estimated variance is then calculated using the obtained x-axis and y-axis. Subsequently, the estimated variance of the overall average housing damage ratio for different combinations is calculated using the above method.
[0056] S42. After gradually increasing the sample size, repeat step S41 to calculate and record the optimal sample combination for each sample size and its estimated variance for the overall housing damage ratio.
[0057] S43. Based on the estimated variance corresponding to different sample sizes, plot a scatter plot of sample size versus estimated variance to obtain the fitting function σ(m) = a·m between sample size m and estimated variance σ(m). b , where a and b are the fitting parameters.
[0058] S44. When the estimated variance corresponding to a certain sample size m satisfies the following formula (5), the optimal sample combination corresponding to the sample size m is taken as the final sampling sample.
[0059] (σ m -σ m+1 ) / σ m <θ (5)
[0060] Where, σ m Let σ be the estimated variance of the m-th sample size. m+1 Let θ be the estimated variance of the (m+1)th sample size, and θ be the threshold.
[0061] In practical applications, the threshold θ is taken as a very small number, specifically 0.01 or 0.05. Specific Implementation
[0063] like Figure 2 As shown, the specific implementation steps of the present invention are as follows:
[0064] Step 1: Collect the average rainfall, maximum rainfall, runoff, proportion of brick and wood structure houses, and proportion of houses damaged as reported by the local authorities for each administrative unit (district, county or township, number n) affected by the rainstorm and flood disaster.
[0065] Step 2: Normalize the average rainfall, maximum rainfall, runoff, and proportion of brick-and-wood houses to the range [0, 1] according to formula (1). Use the normalized data (average rainfall, maximum rainfall, runoff, and proportion of brick-and-wood houses) to construct a four-dimensional attribute space point matrix (denoted as coordinate matrix X, with a size of n×4). Perform logarithmic preprocessing on the locally reported housing damage ratio data. The logarithmic preprocessing formula is shown in formula (2). The logarithmically processed data is denoted as matrix Y (matrix Y has a size of n×1).
[0066]
[0067] Where x i 、x′ i These are the i-th original data and normalized data, respectively, and max(x) and min(x) are the maximum and minimum values of the explanatory variable x, respectively.
[0068] y′ i =log 10 (y i +10 -6 (2)
[0069] In the formula, y i y′ i These are the raw and logarithmic data of the housing damage ratio for the i-th surveyed administrative unit, respectively. A very small number (10-) is added to the formula. 6 This is to avoid the possibility that the damage rate of a house may be 0.
[0070] Step 3: Calculate and generate a semi-variogram scatter plot of the proportion of damaged houses Y as a function of distance within the attribute space X, which consists of four-dimensional variables: average rainfall, maximum rainfall, runoff, and the proportion of brick-and-wood structured houses. Fit the theoretical semi-variogram function. The distance calculation formula between the administrative units to be investigated is shown in equation (3) below.
[0071]
[0072] In the formula, di j x′ is the attribute distance between the i-th and j-th administrative units. i k 、x′ j kThese are the coordinates of the kth (k = 1, ..., 4) of the i-th and j-th administrative units, respectively. The vertical axis is half the square of the corresponding housing damage ratio Y, as shown in equation (4) below.
[0073]
[0074] In the formula, r ij y′ is the semivariance of the housing damage ratio between the i-th and j-th administrative units. i y′ j These are the logarithmic values of the housing damage ratio for the i-th and j-th administrative units, respectively.
[0075] Step 4: Starting with a selected minimum sample size (sample size is set to m, not less than 2), iterate through all possible combinations of m samples in the total administrative unit to be investigated. Using the standard Bock-Kriging model, calculate the estimated variance of the overall average housing damage ratio under different combinations. Select the combination with the smallest variance as the most suitable sample for that sample size. Repeat this step as the sample size gradually increases, recording the most suitable sample combination for each sample size and its theoretical estimated variance of the overall housing damage ratio. Plot a scatter plot of "sample size - estimated variance". The scatter plot in this embodiment is shown below. Figure 3 As shown. Repeat this step until the rate of decrease of the estimated variance in the "Sample Size - Estimated Variance" scatter plot tends to stabilize, then use the sample combination corresponding to that sample size as the final sample. Alternatively, you can obtain a fitting function between the sample size and the estimated variance. According to the fitting function, the influence of the sample size on the estimated variance approaches zero, then use the sample combination corresponding to that sample size as the final sample.
[0076] In practical applications, the fitting function σ(m) = a·m can be obtained between the sample size m and the estimated variance σ(m). b , where a and b are the fitting parameters.
[0077] According to the fitting function, when the estimated variance corresponding to a certain sample size m satisfies the following equation (5), the optimal sample combination corresponding to the sample size m is taken as the final sampling sample.
[0078] (σ m -σ m+1 ) / σ m <θ (5)
[0079] Where, σ m Let σ be the estimated variance of the m-th sample size. m+1 Let θ be the estimated variance of the (m+1)th sample size, and θ be the threshold.
[0080] Comparison shows that, compared with traditional methods, the method of this invention makes the measurement of spatial correlation of samples more reasonable, has stronger interpretability, can reduce the sample size required in actual surveys, helps to make faster and more accurate assessments of losses after rainstorms and floods, and ensures the accuracy of assessment results, making it suitable for widespread application.
[0081] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.
Claims
1. A method for determining the sample for verifying the proportion of house damage caused by rainstorm and flood disasters, characterized in that: It includes the following steps: S1. Collect the average rainfall, maximum rainfall, runoff, proportion of brick-and-wood structure houses, and proportion of damaged houses for each administrative unit to be investigated. Record the number of administrative units to be investigated as n. S2. Normalize the average rainfall, maximum rainfall, runoff and the proportion of brick and wood structure houses to the range of [0,1], and use the normalized data to form a four-dimensional attribute space matrix denoted as X, with the size of X being n×4. Perform logarithmic preprocessing on the house damage proportion data, and the processed house damage proportion matrix is denoted as Y, with the size of Y being n×1. S3. Calculate and generate a semi-variant scatter plot of the house damage ratio matrix Y as a function of distance in the four-dimensional attribute space matrix X, and fit the theoretical semi-variant function. The horizontal axis of the semi-variant scatter plot is the distance between the administrative units to be investigated, and the vertical axis is half of the square of the difference in the house damage ratio of the corresponding administrative units to be investigated. Among them, the distance d between the administrative units to be investigated ij The calculation formula is shown in equation (3) below: Where, d ij It is the distance between the i-th administrative unit and the j-th administrative unit. These are the coordinates of the i-th administrative unit and the k-th administrative unit, respectively, where k = 1, ..., 4; The square of the difference in the proportion of damaged houses in the administrative units to be investigated is r. ij The calculation formula is shown in formula (4): In the formula, r ij y′ is the semivariance of the housing damage ratio between the i-th and j-th administrative units. i y j ′ are the logarithmic values of the housing damage ratio for the i-th and j-th administrative units, respectively; S4. Determine the sampling sample, which includes the following sub-steps: S41. Starting from a selected minimum sample size, let the sample size be m. How many combinations are there of selecting m samples from n administrative units to be investigated? Iterate through these combinations, and each time arbitrarily select a non-repeating combination from these combinations. Calculate the estimated variance of different sample combinations for estimating the overall average housing damage ratio based on the horizontal and vertical coordinates calculated in step S3. Select the sample combination with the smallest estimated variance as the optimal sample combination for this sample size. S42. Repeat step S41 after gradually increasing the sample size, and record the optimal sample combination for each sample size and its estimated variance for the overall housing damage ratio. S43. Based on the estimated variance corresponding to different sample sizes, plot a scatter plot of sample size versus estimated variance to obtain the fitting function σ(m) = a·m between sample size m and estimated variance σ(m). b Where a and b are the fitting parameters; S44. When the estimated variance corresponding to a certain sample size m satisfies the following formula (5), the optimal sample combination corresponding to the sample size m is taken as the final sampling sample. (s m -s m+1 ) / s m <θ (5) Where, σ m Let σ be the estimated variance of the m-th sample size. m+1 Let θ be the estimated variance of the (m+1)th sample size, and θ be the threshold.
2. The method for determining the sample for verifying the proportion of house damage caused by rainstorm and flood disasters according to claim 1, characterized in that: The threshold θ is 0.01 or 0.
05.
3. The method for determining the sample for verifying the proportion of house damage caused by rainstorm and flood disasters according to claim 1, characterized in that: The formula for data normalization in step S2 is shown in equation (1) below: Where x i 、x′ i These are the i-th original data and normalized data, respectively, and max(x) and min(x) are the maximum and minimum values of the explanatory variable x, respectively.
4. The method for determining the sample for verifying the proportion of house damage caused by rainstorm and flood disasters according to claim 1, characterized in that: The formula for logarithmic preprocessing of the housing damage ratio data in step S2 is shown in equation (2) below: and' i =log 10 (and i +10 -6 ) (2) In the formula, y i y′ i These are the raw data and logarithmic data of the housing damage ratio for the i-th survey administrative unit, respectively.
5. The method for determining the sample for verifying the proportion of house damage caused by rainstorm and flood disasters according to claim 1, characterized in that: The specific method for calculating the estimated variance of the overall average housing damage ratio under different sample combinations in step S41 is to use the Block Kriging model to calculate the estimated variance of the overall average housing damage ratio under different combinations.
Citation Information
Patent Citations
Space sampling method oriented to multisource marine environmental monitoring data
CN103678883A
Estimation method and apparatus
US20190197435A1