Methods, devices, equipment and media for constructing dynamic monitoring models for rice waterlogging disasters
By integrating rainfall accumulation, continuous rainy days, and topographic slope index, a dynamic monitoring model for rice waterlogging disasters was constructed, which solved the problems of poor timeliness and limited identification accuracy of existing monitoring results, and achieved high-precision monitoring grouped by growth stage.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA INST OF WATER RESOURCES & HYDROPOWER RES
- Filing Date
- 2025-10-30
- Publication Date
- 2026-06-30
Smart Images

Figure CN121350530B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent analysis technology for agricultural information, and in particular to a method, apparatus, equipment and medium for constructing a dynamic monitoring model for rice waterlogging disasters. Background Technology
[0002] Rice, as a major food crop in my country, is crucial to the national economy and people's livelihood. Waterlogging is a key natural disaster restricting its high and stable yields. Against the backdrop of global climate change, frequent extreme rainfall events have exacerbated the intensity and frequency of waterlogging in major rice-producing areas (such as the middle and lower reaches of the Yangtze River). This region is characterized by low-lying terrain and a dense river network, making it susceptible to widespread waterlogging during the plum rain season. Therefore, achieving high-precision, dynamic monitoring of rice waterlogging is an urgent need to improve agricultural disaster prevention and mitigation capabilities and ensure food security.
[0003] Currently, research on waterlogging disasters largely focuses on dryland crops such as corn and wheat, with relatively weak specialized research on rice. Existing methods for monitoring waterlogging in rice have significant limitations: first, they rely heavily on rainfall as a single indicator, neglecting the impact of topography on water accumulation and drainage, making it difficult to accurately depict the spatial pattern of disasters; second, they often construct monitoring indicators based on meteorological data at a ten-day or monthly scale, leading to delays in disaster identification due to limited temporal resolution; and third, they do not fully consider the differences in flood tolerance at different growth stages of rice, lacking stage-specific monitoring models and thresholds, resulting in limited identification accuracy. Therefore, there is an urgent need to develop a dynamic monitoring technology system for rice waterlogging that integrates multi-source data and differentiates between growth stages. Summary of the Invention
[0004] Based on this, the present invention provides a method, device, equipment and medium for constructing a dynamic monitoring model for rice waterlogging disasters, in order to solve the problems of single indicators, neglect of topographic influence, poor timeliness of monitoring results and failure to distinguish the differences in flood tolerance of rice during its growth period in existing rice waterlogging monitoring, and to realize dynamic monitoring of rice waterlogging disasters on a daily scale.
[0005] In a first aspect, embodiments of the present invention provide a method for constructing a dynamic monitoring model for rice waterlogging disasters, including:
[0006] Acquire historical waterlogging disaster data, historical daily rainfall data, digital elevation model data, and growth period data of the target rice in the target area. The historical waterlogging disaster data includes the disaster level.
[0007] The effective cumulative rainfall index and the continuous rainfall days index are calculated based on the historical daily rainfall data. The slope index is calculated based on the digital elevation model data. The three types of indices after normalization are used as feature variables, and the disaster level is used as a label variable to construct a sample dataset that associates features with disaster level.
[0008] The sample dataset is input into a random forest model to calculate the importance weights of each feature variable, and the importance weights are used as the weight coefficients of the corresponding indices to generate a calculation model for the comprehensive index of rice waterlogging disaster.
[0009] The comprehensive index value of each sample in the sample dataset is solved using the calculation model. The samples are then grouped based on the rice growth period data to construct sample subsets that contain disaster level and comprehensive index value for each growth period.
[0010] Normality tests were performed on each sample subset. For normally distributed sample subsets, the t-distribution interval estimation method was used to determine the comprehensive index grading thresholds corresponding to different disaster levels in each reproductive period.
[0011] By associating and storing the computational model with the graded thresholds for each growth stage, a model that can be directly used for dynamic monitoring of rice waterlogging disasters is formed.
[0012] Secondly, embodiments of the present invention also provide a device for constructing a dynamic monitoring model for rice waterlogging disasters, comprising:
[0013] The historical monitoring data acquisition module is used to acquire historical waterlogging disaster data, historical daily rainfall data, digital elevation model data, and growth period data of the target rice in the target area. The historical waterlogging disaster data includes the disaster level.
[0014] The sample dataset construction module is used to calculate the effective cumulative rainfall index and the continuous rainfall days index based on the historical daily rainfall data, calculate the slope index based on the digital elevation model data, use the three types of indices after normalization as feature variables, and use the disaster level as label variable to construct a sample dataset in which features are associated with disaster level.
[0015] The computational model generation module is used to input the sample dataset into the random forest model to calculate the importance weights of each feature variable, and use the importance weights as the weight coefficients of the corresponding index to generate a computational model for the comprehensive index of rice waterlogging disaster.
[0016] The sample subset construction module is used to solve the comprehensive index value of each sample in the sample dataset using the calculation model, and to group the samples in combination with the rice growth period data to construct sample subsets that contain disaster level and comprehensive index value in each growth period.
[0017] The grading threshold determination module is used to perform normality tests on each sample subset. For normally distributed sample subsets, the t-distribution interval estimation method is used to determine the comprehensive index grading threshold corresponding to different disaster levels in each reproductive period.
[0018] The monitoring model generation module is used to generate a model that can be directly used for dynamic monitoring of rice waterlogging disasters by associating and storing the calculation model with the graded thresholds of each growth stage.
[0019] Thirdly, embodiments of the present invention also provide an electronic device, the electronic device comprising:
[0020] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to execute a method for constructing a dynamic monitoring model for rice waterlogging disasters according to any embodiment of the present invention.
[0021] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing computer instructions, which are used to cause a processor to execute a method for constructing a dynamic monitoring model for rice waterlogging disasters as described in any embodiment of the present invention.
[0022] This invention overcomes the shortcomings of traditional methods that rely on single rainfall data by integrating three key indicators: accumulated rainfall, consecutive rainy days, and topographic slope. This significantly improves the comprehensive judgment ability on the formation mechanism of waterlogging. It sets monitoring thresholds according to rice growth stages, solving the core problem of traditional models ignoring the differences in flood tolerance at different crop stages and greatly improving the accuracy of identification. It uses random forest automatic weighting combined with statistical inference to determine the thresholds, eliminating human experience intervention and ensuring the objectivity and reliability of the model. It constructs an integrated monitoring model to achieve a closed loop from historical modeling to real-time judgment, effectively solving the problem of delayed early warning in traditional methods.
[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is a flowchart of a method for constructing a dynamic monitoring model for rice waterlogging disasters according to Embodiment 1 of the present invention;
[0026] Figure 2This is a flowchart of another method for constructing a dynamic monitoring model for rice waterlogging disasters according to Embodiment 2 of the present invention;
[0027] Figure 3 This is a schematic diagram of a device for constructing a dynamic monitoring model for rice waterlogging disasters according to Embodiment 3 of the present invention;
[0028] Figure 4 This is a schematic diagram of the structure of an electronic device for implementing a method for constructing a dynamic monitoring model for rice waterlogging disasters according to an embodiment of the present invention. Detailed Implementation
[0029] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0030] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0031] Example 1
[0032] Figure 1 This is a flowchart of a method for constructing a dynamic monitoring model for rice waterlogging disasters according to Embodiment 1 of the present invention. This embodiment is applicable to situations where waterlogging disasters are monitored over a large area of rice. The method can be executed by a device for constructing a dynamic monitoring model for rice waterlogging disasters. This device can be implemented in hardware and / or software, and can be configured in a server equipped with a data acquisition interface and waterlogging disaster modeling and monitoring program. Figure 1 As shown, the method includes:
[0033] S110. Obtain historical waterlogging disaster data, historical daily rainfall data, digital elevation model data, and growth period data of the target rice in the target area. The historical waterlogging disaster data includes the disaster level.
[0034] Historical waterlogging disaster data can be obtained through legitimate channels such as the disaster archives of agricultural management departments in the target area, annual disaster survey reports, and field disaster census records. This data must include the disaster level information for each waterlogging disaster event (e.g., it can be divided into three levels: mild, moderate, and severe; specific classification standards can refer to industry-standard norms), as well as the time and spatial information of the disaster to ensure correlation with other data. Historical daily rainfall data can be extracted from the historical databases of meteorological observation stations in the target area. This data records rainfall information on a daily basis, and its temporal coverage should match the time range of historical waterlogging disaster data to ensure that subsequent rainfall characteristics at the time of the disaster can be traced based on the rainfall data. Digital elevation model (DEM) data can be obtained from geographic information data service platforms for the target area. This data reflects the topographic features of the target area, and its spatial coverage should completely include the target rice planting area. The growth period data of the target rice can be obtained through agricultural technology information, crop variety characteristic manuals, or local agricultural technology extension departments. This data needs to clearly define the time intervals of each growth stage (such as sowing period, tillering period, heading period, etc.) of the target rice variety in the target area, so as to provide a basis for subsequent grouping by growth period.
[0035] S120. Calculate the effective cumulative rainfall index and the continuous rainfall days index based on the historical daily rainfall data, calculate the slope index based on the digital elevation model data, use the three types of indices after normalization as feature variables, use the disaster level as label variables, and construct a sample dataset that associates features with disaster level.
[0036] This embodiment generates a sample dataset for model training based on the acquired basic data. Normalization refers to normalizing the three types of indices mentioned above to eliminate the influence of differences in the units of measurement between different indices. Sample construction involves using the normalized effective cumulative rainfall index, consecutive rainfall days index, and slope index as feature variables, and the disaster level from historical flooding disaster data as label variables, constructing the sample dataset through spatiotemporal matching. For example, associating the three types of index values corresponding to a disaster event with the moderate disaster level label of that event forms a sample dataset.
[0037] Optionally, calculating the effective cumulative rainfall index and the consecutive rainfall days index based on the historical daily rainfall data, and calculating the slope index based on the digital elevation model data, may include:
[0038] Based on historical daily rainfall data, the effective cumulative rainfall is calculated iteratively day by day according to the time series. The calculation formula is as follows: ,in, For the first Daily effective cumulative rainfall, For the first Daily rainfall, Let be the rainfall attenuation coefficient, and Initialize to 0 in the first iteration. After completing the full time series iteration, use the obtained effective cumulative rainfall as the effective cumulative rainfall index.
[0039] Based on historical daily rainfall data, the system sequentially determines whether the daily rainfall is not less than a preset rainfall threshold. If so, the number of consecutive rainy days on that day is counted as the number of consecutive rainy days on the previous day plus one. If not, the number of consecutive rainy days on that day is reset to 0. The consecutive rainfall days index is obtained by summing the consecutive days in the time series.
[0040] Based on digital elevation model data, for each pixel in the target area, the elevation change rate in the x and y directions is calculated according to the elevation difference between the current pixel and its neighboring pixels. Then, by substituting the elevation change rates in the x and y directions into the preset slope calculation formula, the slope value of each pixel is calculated. The set of slope values of all pixels is the slope index.
[0041] Based on the time series of historical daily rainfall data, calculations are performed day by day starting from the first date. First iteration ( When ), the effective cumulative rainfall of the previous day ( Initialize to 0, the effective cumulative rainfall for the day ( , that is ;Day 2 ( )hour,( This process is repeated until the entire time series is iterated. The value of the rainfall attenuation coefficient W needs to be determined in conjunction with the flood tolerance characteristics of the rice varieties in the target area. After completing the full time series iteration, the obtained daily effective cumulative rainfall constitutes a continuous series, such as... to This sequence is the effective cumulative rainfall index, which can be directly used for subsequent time matching with disaster events.
[0042] The preset rainfall threshold needs to be set based on the minimum daily rainfall that can affect rice growth. In a specific example, if the rainfall threshold is set to 5mm (a day with rainfall ≥ 5mm is considered a valid rainfall day; < 5mm is considered no valid rainfall), the following criteria are applied sequentially according to the time sequence: if the rainfall on day i is ≥ 5mm, and the number of consecutive rainfall days on day i-1 is n, then the number of consecutive rainfall days on day i is counted as n+1; if the rainfall on day i is < 5mm, then regardless of the previous day's count, the count for that day is reset to 0. Example: A 5-day rainfall sequence is [6mm, 8mm, 3mm, 7mm, 9mm]. The corresponding consecutive rainfall days are calculated as follows: Day 1 = 1 (first valid rainfall), Day 2 = 2 (consecutive valid rainfall), Day 3 = 0 (rainfall < 5mm), Day 4 = 1 (valid rainfall resumes), Day 5 = 2 (consecutive valid rainfall). This sequence is the consecutive rainfall days index, and the value on the day of the disaster can be extracted as a sample feature.
[0043] slope Slope refers to the degree of inclination of a point on the Earth's surface relative to the horizontal plane. It is a key indicator for quantifying topographic relief and determines the velocity of surface runoff; the steeper the slope, the faster the drainage. The slope calculation formula for a given pixel in the elevation is as follows: ,in, and This represents the rate of change of elevation in the x and y directions.
[0044] Furthermore, by using the three normalized indices as feature variables and the disaster level as a label variable, a sample dataset relating features to disaster levels can be constructed, which may include:
[0045] Read each disaster event record from the historical flood disaster data and extract the occurrence time, location, and disaster level of each disaster event record;
[0046] Based on the occurrence time, extract the target effective cumulative rainfall index value at the current occurrence time from the effective cumulative rainfall index, and extract the target continuous rainfall day index value at the current occurrence time from the continuous rainfall day index;
[0047] Based on the location, extract the target slope index value for the current location from the set of pixel values of the slope index;
[0048] The extracted target effective cumulative rainfall index, target continuous rainfall days index, and target slope index are combined into a feature vector, and this feature vector is associated with the disaster level, the time and location of the disaster event to form a complete sample data.
[0049] By iterating through all disaster event records in the historical flood disaster data, and repeating the above steps for each disaster event record to obtain all sample data, a sample dataset with features correlated with disaster level is constructed.
[0050] If historical flooding disaster data is in paper form, it must first be digitized into structured data, with fields including at least: disaster number, occurrence time, location, and disaster level. If it is electronic data, data integrity must be checked. The occurrence time must be accurate to the day. If the original record is only accurate to the month, the specific date must be calculated by combining rainfall data from the same period and field survey records; if it cannot be accurate to the day, the record must be excluded to avoid excessive time matching errors. Based on the occurrence time of the disaster event, extract the corresponding date values from the effective cumulative rainfall index sequence and the continuous rainfall days index sequence. For example, for a disaster that occurred on July 15, 2021, directly extract the effective cumulative rainfall index value and the continuous rainfall days index value for that day. Based on the latitude and longitude coordinates of the disaster location, locate the corresponding cell in the slope index raster data through spatial interpolation (such as nearest neighbor interpolation), and extract the slope value of that cell as the target slope index value. The three types of index values (e.g., [effective cumulative rainfall index, consecutive rainfall days index, slope index]) are combined in a fixed order to form a feature vector with a dimension of 3. Each sample data point must contain "feature vector + label + spatiotemporal identifier," that is: the feature vector is the three types of index values, the label is the disaster level, and the spatiotemporal identifier is the time and location of occurrence. A programming script is used to automatically traverse all disaster event records, repeating the above extraction, matching, and combination steps for each record. The final sample dataset can be stored in CSV format or a database table, with fields including "effective cumulative rainfall index, consecutive rainfall days index, slope index, disaster level, time of occurrence, and location," facilitating subsequent reading and filtering.
[0051] S130. Input the sample dataset into the random forest model to calculate the importance weight of each feature variable, and use the importance weight as the weight coefficient of the corresponding index to generate a calculation model for the comprehensive index of rice waterlogging disaster.
[0052] The random forest model was chosen because it effectively handles the nonlinear relationships among multiple feature variables (rice waterlogging disasters are influenced by both rainfall and topography, and these influences are not simply linearly additive), and it can automatically calculate feature importance. Compared to models like linear regression, it is more robust to outliers (suitable for a small number of possible anomalies in historical disaster data). Key points for weight calculation include: clearly distinguishing between feature variables and label variables when inputting the sample dataset (feature variables are three-class normalized indices, and label variables are disaster levels); ensuring a balanced sample size for each disaster level during model training; and validating the weight calculation results.
[0053] S140. Using the calculation model, solve for the comprehensive index value of each sample in the sample dataset, and group the samples according to the rice growth period data to construct sample subsets that contain disaster level and comprehensive index value in each growth period.
[0054] Rice exhibits significantly different tolerances to waterlogging at different growth stages. For example, during the tillering stage, the root system is shallow, and even mild waterlogging can lead to reduced tillering. Waterlogging during the grain-filling stage directly impacts grain filling, resulting in a higher risk of yield reduction. If samples are not grouped by growth stage, a uniform threshold can lead to misclassification of severe waterlogging during the tillering stage and underclassification during the grain-filling stage. Therefore, it is essential to separate samples by growth stage. A growth stage time reference table must first be established, and then the date of the disaster for each sample must be matched against the table. If the date of the disaster falls at the boundary between two growth stages, field records must be consulted to determine whether the plants have entered the jointing stage, thus avoiding errors caused by mechanically grouping by date.
[0055] S150. Perform a normality test on each sample subset. For normally distributed sample subsets, use the t-distribution interval estimation method to determine the comprehensive index grading threshold corresponding to different disaster levels in each reproductive period.
[0056] The t-distribution interval estimation method assumes that the data conforms to a normal distribution. If the sample data is skewed, directly using the t-distribution will lead to bias in the threshold calculation. Therefore, it is necessary to first screen out a sample subset that meets the conditions to ensure the statistical validity of the threshold calculation. In a specific example, taking the threshold for a minor disaster as an example, when using the t-distribution interval estimation method, it is necessary to calculate the lower limit of the 95% confidence interval of the comprehensive index of the minor disaster sample. This is because the lower limit can cover the comprehensive index range of most minor disasters, avoiding misjudging normal situations close to minor as disasters. The thresholds for different disaster levels must meet the logical order of "minor threshold < moderate threshold < severe threshold". If threshold overlap occurs, it is necessary to check whether there are anomalies in the sample data.
[0057] S160. By associating and storing the calculation model with the graded thresholds of each growth stage, a model that can be directly used for dynamic monitoring of rice waterlogging disasters is formed.
[0058] The comprehensive index calculation model can be encapsulated into a callable algorithm module using an algorithm + database storage structure. The threshold values for each growth stage are stored in a database table, and the module is associated with the database through the name of the growth stage. The monitoring model can be directly used for real-time monitoring because it integrates the entire calculation and judgment process. During real-time monitoring, only real-time rainfall data, real-time plot slope data, and the current growth stage need to be input to automatically calculate the comprehensive index, match the corresponding thresholds, and output the disaster level. This solves the efficiency problem of traditional methods that require repeated manual calculations and judgments. Furthermore, when covering a large area, only batch input of data from different plots is needed, meeting the needs of large-scale monitoring.
[0059] Optionally, after storing the computational model in association with the grading thresholds for each growth stage to form a model that can be directly used for monitoring the level of rice waterlogging disasters, it may further include:
[0060] Real-time daily rainfall data is obtained through regional meteorological monitoring stations, pre-processed digital elevation model data of the target area is retrieved through geographic information systems, and the current growth stage information of the target rice is obtained through field IoT equipment monitoring or planting record query.
[0061] Based on the real-time daily rainfall data, the real-time effective cumulative rainfall index and the continuous rainfall days index are calculated. Based on the digital elevation model data of the target area, the real-time slope index is calculated.
[0062] After normalizing the real-time effective cumulative rainfall index, continuous rainfall days index, and slope index, the data are input into the dynamic monitoring model for rice waterlogging disaster to obtain the current real-time comprehensive index value of rice waterlogging disaster.
[0063] From the set of comprehensive index grading thresholds for each growth stage, which are stored in association with the dynamic monitoring model of rice waterlogging disaster and sorted from low to high according to the disaster level, the target grading threshold corresponding to the target growth stage of the target rice is determined.
[0064] The real-time comprehensive index value of rice waterlogging disaster is matched with the determined target classification threshold, and the waterlogging disaster level of the target rice in the target area is determined according to the threshold interval it falls into.
[0065] Real-time daily rainfall data is obtained from regional meteorological monitoring stations in the target area. The data must include the actual measured rainfall for the day, and the acquisition frequency can be once daily to ensure the data's timeliness meets real-time calculation requirements. Digital elevation model (DEM) data is retrieved from the target area via a geographic information system. This data must be pre-processed raster data, and its spatial extent must completely cover the target monitoring area. Current growth stage information can be obtained in two ways: first, through real-time monitoring of plant growth status using field IoT devices (such as crop growth sensors); second, by deducing it from the target rice's planting records (such as sowing date and variety characteristics records). The hierarchical thresholds associated with the dynamic monitoring model for rice waterlogging disasters are organized in a three-dimensional structure of "growth stage-disaster level-threshold" (e.g., for the Jiangsu demonstration area, during the tillering stage: mild waterlogging 0.2–0.32; moderate waterlogging 0.32–0.48; severe waterlogging >0.48; for the Hubei demonstration area, during the tillering stage: mild waterlogging 0.27–0.38; moderate waterlogging 0.38–0.55; severe waterlogging >0.55). The hierarchical thresholds for each growth stage are stored in ascending order of disaster level. Based on the current target growth stage of the target rice variety, the corresponding target hierarchical threshold set is retrieved and extracted from the above storage structure.
[0066] This invention overcomes the shortcomings of traditional methods that rely on single rainfall data by integrating three key indicators: accumulated rainfall, consecutive rainy days, and topographic slope. This significantly improves the comprehensive judgment ability on the formation mechanism of waterlogging. It sets monitoring thresholds according to rice growth stages, solving the core problem of traditional models ignoring the differences in flood tolerance at different crop stages and greatly improving the accuracy of identification. It uses random forest automatic weighting combined with statistical inference to determine the thresholds, eliminating human experience intervention and ensuring the objectivity and reliability of the model. It constructs an integrated monitoring model to achieve a closed loop from historical modeling to real-time judgment, effectively solving the problem of delayed early warning in traditional methods.
[0067] Example 2
[0068] Figure 2 This is a flowchart of another method for constructing a dynamic monitoring model for rice waterlogging disasters provided in Embodiment 2 of the present invention. This embodiment is a refinement based on Embodiment 1, specifically as follows: Figure 2 As shown, the method includes:
[0069] S210. Obtain historical waterlogging disaster data, historical daily rainfall data, digital elevation model data, and growth period data of the target rice in the target area. The historical waterlogging disaster data includes the disaster level.
[0070] S220. Calculate the effective cumulative rainfall index and the continuous rainfall days index based on the historical daily rainfall data, calculate the slope index based on the digital elevation model data, use the three types of indices after normalization as feature variables, use the disaster level as label variables, and construct a sample dataset that associates features with disaster level.
[0071] S230. Input the sample dataset into the random forest model to calculate the importance weight of each feature variable, and use the importance weight as the weight coefficient of the corresponding index to generate a calculation model for the comprehensive index of rice waterlogging disaster.
[0072] Optionally, the sample dataset is input into a random forest model to calculate the importance weights of each feature variable, and the importance weights are used as the weight coefficients of the corresponding indices to generate a calculation model for the comprehensive index of rice waterlogging disaster. This model may include:
[0073] The sample dataset is split into training data and independent validation data, and input into a preset random forest model for training. When the prediction accuracy of the random forest model on the training data reaches a preset accuracy threshold, and the prediction accuracy on the independent validation data stably reaches a preset validation threshold, the model training is deemed complete, and the trained random forest model is obtained.
[0074] The built-in feature importance assessment function in the trained random forest model is invoked, and the importance scores of the three feature variables—effective cumulative rainfall index, consecutive rainfall days index, and slope index—are calculated using the average impurity reduction method.
[0075] The importance scores are normalized so that the sum of the weight coefficients is 1. The final weight coefficients corresponding to the effective cumulative rainfall index, the consecutive rainfall days index, and the slope index are then obtained, denoted as […]. , , ;
[0076] Based on the final weighting coefficient A calculation model for the comprehensive index of rice waterlogging disaster was constructed, denoted as . The expression for the computational model is: in, , , These represent the normalized effective cumulative rainfall index, the consecutive rainfall days index, and the slope index, respectively.
[0077] Stratified random sampling was used to split the sample dataset, with stratification based on disaster level. This ensured that the proportions of mild, moderate, and severe disaster samples in the training and validation sets were completely consistent, avoiding model bias towards predicting disaster levels with a high proportion of samples due to uneven sample distribution. The calculated importance scores were normalized using a sum-of-the-parts normalization method, dividing the original score of each feature by the sum of the scores of the three features, so that the sum of the processed weight coefficients was 1. The final weight coefficients were then used as the basis for the normalization. , , A calculation model for the comprehensive index of rice waterlogging disaster was constructed, in which: This represents the normalized effective cumulative rainfall index; Represents the normalized index of consecutive rainfall days; This represents the normalized slope index; the RIWI calculation result ranges from 0 to 1 (a higher value indicates a higher risk of waterlogging). In the random forest model, each decision tree reduces node impurity through feature splitting. The importance score of a feature = the sum of the differences in impurity of that feature before and after splitting across all decision trees ÷ the number of decision trees. A higher score indicates a greater contribution of that feature to the disaster level classification. Assuming there are 100 decision trees, the effective cumulative rainfall index (... The total impurity reduction was 42, and the continuous rainfall days index ( The slope index is 31. If the score is 27, then the original importance scores for the three factors are 0.42, 0.31, and 0.27, respectively.
[0078] S240. Traverse each sample data in the sample dataset, and input the normalized effective cumulative rainfall index, continuous rainfall days index and slope index of each sample data into the calculation model of the comprehensive index of rice waterlogging disaster, and solve to obtain the comprehensive index value corresponding to each sample data.
[0079] Extract the normalized effective cumulative rainfall index from each sample data point. ), consecutive rainfall days index ( ) and slope index ( The three index values mentioned above are input into the comprehensive index calculation model for rice waterlogging disaster. The comprehensive index value corresponding to the sample data is calculated by the model and stored in association with the original sample data.
[0080] S250. Read the rice growth period data of the target rice variety and obtain the time range of each growth stage of the target rice variety in the target area.
[0081] Read the rice growth period data of the target rice variety. This data includes the division information of each growth stage (such as tillering stage, jointing stage, grain filling stage, etc.) of the variety from sowing to maturity in the target area. Extract the time range corresponding to each growth stage from the growth period data to form a "growth period-time range" correspondence table.
[0082] S260. Extract the time of occurrence of the disaster event associated with each sample data, and determine the reproductive period to which the current sample data belongs by matching the time range.
[0083] Extract the occurrence time of the disaster event associated with each sample data, match the occurrence time with the "fertility period-time range" correspondence table to determine the fertility period to which the sample data belongs (e.g., if the occurrence time is May 20, then match to the tillering stage).
[0084] S270. Group the sample dataset according to the reproductive period, summarize the comprehensive index value and corresponding disaster level in each group of sample data, and construct a sample subset in which the disaster level and comprehensive index value are associated in each reproductive period.
[0085] Based on the matching results above, the sample dataset is grouped according to the reproductive period. After summarizing the sample data of each group, a sample subset containing two core pieces of information, namely the comprehensive index value and the corresponding disaster level, is formed. The sample subsets of all reproductive periods together constitute the sample set in which the disaster level and the comprehensive index value are associated within each reproductive period, which is used for subsequent calculation of the classification threshold.
[0086] S280. Perform a normality test on each sample subset. For the normally distributed sample subset, use the t-distribution interval estimation method to determine the comprehensive index grading threshold corresponding to different disaster levels in each reproductive period.
[0087] Furthermore, a normality test is performed on each sample subset. For normally distributed sample subsets, the t-distribution interval estimation method is used to determine the comprehensive index grading thresholds corresponding to different disaster levels within each reproductive period. This may include:
[0088] Each sample subset was grouped separately according to the disaster level, resulting in a set of comprehensive index values that correspond one-to-one with the disaster level within the same reproductive period;
[0089] For each group of composite index values within the current reproductive period, a normality test is performed. If any group of target composite index values does not conform to a normal distribution, then all composite index values within the target composite index value group are mathematically transformed until the target composite index value group conforms to a normal distribution.
[0090] For each group of comprehensive index values that conforms to a normal distribution within the current reproductive period, the t-distribution interval estimation method is used to calculate its confidence interval at a given confidence level. The lower limit of each confidence interval is used as the comprehensive index classification threshold for the corresponding disaster level within the current reproductive period.
[0091] By sorting all comprehensive index grading thresholds within the same reproductive period in order of disaster level from low to high, a set of comprehensive index grading thresholds for the current reproductive period is constructed.
[0092] For each subset of samples from different reproductive stages, separate groups are formed according to the severity of the disaster (e.g., mild, moderate, severe). For example, the tillering stage sample subset contains 100 samples, of which 30 are mild, 50 are moderate, and 20 are severe. After grouping, three independent sets of comprehensive index values are formed: "Tillering Stage - Mild," "Tillering Stage - Moderate," and "Tillering Stage - Severe." Grouping is performed using data processing tools, filtering by the dual key fields of reproductive stage and disaster severity to ensure that each comprehensive index value group contains only sample data from the same reproductive stage and the same disaster severity, and that the data format within each group is consistent. The Shapiro-Wilke test is used to test the normality of each comprehensive index value group. The test statistic W ranges from 0 to 1, and a significance level α = 0.05 is set. If the p-value > 0.05, the data in that group is considered to conform to a normal distribution; if the p-value ≤ 0.05, it is considered not to conform to a normal distribution. For target composite index value groups that do not conform to a normal distribution, mathematical transformation methods are used until they conform to a normal distribution. The Shapiro-Wilke test must be repeated after each transformation. For composite index value groups that conform to a normal distribution, the t-distribution interval estimation method is used to calculate the confidence interval. If a given confidence level of 95% is set, the lower limit of the calculated 95% confidence interval is extracted as the composite index classification threshold for the corresponding disaster level. For example, the confidence interval for the "tillering stage - mild" group is [0.28, 0.35], then the mild disaster classification threshold is 0.28. All composite index classification thresholds within the same reproductive period are sorted in ascending order of disaster level to form a set of composite index classification thresholds for that reproductive period. The final constructed set of classification thresholds for each reproductive period needs to be stored as structured data, with fields including: reproductive period name, mild threshold, moderate threshold, severe threshold, sample size, and confidence level, for easy retrieval by the subsequent monitoring model.
[0093] S290. By associating and storing the calculation model with the grading thresholds of each growth stage, a rice waterlogging disaster level monitoring model that can be directly used for real-time monitoring is formed.
[0094] This invention addresses the problems of subjective weighting and unclear multi-factor fusion logic in existing technologies by objectively calculating and normalizing the importance weights of three indices: effective cumulative rainfall, number of consecutive rainy days, and slope using a random forest model. This enables the constructed comprehensive index calculation model to scientifically quantify the contribution of each factor and reduce human error. By associating the sample comprehensive index value with the growth period and grouping the samples into subsets, this invention overcomes the shortcomings of existing technologies that do not consider the differences in flood resistance of rice at different growth periods and lack correlation between samples and growth periods. This lays the foundation for determining accurate thresholds for subsequent growth periods and avoids bias caused by a one-size-fits-all assessment.
[0095] Example 3
[0096] Figure 3 This is a schematic diagram of a device for constructing a dynamic monitoring model for rice waterlogging disasters, provided in Embodiment 3 of the present invention. Figure 3 As shown, the device includes:
[0097] The historical monitoring data acquisition module 310 is used to acquire historical waterlogging disaster data, historical daily rainfall data, digital elevation model data, and growth period data of the target rice in the target area. The historical waterlogging disaster data includes the disaster level.
[0098] The sample dataset construction module 320 is used to calculate the effective cumulative rainfall index and the continuous rainfall days index based on the historical daily rainfall data, calculate the slope index based on the digital elevation model data, use the three types of indices after normalization as feature variables, and use the disaster level as a label variable to construct a sample dataset in which features are associated with disaster level.
[0099] The calculation model generation module 330 is used to input the sample dataset into the random forest model to calculate the importance weight of each feature variable, and use the importance weight as the weight coefficient of the corresponding index to generate a calculation model for the comprehensive index of rice waterlogging disaster.
[0100] The sample subset construction module 340 is used to solve the comprehensive index value of each sample in the sample dataset using the calculation model, and to group the samples in combination with the rice growth period data to construct sample subsets that contain disaster level and comprehensive index value in each growth period.
[0101] The grading threshold determination module 350 is used to perform normality tests on each sample subset. For normally distributed sample subsets, the t-distribution interval estimation method is used to determine the comprehensive index grading threshold corresponding to different disaster levels in each reproductive period.
[0102] The monitoring model generation module 360 is used to form a model that can be directly used for dynamic monitoring of rice waterlogging disasters by associating and storing the calculation model with the graded thresholds of each growth stage.
[0103] This invention overcomes the shortcomings of traditional methods that rely on single rainfall data by integrating three key indicators: accumulated rainfall, consecutive rainy days, and topographic slope. This significantly improves the comprehensive judgment ability on the formation mechanism of waterlogging. It sets monitoring thresholds according to rice growth stages, solving the core problem of traditional models ignoring the differences in flood tolerance at different crop stages and greatly improving the accuracy of identification. It uses random forest automatic weighting combined with statistical inference to determine the thresholds, eliminating human experience intervention and ensuring the objectivity and reliability of the model. It constructs an integrated monitoring model to achieve a closed loop from historical modeling to real-time judgment, effectively solving the problem of delayed early warning in traditional methods.
[0104] Optionally, based on the above embodiments, the sample dataset construction module 320 may include:
[0105] The effective cumulative rainfall index generation unit is used to calculate the effective cumulative rainfall iteratively day by day based on historical daily rainfall data, using the following formula: ,in, For the first Daily effective cumulative rainfall, For the first Daily rainfall, Let be the rainfall attenuation coefficient, and Initialize to 0 in the first iteration. After completing the full time series iteration, use the obtained effective cumulative rainfall as the effective cumulative rainfall index.
[0106] The continuous rainfall days index generation unit is used to determine, in chronological order, whether the rainfall of the day is not less than a preset rainfall threshold based on historical daily rainfall data. If so, the continuous rainfall days of the day are counted as the number of continuous rainfall days of the previous day plus one. If not, the continuous rainfall days of the day are reset to 0. The continuous rainfall days index is obtained by summing the number of consecutive days in the time series.
[0107] The slope index generation unit is used to calculate the elevation change rate in the x and y directions of each pixel in the target area based on the elevation difference between the current pixel and its neighboring pixels, according to the digital elevation model data. Then, by substituting the elevation change rates in the x and y directions into a preset slope calculation formula, the slope value of each pixel is calculated. The set of slope values of all pixels is the slope index.
[0108] Optionally, based on the above embodiments, the sample dataset construction module 320 may further include:
[0109] The disaster event record reading unit is used to read each disaster event record recorded in the historical flood disaster data and extract the occurrence time, location and disaster level of each disaster event record;
[0110] The first type of index value extraction unit is used to extract the target effective cumulative rainfall index value at the current occurrence time from the effective cumulative rainfall index based on the occurrence time, and to extract the target continuous rainfall days index value at the current occurrence time from the continuous rainfall days index.
[0111] The target second-type index value extraction unit is used to extract the target slope index value of the current location from the set of pixel values of the slope index based on the current location.
[0112] The sample data construction unit is used to combine the extracted target effective cumulative rainfall index value, target continuous rainfall days index value, and target slope index value into a feature vector, and associate the feature vector with the disaster level, the occurrence time and location of the disaster event to form a complete sample data;
[0113] The disaster event record traversal unit is used to traverse all disaster event records in the historical flood disaster data. By repeating the above steps for each disaster event record, all sample data are obtained, and a sample dataset with features associated with disaster level is constructed.
[0114] Optionally, based on the above embodiments, the calculation model generation module 330 may include:
[0115] The random forest model training unit is used to split the sample dataset into training data and independent validation data, input them into a preset random forest model for training, and determine that the model training is complete when the prediction accuracy of the random forest model on the training data reaches a preset accuracy threshold and the prediction accuracy on the independent validation data stably reaches a preset validation threshold, thus obtaining the trained random forest model.
[0116] The importance score calculation unit is used to call the feature importance evaluation function built into the trained random forest model and calculate the importance scores of the three feature variables, namely the effective cumulative rainfall index, the consecutive rainfall days index, and the slope index, respectively, using the average impurity reduction method.
[0117] The weighting coefficient determination unit is used to normalize the importance score so that the sum of the processed weighting coefficients is 1, and to obtain the final weighting coefficients corresponding to the effective cumulative rainfall index, the continuous rainfall days index, and the slope index, respectively, denoted as . , , ;
[0118] The rice waterlogging disaster comprehensive index calculation model construction unit is used to construct the model based on the final weighting coefficients. A calculation model for the comprehensive index of rice waterlogging disaster was constructed, denoted as . The expression for the computational model is: in, , , These represent the normalized effective cumulative rainfall index, the consecutive rainfall days index, and the slope index, respectively. Optionally, based on the above embodiments, the sample subset construction module 340 may include:
[0119] The comprehensive index value solving unit is used to traverse each sample data in the sample dataset, input the normalized effective cumulative rainfall index, continuous rainfall days index and slope index of each sample data into the calculation model of the comprehensive index of rice waterlogging disaster, and solve to obtain the comprehensive index value corresponding to each sample data.
[0120] The growth stage acquisition unit is used to read rice growth period data for the target rice variety and obtain the time range of each growth stage of the target rice variety within the target area.
[0121] The reproductive period determination unit is used to extract the occurrence time of the disaster event associated with each sample data, and determine the reproductive period to which the current sample data belongs by matching the time range;
[0122] The comprehensive index value and disaster level association unit is used to group the sample dataset according to the reproductive period, summarize the comprehensive index value and corresponding disaster level in each group of sample data, and construct a sample subset in which the disaster level and comprehensive index value are associated in each reproductive period.
[0123] Optionally, based on the above embodiments, the grading threshold determination module 350 may include:
[0124] The sample subset grouping unit is used to group each sample subset separately according to the disaster level, so as to obtain a set of comprehensive index values that correspond one-to-one with the disaster level within the same reproductive period;
[0125] The normality test unit is used to perform a normality test on each composite index value group within the current reproductive period. If any target composite index value group does not conform to the normal distribution, then all composite index values within the target composite index value group are mathematically transformed until the target composite index value group conforms to the normal distribution.
[0126] The comprehensive index grading threshold determination unit is used to calculate the confidence interval of each comprehensive index value group that conforms to a normal distribution within the current fertility period using the t-distribution interval estimation method at a given confidence level, and to use the lower limit of each confidence interval as the comprehensive index grading threshold of the corresponding disaster level within the current fertility period;
[0127] The comprehensive index grading threshold set construction unit is used to sort all comprehensive index grading thresholds within the same reproductive period in order of disaster level from low to high, and construct the comprehensive index grading threshold set for the current reproductive period.
[0128] Optionally, based on the above embodiments, it may also include: a real-time waterlogging disaster level determination unit, which is used to obtain real-time daily rainfall data through regional meteorological monitoring stations, retrieve pre-processed digital elevation model data of the target area through geographic information system, and obtain the current growth stage information of the target rice through field IoT equipment monitoring or planting record query.
[0129] Based on the real-time daily rainfall data, the real-time effective cumulative rainfall index and the continuous rainfall days index are calculated. Based on the digital elevation model data of the target area, the real-time slope index is calculated.
[0130] After normalizing the real-time effective cumulative rainfall index, continuous rainfall days index, and slope index, the data are input into the dynamic monitoring model for rice waterlogging disaster to obtain the current real-time comprehensive index value of rice waterlogging disaster.
[0131] From the set of comprehensive index grading thresholds for each growth stage, which are stored in association with the dynamic monitoring model of rice waterlogging disaster and sorted from low to high according to the disaster level, the target grading threshold corresponding to the target growth stage of the target rice is determined.
[0132] The real-time comprehensive index value of rice waterlogging disaster is matched with the determined target classification threshold, and the waterlogging disaster level of the target rice in the target area is determined according to the threshold interval it falls into.
[0133] The rice waterlogging disaster dynamic monitoring model construction device provided in this embodiment of the invention can execute the rice waterlogging disaster dynamic monitoring model construction method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0134] Example 4
[0135] Figure 4A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0136] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0137] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0138] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, central processing unit (CPU), graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as a method for constructing a dynamic monitoring model for rice waterlogging disasters.
[0139] That is: to obtain historical waterlogging disaster data, historical daily rainfall data, digital elevation model data, and growth period data of the target rice in the target area, wherein the historical waterlogging disaster data includes the disaster level;
[0140] The effective cumulative rainfall index and the continuous rainfall days index are calculated based on the historical daily rainfall data. The slope index is calculated based on the digital elevation model data. The three types of indices after normalization are used as feature variables, and the disaster level is used as a label variable to construct a sample dataset that associates features with disaster level.
[0141] The sample dataset is input into a random forest model to calculate the importance weights of each feature variable, and the importance weights are used as the weight coefficients of the corresponding indices to generate a calculation model for the comprehensive index of rice waterlogging disaster.
[0142] The comprehensive index value of each sample in the sample dataset is solved using the calculation model. The samples are then grouped based on the rice growth period data to construct sample subsets that contain disaster level and comprehensive index value for each growth period.
[0143] Normality tests were performed on each sample subset. For normally distributed sample subsets, the t-distribution interval estimation method was used to determine the comprehensive index grading thresholds corresponding to different disaster levels in each reproductive period.
[0144] By associating and storing the computational model with the graded thresholds for each growth stage, a model that can be directly used for dynamic monitoring of rice waterlogging disasters is formed.
[0145] In some embodiments, a method for constructing a dynamic monitoring model for rice waterlogging disasters can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the method for constructing a dynamic monitoring model for rice waterlogging disasters described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to execute a method for constructing a dynamic monitoring model for rice waterlogging disasters by any other suitable means (e.g., by means of firmware).
[0146] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0147] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0148] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0149] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0150] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0151] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0152] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0153] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for constructing a dynamic monitoring model for rice waterlogging disasters, characterized in that, include: Acquire historical waterlogging disaster data, historical daily rainfall data, digital elevation model data, and growth period data of the target rice in the target area. The historical waterlogging disaster data includes the disaster level. The effective cumulative rainfall index and the continuous rainfall days index are calculated based on the historical daily rainfall data. The slope index is calculated based on the digital elevation model data. The three types of indices after normalization are used as feature variables, and the disaster level is used as a label variable to construct a sample dataset that associates features with disaster level. The sample dataset is input into a random forest model to calculate the importance weights of each feature variable, and the importance weights are used as the weight coefficients of the corresponding indices to generate a calculation model for the comprehensive index of rice waterlogging disaster. The comprehensive index value of each sample in the sample dataset is solved using the calculation model. The samples are then grouped based on the rice growth period data to construct sample subsets that contain disaster level and comprehensive index value for each growth period. Normality tests were performed on each sample subset. For normally distributed sample subsets, the t-distribution interval estimation method was used to determine the comprehensive index grading thresholds corresponding to different disaster levels in each reproductive period. By associating and storing the computational model with the graded thresholds for each growth stage, a model that can be directly used for dynamic monitoring of rice waterlogging disasters is formed. The comprehensive index value of each sample in the sample dataset is calculated using the aforementioned calculation model. The samples are then grouped based on the rice growth period data to form subsets for each growth period, each containing a disaster level and its corresponding comprehensive index value. These subsets include: Iterate through each sample data in the sample dataset, and input the normalized effective cumulative rainfall index, continuous rainfall days index, and slope index of each sample data into the calculation model of the comprehensive index of rice waterlogging disaster, and solve to obtain the comprehensive index value corresponding to each sample data. Read rice growth period data for the target rice variety to obtain the time range of each growth stage of the target rice variety within the target region; Extract the occurrence time of the disaster event associated with each sample data, and determine the reproductive period to which the current sample data belongs by matching the time range; The sample dataset is grouped according to the reproductive period, and the comprehensive index value and corresponding disaster level in each group of sample data are summarized to construct a sample subset in which the disaster level and comprehensive index value are associated within each reproductive period.
2. The method according to claim 1, characterized in that, The effective cumulative rainfall index and the consecutive rainfall days index are calculated based on the historical daily rainfall data, and the slope index is calculated based on the digital elevation model data, including: Based on historical daily rainfall data, the effective cumulative rainfall is calculated iteratively day by day according to the time series. The calculation formula is as follows: ,in, For the first Daily effective cumulative rainfall, For the first Daily rainfall, Let be the rainfall attenuation coefficient, and Initialize to 0 in the first iteration. After completing the full time series iteration, use the obtained effective cumulative rainfall as the effective cumulative rainfall index. Based on historical daily rainfall data, the system sequentially determines whether the daily rainfall is not less than a preset rainfall threshold. If so, the number of consecutive rainy days on that day is counted as the number of consecutive rainy days on the previous day plus one. If not, the number of consecutive rainy days on that day is reset to 0. The consecutive rainfall days index is obtained by summing the consecutive days in the time series. Based on digital elevation model data, for each pixel in the target area, the elevation change rate in the x and y directions is calculated according to the elevation difference between the current pixel and its neighboring pixels. Then, by substituting the elevation change rates in the x and y directions into the preset slope calculation formula, the slope value of each pixel is calculated. The set of slope values of all pixels is the slope index.
3. The method according to claim 2, characterized in that, Using the three normalized indices as feature variables and the disaster level as a label variable, a sample dataset correlated with the features and the disaster level is constructed, including: Read each disaster event record from the historical flood disaster data and extract the occurrence time, location, and disaster level of each disaster event record; Based on the occurrence time, extract the target effective cumulative rainfall index value at the current occurrence time from the effective cumulative rainfall index, and extract the target continuous rainfall day index value at the current occurrence time from the continuous rainfall day index; Based on the location, extract the target slope index value for the current location from the set of pixel values of the slope index; The extracted target effective cumulative rainfall index, target continuous rainfall days index, and target slope index are combined into a feature vector, and this feature vector is associated with the disaster level, the time and location of the disaster event to form a complete sample data. By iterating through all disaster event records in the historical flood disaster data, and repeating the above steps for each disaster event record to obtain all sample data, a sample dataset with features correlated with disaster level is constructed.
4. The method according to claim 1, characterized in that, The sample dataset is input into a random forest model to calculate the importance weights of each feature variable. These importance weights are then used as weight coefficients for the corresponding indices to generate a calculation model for the comprehensive index of rice waterlogging disaster, including: The sample dataset is split into training data and independent validation data, and input into a preset random forest model for training. When the prediction accuracy of the random forest model on the training data reaches a preset accuracy threshold, and the prediction accuracy on the independent validation data stably reaches a preset validation threshold, the model training is deemed complete, and the trained random forest model is obtained. The built-in feature importance assessment function in the trained random forest model is invoked, and the importance scores of the three feature variables—effective cumulative rainfall index, consecutive rainfall days index, and slope index—are calculated using the average impurity reduction method. The importance scores are normalized so that the sum of the weight coefficients is 1. The final weight coefficients corresponding to the effective cumulative rainfall index, the consecutive rainfall days index, and the slope index are then obtained, denoted as […]. , , ; Based on the final weighting coefficient , , A calculation model for the comprehensive index of rice waterlogging disaster was constructed, denoted as . The expression for the computational model is: in, , , These represent the normalized effective cumulative rainfall index, the consecutive rainfall days index, and the slope index, respectively.
5. The method according to claim 1, characterized in that, Normality tests were performed on each sample subset. For normally distributed sample subsets, the t-distribution interval estimation method was used to determine the comprehensive index grading thresholds corresponding to different disaster levels within each reproductive period, including: Each sample subset was grouped separately according to the disaster level, resulting in a set of comprehensive index values that correspond one-to-one with the disaster level within the same reproductive period; For each group of composite index values within the current reproductive period, a normality test is performed. If any group of target composite index values does not conform to a normal distribution, then all composite index values within the target composite index value group are mathematically transformed until the target composite index value group conforms to a normal distribution. For each group of comprehensive index values that conforms to a normal distribution within the current reproductive period, the t-distribution interval estimation method is used to calculate its confidence interval at a given confidence level. The lower limit of each confidence interval is used as the comprehensive index classification threshold for the corresponding disaster level within the current reproductive period. By sorting all comprehensive index grading thresholds within the same reproductive period in order of disaster level from low to high, a set of comprehensive index grading thresholds for the current reproductive period is constructed.
6. The method according to claim 1, characterized in that, After developing a model that can be directly used for dynamic monitoring of waterlogging disasters in rice paddies, the following steps are also included: Real-time daily rainfall data is obtained through regional meteorological monitoring stations, pre-processed digital elevation model data of the target area is retrieved through geographic information systems, and the current growth stage information of the target rice is obtained through field IoT equipment monitoring or planting record query. Based on the real-time daily rainfall data, the real-time effective cumulative rainfall index and the continuous rainfall days index are calculated. Based on the digital elevation model data of the target area, the real-time slope index is calculated. After normalizing the real-time effective cumulative rainfall index, continuous rainfall days index, and slope index, the data are input into the dynamic monitoring model for rice waterlogging disaster to obtain the current real-time comprehensive index value of rice waterlogging disaster. From the set of comprehensive index grading thresholds for each growth stage, which are stored in association with the dynamic monitoring model of rice waterlogging disaster and sorted from low to high according to the disaster level, the target grading threshold corresponding to the target growth stage of the target rice is determined. The real-time comprehensive index value of rice waterlogging disaster is matched with the determined target classification threshold, and the waterlogging disaster level of the target rice in the target area is determined according to the threshold interval it falls into.
7. A device for constructing a dynamic monitoring model for rice waterlogging disasters, characterized in that, The device includes: The historical monitoring data acquisition module is used to acquire historical waterlogging disaster data, historical daily rainfall data, digital elevation model data, and growth period data of the target rice in the target area. The historical waterlogging disaster data includes the disaster level. The sample dataset construction module is used to calculate the effective cumulative rainfall index and the continuous rainfall days index based on the historical daily rainfall data, calculate the slope index based on the digital elevation model data, use the three types of indices after normalization as feature variables, and use the disaster level as label variable to construct a sample dataset in which features are associated with disaster level. The computational model generation module is used to input the sample dataset into the random forest model to calculate the importance weights of each feature variable, and use the importance weights as the weight coefficients of the corresponding index to generate a computational model for the comprehensive index of rice waterlogging disaster. The sample subset construction module is used to solve the comprehensive index value of each sample in the sample dataset using the calculation model, and to group the samples in combination with the rice growth period data to construct sample subsets that contain disaster level and comprehensive index value in each growth period. The grading threshold determination module is used to perform normality tests on each sample subset. For normally distributed sample subsets, the t-distribution interval estimation method is used to determine the comprehensive index grading threshold corresponding to different disaster levels in each reproductive period. The monitoring model generation module is used to form a model that can be directly used for dynamic monitoring of rice waterlogging disasters by associating and storing the calculation model with the graded thresholds of each growth stage; The sample subset construction module includes: The comprehensive index value solving unit is used to traverse each sample data in the sample dataset, input the normalized effective cumulative rainfall index, continuous rainfall days index and slope index of each sample data into the calculation model of the comprehensive index of rice waterlogging disaster, and solve to obtain the comprehensive index value corresponding to each sample data. The growth stage acquisition unit is used to read rice growth period data for the target rice variety and obtain the time range of each growth stage of the target rice variety within the target area. The reproductive period determination unit is used to extract the occurrence time of the disaster event associated with each sample data, and determine the reproductive period to which the current sample data belongs by matching the time range; The comprehensive index value and disaster level association unit is used to group the sample dataset according to the reproductive period, summarize the comprehensive index value and corresponding disaster level in each group of sample data, and construct a sample subset in which the disaster level and comprehensive index value are associated in each reproductive period.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the method for constructing a dynamic monitoring model for rice waterlogging disasters as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that are used to cause a processor to execute the method for constructing a dynamic monitoring model for rice waterlogging disasters according to any one of claims 1-6.
Citation Information
Patent Citations
System and method for monitoring berry tea meteorological disasters
CN118071194A
Flood disaster situation prediction method and device, computer equipment, storage medium and product
CN119917861A