Multi-source data fusion-based summer corn flowering phase high-temperature disaster-caused evaluation model and disaster damage estimation system
Through the summer corn flower high-temperature disaster-causing evaluation model and disaster-causing estimation system based on multi-source data fusion, the problem of insufficient accuracy and comprehensiveness of summer corn high-temperature disaster-causing evaluation and disaster-causing estimation in the existing technology is solved, and more accurate high-temperature disaster-causing evaluation and disaster-causing estimation is achieved, meeting the needs of precise agriculture.
Patent Information
- Application Number
- CN202510053988.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-06
AI Technical Summary
The existing technology lacks accuracy and comprehensiveness in the evaluation of high temperature disasters caused by summer corn and disaster damage estimation, and lacks an evaluation system that integrates meteorological conditions and crop growth parameters, making it difficult to meet the needs of precision agriculture.
The summer popcorn period high temperature disaster assessment model and disaster loss estimation system based on multi-source data fusion are adopted, including data source and processing module, feature factor screening module, model construction and evaluation module, and disaster loss estimation and analysis module. By obtaining and processing meteorological data, remote sensing data and disaster investigation data, screening characteristic factors, and constructing multiple linear regression models, disaster loss estimation and analysis are realized.
It improves the accuracy and comprehensiveness of high-temperature disaster assessment, can more accurately reflect the degree of high-temperature disasters in summer corn, enhances the accuracy and mechanism of disaster monitoring, and meets the needs of precision agriculture.
Smart Images

Figure CN119941037A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of agricultural disaster monitoring and assessment, and more specifically, to a summer corn flowering period high temperature disaster assessment model and a disaster loss estimation system based on multi-source data fusion. Background Art
[0002] As one of the most important food crops in the world, the stability of summer corn yield and quality is of great significance to ensuring global food security. However, with the intensification of global climate change, extreme high temperature events occur frequently, posing a serious threat to the growth and yield of summer corn. High temperature not only affects the photosynthesis, respiration and water balance of corn, but also accelerates physiological and biochemical reactions, shortens the growth cycle, and ultimately leads to reduced yield or even total crop failure. Therefore, the construction of a summer corn high temperature disaster assessment model and the study of disaster loss estimation are of great significance for formulating effective disaster prevention and mitigation measures and ensuring food production safety.
[0003] At present, the research on summer corn heat damage mainly focuses on the disaster mechanism, disaster identification and disaster monitoring and evaluation. In terms of disaster mechanism, studies have shown that different growth and development stages of corn have different sensitivities to high temperatures, especially during the flowering and pollination period, when high temperatures can significantly reduce pollen vitality and affect pollination and fruiting. In addition, high temperatures can also lead to decreased photosynthetic protease activity, damaged chloroplast structure, increased leaf temperature, and reduced transpiration rate, thereby affecting photosynthesis and dry matter accumulation. In terms of disaster identification, different varieties of corn have different tolerance to high temperatures, and soil moisture conditions and agricultural management measures can also affect corn's tolerance to high temperatures. At present, most studies identify and grade high temperature heat damage based on high temperature intensity, duration, and accumulated temperature at different temperature thresholds, but the temperature thresholds and judgment criteria used in the research on high temperature heat damage of corn at different development stages in different regions are different. In terms of disaster monitoring and evaluation, ground meteorological station data were previously used, combined with crop development period data, to analyze the number and spatial distribution of high temperature heat damage. With the continuous development of remote sensing technology, it has been widely used in crop high temperature stress monitoring and evaluation due to its advantages such as real-time and regionality. However, all of them use the land surface temperature inverted by remote sensing to study the high temperature heat damage of corn. However, the occurrence of high temperature heat damage is not only related to temperature, but also closely related to soil moisture, crop varieties, management measures, etc. In addition, high temperature heat damage has certain concealment and accumulation. Therefore, dynamic monitoring of corn growth process during and after high temperature process, finding sensitive growth parameters affected by disasters and determining the index threshold of the degree of disaster, and constructing meteorological and remote sensing indicators of corn stress resistance under high temperature weather process stress can enhance the accuracy and mechanism of disaster monitoring. Through remote sensing technology, the impact of high temperature stress on summer corn physiological parameters, such as chlorophyll content, leaf temperature, transpiration rate, etc., can be inverted and quantified. The changes in these physiological parameters can directly reflect the growth state of summer corn under high temperature stress, providing an important basis for monitoring and early warning of high temperature disasters. In addition, combining meteorological, soil, crop growth and other information, using remote sensing technology to construct a high temperature disaster evaluation model has become a hot spot and trend in current research.
[0004] Although some progress has been made in the evaluation of summer corn heat damage and the estimation of damage, there are still some shortcomings. First, the current standards for identifying corn heat damage are generally based on daily temperature values, lacking the application of more accurate hourly temperature data; and the existing models mostly focus on using temperature as a single indicator, lacking the construction of an evaluation system that integrates meteorological conditions and crop growth parameters, resulting in the accuracy and comprehensiveness of monitoring and evaluation results being limited. Secondly, when constructing high temperature disaster evaluation models, existing studies often ignore the differences in the sensitivity of corn to high temperatures at different growth and development stages, resulting in the need to improve the applicability and accuracy of the model. Finally, existing studies rely more on linear empirical models in terms of damage estimation, lack in-depth research based on remote sensing technology and machine learning methods, and are difficult to meet the needs of precision agriculture.
[0005] In order to solve the above problems, a technical solution is now provided. Summary of the invention
[0006] In order to overcome the above-mentioned defects of the prior art, the embodiments of the present invention provide a summer corn flowering period high temperature disaster assessment model and a disaster loss estimation system based on multi-source data fusion to solve the problems raised in the above-mentioned background technology.
[0007] To achieve the above object, the present invention provides the following technical solutions:
[0008] A high temperature disaster assessment model and disaster loss estimation system for summer corn flowering period based on multi-source data fusion, including data source and processing module, characteristic factor screening module, model construction and evaluation module, and disaster loss estimation and analysis module;
[0009] The data source and processing module is used to obtain meteorological data, remote sensing data and disaster investigation data, and process the meteorological data, remote sensing data and disaster investigation data respectively;
[0010] The feature factor screening module uses the correlation coefficient method, the feature selection algorithm based on random forest and the feature selection method based on model for screening;
[0011] The model building and evaluation module is used for model building and model evaluation;
[0012] The disaster loss estimation and analysis module uses the constructed model to perform disaster loss estimation and analysis.
[0013] In a preferred embodiment, the data source and processing module operation specifically include the following:
[0014] Meteorological data include the daily maximum temperature, daily average temperature, hourly temperature, maximum soil relative humidity, minimum soil relative humidity and average daily value data during the flowering period of summer corn;
[0015] Remote sensing data include surface temperature products, leaf area index products, potential evapotranspiration products, vegetation index products and surface reflectance products during the summer corn flowering period and the following 10 days.
[0016] In a preferred embodiment, the data source and processing module operation specifically further includes the following:
[0017] The meteorological data is processed to construct a corn high temperature disaster index to reflect the impact of extreme high temperature events on summer corn. The specific calculation formula is as follows:
[0018]
[0019] HI 综 =HI 32 ×0.4+HI 35 ×0.6;
[0020] Among them, HI 32 It refers to the high temperature index when the threshold is 32℃;
[0021] DH 32 It refers to the number of high temperature hours with a threshold of 32°C;
[0022] DH min and DH max They refer to the minimum and maximum values of DH during the flowering period of summer maize in 20 years, respectively;
[0023] AH 32 It refers to the accumulated heat when the threshold is 32°C, that is, the sum of temperatures greater than 32°C;
[0024] AH min and AH max They refer to the minimum and maximum values of AH during the flowering period of summer corn in 20 years;
[0025] HI ranges from 0 to 1, with larger values indicating more severe heat damage;
[0026] HI 综 It is a comprehensive high temperature disaster index, taking into account HI 32 and HI 35 , calculated by weighted average.
[0027] In a preferred embodiment, the data source and processing module operation specifically further includes the following:
[0028] The remote sensing data were processed, and according to the influence mechanism of high temperature stress and the remote sensing inversion principle of plant physiological and biochemical parameters, 9 remote sensing factors were selected as the disaster-causing factors of high temperature stress during the flowering period of summer corn. Each factor included data products for three different time periods: the early flowering period, the late flowering period, and the post-flowering period.
[0029] The actual evapotranspiration ET and potential evapotranspiration PET are obtained, and the crop water stress index CWSI is calculated using the following formula:
[0030] CWSI=1-ET / PET.
[0031] In a preferred embodiment, the data source and processing module operation specifically further includes the following:
[0032] The disaster survey data was processed and the number of grains per plant was calculated to characterize the impact of high temperature heat damage on corn. The number of grains per plant was used as the modeling dependent variable, and its rate of change characterized the degree of damage.
[0033] In a preferred embodiment, the operation of the characteristic factor screening module specifically includes the following contents:
[0034] Normal distribution test The Shapiro-Wi lk test was used for normality test, and the data distribution was judged comprehensively with the graphs. The factors that were in line with or close to normal distribution and had continuous data were analyzed later;
[0035] Feature selection first uses the correlation coefficient method to remove features with low correlation with the target variable, then uses the random forest-based feature selection algorithm to verify the optimization, and finally uses the model-based feature selection method for further screening.
[0036] In a preferred embodiment, the model building and evaluation module operation includes the following specific contents:
[0037] Construct multiple linear regression, ridge regression, LASSO regression, random forest regression, and XGBoost regression models;
[0038] The models were evaluated using the coefficient of determination, root mean square error, and mean absolute error.
[0039] In a preferred embodiment, the operation of the disaster loss estimation and analysis module includes the following specific contents:
[0040] The constructed optimal model was used to predict and compare the number of grains per plant in typical years with high temperature heat damage during the flowering period of summer corn and years with no obvious disasters, calculate the change rate of the number of grains per plant, divide the degree of yield reduction, and realize disaster loss estimation and analysis.
[0041] Technical effects and advantages of the summer corn flowering period high temperature disaster assessment model and disaster loss estimation system based on multi-source data fusion of the present invention:
[0042] Based on hourly temperature data, an innovative high temperature disaster index was constructed, and a high temperature heat damage evaluation system for summer corn during the flowering period was established by combining meteorological, soil, crop growth and other information. An evaluation model was constructed using methods such as multiple linear regression (MLR), ridge regression, LASSO regression, random forest regression (RF) and XGBoost regression to achieve accurate assessment of high temperature heat damage to summer corn and disaster loss estimation, thereby improving the comprehensiveness and reliability of the evaluation results. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is the Yield histogram and QQ plot.
[0044] Figure 2 HTHours 35 Histograms and QQ plots.
[0045] Figure 3 LAI histogram and QQ plot.
[0046] Figure 4 This is a bar chart of Pearson correlation analysis of disaster factors.
[0047] Figure 5 This is the heat map of factors with Pearson correlation coefficient > 0.3.
[0048] Figure 6 This is a graph of feature importance analysis for the random forest method.
[0049] Figure 7 This is a comparison chart of model accuracy.
[0050] Figure 8 Taylor diagram for model comparison.
[0051] Figure 9-13 This is the estimation effect of the MLRRidgeLASSORFXGB model on the number of grains per plant in summer corn in Henan Province.
[0052] Fig.14 For 2016 HI 35 Spatial distribution map.
[0053] Fig.15 This is the RF prediction result diagram.
[0054] Fig.16 This is the XGBoost prediction result graph. DETAILED DESCRIPTION
[0055] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0056] Example 1
[0057] The present invention provides a summer corn flowering period high temperature disaster evaluation model and disaster loss estimation system based on multi-source data fusion, including a data source and processing module, a characteristic factor screening module, a model building and evaluation module, and a disaster loss estimation and analysis module;
[0058] The data source and processing module is used to obtain meteorological data, remote sensing data and disaster investigation data, and process the meteorological data, remote sensing data and disaster investigation data respectively;
[0059] The feature factor screening module uses the correlation coefficient method, the feature selection algorithm based on random forest and the feature selection method based on model to screen the feature factors;
[0060] The model building and evaluation module is used for model building and model evaluation;
[0061] The disaster loss estimation and analysis module uses the constructed model to perform disaster loss estimation and analysis.
[0062] The data source and processing module operation specifically include the following:
[0063] Meteorological data include the daily maximum temperature, daily average temperature, hourly temperature, maximum soil relative humidity, minimum soil relative humidity and average daily value data during the flowering period of summer corn;
[0064] Remote sensing data include surface temperature products, leaf area index products, potential evapotranspiration products, vegetation index products and surface reflectance products during the summer corn flowering period and the following 10 days.
[0065] It should be added that the meteorological data used in the study comes from the meteorological big data cloud platform "Tianqing" of the National Meteorological Confidence Center, including the daily maximum temperature, daily average temperature, hourly temperature, maximum soil relative humidity, minimum soil relative humidity and average daily value data during the summer corn flowering period in Henan Province from 2005 to 2024.
[0066] The growth and development data of summer corn, including key growth periods such as sowing period, flowering period, and tasseling and silking period, are derived from the statistical data and historical records of the Henan Provincial Meteorological Department. These data are used to determine the key period for high temperature heat damage assessment.
[0067] The disaster survey data comes from the disaster surveys conducted by the Henan Provincial Meteorological Bureau at the provincial, municipal and county levels in the province’s meteorological systems in 2016 and 2017, which were typical years of high temperatures during the summer corn flowering period. A total of 88 survey plots were obtained, and 76 plots were obtained after removing outliers.
[0068] The data source and processing module operation also includes the following contents:
[0069] The meteorological data is processed to construct a corn high temperature disaster index to reflect the impact of extreme high temperature events on summer corn. The specific calculation formula is as follows:
[0070]
[0071] HI 综 =HI 32 ×0.4+HI 35 ×0.6;
[0072] Among them, HI 32 It refers to the high temperature index when the threshold is 32℃;
[0073] DH 32 It refers to the number of high temperature hours with a threshold of 32°C;
[0074] DH min and DH max They refer to the minimum and maximum values of DH during the flowering period of summer maize in 20 years, respectively;
[0075] AH 32 It refers to the accumulated heat when the threshold is 32°C, that is, the sum of temperatures greater than 32°C;
[0076] AH min and AH max They refer to the minimum and maximum values of AH during the flowering period of summer corn in 20 years;
[0077] HI ranges from 0 to 1, with larger values indicating more severe heat damage;
[0078] HI 综 It is a comprehensive high temperature disaster index, taking into account HI 32 and HI 35 , calculated by weighted average.
[0079] It should be noted that different growth and development stages of corn have different sensitivities to high temperatures. The flowering and pollination period is particularly sensitive to high temperatures, which can significantly reduce pollen vitality and affect pollination and fruiting. The greater the intensity and the longer the duration of high temperatures, the more serious the damage to corn and the more difficult it is to recover. Therefore, based on the growth and development of summer corn and hourly temperature data, a new corn high temperature disaster index (HI) is constructed to reflect the impact of extreme high temperature events on summer corn.
[0080] According to the actual impact of the thresholds of 32℃ and 35℃ on summer corn, weights of 0.4 and 0.6 were assigned respectively to calculate HI 综 All meteorological elements were interpolated to 1 km × 1 km grid points using the Kriging method. After processing, the meteorological factor variables in Table 1 were obtained for further factor analysis and modeling.
[0081] Table 1 Meteorological factors causing high temperature stress in summer corn
[0082]
[0083]
[0084] The data source and processing module operation also includes the following contents:
[0085] The remote sensing data were processed, and according to the influence mechanism of high temperature stress and the remote sensing inversion principle of plant physiological and biochemical parameters, 9 remote sensing factors were selected as the disaster-causing factors of high temperature stress during the flowering period of summer corn. Each factor included data products for three different time periods: the early flowering period, the late flowering period, and the post-flowering period.
[0086] The actual evapotranspiration ET and potential evapotranspiration PET are obtained, and the crop water stress index CWSI is calculated using the following formula:
[0087] CWSI=1-ET / PET.
[0088] It should be noted that high temperature will destroy the chloroplast structure and reduce the activity of photosynthetic proteases, thereby accelerating the degradation of chlorophyll and decreasing its content. Chlorophyll is the main pigment for photosynthesis in plants, and a reduction in its content will directly affect the photosynthetic capacity of plants. Under high temperature stress, the temperature of plant leaves will rise significantly, which is due to the weakening of transpiration and the decrease in the heat dissipation capacity of the plant body. High temperature will also increase the respiration of leaves, further aggravating energy consumption and temperature increase. High temperature will cause the stomata to close, thereby reducing the amount of water evaporation required for transpiration and reducing the transpiration rate. The decrease in transpiration rate will affect the water balance and heat dissipation capacity of the plant, further aggravating high temperature stress. Based on the above high temperature stress mechanism, the physiological and biochemical parameters that can reflect the possible high temperature stress of summer corn include:
[0089] (1) Chlorophyll content, estimated by analyzing the spectral reflectance characteristics of vegetation;
[0090] (2) Leaf temperature, using the thermal infrared band to invert the surface temperature;
[0091] (3) transpiration rate, which is estimated indirectly by monitoring other relevant parameters (such as leaf temperature, soil moisture content, etc.);
[0092] (4) Crop Water Stress Index (CWS I), which combines the difference between canopy temperature and air temperature and can reflect the degree of water stress of plants. Under high temperature stress, CWS I usually increases;
[0093] (5) Photosynthetic rate can be estimated indirectly by monitoring parameters such as chlorophyll content and leaf temperature.
[0094] Surface temperature product (MOD11A2), leaf area index product (MCD15A3H), potential evapotranspiration product (MOD16A2), vegetation index product (MOD13Q1) and surface reflectance product (MOD09Q1), Fpar and LAI are obtained from MCD15A3H product, LST is obtained from MOD11A2 product, B1 and B2 are obtained from MOD09Q1 product, NDVI is obtained from MOD13Q1 product, PET and ET are obtained from MOD16A2 product.
[0095] The actual evapotranspiration (ET) reflects the amount of water actually consumed by plants, while the potential evapotranspiration (PET) is the amount of water that plants may consume under ideal conditions. The closer the value of CWS I is to 1, the more severe the water stress on the plants. All remote sensing factors are processed into products with a temporal resolution of 8 days and a spatial resolution of 1 km.
[0096] Table 2 Remote sensing factors of summer maize high temperature stress
[0097]
[0098]
[0099] The data source and processing module operation also includes the following contents:
[0100] The disaster survey data was processed and the number of grains per plant was calculated to characterize the impact of high temperature heat damage on corn. The number of grains per plant was used as the modeling dependent variable, and its rate of change characterized the degree of damage.
[0101] It should be noted that 201 data on summer corn yield structure analysis were collected from 22 agricultural meteorological observation stations from 2005 to 2024. Correlation analysis showed that the correlation between high temperature index and the number of grains per plant and grain weight per plant was strong, reaching a very significant correlation at the p<0.001 level, and the difference was not large. At the same time, Zhang Chuan et al. showed that high temperature stress significantly reduced the dry matter accumulation, fruit setting rate and number of grains per ear of corn plants, ultimately leading to a decrease in grain yield. Considering that after the high temperature heat damage process, the 100-grain weight can usually be obtained after harvest, so in order to efficiently and timely predict the impact of high temperature heat damage on summer corn, this study used the rate of change in grain number per plant (RCGNP) to characterize the impact of high temperature heat damage on corn, that is, to construct a prediction model for the number of grains per plant (GNP), and to characterize the degree of damage by calculating the rate of change in grain number per plant.
[0102] The GNP of each survey point of disaster survey data is calculated as follows:
[0103] GNP = (number of grains per plant 1 + number of grains per plant 2 + number of grains per plant 3 + number of grains per plant 4 + number of grains per plant 5) / 5
[0104] RCGNP = (GNP of the current year - GNP of the year without disaster or normal year) / GNP of the year without disaster or normal year × 100%
[0105] RCGNP represents the extent of disaster losses compared with non-disaster years or normal years.
[0106] The operation of the characteristic factor screening module specifically includes the following contents:
[0107] Normal distribution test The Shapiro-Wilk test was used for normality test, and the data distribution was judged comprehensively by combining graphics. The factors that met or were close to normal distribution and had continuous data were analyzed later;
[0108] Feature selection first uses the correlation coefficient method to remove features with low correlation with the target variable, then uses the random forest-based feature selection algorithm to verify the optimization, and finally uses the model-based feature selection method for further screening.
[0109] It should be noted that the normality test is performed using the Shapiro-Wilk test. When the p-value is greater than the set significance level, the data can be considered to conform to the normal distribution. In statistics, the commonly used significance level is 0.05. Therefore, if the p-value of the Shapiro-Wilk test is greater than 0.05, it is generally believed that the data conforms to the normal distribution. The QQ graph determines whether the data conforms to the normal distribution by comparing the sample quantile and the theoretical normal quantile. If the data points in the QQ graph are distributed roughly along the diagonal, it indicates that the data may conform to the normal distribution.
[0110] The correlation coefficient method selects highly correlated features by calculating the correlation coefficient between the features and the target variable. It can quickly identify features that are linearly correlated with the target variable, but may not capture nonlinear relationships and interactions between features, and may not be able to effectively distinguish multicollinear features.
[0111] The feature selection algorithm based on random forest evaluates the importance of features by creating a "shadow" feature for each feature. Shadow features are features with the same distribution as the original features but arranged randomly. Important features are selected by comparing the importance scores of the original features and the shadow features. It can handle nonlinear relationships and interactions between features, and has good discrimination ability for multicollinear features, but the amount of calculation is large, and it is necessary to reasonably set the number of iterations and the threshold of the importance score.
[0112] The embedding method embeds feature selection into the model training process and selects features through regularization or feature importance scores. For example, Lasso regression and decision tree-based models. It can perform feature selection and model training at the same time to improve efficiency, but the effect of the embedding method depends on the selected model and regularization parameters.
[0113] In order to ensure both the importance factor and avoid too many cross-correlation factors, the accuracy and reliability of the corn high temperature heat damage prediction model are improved. The present invention first uses the correlation coefficient method to remove the features with low correlation with the target variable. This step can quickly reduce the number of features and improve the efficiency of subsequent processing. Then use the feature selection algorithm based on random forest or other advanced feature selection algorithms for verification and optimization. This step can ensure the stability and reliability of the selected features and further improve the prediction performance of the model. Finally, a model-based feature selection method (Lasso regression and decision tree-based model, etc.) is used for further selection. These methods can more accurately evaluate the importance of features and take into account nonlinear relationships and interactions between features.
[0114] The model building and evaluation module operation includes the following specific contents:
[0115] Construct multiple linear regression, ridge regression, LASSO regression, random forest regression, and XGBoost regression models;
[0116] The models were evaluated using the coefficient of determination, root mean square error, and mean absolute error.
[0117] It should be noted that multiple linear regression (MLR) is a statistical method used to analyze the linear relationship between multiple independent variables and a dependent variable. It can comprehensively consider multiple factors, effectively improve the prediction accuracy by introducing multiple independent variables, quantify the influence of each independent variable on the dependent variable, and has strong interpretability. However, when there is a high correlation between independent variables, it will lead to multicollinearity problems, making the estimation of regression coefficients unstable and reducing interpretability.
[0118] Ridge regression effectively reduces the impact of multicollinearity on the model by introducing L2 regularization terms, improving the stability of the model and the prediction accuracy; however, its disadvantage is that the introduction of regularization terms may cause the model to be too smooth in some cases, thereby losing some information and affecting the interpretability of the model. Compared with OLS regression, the coefficients of ridge regression no longer have intuitive interpretation significance.
[0119] LASSO regression is a linear model construction method for feature selection and parameter estimation. By adding the L1 regularization term, some unimportant feature coefficients can be compressed to zero, thereby achieving feature selection and improving the interpretability and generalization ability of the model. It is particularly suitable for high-dimensional data sets and can screen out the variables that have the greatest impact on the dependent variable when the number of features is much larger than the number of samples. It performs well in dealing with multicollinearity problems and can stably select important features. The regularization parameter α of LASSO needs to be carefully selected through cross-validation and other methods. Improper selection may affect model performance and is more sensitive to outliers.
[0120] Random forest regression (RF) is an ensemble learning-based algorithm that performs regression tasks by building multiple decision trees and integrating their predictions. When building each decision tree, random forests randomly select data subsets and feature subsets. This randomness helps reduce the risk of overfitting. Random forest regression can evaluate the importance of each feature to the model, which is very helpful for feature selection and interpretation of the model. Since the decision trees in random forests are independent of each other, they can be easily parallelized to increase training speed. The performance of random forest regression is sensitive to parameter settings (such as the number of decision trees, the size of feature subsets, etc.), and appropriate parameter tuning is required to achieve optimal performance. Building a large number of decision trees and integrating their results may require more computing resources, which may become a bottleneck when processing large-scale data sets. When the random forest regression model involves multiple features and complex interactions, its output results may be difficult to interpret.
[0121] XGBoost regression is a powerful machine learning algorithm. It is based on the gradient boosting algorithm and combines a variety of optimization techniques to achieve excellent performance when dealing with complex nonlinear data. XGBoost uses a range of optimization techniques, such as parallel processing, cache optimization, and approximation algorithms, to make it faster to train on large-scale datasets. The importance of each feature can be evaluated based on the constructed decision tree model to help feature selection and feature engineering. XGBoost can automatically handle missing values without additional preprocessing. XGBoost has many parameters that need to be adjusted, and different parameter adjustments are required for different datasets, which increases the difficulty of use.
[0122] The model evaluation index used in this invention is R 2 , RMSE and MAE. 2 Also known as the coefficient of determination, it is an important indicator for measuring the goodness of fit of a regression model. It indicates the proportion of the variation of the dependent variable that can be explained by the independent variable. RMSE, or root mean square error, is one of the commonly used indicators for measuring the accuracy of a prediction model. It is the degree of deviation between the predicted value and the true value. MAE, or mean absolute error, is the average of the absolute values of the prediction errors of each sample. It is a statistical indicator for measuring the average difference between the predicted value and the actual value. Finally, the Taylor diagram is drawn using the correlation coefficient R, standard deviation (SDEV), and root mean square error (RMSD) to compare the prediction models.
[0123] The operation of the disaster loss estimation and analysis module includes the following specific contents:
[0124] The constructed optimal model was used to predict and compare the number of grains per plant in typical years with high temperature heat damage during the flowering period of summer corn and years with no obvious disasters, calculate the change rate of the number of grains per plant, divide the degree of yield reduction, and realize disaster loss estimation and analysis.
[0125] Result analysis:
[0126] Normal distribution test
[0127] from Figure 1-Figure 3 It can be seen that most of the factor data conform to or are close to normal distribution and the data are continuous, so Pearson correlation analysis is used.
[0128] Pearson correlation analysis
[0129] Pearson correlation analysis, also known as Pearson correlation coefficient, is a statistical analysis method used to measure the strength and direction of the linear relationship between two variables. The strength of the correlation can be determined based on the absolute value of the correlation coefficient, such as Figure 4 shown.
[0130] Heatmap is a powerful data visualization tool that can help us intuitively display the strength and direction of linear correlation between multiple variables, providing strong support for data analysis, decision making and scientific research. 14 disaster factors with Pearson correlation coefficient>0.3 are selected to make heatmap, such as Figure 5 shown.
[0131] Feature selection based on random forest
[0132] Based on the correlation analysis results, the random forest method was used to select features for factors with correlation coefficients greater than 0.3 ( Figure 6 ). Refer to the heat map ( Figure 5 ), select one of the two factors with strong correlation that has strong representativeness to avoid strong mutual correlation between factors. 35 Although the correlation with HI is strong, considering their importance to high temperature heat damage during corn flowering, they are temporarily retained. When modeling, the feature selection method based on the model (Lasso regression and decision tree-based model, etc.) will be considered for further selection. The current factor selection results are: AH 35 , LST PostFL , E.T. PreFL ,HI,T avg , T max , PET AfterFL .
[0133] Model building
[0134] MLR, Ridge, Lasso, RF and XGBoost methods were used to construct the high temperature disaster model during the flowering period of summer corn. The model performance is as follows: Figure 7 As shown. Training set R 2 The average MAE is 491.8 and the average RMSE is 618.2, which is between 0.277 and 0.722. 2 The maximum is 0.722, while MAE and RMSE are the smallest, followed by RF. 2 The mean of MAE is 617.4 and the mean of RMSE is 692.2, which is between 0.228 and 0.444. 2 The maximum is 0.444, while MAE and RMSE are the smallest, followed by XGBoost. Although the XGBoost model fits the training data well, it may be at risk of overfitting; the RF model with the best performance on the test set has better generalization ability and is more likely to perform well in practical applications. Figure 8The model comparison in the figure clearly shows that the std of the RF model is closer to the measured value, the RMSE is smaller, and the correlation coefficient is closest to 0.7. Taking all factors into consideration, the RF model is significantly better than other models, followed by the XGBoost model.
[0135] In general, the prediction accuracy of different regression models for the degree of high temperature disaster during the flowering period of summer corn generally shows a trend of RF>XGBoost>Ridge>Lasso>MLR. Compared with the traditional linear regression method, the two machine learning methods RF and XGBoost showed superior performance. This result shows that the relationship between the high temperature disaster factor and yield loss of summer corn may be extremely complex and difficult to accurately describe by a simple linear model. Therefore, it may be more appropriate to use a machine learning method that can capture complex nonlinear relationships to reveal the deep laws hidden behind the data.
[0136] Comparison of different models for estimating damage caused by high temperature stress during summer corn flowering period in Henan Province
[0137] After comparing and analyzing the prediction performance of MLR, Ridge, LASSO, RF and XGBoost models for the number of grains per plant in summer corn in Henan Province in 2016 (such as Figure 9-13 It is found that the prediction trends of these models show a high degree of consistency, although there are slight differences in the specific proportion distribution of each level. These prediction results are compared with the spatial distribution map of the high temperature disaster index in the same year ( Fig.14 ) are superimposed and compared, and it can be clearly observed that the spatial distribution patterns between them are generally consistent, which verifies the effectiveness of the model prediction. It is particularly worth mentioning that the RF model has shown a more outstanding ability in capturing local detail features. For example, in the Zhoukou area, the prediction accuracy of the RF model is significantly better than other models. This finding highlights its advantage in capturing complex spatial variability. Further in-depth analysis found that the RF model can not only accurately reflect the overall trend of high temperature disasters during the flowering period of summer corn, but also perform particularly well in revealing the spatial heterogeneity of disasters. The generated prediction image texture is richer and more delicate, and contains more detailed information about the occurrence and distribution of disasters. This feature shows that the RF model is extremely sensitive to identifying and analyzing the subtle differences in the impact of high temperature disasters on summer corn production, providing strong decision-making support for precision agricultural management and disaster warning.
[0138] Analysis of high temperature damage during summer corn flowering period
[0139] The optimal model constructed by the present invention is used to estimate the grains of the summer corn plant in the typical year of high temperature heat damage during flowering, and compared with the year without obvious disasters during the summer corn production season, so as to obtain the damage of high temperature heat damage during flowering period of summer corn. 2016 is a typical year of high temperature heat damage during flowering period of summer corn, and there is no obvious disaster during the summer corn production season of 2015. Therefore, the present invention takes these two years as examples to analyze the damage of high temperature damage during flowering period of summer corn.
[0140] The RF and XGBoost models were used to predict the number of grains per unit area in 2016 and 2015, and then raster calculations were performed in ArcGIS:
[0141] Change rate of number of seeds per plant = (2016 predicted value - 2015 predicted value) / 2015 predicted value × 100%, see Fig.15 and Fig.16 In the present invention, <-10% is a severe reduction in production, -10% to 0% is a slight reduction in production, and >0% is a flat or increased production.
[0142] in conclusion
[0143] Aiming at the problem of high temperature disasters in summer corn, the present invention constructs a high temperature disaster evaluation model based on multi-source data fusion, and estimates the damage of high temperature disasters in summer corn in Henan Province. By integrating remote sensing data, meteorological data and summer corn growth and development data, a comprehensive high temperature disaster index was successfully established. The index comprehensively considers the impact of high temperature intensity, duration and high temperature on different growth and development stages of corn, and can more accurately reflect the degree of high temperature disasters in summer corn. By introducing information such as meteorology, soil, and crop growth, remote sensing technology was used to invert the remote sensing products related to the physiological parameters (such as chlorophyll content, leaf temperature, etc.) of summer corn under high temperature stress, and a relationship model of the number of grains per plant with these parameters significantly correlated with the degree of high temperature disaster was established. The research results show that the random forest (RF) and extreme gradient boosting (XGBoost) models perform well in the prediction of high temperature disasters in summer corn, especially the RF model, which shows unique advantages in capturing spatial variability and improving prediction accuracy. The model prediction results are highly consistent with the actual spatial distribution map of yield reduction, which verifies the effectiveness and reliability of the model. In addition, the prediction accuracy of different regression models for the degree of high temperature disaster during summer corn flowering period showed a trend of RF>XGBoost>
[0144] The trend of Ridge>Lasso>MLR emphasizes the superiority of machine learning methods in dealing with complex nonlinear relationships.
[0145] In terms of disaster-causing factors selection, previous studies have mostly evaluated high temperature disasters based on a single meteorological factor (such as the daily maximum temperature), ignoring the differences in the sensitivity of corn growth and development stages to high temperatures and the impact of other factors such as soil moisture and agricultural management measures. The present invention comprehensively considers multiple meteorological factors and remote sensing factors, and screens out key factors that have a significant impact on high temperature disasters through feature selection algorithms, such as remote sensing parameters that are closely related to chlorophyll content and leaf temperature, thereby improving the accuracy and applicability of the model.
[0146] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0147] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A high temperature disaster assessment model and disaster loss estimation system for summer corn flowering period based on multi-source data fusion, characterized in that: It includes data source and processing module, characteristic factor screening module, model building and evaluation module and disaster loss estimation and analysis module; The data source and processing module is used to obtain meteorological data, remote sensing data and disaster investigation data, and process the meteorological data, remote sensing data and disaster investigation data respectively; The feature factor screening module uses the correlation coefficient method, the feature selection algorithm based on random forest and the feature selection method based on model to screen the feature factors; The model building and evaluation module is used for model building and model evaluation; The disaster loss estimation and analysis module uses the constructed model to perform disaster loss estimation and analysis.
2. The summer corn flowering period high temperature disaster assessment model and disaster loss estimation system based on multi-source data fusion according to claim 1 is characterized by: The data source and processing module operation specifically include the following: Meteorological data include the daily maximum temperature, daily average temperature, hourly temperature, maximum soil relative humidity, minimum soil relative humidity and average daily value data during the flowering period of summer corn; Remote sensing data include surface temperature products, leaf area index products, potential evapotranspiration products, vegetation index products and surface reflectance products during the summer corn flowering period and the following 10 days.
3. The summer corn flowering period high temperature disaster assessment model and disaster loss estimation system based on multi-source data fusion according to claim 2 is characterized by: The data source and processing module operation also includes the following contents: The meteorological data is processed to construct a corn high temperature disaster index to reflect the impact of extreme high temperature events on summer corn. The specific calculation formula is as follows: HI 综 HI 32 ×0.4+HI 35 ×0.6; Among them, HI 32 It refers to the high temperature index when the threshold is 32℃; DH 32 It refers to the number of high temperature hours with a threshold of 32°C; DH min and DH max They refer to the minimum and maximum values of DH during the flowering period of summer maize in 20 years, respectively; AH 32 It refers to the accumulated heat when the threshold is 32°C, that is, the sum of temperatures greater than 32°C; AH min and AH max They refer to the minimum and maximum values of AH during the flowering period of summer corn in 20 years; HI ranges from 0 to 1, with larger values indicating more severe heat damage; HI 综 It is a comprehensive high temperature disaster index, taking into account HI 32 and HI 35 , calculated by weighted average.
4. The summer corn flowering period high temperature disaster assessment model and disaster loss estimation system based on multi-source data fusion according to claim 3 is characterized by: The data source and processing module operation also includes the following contents: The remote sensing data were processed, and according to the influence mechanism of high temperature stress and the remote sensing inversion principle of plant physiological and biochemical parameters, 9 remote sensing factors were selected as the disaster-causing factors of high temperature stress during the flowering period of summer corn. Each factor included data products for three different time periods: the early flowering period, the late flowering period, and the post-flowering period. The actual evapotranspiration ET and potential evapotranspiration PET are obtained, and the crop water stress index CWSI is calculated using the following formula: CWSI=1-ET / PET.
5. The summer corn flowering period high temperature disaster assessment model and disaster loss estimation system based on multi-source data fusion according to claim 4 is characterized by: The data source and processing module operation also includes the following contents: The disaster survey data was processed and the number of grains per plant was calculated to characterize the impact of high temperature heat damage on corn. The number of grains per plant was used as the modeling dependent variable, and its rate of change characterized the degree of damage.
6. The summer corn flowering period high temperature disaster assessment model and disaster loss estimation system based on multi-source data fusion according to claim 5 is characterized by: The operation of the characteristic factor screening module specifically includes the following contents: Normal distribution test The Shapiro-Wilk test was used to test the normality, and the data distribution was judged comprehensively with the graphs. The factors that met or were close to the normal distribution and had continuous data were subsequently analyzed; Feature selection first uses the correlation coefficient method to remove features with low correlation with the target variable, then uses the random forest-based feature selection algorithm to verify the optimization, and finally uses the model-based feature selection method for further screening.
7. The summer corn flowering period high temperature disaster assessment model and disaster loss estimation system based on multi-source data fusion according to claim 6 is characterized by: The model building and evaluation module operation includes the following specific contents: Construct multiple linear regression, ridge regression, LASSO regression, random forest regression, and XGBoost regression models; The models were evaluated using the coefficient of determination, root mean square error, and mean absolute error.
8. The summer corn flowering period high temperature disaster assessment model and disaster loss estimation system based on multi-source data fusion according to claim 7 is characterized by: The operation of the disaster loss estimation and analysis module includes the following specific contents: The optimal model constructed was used to predict and compare the number of grains per plant in typical years with high temperature heat damage during the flowering period of summer corn and years with no obvious disasters, calculate the change rate of the number of grains per plant, divide the degree of yield reduction, and realize disaster loss estimation and analysis.