Method and device for constructing farmland weather forecast correction model
By constructing a farmland meteorological forecast correction model, using a variety of machine learning algorithms and stacked learning methods, the problem of insufficient accuracy of the existing meteorological forecast model when considering specific factors in farmland is solved, and higher accuracy of meteorological forecasts and reliability of agricultural production decisions are achieved.
Patent Information
- Application Number
- CN202510218455.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-27
AI Technical Summary
The existing numerical meteorological forecasting model has systematic deviations when considering factors such as farmland crop coverage, water level height and topography, resulting in insufficient accuracy of meteorological forecasting and difficult to meet the needs of modern agricultural production.
By obtaining historical data of the target farmland, using light gradient enhancement algorithm, extreme gradient enhancement algorithm and adaptive enhancement algorithm for fitting training, optimizing the prediction model, and using stacking learning methods to combine the results of multiple prediction models, a farmland meteorological forecast correction model is constructed.
It significantly improves the accuracy of farmland meteorological forecasts, reduces errors, and ensures the reliability of long-term meteorological forecasts, thereby improving the quality of decision-making in agricultural production and reducing economic losses caused by false positives.
Smart Images

Figure CN120216985A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of meteorological prediction, and particularly to a method and device for constructing a farmland meteorological forecast correction model. Background Art
[0002] Numerical weather prediction is a technology for predicting the future atmospheric motion state and weather phenomena by solving the equations of weather dynamics and thermodynamics under given initial conditions and boundary conditions. However, existing numerical meteorological forecast models generally have systematic biases. These models usually fail to fully consider key factors such as crop coverage, water level height, and terrain in farmland, especially paddy fields, which have a significant impact on the accuracy of meteorological forecasts. Therefore, the models of the prior art are difficult to meet the requirements of modern agricultural production for high-quality meteorological forecasts. Summary of the Invention
[0003] An object of an embodiment of the present invention is to provide a method and device for constructing a farmland meteorological forecast correction model, and the construction method improves the accuracy of the farmland meteorological forecast correction model.
[0004] To achieve the above object, an embodiment of the present invention provides a method for constructing a farmland meteorological forecast correction model, the method comprising:
[0005] Obtaining historical data of a target farmland;
[0006] Respectively using a light gradient boosting algorithm, an extreme gradient boosting algorithm, and an adaptive boosting algorithm to perform fitting training on the historical data to obtain a first prediction model, a second prediction model, and a third prediction model;
[0007] At different forecast lead times, respectively optimizing the three prediction models according to the convergence of the first prediction model, the second prediction model, and the third prediction model, and respectively obtaining the prediction results of the optimized three prediction models and the observed data corresponding to the prediction results;
[0008] Using the prediction results as feature data and the observed data as label data, performing stacked learning training and verification on the feature data and the label data according to a predetermined ratio to obtain a farmland meteorological forecast correction model.
[0009] Optionally, the historical data includes observed data and forecast data of temperature, relative humidity, wind speed, and air pressure, as well as water level data, elevation data, and normalized difference vegetation index data.
[0010] Optionally, using a light gradient boosting algorithm to perform fitting training on the historical data to obtain a first prediction model, comprising:
[0011] Use the decision tree splitting algorithm with a histogram and the leaf-first growth strategy to perform fitting training on the historical data to obtain a first prediction model.
[0012] Optionally, use the extreme gradient boosting algorithm to perform fitting training on the historical data to obtain a second prediction model, including:
[0013] Perform second-order approximation, regularization, sparse processing, distributed computing, and training on the historical data through the gradient boosting algorithm, regularization term, and weighted splitting strategy to obtain a second prediction model.
[0014] Optionally, use the adaptive boosting algorithm to perform fitting training on the historical data to obtain a third prediction model, including:
[0015] Iteratively train multiple weak classifiers based on the historical data, and adjust the weights of the misclassified samples in the multiple weak classifiers;
[0016] Combine the multiple weak classifiers through weighted voting to obtain a third prediction model, and the third prediction model is a strong classifier.
[0017] Optionally, under the prediction time limit, optimize the three prediction models according to the convergence of the first prediction model, the second prediction model, and the third prediction model respectively, including:
[0018] Use the K-fold cross-validation method and the grid search method to determine the training samples and validation samples of the first prediction model, the second prediction model, and the third prediction model respectively;
[0019] Determine the convergence and accuracy of the three models according to the training samples and validation samples;
[0020] Determine the optimized first prediction model, second prediction model, and third prediction model according to the convergence and accuracy.
[0021] Optionally, use the prediction result as the feature data and the observed data as the label data, and perform stacked learning training and validation on the feature data and the label data according to a predetermined ratio to obtain a farmland weather forecast correction model, including:
[0022] The prediction results include a first prediction result, a second prediction result, and a third prediction result;
[0023] Adopt the regression method to combine the first prediction result, the second prediction result, and the third prediction result to obtain the feature data, and combine the observed data at the corresponding moments of the first prediction result, the second prediction result, and the third prediction result to obtain the label data;
[0024] Perform cyclic training on the feature data and the label data according to a preset ratio using the grid search method and the K-fold cross-validation method to obtain training data;
[0025] Verify the training data to obtain a farmland weather forecast correction model for each forecast time range;
[0026] The parameters of the stacked learning training include the regularization strength and the regularization method.
[0027] Optionally, the method further includes: evaluating the farmland weather forecast correction model:
[0028] Obtain the difference between the predicted value and the actual data of the weather of the target farmland by using the root mean square error, and obtain the model with the smallest error as the final farmland weather forecast correction model.
[0029] Optionally, the method further includes:
[0030] Preprocess the historical data;
[0031] The preprocessing includes outlier identification, outlier removal, and data filling.
[0032] On the other hand, the present invention also proposes a device for constructing a farmland weather forecast correction model, and the device includes:
[0033] An acquisition module, configured to acquire historical data of a target farmland;
[0034] A first processing module, configured to respectively use a light gradient boosting algorithm, an extreme gradient boosting algorithm, and an adaptive boosting algorithm to perform fitting training on the historical data to obtain a first prediction model, a second prediction model, and a third prediction model;
[0035] A second processing module, configured to optimize the three prediction models according to the convergence of the first prediction model, the second prediction model, and the third prediction model respectively at different forecast time ranges, and respectively obtain the prediction results of the optimized three prediction models and the actual data corresponding to the prediction results;
[0036] A third processing module, configured to perform stacked learning training and verification on the feature data and the label data in a predetermined proportion with the prediction results as the feature data and the actual data as the label data, to obtain a farmland weather forecast correction model.
[0037] On the other hand, the present invention also proposes a method for predicting farmland weather, and the method uses the farmland weather forecast correction model constructed by the method for constructing a farmland weather forecast correction model described above to predict the weather of the target farmland.
[0038] A construction method of a farmland meteorological forecast correction model of the present invention includes: obtaining historical data of a target farmland; respectively using a light gradient boosting algorithm, an extreme gradient boosting algorithm, and an adaptive boosting algorithm to fit and train the historical data to obtain a first prediction model, a second prediction model, and a third prediction model; at different forecast lead times, respectively optimize the three prediction models according to the convergence of the first prediction model, the second prediction model, and the third prediction model, and respectively obtain the prediction results of the optimized three prediction models and the observed data corresponding to the prediction results; using the prediction results as feature data and the observed data as label data, perform stacked learning training and verification on the feature data and the label data according to a predetermined ratio to obtain a farmland meteorological forecast correction model. This construction method trains independent models for different forecast lead times and performs fine corrections, which can significantly reduce errors, ensure the accuracy of long-term meteorological forecasts, thereby improving the decision-making quality in agricultural production and reducing economic losses caused by false alarms.
[0039] Other features and advantages of the embodiments of the present invention will be described in detail in the following specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification, and are used together with the following specific implementation to explain the embodiments of the present invention, but do not constitute a limitation to the embodiments of the present invention. In the drawings:
[0041] Figure 1 is a schematic flowchart of a construction method of a farmland meteorological forecast correction model of the present invention;
[0042] Figure 2 is a schematic diagram of a construction device of a farmland meteorological forecast correction model of the present invention.
[0043] DESCRIPTION OF THE REFERENCE NUMERALS
[0044] 100 - Construction device of the farmland meteorological forecast correction model;
[0045] 200 - Acquisition module;
[0046] 300 - First processing module;
[0047] 400 - Second processing module;
[0048] 500 - Third processing module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] The following details the specific implementation of the embodiments of the present invention with reference to the drawings. It should be understood that the specific implementation described herein is only for explaining and understanding the embodiments of the present invention, and is not used to limit the embodiments of the present invention.
[0050] It should be noted that in the technical solution of this application, the acquisition, transmission, storage, use, processing, etc. of data all comply with the relevant provisions of national laws and regulations. In the embodiments of this application, some existing solutions in the industry such as certain software, components, models, etc. may be mentioned. They should be regarded as exemplary, and their purpose is only to illustrate the feasibility in the implementation of the technical solution of this application, but it does not mean that the applicant has already or necessarily used this solution.
[0051] Figure 1 It is a schematic flowchart of the construction method of a farmland meteorological forecast correction model of the present invention. As Figure 1 shown, a construction method of a farmland meteorological forecast correction model for predicting farmland meteorology of the present invention includes: step S101 is to obtain historical data of the target farmland.
[0052] According to a specific implementation manner, the historical data includes actual data and forecast data of temperature, relative humidity, wind speed, and air pressure, as well as water level data, elevation data, and normalized difference vegetation index (NDVI) data. The farmland can be wheat fields, paddy fields, and other crop fields, preferably paddy fields.
[0053] The historical observation data of the farmland meteorological station includes key meteorological variables such as air temperature, relative humidity, wind speed, and air pressure, which are collected from meteorological monitoring equipment installed in the paddy field, with a time span of nearly 5 years and hourly records. The historical forecast data is the CMA - MESO regional forecast product, which has a resolution of 3 km, a forecast period of 72 hours, and a time resolution of hourly. Data with the same time range and meteorological element variables as the observation data of the farmland meteorological station are collected, including 2m temperature, 2m relative humidity, 10m wind speed, air pressure, etc. The historical farmland water level data is recorded by water level gauges distributed in different fields to obtain hourly average water level data, and the time range is the same as that of the observation data of the farmland meteorological station. Historical vegetation index (NDVI) data: Since the growth of crops changes little within a day, the daily average NDVI data collected by a remote - sensing drone at a height of 120m can be used as the hourly NDVI value for the day. The farm elevation data is static elevation data, and the elevation of the farm is determined according to the longitude and latitude information of the farm meteorological station.
[0054] The method further includes: pre - processing the historical data; the pre - processing includes outlier identification, outlier removal, and data filling.
[0055] Specifically, the preprocessing is mainly carried out on historical observation data and forecast data. The outlier removal is to remove data points that significantly exceed the climate norm in all variables. The outlier identification is to identify outliers in continuous data such as temperature, relative humidity, and wind speed using the Z-score method (the distance of each data point from the standard deviation). Data points with a Z-score greater than 3 or less than -3 are regarded as outliers. The data filling is to fill in missing or abnormal data points using the linear interpolation method.
[0056] The hourly historical observation data of the farmland meteorological station formed after data preprocessing is used as the feature data label dataset (i.e., the true value), and other data combinations become the feature data. Since the CMA-MESO model has different systematic errors for forecasts of different lead times, different models need to be constructed for different lead times. That is, when the lead time is 72 hours, 72 feature data and feature labels are constructed respectively according to the lead time. The feature dataset and the label dataset can be divided into a training dataset and a test dataset in a ratio of 8:2.
[0057] Step S102 is to respectively use the light gradient boosting algorithm, the extreme gradient boosting algorithm, and the adaptive boosting algorithm to fit and train the historical data to obtain the first prediction model, the second prediction model, and the third prediction model.
[0058] According to a specific implementation manner, using the light gradient boosting algorithm to fit and train the historical data to obtain the first prediction model includes: using the decision tree splitting algorithm of the histogram and the leaf-first growth strategy to fit and train the historical data to obtain the first prediction model.
[0059] Specifically, assume that the objective function of the first prediction model is:
[0060]
[0061] where: y i is the true label of the i-th sample, and f(x i ) is the predicted value of the model (i.e., the output of the tree). L(y i , f(x i ) is the loss function (such as mean squared error, cross entropy, etc.), and Ω(f) is the regularization term, which is used to control the complexity of the model and prevent overfitting.
[0062] After giving a target function L(θ), in each round of iteration t, a new tree h t (x) will be added to the current model, so that the new prediction result becomes:
[0063] f t (x) = f t-1 (x) + ηht (x)
[0064] where: f t-1 (x) is the model prediction value of the previous iteration, η is the learning rate (usually a positive number less than 1, controlling the step size). h_t(x) is the newly added decision tree at the t-th round.
[0065] In each iteration, the first prediction model updates the model by calculating the gradient and the second-order gradient. For each tree t, the negative gradient based on the loss function is calculated, and these gradients are used to fit a new decision tree. The calculation formula is:
[0066] The first-order gradient (g i ):
[0067] The second-order gradient (h i ):
[0068] The first prediction model splits the decision tree through an optimized "leaf first" strategy, selecting the split point that can minimize the loss function. By optimizing the objective function, the gain of each node split is calculated:
[0069]
[0070] where: ∑g i is the sum of the first-order gradients of all samples in the current node. ∑h i is the sum of the second-order gradients of all samples in the current node. λ is the regularization term, controlling the complexity of the leaf nodes. In this way, the first prediction model optimizes the split of each tree during training, making each split maximize the performance improvement of the model.
[0071] To control the complexity of the model, the first prediction model introduces a regularization term:
[0072]
[0073] where: γ is the complexity coefficient of each tree, T is the depth of the tree. λ is the L2 regularization parameter, ∥w k ∥ is the weight of each leaf node.
[0074] After multiple iterations, the final prediction result is the weighted output sum of all trees:
[0075]
[0076] where, T is the total number of trees, η is the learning rate, h t (x) is the output of the t-th tree.
[0077] The historical data is fitted and trained using the extreme gradient boosting algorithm to obtain a second prediction model, including: performing second-order approximation, regularization, sparse processing, distributed calculation, and training on the historical data through the gradient boosting algorithm, regularization term, and weighted splitting strategy to obtain the second prediction model.
[0078] Specifically, define the objective function:
[0079] is the loss function, which is used to measure the error between the predicted value and the actual value of the model. Ω(f) is the regularization term, which is used to control the model complexity and prevent overfitting.
[0080] Initialize the model: The initial prediction F0(x) of the model is calculated as:
[0081]
[0082] where c is a constant value, is the loss function
[0083] Calculate the gradient and second-order gradient. In each round of iteration, the second prediction model guides the model update by calculating the first-order and second-order derivatives of the loss function. For each sample i, calculate the gradient and Hessian (i.e., second-order derivative):
[0084] First-order gradient (Gradient):
[0085] Second-order gradient (Hessian):
[0086] Incremental training: Assume the current predicted value is F t-1 (x). We hope to update the predicted value by adding the tree h t (x) in the t-th round. The new predicted value is:
[0087] F t (x) = F t-1 (x) + ηh t (x)
[0088] where η is the learning rate and h t (x) is the newly added decision tree.
[0089] Gain calculation. The training objective of each tree is to minimize the objective function, especially by splitting nodes to minimize the loss function. The second prediction model calculates the gain of each split point, that is, the reduction in the loss of the split node. The formula for calculating the gain is:
[0090]
[0091] IL and I R are the sample index sets of the left and right subtrees after the current split, respectively. and are the sums of the gradients and Hessians of the left subtree, respectively. The same applies to the right subtree. λ is the regularization parameter used to prevent overfitting.
[0092] Update the tree structure and leaf nodes. According to the calculation result of the gain, the second prediction model selects the optimal split point and constructs the tree. The output w of each leaf node k is:
[0093]
[0094] where I k is the sample index of the k-th leaf node, g i and h i are the gradient and Hessian of the sample.
[0095] Regularization. The second prediction model introduces a regularization term Ω(f) to control the complexity of the tree. The regularization term consists of two parts:
[0096] Number of leaves of the tree: T is the number of leaves of the tree.
[0097] Sum of the squares of the leaf node weights: where w k is the weight of the leaf node.
[0098] The total regularization term is:
[0099]
[0100] The purpose of the regularization term is to limit the complexity of the model and prevent overfitting.
[0101] Update the objective function and the final model. After each iteration is completed, the second prediction model updates the objective function of the model. The final predicted value is the sum of the weighted outputs of all the trees:
[0102]
[0103] where F(x) is the final prediction result, T is the total number of trees, η is the learning rate, and h t (x) is the predicted value of the t-th tree.
[0104] For prediction of new samples, the final prediction is the weighted sum of the outputs of all the trees:
[0105]
[0106] The output h of each tree t(x) is calculated based on the leaf node weights learned during the training process. Using the AdaBoost algorithm to perform fitting training on the historical data to obtain a third prediction model, including: iteratively training multiple weak classifiers according to the historical data, and adjusting the weights of the misclassified samples in the multiple weak classifiers; combining the multiple weak classifiers through weighted voting to obtain a third prediction model, and the third prediction model is a strong classifier.
[0107] Specifically, initialize the sample weights:
[0108]
[0109] where N is the number of training samples.
[0110] Train weak classifiers. The third prediction model is iteratively trained through a series of weak classifiers, usually using decision stumps (a tree with only a single decision node). Each weak classifier is trained based on the weights of the current samples. In the t-th round of training, the third prediction model trains a weak classifier h t (x), and calculates the error rate ∈ t .
[0111] The error rate ∈ t is calculated as
[0112]
[0113] where: I(h t (x i )≠y i ) is an indicator function. If the classifier h t (x i ) misclassifies sample i, the value is 1; otherwise it is 0. is the weight of sample i in the t-th round.
[0114] Calculate the classifier weights: The weight α t of each weak classifier h t (x) reflects its importance in the overall model. The weight α t of the classifier is related to its error rate ∈ t . The calculation formula is:
[0115]
[0116] where: when ∈ t →0, α t →∞, indicating that the classifier is almost perfect and gives it a large weight. When ∈ t →0.5, α t →0, indicating that the classifier can hardly distinguish positive and negative samples and has a small weight.
[0117] Update the weights of the samples: The key to the third type of prediction model is to assign higher weights to the misclassified samples, prompting subsequent classifiers to pay more attention to these difficult-to-classify samples.
[0118] The updated sample weights The calculation formula is as follows:
[0119]
[0120] Where: y i h t (x i ) is the prediction result of the classifier. If the classification is correct, then y i h t (x i ) = +1. If the classification is incorrect, then y i h t (x i ) = -1. The weights of misclassified samples will increase, while the weights of correctly classified samples will decrease. To avoid excessive weights, the weights of all samples are usually normalized so that the sum of all weights is 1:
[0121]
[0122] In this application, three machine learning frameworks, namely the Light Gradient Boosting Algorithm, the Extreme Gradient Boosting Algorithm, and the Adaptive Boosting Algorithm, are used to fit and train the training dataset. K-fold cross-validation (CV) is adopted to further improve the model accuracy, that is, the training dataset is divided into K parts (K is generally 10). One part is taken as the validation set in turn to test the accuracy. Finally, the average accuracy of K tests is taken, and then the trained model is used to test the test dataset. Finally, a correction model with good convergence and relatively high prediction accuracy is selected under each model framework at each forecast time.
[0123] During model training, Grid Search is used to tune the parameters of the three machine learning models. Within the specified parameter range, the parameters are cyclically adjusted at the specified step size and the learner is trained to find the parameters with the highest accuracy on the validation set within the set parameter range. The parameters for Grid Search include learning_rate, max_depth, n_estimators, etc.
[0124] Step S103 is to optimize the three prediction models according to the convergence of the first prediction model, the second prediction model, and the third prediction model respectively at different forecast times, and obtain the prediction results of the three optimized prediction models and the corresponding observed data of the prediction results respectively.
[0125] According to a specific embodiment, under the prediction time limit, the three prediction models are optimized according to the convergence of the first prediction model, the second prediction model, and the third prediction model respectively, including: determining the accuracy of the first prediction model, the second prediction model, and the third prediction model respectively by using the K-fold cross-validation method; determining the convergence and accuracy rate of the three models according to the accuracy; and determining the optimized first prediction model, second prediction model, and third prediction model according to the convergence and accuracy rate.
[0126] Step S104 is to use the prediction result as the feature data and the actual data as the label data, and perform stacked learning training and verification on the feature data and the label data according to a predetermined ratio to obtain a farmland meteorological forecast correction model.
[0127] According to a specific embodiment, using the prediction result as the feature data and the actual data as the label data, and performing stacked learning training and verification on the feature data and the label data according to a predetermined ratio to obtain a farmland meteorological forecast correction model, including: the prediction result includes a first prediction result, a second prediction result, and a third prediction result; the actual data at the corresponding moments of the first prediction result, the second prediction result, and the third prediction result is used as the label data (the actual data corresponding to the first prediction result, the second prediction result, and the third prediction result is the same data); performing cyclic training on the feature data and the label data according to a preset ratio by using the grid search method and the K-fold cross-validation method to obtain training data; verifying the training data to obtain a farmland meteorological forecast correction model for each prediction time limit; the parameters of the stacked learning training include the regularization strength and the regularization method.
[0128] The present invention uses the model stacking (Stacking) method to integrate the above three prediction models (base learners). The main idea of the stacking method is to use the prediction results of the three base learners as new features, use the meteorological actual observation value at the predicted moment as the label, and then combine the outputs of the three base learners through a meta-learner to generate a more accurate final prediction result.
[0129] The prediction results of the three base learners for each sample are spliced together to form a new feature vector. For n samples and k = 3 base learners, the prediction results of the base learners can be represented as an n×k matrix, and each column corresponds to the predicted value of a base learner. The true label of the meta-learner is still the true value of the original problem, that is, the actual observation value of the farmland meteorological station (such as 2m temperature, 2m relative humidity, etc.). Therefore, the meta-learner needs to fit the true observation value based on the above n×k prediction result features.
[0130] The present invention selects LASSO regression (Least Absolute Shrinkage and Selection Operator) as the meta-learner, and trains the model by minimizing the objective function with an L1 regularization term. This can obtain more sparse and interpretable parameters while maintaining accuracy.
[0131] For n training samples, each sample is represented by x i ∈R k , that is, the vector composed of the predicted values of three base learners for this sample, and the true observed value is y i . In LASSO regression, the following objective function needs to be solved: The first term: is the term for minimizing the mean squared error (MSE), which is used to ensure the fitting accuracy. The second term (regularization term): λ∥ω∥1 = λ∑ j |ω j | is the L1 norm penalty term, which can make some coefficients tend to zero, thus achieving the effects of feature selection and overfitting suppression. λ is the regularization strength (hyperparameter), which is used to control the balance between model complexity and sparsity. By optimizing the above objective function, the meta-learner can make a linear combination between the outputs of different base learners and perform feature selection, making the final prediction performance more robust. When stacking learning is trained, the feature data and label data are divided into a training data set and a test data set according to a ratio of 7:3.
[0132] The training set is used to train three base learners and generate predicted outputs, and then further train the meta-learner (LASSO). The test set is used to finally verify the overall prediction performance after stacking. When training the meta-learner (LASSO), K-fold cross-validation (such as K = 10) is also introduced to reduce the overfitting risk and improve the model accuracy. In each round, different folds are selected as the validation set, and the remaining folds are used as the training set. The validation results of each round are statistically analyzed and averaged as the evaluation index of the model performance.
[0133] Stacking learning also uses grid search to select the optimal LASSO hyperparameters, including: the value range of the regularization strength (λ); the regularization form (L1 / L2 or Elastic Net, etc., different penalty forms can be selected according to actual needs). The optimal combination is selected through the evaluation index on the validation set, and then the combination is used for the final training on the entire training set. After the meta-learner is trained, it is evaluated on the test data set, and the final stacked correction model prediction results are obtained at each forecast time limit to compare the improvement in its effect compared with the prediction of a single base learner.
[0134] The method further includes: evaluating the farmland meteorological forecast correction model: obtaining the difference between the predicted value and the actual data of the meteorology of the target farmland by using the root mean square error, and obtaining the model with the smallest error as the final farmland meteorological forecast correction model.
[0135] According to a specific implementation manner, in the base learner training and stacking model training, the root mean square error (RMSE) is used to measure the difference between the predicted value and the actual observed value. For the predicted value and the true observed value y i (i = 1, …, n), its calculation formula is:
[0136]
[0137] During the model training process, the model with the smallest RMSE is usually regarded as the model with the optimal performance. Each base learner and stacking model are evaluated on the cross-validation and test sets through the RMSE index, and finally the optimal model parameters and structure are selected.
[0138] The future meteorological data is forecast and corrected through the trained farmland meteorological forecast correction model. For each new prediction moment, the meteorological data is corrected. With the new CMA-MESO data and multi-dimensional features such as water level, NDVI, and farm elevation as the input, the output of each prediction includes the corrected values of meteorological variables such as temperature, relative humidity, wind speed, and air pressure. The model will generate hourly corrected meteorological data and can provide more accurate meteorological information for the farm.
[0139] The present invention realizes the accurate prediction of the meteorology of paddy fields through the integration and comprehensive analysis of multi-data sources. Specifically, by combining multi-source data such as hourly historical observation data, historical forecast data, farmland water level data, NDVI (vegetation index) data, and farm elevation data of farmland meteorological stations, a high-dimensional feature data set is constructed. Through data integration, various factors affecting the accuracy of paddy field meteorological forecasts can be captured more comprehensively, overcoming the limitations of a single meteorological data source or a single model. This method can more accurately reflect the climate change of the environment around the paddy field, especially when it comes to complex agro-meteorological factors (such as water level, crop growth, farm terrain, etc.), improving the accuracy of meteorological forecasts.
[0140] The present invention constructs different models for the correction models with different forecast lead times to perform forecast lead time correction. This time-sensitive model construction method effectively solves the problem of large error differences under different forecast lead times. By training independent models for different forecast lead times and performing fine corrections, the error can be significantly reduced, ensuring the accuracy of long-term meteorological forecasts, thereby improving the decision-making quality in agricultural production and reducing economic losses caused by false alarms.
[0141] The present invention also proposes a device for constructing a farmland weather forecast correction model, such as Figure 2 As shown, the construction device 100 of the farmland meteorological forecast correction model includes: an acquisition module 200, which is used to obtain historical data of the target farmland; a first processing module 300, which is used to fit the historical data using a mild gradient boosting algorithm, an extreme gradient boosting algorithm and an adaptive enhancement algorithm to obtain a first prediction model, a second prediction model and a third prediction model; a second processing module 400, which is used to optimize the three prediction models according to the convergence of the first prediction model, the second prediction model and the third prediction model under different forecast time limits, and obtain the prediction results of the three optimized prediction models and the actual data corresponding to the prediction results; a third processing module 500, which is used to use the prediction results as feature data and the actual data as label data, and to stack the feature data and the label data in a predetermined ratio for learning, training and verification, so as to obtain the farmland meteorological forecast correction model.
[0142] The corrected meteorological variables of the constructed device include temperature, relative humidity, wind speed, and air pressure, which are used to solve the problem that the current general weather forecast for rice fields is not accurate enough. The observation data of the farmland meteorological station, historical forecast data, historical farmland average water level, average vegetation index (NDVI) data, and elevation data of the location of the farm meteorological station are collected, and the data are preprocessed to remove outliers, thereby obtaining feature variables, feature data sets, and label data sets. Then, different model frameworks are used to train the correction model, and the model stacking method is used to integrate the model to further improve the generalization ability of the model, and then the new prediction data and correction model are used to realize the correction of the weather forecast.
[0143] On the other hand, the present invention also proposes a method for predicting farmland weather, which uses the evaluation model constructed according to the above-mentioned method for constructing the farmland weather forecast correction model to predict the weather of the target farmland.
[0144] A method for constructing a correction model for farmland weather forecasting of the present invention includes: obtaining historical data of a target farmland; respectively using a light gradient boosting algorithm, an extreme gradient boosting algorithm, and an adaptive boosting algorithm to perform fitting training on the historical data to obtain a first prediction model, a second prediction model, and a third prediction model; at different forecast lead times, respectively optimizing the three prediction models according to the convergence of the first prediction model, the second prediction model, and the third prediction model, and respectively obtaining the prediction results of the optimized three prediction models and the observed data corresponding to the prediction results; using the prediction results as feature data and the observed data as label data, performing stacked learning training and verification on the feature data and the label data according to a predetermined ratio to obtain a correction model for farmland weather forecasting. This construction method trains independent models for different forecast lead times and performs fine correction, which can significantly reduce errors, ensure the accuracy of long-term weather forecasting, thereby improving the decision-making quality in agricultural production and reducing economic losses caused by false alarms.
[0145] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0146] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 or multiple flows and / or blocks.
[0147] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in Figure 1 one or more flows and / or blocks Figure 1 or multiple flows and / or blocks.
[0148] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 steps of the functions specified in one block or multiple blocks.
[0149] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0150] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0151] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0152] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of another identical element in the process, method, commodity or device comprising the element.
[0153] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A method for constructing a farmland weather forecast correction model, characterized in that: The method includes: Obtain historical data of target farmland; Using a mild gradient boosting algorithm, an extreme gradient boosting algorithm, and an adaptive boosting algorithm to perform fitting training on the historical data, respectively, to obtain a first prediction model, a second prediction model, and a third prediction model; Under different forecast time effectiveness, respectively optimizing the three prediction models according to the convergence of the first prediction model, the second prediction model and the third prediction model, respectively obtaining the prediction results of the three optimized prediction models and the actual data corresponding to the prediction results; The prediction results are used as feature data and the actual data are used as label data. The feature data and label data are stacked for learning, training and verification in a predetermined ratio to obtain a farmland weather forecast correction model.
2. The construction method according to claim 1, characterized in that: The historical data include actual data and forecast data of temperature, relative humidity, wind speed, air pressure, water level data, elevation data and normalized difference vegetation index data.
3. The construction method according to claim 1, characterized in that: The historical data is fit trained using a mild gradient boosting algorithm to obtain a first prediction model, including: The historical data is fitted and trained using a histogram decision tree splitting algorithm and a leaf-first growth strategy to obtain a first prediction model.
4. The construction method according to claim 1, characterized in that: The historical data is fit trained using an extreme gradient boosting algorithm to obtain a second prediction model, including: The second prediction model is obtained by performing second-order approximation, regularization, sparse processing, distributed computing and training on the historical data through a gradient boosting algorithm, a regularization term and a weighted splitting strategy.
5. The construction method according to claim 1, characterized in that: The historical data is fit trained using an adaptive enhancement algorithm to obtain a third prediction model, including: Iteratively training a plurality of weak classifiers according to the historical data, and adjusting the weights of misclassified samples in the plurality of weak classifiers; The multiple weak classifiers are combined by weighted voting to obtain a third prediction model, where the third prediction model is a strong classifier.
6. The construction method according to claim 1, characterized in that: The three prediction models are optimized according to the convergence of the first prediction model, the second prediction model and the third prediction model under the prediction timeliness, including: Using K-fold cross validation method and grid search method to determine the training samples and verification samples of the first prediction model, the second prediction model and the third prediction model respectively; Determining the convergence and accuracy of the three models based on the training samples and the validation samples; The optimized first prediction model, second prediction model and third prediction model are determined according to the convergence and accuracy.
7. The construction method according to claim 1, characterized in that: The prediction result is used as feature data, the actual data is used as label data, and the feature data and label data are stacked, trained and verified in a predetermined ratio to obtain a farmland meteorological forecast correction model, including: The prediction results include a first prediction result, a second prediction result and a third prediction result; Using a regression method, combining the first prediction result, the second prediction result, and the third prediction result to obtain feature data, and combining the actual data at the corresponding moments of the first prediction result, the second prediction result, and the third prediction result to obtain label data; The feature data and label data are subjected to cyclic training using a grid search method and a K-fold cross validation method according to a preset ratio to obtain training data; Verifying the training data to obtain a farmland meteorological forecast correction model for each forecast time period; The parameters of the stacked learning training include regularization strength and regularization method.
8. The construction method according to claim 1, characterized in that: The method further includes: evaluating the farmland meteorological forecast correction model: The root mean square error is used to obtain the difference between the predicted value and the actual data of the weather of the target farmland, and the model with the smallest error is obtained as the final farmland weather forecast correction model.
9. The construction method according to claim 1, characterized in that: The method further includes: Preprocessing the historical data; The preprocessing includes outlier identification, outlier removal and data filling.
10. A device for constructing a farmland weather forecast correction model, characterized in that: The device includes: An acquisition module is used to obtain historical data of the target farmland; A first processing module is used to perform fitting training on the historical data using a mild gradient boosting algorithm, an extreme gradient boosting algorithm, and an adaptive boosting algorithm to obtain a first prediction model, a second prediction model, and a third prediction model; The second processing module is used to optimize the three prediction models according to the convergence of the first prediction model, the second prediction model and the third prediction model under different forecast time effectiveness, and obtain the prediction results of the three optimized prediction models and the actual data corresponding to the prediction results respectively; The third processing module is used to use the prediction results as feature data and the actual data as label data, and to perform stacking learning training and verification on the feature data and label data in a predetermined ratio to obtain a farmland meteorological forecast correction model.
11. A method for predicting farmland weather, characterized in that: The method predicts the weather of the target farmland using the farmland weather forecast correction model constructed according to the method for constructing the farmland weather forecast correction model according to any one of claims 1 to 9.
Citation Information
Cited By
Credibility quantification method, system and equipment based on vertical atmospheric prediction
CN121071435A