Forecasting method for borer occurrence based on multi-source meteorological lag characteristics and ensemble learning

By constructing a multi-source meteorological lag feature and ensemble learning method, the problem of simple feature types in existing technologies is solved, and the accuracy and stability of soybean pod borer prediction are improved, providing early warning and precise control capabilities.

CN122132691APending Publication Date: 2026-06-02JILIN AGRICULTURAL UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JILIN AGRICULTURAL UNIV
Filing Date
2026-05-07
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing methods for predicting soybean pod borers rely on overly simplistic feature types built through feature engineering, and fail to delve into the multivariate relationships between meteorological factors and pest occurrence over time, resulting in limited prediction accuracy.

Method used

We employ a multi-source meteorological lag feature and ensemble learning approach to construct a multi-scale lag feature set. We then combine various machine learning algorithms to compare and fuse models, forming an ensemble prediction model. This model includes data acquisition, feature selection and dimensionality reduction, comparison of multiple algorithm models, and ensemble learning fusion.

Benefits of technology

It significantly improved the accuracy and stability of soybean pod borer prediction, increasing the R² of the daily pod population prediction model from 0.12 to 0.65, and reducing the average absolute error of the occurrence period prediction to 1.7 days, thus achieving early warning and precise control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122132691A_ABST
    Figure CN122132691A_ABST
Patent Text Reader

Abstract

This method for predicting the occurrence of soybean pod borer, based on multi-source meteorological lag features and ensemble learning, belongs to the field of soybean pest forecasting technology. It addresses the technical problem that existing soybean pod borer forecasting methods often employ overly simplistic feature types in feature engineering, failing to delve into the multivariate relationships between meteorological factors and pest occurrence over time. The method is generally divided into five parts: S1, data acquisition and preprocessing; S2, multi-scale lag feature construction; S3, feature selection and dimensionality reduction; S4, multi-algorithm model comparison and optimization; and S5, ensemble learning and prediction. By constructing a multi-scale lag meteorological feature set, comparing and optimizing multiple algorithm models, and performing ensemble learning and prediction, it achieves high-precision, early warning prediction of the soybean pod borer occurrence period and population size.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of soybean pest prediction technology, specifically involving a method for predicting the occurrence of the pod borer based on multi-source meteorological lag characteristics and ensemble learning. Background Technology

[0002] The soybean pod borer is one of the major pests in my country's main soybean-producing areas, such as the Huang-Huai-Hai Plain and Northeast China. Its larvae bore into soybean grains, leading to yield losses and reduced quality. In years with severe infestations, the infestation rate can reach 20%-30%, and in severe cases, even exceed 50%, causing significant economic losses to soybean production. Therefore, accurately predicting the occurrence period and dynamics of the soybean pod borer is crucial for developing scientific control strategies, reducing pesticide use, and ensuring safe soybean production.

[0003] Traditional methods for predicting soybean pod borers can be mainly categorized as follows: (1) Field survey method: The number of adult insects and the larval borer rate are counted by trapping them with sex pheromones or taking samples in the field. This method requires professional personnel to stay in the field for a long time, which is time-consuming and labor-intensive, and cannot achieve early warning.

[0004] (2) Experience-based prediction method: Based on historical occurrences and meteorological conditions, qualitative judgments are made relying on expert experience. This method is highly subjective, and the accuracy of predictions is greatly affected by personal experience, making it difficult to promote and apply.

[0005] (3) Accumulated temperature model method: This method uses the effective accumulated temperature rule to predict the development progress of pests. This method only considers temperature as a single factor and ignores the influence of other important meteorological factors such as humidity and precipitation, resulting in limited prediction accuracy.

[0006] (4) Simple regression model: Establish a linear regression relationship between meteorological factors and pest occurrence. This method cannot capture the nonlinear relationship between meteorological factors and pest occurrence, and does not consider the lag effect of meteorological factors, resulting in unsatisfactory prediction results.

[0007] To address the aforementioned issues, Chinese invention patent application "A Machine Learning-Based Method for Predicting and Forecasting Soybean Harvester Borer" (Publication No. CN121660484A) discloses a machine learning-based method for predicting and forecasting soybean harvester borer, solving the technical problems of low prediction efficiency and unsatisfactory prediction results in traditional methods. However, its feature engineering construction uses traditional feature indicators such as continuous phenological indicators (based on effective accumulated temperature), daily air temperature, relative humidity, and historical data on total trapping. The constructed feature types are too simple and do not delve into the multivariate relationship between meteorological factors and pest occurrence over time. Summary of the Invention

[0008] To address the technical problem that existing methods for predicting and forecasting soybean pod borers have overly simplistic feature types constructed through feature engineering and have not delved into the multivariate relationships between meteorological factors and pest occurrence over time, this invention provides a pod borer occurrence prediction method based on multi-source meteorological lag features and ensemble learning.

[0009] The method includes the following steps: S1. Data Acquisition and Preprocessing: Collect meteorological and insect data for the field to be predicted during the monitoring period, and perform outlier detection, missing value imputation, and data standardization. S2. Construction of multi-scale lag features: Construct a multi-scale lag feature set for the preprocessed data; S3. Feature selection and dimensionality reduction: Evaluate the nonlinear correlation between each feature and insect infestation, select the feature subset with the strongest predictive ability, and construct a dimensionality-reduced multi-scale lag feature set. S4. Comparison and Optimization of Multiple Algorithm Models: The one-year leave-one-year cross-validation method is used to evaluate the predictive performance of each prediction model for daily insect population and occurrence period; S5. Ensemble Learning and Prediction: Based on the model comparison results, the three prediction models with the best performance in predicting daily insect population and occurrence period are selected and weighted and fused to construct an ensemble prediction model for daily insect population and an ensemble prediction model for occurrence period. Based on the feature selection and dimensionality reduction dataset, daily insect population and occurrence period prediction are performed.

[0010] Furthermore, the meteorological data monitoring period is from April 1st to September 30th each year, covering the entire occurrence period of the soybean pod borer; Insect data were obtained using the sex pheromone trapping method. The monitoring indicators included the daily number of adult insects and the cumulative number of insects. The monitoring period for insect data was from July 20 to September 30 each year.

[0011] Furthermore, outlier detection: The 3σ criterion is used to identify and remove outliers in meteorological data. Data points that exceed the mean ± 3 times the standard deviation are marked and corrected. Missing value imputation: For missing meteorological data, linear interpolation is used to impute them. For data that is missing for more than 3 consecutive days, the average value of the same period over many years is used as a substitute. Data standardization: Meteorological features are Z-score standardized to eliminate the influence of dimensional differences on the model.

[0012] Furthermore, the multi-scale lag feature set specifically includes: Real-time characteristics: average temperature, daily minimum temperature, daily maximum temperature, daily temperature range, daily relative humidity, precipitation, effective accumulated temperature, daily maximum wind speed, temperature and humidity index, and heat stress index; Lag characteristics: Average temperature, maximum temperature, minimum temperature, relative humidity, precipitation, effective accumulated temperature and wind speed with lags of 1, 2, 3, 5, 7, 10 and 14 days respectively were constructed; Rolling statistical characteristics: Rolling average temperature, rolling temperature standard deviation, rolling average humidity, rolling cumulative precipitation, and rolling number of precipitation days are constructed for rolling windows of 3, 5, 7, 10, 14, and 21 days, respectively. Cumulative characteristics: cumulative precipitation, cumulative number of high-temperature days, and cumulative number of precipitation days during the meteorological data monitoring period; Interaction characteristics: temperature-precipitation interaction term and humidity-precipitation interaction term; Phenological characteristics: annual accumulated days, whether it occurs in May-June, whether it occurs in July-August, extreme high-temperature days, and rainstorm days.

[0013] Furthermore, the specific process of feature selection and dimensionality reduction is as follows: Calculate the mutual information value between each feature and the target variable; Sort the features in descending order of mutual information value and select the top 40 features to form the optimal feature subset; The 40 selected features include: instantaneous temperature features, 1-14 day hysteresis temperature features, 3-21 day rolling temperature features, and cumulative features.

[0014] Furthermore, the prediction models include random forest model, gradient boosting tree model, extreme random tree model, support vector regression model, K-nearest neighbor regression model, ridge regression model, Lasso regression model and elastic network regression model.

[0015] Furthermore, for the prediction of daily insect population, the constructed integrated prediction model for daily insect population is as follows: extreme random tree model weight 0.4, support vector regression model weight 0.35, and Lasso regression model weight 0.25. For occurrence period prediction, the constructed occurrence period ensemble prediction model is as follows: Random Forest model weight 0.4, Extreme Random Tree model weight 0.4, and K Nearest Neighbor Regression model weight 0.2.

[0016] The beneficial effects of the method described in this invention are as follows: A method for constructing multi-scale lag meteorological features: This invention, for the first time, systematically constructs a multi-scale meteorological feature set including 1-14 day lags, 3-21 day rolling statistics, and cumulative data throughout the entire growth period, fully considering the lag and cumulative effects of meteorological conditions on the occurrence of soybean pod borer. Compared with existing methods that only use meteorological data from the same period, the feature set of this invention can capture the impact of early meteorological conditions on pest development, reproduction, and survival, significantly improving prediction accuracy. Experimental results show that after adding lag features, the R² of the daily pest population prediction model increased from 0.12 to 0.65, an improvement of 428.4%.

[0017] Multi-algorithm comparison and optimization combined with ensemble learning strategy: This invention innovatively employs a multi-algorithm comparison and optimization strategy, systematically comparing the performance of eight machine learning algorithms on the soybean codling moth prediction task, selecting the most suitable combination of base models, and then using ensemble learning for fusion prediction. Compared with single models, the ensemble model can combine the advantages of each base model, reduce prediction variance, and improve prediction stability. Experimental results show that the prediction performance of the ensemble model is superior to any single base model, with the mean absolute error of the occurrence period prediction reduced to 1.7 days, an improvement of approximately 12% compared to the best single model.

[0018] A Phenological Period-Based Annual Forecasting Method: This invention innovatively proposes a phenological period-based annual forecasting method. Using meteorological data from April to July, it can predict the occurrence period and quantity of pests 1-2 months before their emergence, achieving true early warning. Analysis reveals a strong negative correlation between total precipitation in July and the peak occurrence period (r=-0.81), and a positive correlation between average temperature in July and the peak occurrence period (r=0.62). These key meteorological factors can serve as early warning indicators. This method provides ample time for decision-making preparation for the precise control of soybean pod borers, contributing to reduced pesticide use and green pest control.

[0019] A complete technical methodology and system implementation: Based on the proposed effective prediction method, this invention also constructs a complete technical system from data acquisition, feature engineering, model training to prediction applications, forming a scalable and replicable prediction system. This system features modular design, configurable parameters, and visualized results, facilitating its application in pest prediction across different regions and crops. Attached Figure Description

[0020] Figure 1 This is a flowchart of the method described in an embodiment of the present invention; Figure 2 This is a comparison chart of the predicted total insect population and the actual value in an embodiment of the present invention; Figure 3 This is a comparison chart of the predicted occurrence period and the actual value in an embodiment of the present invention. Detailed Implementation

[0021] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.

[0022] Example 1 This embodiment provides a method for predicting the occurrence of codling moths based on multi-source meteorological lag features and ensemble learning. The overall process of the method is as follows: Figure 1 As shown in the attached diagram, the scheme will be explained in detail.

[0023] 1. Data Acquisition and Preprocessing The data used were derived from continuous monitoring data of the experimental field from 2020 to 2025. Meteorological data were collected by automatic weather stations installed at the experimental site, and the observed elements included: daily average temperature (°C), daily minimum temperature (°C), daily maximum temperature (°C), daily maximum wind speed (m / s), daily relative humidity (%), precipitation (mm), and effective accumulated temperature (°C·d). The data acquisition time resolution was 30 minutes, the statistics were performed on a daily scale, and the monitoring period was from April 1 to September 30 each year, covering the complete occurrence period of the soybean pod borer.

[0024] Insect data were obtained using a pheromone trapping method. A 20-meter protective row was set up around the perimeter of the experimental site. Three water-basin-shaped traps were placed every 25 meters along the same ridge. The traps were placed on iron frames and secured with wire. A pheromone lure was suspended inside the trap, and approximately one-third of the container was filled with detergent water. The number of soybean pod borers collected was checked daily at 8:00 AM, and the trapped soybean pod borers were collected. The lure was replaced every 15 days. Monitoring indicators included: daily adult emergence (heads / day) and cumulative insect population (heads). (The monitoring period was from July 20th to September 30th each year, covering the complete occurrence period of the soybean pod borer.) Data preprocessing includes the following steps: (1) Outlier detection: The 3σ criterion is used to identify and remove outliers in meteorological data. Data points that exceed the mean ± 3 times the standard deviation are marked and corrected.

[0025] (2) Missing value imputation: For missing meteorological data, linear interpolation is used to fill in the missing data; for data missing for more than 3 consecutive days, the average value of the same period over many years is used as a substitute.

[0026] (3) Data standardization: The meteorological characteristics are Z-score standardized to eliminate the influence of dimensional differences on the model.

[0027] 2. Construction of multi-scale hysteresis features Based on the preprocessed meteorological data, a multi-scale lag feature set is constructed according to the method described in this invention, specifically including: (1) Instantaneous characteristics (10): average temperature, daily minimum temperature, daily maximum temperature, daily temperature difference, daily relative humidity, precipitation, effective accumulated temperature, daily maximum wind speed, temperature and humidity index, and heat stress index.

[0028] (2) Lag characteristics (49): The average temperature, maximum temperature, minimum temperature, relative humidity, precipitation, effective accumulated temperature and wind speed with lags of 1, 2, 3, 5, 7, 10 and 14 days are constructed respectively, for a total of 7×7=49 characteristics.

[0029] (3) Rolling statistical features (30): Rolling average temperature, rolling temperature standard deviation, rolling average humidity, rolling cumulative precipitation, and rolling number of precipitation days for rolling windows of 3, 5, 7, 10, 14, and 21 days are constructed respectively, for a total of 6×5=30 features.

[0030] (4) Cumulative characteristics (3): Construct the cumulative precipitation, cumulative number of high temperature days, and cumulative number of precipitation days from April 1 to September 30.

[0031] (5) Interaction features (2): temperature-precipitation interaction term, humidity-precipitation interaction term.

[0032] (6) Phenological characteristics (4): annual accumulated days, whether it is May-June, whether it is July-August, extreme high temperature days, and rainstorm days.

[0033] The final feature set contains 101 features, providing rich input information for subsequent model training and prediction.

[0034] 3. Feature Selection and Dimensionality Reduction The mutual information method was used to assess the nonlinear correlation between various features and pest infestations, and the subset of features with the strongest predictive power was selected. The mutual information method can capture any statistical dependence between features and the target variable, and is suitable for handling the complex nonlinear relationship between meteorological factors and pest occurrence.

[0035] The feature selection process is as follows: (1) Calculate the mutual information value between each feature and the target variable (daily insect population); (2) Sort by mutual information value in descending order and select the top 40 features to form the optimal feature subset; (3) The 40 selected features include: instantaneous temperature features, 1-14 day lag temperature features, 3-21 day rolling temperature features, cumulative features, and phenological period features, as detailed below: Real-time characteristics (5 in total): average temperature, daily minimum temperature, effective accumulated temperature, temperature and humidity index, and annual accumulated days; Lag characteristics 1-14 days (18 in total): Lagging average temperature by 1 day, lag by 2 days, lag by 3 days, lag by 5 days, lag by 7 days, lag by 10 days, and lag by 14 days; Lagging minimum temperature by 1 day, lag by 2 days, lag by 3 days, lag by 5 days, lag by 7 days, lag by 10 days, and lag by 14 days; Lagging effective accumulated temperature by 1 day, lag by 2 days, lag by 3 days, lag by 5 days, lag by 7 days, lag by 10 days, and lag by 14 days. Rolling characteristics for 3-21 days (14 in total): Rolling 3-day average temperature, Rolling 5-day average temperature, Rolling 7-day average temperature, Rolling 10-day average temperature, Rolling 14-day average temperature, Rolling 21-day average temperature; Rolling 7-day relative humidity, Rolling 10-day relative humidity, Rolling 14-day relative humidity, Rolling 21-day relative humidity; Cumulative characteristics (3 in total): cumulative precipitation, cumulative number of high-temperature days, and cumulative number of precipitation days.

[0036] After feature selection, the feature dimension was reduced from 101 to 40, significantly reducing the model complexity while retaining the main prediction information.

[0037] 4. Optimization through comparison of multiple algorithm models The predictive performance of each model is evaluated using a leave-one-year cross-validation method, which uses five years of data as the training set and the remaining year of data as the test set. The models are validated six times in turn to ensure that each year is used as the test set once. Finally, the average result is taken.

[0038] The machine learning algorithms compared include: Random Forest, Gradient Boosting, Extra Trees, Support Vector Regression (SVR), K-Nearest Neighbors Regression (KNN), Ridge Regression, Lasso Regression, and ElasticNet Regression.

[0039] Table 1 shows the comparison results of the daily insect population prediction models: Table 1. Comparison of Daily Insect Population Prediction Models (after logarithmic transformation)

[0040] As shown in Table 1, Extremely Random Trees performed best, with a coefficient of determination (R²) of 0.67, followed by Support Vector Regression (R²=0.65) and Lasso Regression (R²=0.63). Overall, ensemble learning methods (Random Forest, Gradient Boosting Tree, Extremely Random Trees) outperformed linear and single-model methods.

[0041] The comparison results of the prediction models for the occurrence period are shown in Table 2: Table 2 Comparison Results of Occurrence Prediction Models

[0042] As shown in Table 2, K-Nearest Neighbors Regression performed best, with a mean absolute error (MAE) of 4.60 days, followed by Random Forest (4.64 days) and Extremely Random Tree (4.71 days). Ensemble learning methods also performed excellently on the occurrence prediction task.

[0043] 5. Ensemble learning and prediction Based on the model comparison results, the three best-performing models—random forest, extreme random tree, and K-nearest neighbor regression—were selected and weighted together to construct an integrated prediction model.

[0044] Weighting strategy: A dynamic weighting method is adopted based on the performance of each model in cross-validation. For daily insect population prediction, the weight of Extreme Random Tree is 0.4, the weight of Support Vector Regression is 0.35, and the weight of Lasso Regression is 0.25; for occurrence period prediction, the weight of Random Forest is 0.4, the weight of Extreme Random Tree is 0.4, and the weight of K-Nearest Neighbor Regression is 0.2.

[0045] ensemble model prediction performance: (1) Daily insect population prediction: The coefficient of determination of the ensemble model was 0.68, and the root mean square error was 1.12, which is about 2% higher than the best single model; (2) Prediction of occurrence period: The average absolute error of the integrated model is 4.90 days, which is basically the same as the best single model, but the prediction stability is better.

[0046] 6. Annual Forecast Results and Analysis Based on meteorological data from April to July, the occurrence period and total number of soybean pod borers were predicted for each year. The results are shown in Table 3. Table 3 Summary of Annual Forecast Results

[0047] As can be seen from Table 3: (1) High accuracy in predicting the occurrence period: The prediction error for each year is within 3 days, with an average absolute error of 1.7 days and a maximum error of 2.3 days (2020). The prediction accuracy meets the actual needs of agricultural production and can provide a reliable basis for determining the prevention and control period.

[0048] (2) The overall insect population prediction trend is accurate: Although the prediction deviation is large in some years (such as 2021), the overall trend prediction is accurate. It can distinguish between years with severe outbreaks (2020, 2022) and years with mild outbreaks (2021, 2024), providing a basis for the classification of the severity of the outbreak.

[0049] A more intuitive comparison chart of the predicted and actual total insect population is shown below. Figure 2 As shown in the figure, the comparison between the predicted and actual values ​​for the occurrence period is as follows: Figure 3 As shown.

[0050] 7. Single-variable experiment The main purpose of this section is to verify the impact of adding lag features on the predictive performance of soybean pod borer daily populations. The R² value of the prediction model using the traditional method is used to illustrate that adding lag features significantly improves the explanatory power and generalization ability of the prediction model.

[0051] (1) Variables without lag characteristics (11): average temperature, daily minimum temperature, daily maximum temperature, daily maximum wind speed, daily relative humidity, precipitation, temperature-humidity ratio, temperature difference, effective accumulated temperature, daily temperature difference, temperature-humidity index, and annual accumulated days.

[0052] Indicator Calculation Explanation: Temperature-humidity ratio: the ratio of temperature to humidity; Effective accumulated temperature: max(0, daily average temperature -10℃); Daily temperature range: the difference between the highest and lowest temperatures of the day; Temperature and humidity index: average temperature × daily relative humidity / 100; Yearly accumulated days: The cumulative number of days from January 1st of the current year to that day.

[0053] (2) Variables with lagged characteristics (53): Variables without lagged characteristics (11) plus lagged characteristics (42), of which the lagged characteristics (42) are the average temperature, maximum temperature, minimum temperature, relative humidity, precipitation and effective accumulated temperature lagged by 1, 2, 3, 5, 7, 10 and 14 days respectively.

[0054] (3) Model used: Lasso regression (alpha=0.1) (4) The experimental comparison results are shown in Table 4: Table 4

[0055] Conclusion: After adding the lag feature, the R² of the daily insect population prediction model increased from 0.12 to 0.65, an increase of 428.4%.

[0056] Example 2 This embodiment further defines Embodiment 1, conducts key meteorological factor analysis, and promotes the system application and application, thereby improving and verifying the technical solution in Embodiment 1.

[0057] 1. Analysis of key meteorological factors By analyzing the importance of features, key meteorological factors influencing the occurrence of soybean pod borer were identified: (1) Total precipitation in July: It is strongly negatively correlated with the peak occurrence period (r = -0.81) and is the most important predictor. More precipitation in July is conducive to adult emergence and oviposition, leading to an earlier peak occurrence period.

[0058] (2) Total precipitation from April to July: It is negatively correlated with the peak period of occurrence (r = -0.63). More precipitation in the early period is conducive to the survival and development of overwintering larvae.

[0059] (3) Average temperature in July: positively correlated with the peak period of occurrence (r=0.62). High temperature will delay the development of pests and postpone the peak period of occurrence.

[0060] (4) Number of high-temperature days: The number of days with the highest daily temperature > 30℃ is negatively correlated with the total number of insects (r = -0.41), indicating that high temperature has an inhibitory effect on pests.

[0061] The aforementioned key meteorological factors can serve as early warning indicators, allowing for a preliminary assessment of the annual trend based on meteorological conditions before the end of July.

[0062] 2. System Application and Promotion The prediction method described in this invention has been developed into a complete software system, featuring modular design, configurable parameters, and result visualization. The system application process is as follows: (1) Data input: Users input meteorological observation data from April 1 to July 31 of the current year, including daily average temperature, daily maximum temperature, daily minimum temperature, daily relative humidity, precipitation, effective accumulated temperature and other indicators.

[0063] (2) Feature calculation: The system automatically calculates multi-scale lag features and annual meteorological features.

[0064] (3) Model prediction: Call the trained prediction model and output the predicted date of the peak period and the predicted total number of insects.

[0065] (4) Results presentation: The prediction results are presented in the form of charts and reports, including the judgment of the occurrence trend, the degree of occurrence, and prevention and control recommendations.

[0066] (5) Decision support: Based on the forecast results, agricultural technicians can formulate prevention and control plans in advance and carry out chemical prevention and control 7-10 days before the predicted peak period to achieve precise prevention and control.

[0067] The above system has the following beneficial effects: The prediction accuracy has been significantly improved: the constructed daily insect population prediction model has an R² of 0.68, the average absolute error of the occurrence period prediction is only 1.7 days, and the RMSE of the total insect population prediction is 0.54 (log scale). All performance indicators are superior to the existing technology. Early warning can be achieved: by using meteorological data from April to July, the occurrence of soybean pod borer can be predicted before the end of July, which is 1-2 months earlier than the actual occurrence of the pest, allowing sufficient time for prevention and control decisions. Quantitative prediction results: Provides quantitative prediction results, including the specific date of the peak occurrence and the predicted total insect population, which can be directly used to guide the calculation of pesticide dosage and the determination of the control period, so as to achieve precise control; Simple to operate and low cost: Forecasts can be completed by simply inputting routine meteorological observation data. There is no need for complicated field surveys and expensive monitoring equipment, which can significantly reduce forecasting costs and is suitable for application by grassroots agricultural technology extension departments. The method is universal and scalable: the technical methodology has good versatility and can be extended to the prediction and forecasting of other crop pests by adjusting the feature parameters and model parameters.

Claims

1. A method for predicting the occurrence of codling moths based on multi-source meteorological lag characteristics and ensemble learning, characterized in that, The method includes the following steps: S1. Data Acquisition and Preprocessing: Collect meteorological and insect data for the field to be predicted during the monitoring period, and perform outlier detection, missing value imputation, and data standardization. S2. Construction of multi-scale lag features: Construct a multi-scale lag feature set for the preprocessed data; S3. Feature selection and dimensionality reduction: Evaluate the nonlinear correlation between each feature and insect infestation, select the feature subset with the strongest predictive ability, and construct a dimensionality-reduced multi-scale lag feature set. S4. Comparison and Optimization of Multiple Algorithm Models: The one-year leave-one-year cross-validation method is used to evaluate the predictive performance of each prediction model for daily insect population and occurrence period; S5. Ensemble Learning and Prediction: Based on the model comparison results, the three prediction models with the best performance in predicting daily insect population and occurrence period are selected and weighted and fused to construct an ensemble prediction model for daily insect population and an ensemble prediction model for occurrence period. Based on the feature selection and dimensionality reduction dataset, daily insect population and occurrence period prediction are performed.

2. The method for predicting the occurrence of codling moths based on multi-source meteorological lag features and ensemble learning according to claim 1, characterized in that, The meteorological data monitoring period is from April 1 to September 30 each year, covering the entire occurrence period of the soybean pod borer; Insect data were obtained using the sex pheromone trapping method. The monitoring indicators included the daily number of adult insects and the cumulative number of insects. The monitoring period for insect data was from July 20 to September 30 each year.

3. The method for predicting the occurrence of codling moths based on multi-source meteorological lag features and ensemble learning according to claim 1, characterized in that, Outlier detection: Outliers in meteorological data are identified and removed using the 3σ criterion. Data points exceeding the mean ± 3 standard deviations are marked and corrected. Missing value imputation: Missing meteorological data are filled using linear interpolation. For data missing for more than 3 consecutive days, the average value of the same period over many years is used as a substitute. Data standardization: Meteorological features are Z-score standardized to eliminate the influence of dimensional differences on the model.

4. The method for predicting the occurrence of codling moths based on multi-source meteorological lag features and ensemble learning according to claim 3, characterized in that, Multi-scale lag feature sets, specifically including: Real-time characteristics: average temperature, daily minimum temperature, daily maximum temperature, daily temperature range, daily relative humidity, precipitation, effective accumulated temperature, daily maximum wind speed, temperature and humidity index, and heat stress index; Lag characteristics: Average temperature, maximum temperature, minimum temperature, relative humidity, precipitation, effective accumulated temperature and wind speed with lags of 1, 2, 3, 5, 7, 10 and 14 days respectively were constructed; Rolling statistical characteristics: Rolling average temperature, rolling temperature standard deviation, rolling average humidity, rolling cumulative precipitation, and rolling number of precipitation days were constructed for rolling windows of 3, 5, 7, 10, 14, and 21 days, respectively. Cumulative characteristics: cumulative precipitation, cumulative number of high-temperature days, and cumulative number of precipitation days during the meteorological data monitoring period; Interaction characteristics: temperature-precipitation interaction term and humidity-precipitation interaction term; Phenological characteristics: annual accumulated days, whether it occurs in May-June, whether it occurs in July-August, extreme high-temperature days, and rainstorm days.

5. The method for predicting the occurrence of codling moths based on multi-source meteorological lag features and ensemble learning according to claim 4, characterized in that, The specific process of feature selection and dimensionality reduction is as follows: Calculate the mutual information value between each feature and the target variable; Sort the features in descending order of mutual information value and select the top 40 features to form the optimal feature subset; The 40 selected features include: instantaneous temperature features, 1-14 day hysteresis temperature features, 3-21 day rolling temperature features, and cumulative features.

6. The method for predicting the occurrence of codling moths based on multi-source meteorological lag features and ensemble learning according to claim 5, characterized in that, Prediction models include random forest, gradient boosting tree, extreme random tree, support vector regression, K-nearest neighbor regression, ridge regression, Lasso regression, and elastic network regression.

7. The method for predicting the occurrence of codling moths based on multi-source meteorological lag features and ensemble learning according to claim 6, characterized in that, For daily insect population prediction, the constructed integrated prediction model for daily insect population is as follows: extreme random tree model weight 0.4, support vector regression model weight 0.35, and Lasso regression model weight 0.

25. For occurrence prediction, the constructed occurrence prediction ensemble model is as follows: Random Forest model weight 0.4, Extreme Random Tree model weight 0.4, and K-Nearest Neighbor Regression model weight 0.

2.

8. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-7.

9. A computer-readable storage medium for storing computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the steps of the method according to any one of claims 1-7.