Machine learning algorithm-based remote sensing estimation system and method for grass yield of grassland

Through the remote sensing estimation system of grassland production and grass volume based on machine learning algorithms, a variety of influencing factors are considered to be comprehensively considered to build a grassland production and grass volume estimation model, which solves the problem of insufficient estimation accuracy in the existing technology, and achieves high-precision grass volume estimation.

CN119961594APending Publication Date: 2025-05-09青海省气象科学研究所
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510022538.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The estimation accuracy of the existing grassland grassland estimation model is insufficient and multiple influencing factors cannot be effectively considered.

Method used

A remote sensing estimation system for grassland production and grass volume based on machine learning algorithms is adopted, and a multi-source data such as vegetation index, meteorological factors, and topographic factors are comprehensively considered, and a grassland production and grass volume estimation model is constructed through various machine learning algorithms.

Benefits of technology

The accuracy of grassland production and grass volume estimation has been significantly improved, and efficient estimation of grassland production has been achieved, which is suitable for promotion and application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961594A_ABST
    Figure CN119961594A_ABST
Patent Text Reader

Abstract

The invention discloses a machine learning algorithm-based remote sensing estimation system and method for the grass yield of a grassland. The system comprises a data acquisition module, a data preprocessing module, a feature selection module, a machine learning algorithm module, a model training and verification module and a grass yield estimation module of the grassland. The invention further discloses a remote sensing estimation method for the grass yield of the grassland based on the machine learning algorithm. According to the technical scheme, multi-source data such as the vegetation index, the meteorological factor and the topographic factor are comprehensively considered, the grass land grass yield estimation model is constructed through multiple machine learning algorithms, and the estimation precision is remarkably improved. Meanwhile, through comprehensive analysis of various data such as weather, vegetation and terrain and prediction by using a machine learning model, efficient estimation of the grassland yield is realized, and the method is suitable for popularization and application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer technology and relates to a grass production remote sensing estimation system and method based on a machine learning algorithm. Background Art

[0002] Grassland biomass (grass production) can characterize the soil fertility of grassland, livestock carrying capacity and maintain the stability of the carbon cycle in the region, and occupies an important position in the ecosystem. Accurate inversion of grassland forage production has always been the research focus of terrestrial ecology. Grassland ecosystems play a very important role in maintaining regional economic development, ensuring plateau water conservation, maintaining biodiversity, and maintaining carbon fixation. In view of the important role of grassland in ecological environment and global climate change, as well as the fragile characteristics of the ecological environment in some regions, accurate monitoring of vegetation growth is of great significance. Constructing an accurate grassland forage production estimation model is of great scientific significance for grassland management, grass-livestock balance, grassland growth status assessment and ecological environment protection.

[0003] The grassland biomass estimation model currently used in business is a parameter regression model based on a single vegetation index for different grassland types. However, grassland biomass is affected by climate factors such as light, temperature and precipitation, soil factors such as soil nutrients, soil structure and fertility, biological factors such as grassland type, species richness and distribution of poisonous weeds, and management factors such as grazing, fencing and rotational grazing. If only a single vegetation index is used to build a model, its estimation accuracy is often limited. In view of the above problems, it is urgent to build a high-precision grassland biomass estimation model. Summary of the invention

[0004] In view of the problem of insufficient accuracy in grassland forage yield estimation in the prior art, the purpose of the present invention is to provide a grassland forage yield remote sensing estimation system and method based on machine learning algorithms. The system comprehensively considers multi-source data such as vegetation index, meteorological factors, terrain factors, etc., and constructs a grassland forage yield estimation model through a variety of machine learning algorithms, which significantly improves the estimation accuracy.

[0005] The technical solution is as follows:

[0006] A grassland grass yield remote sensing estimation system based on machine learning algorithm, comprising a data acquisition module, a data preprocessing module, a feature selection module, a machine learning algorithm module, a model training and verification module and a grassland grass yield estimation module, wherein the data acquisition module, the data preprocessing module, the feature selection module, the machine learning algorithm module, the model training and verification module and the grassland grass yield estimation module are connected in sequence.

[0007] The data acquisition module is responsible for collecting remote sensing data and related ground measured data, and outputs: raw data. It provides basic data input for the entire system and is connected to the data preprocessing module.

[0008] The data preprocessing module is used to perform preprocessing operations such as interpolation and normalization on the raw data provided by the data acquisition module. Output: standardized data that can be used for feature selection and modeling. The preprocessed data is passed to the feature selection module.

[0009] The feature selection module is used to extract key features (such as vegetation index, meteorological factors, etc.) that are highly correlated with grassland production from the preprocessed data. Output: feature variable set. The extracted feature variables are input into the machine learning algorithm module.

[0010] The machine learning algorithm module selects and constructs a machine learning model (such as random forest, support vector machine, etc.) suitable for estimating grass yield based on the features provided by the feature selection module. Output: machine learning model. The constructed model is provided to the model training and verification module.

[0011] The model training and validation module uses the training data to train the machine learning model and uses the validation data to evaluate the model's performance. Output: The optimized machine learning model. The validated model is provided to the grassland yield estimation module.

[0012] The grassland yield estimation module uses the optimized model to estimate the grassland yield in the target area. Output: Grass yield estimation results. The final output provides data for decision support.

[0013] Specifically, the data acquisition module is used to obtain data from remote sensing satellite data sources, meteorological observation stations and terrain data platforms. The data are processed and used as input features of the model to construct a grassland production estimation model.

[0014] The data preprocessing module is used to preprocess multi-source data to ensure the quality and consistency of the data.

[0015] The feature selection module is used to eliminate the influence of multicollinearity between different variables on the estimation accuracy of the model, and to reduce the complexity of the model and avoid overfitting problems through feature selection.

[0016] The machine learning algorithm module applies a variety of machine learning algorithms, each algorithm is used to build a grassland production estimation model, and compare the performance of different models to screen out the optimal model.

[0017] The model training and verification module uses historical data to train the machine learning model and evaluates the model by dividing the data into a training set and a verification set. Model training includes data input, algorithm selection, and parameter optimization. After training, the model will output the prediction results for the verification set and use R 2 The accuracy of the prediction results is evaluated by using indicators such as value and root mean square error to select the optimal model.

[0018] The grassland yield estimation module is used to estimate the grassland yield using new remote sensing data and meteorological data after the optimal model training is completed. The system can estimate the grassland yield in the area and generate a corresponding grassland yield spatial distribution map based on the input characteristics of the target area provided by the user.

[0019] Furthermore, the data acquired by the data acquisition module include Normalized Difference Vegetation Index (NDVI), Enhanced Vegetation Index (EVI), temperature, precipitation, sunshine hours and terrain factors.

[0020] Furthermore, the data preprocessing in the data preprocessing module includes data missing value filling, outlier detection and normalization processing. For vegetation index data, the maximum value synthesis method is also used to reduce the interference of clouds on remote sensing data.

[0021] Furthermore, the feature selection module uses the Shapley Additive Explanations (SHAP) method to calculate the relative importance of each input factor, and then screens out the target independent variables that have a significant impact on the grass yield.

[0022] Furthermore, the machine learning algorithm module applies multiple machine learning algorithms, including support vector machine (SVM), random forest (RF) and multivariate linear regression.

[0023] The method for estimating grass yield by remote sensing based on a machine learning algorithm of the present invention comprises the following steps:

[0024] Step 1: Obtain meteorological data, MODIS vegetation index data (NDVI, EVI), and digital elevation model (DEM) data in the observation area and perform data preprocessing;

[0025] Step 2: interpolate the acquired meteorological data and crop it to the specified range according to the provided shp file;

[0026] Step 3: Synthesize the maximum value of vegetation index data (NDVI, EVI) and use shp file to clip the data;

[0027] Step 4: Use DEM data to obtain the terrain features of the area and clip it to the boundaries defined by the shp file;

[0028] Step 5: Based on the acquired meteorological data, vegetation index and DEM data, machine learning models such as random forest, SVM and decision tree are trained, and the features with greater impact on grassland yield are screened out by combining the SHAP method;

[0029] Step 6: Using the selected features as input, the trained model is used to predict the grassland yield in the target year;

[0030] Step 7: Generate a spatial distribution tif file of the prediction results, and perform statistics on grassland yields in different areas, and calculate the average and total values ​​within the area;

[0031] Step 8: Calculate the deviation percentage between the predicted grassland yield and the yield of the previous year, the previous five years and the previous ten years, and generate the corresponding deviation percentage change tif file.

[0032] After the data resampling is completed, preferably, the preprocessed data is clipped and processed, including the following steps:

[0033] Interpolation and clipping of meteorological data:

[0034] First, meteorological observation data (such as temperature, precipitation, sunshine, etc.) are read and interpolated. The griddata function in the scipy library is used to interpolate irregular observation point data into regular grid data to form a meteorological data grid with a specified resolution (such as 250 meters). Then, by using the geographic boundary information in the shapefile and combining the crop_to_shape function, the interpolated meteorological data is spatially cropped to ensure that the cropped data is consistent with the geographic boundary of the study area. The cropping process defines the spatial range of the data according to the boundaries in the shapefile file (such as administrative division boundaries), and finally generates a meteorological data grid consistent with the shapefile boundaries.

[0035] Processing and maximum value synthesis of MODIS vegetation index (NDVI and EVI):

[0036] Obtain the vegetation index (NDVI and EVI) data provided by MODIS and read it as raster data through the gdal.Open function. By using the maximum function in numpy, the NDVI and EVI data are synthesized at the maximum value to generate a new raster data representing the maximum value of these two types of vegetation indices. This step ensures that the comprehensive effect of the vegetation index is maximized to better reflect the vegetation growth. After that, using the boundary information in the shapefile file, the maximum value synthesized raster data is spatially clipped through the clipping function to retain only the vegetation index data in the target area. This process ensures that the vegetation data coincides with the boundary of the observation area, which facilitates subsequent analysis and processing.

[0037] Machine learning model training and feature screening:

[0038] A machine learning model is constructed using processed meteorological data, vegetation index data, and DEM data. Specifically, classic algorithms such as random forest, support vector machine (SVM), and decision tree are used for model training. The input data of the model includes various environmental factors (temperature, precipitation, sunshine, terrain, etc.), and the output is grassland yield data. In order to further improve the prediction accuracy of the model, the SHAP (Shapley Additive Explanations) method is introduced in the code. The contribution of each feature in the model is calculated by the SHAP value, and the features are screened. Through this method, several key features that have the greatest impact on grassland yield are screened to ensure the interpretability and usability of the model.

[0039] Grassland production forecast:

[0040] The selected important features are used as input, and the trained optimal model (such as random forest or SVM) is used to predict the grassland yield in the target year. The input data of the model are the meteorological data and other environmental factors for the next year, and the output is the predicted value of grassland yield. Through the prediction results, a spatially distributed tif file is generated. The tif file represents the spatial variation of grassland yield in the entire region. The data can be further used to analyze the spatial distribution characteristics of grassland productivity.

[0041] Zonal statistics calculations:

[0042] The zonal_stats function in the rasterstats library is used to perform statistical analysis on the prediction results in combination with the regional boundary information in the shapefile file. The specific steps include calculating the average and total grassland yield values ​​for each zone and saving these statistical results as Excel files. This step provides users with detailed yield information based on geographical regions, which facilitates more detailed decisions in subsequent grassland management.

[0043] Calculation of deviation percentage:

[0044] In order to analyze the interannual changes in grassland yield, the code implements the calculation of the deviation percentage. By comparing the predicted grassland yield with historical data (such as the previous year, the previous five years, and the previous ten years), the deviation percentage is calculated, which can quantify the change in grassland yield relative to the historical level. Each deviation percentage result will generate a corresponding tif file, which can intuitively understand the temporal and spatial changes and growth trends of grassland yield.

[0045] Beneficial effects of the present invention:

[0046] The technical solution of the present invention comprehensively considers multiple source data such as vegetation index, meteorological factors, and terrain factors, and constructs a grassland yield estimation model through multiple machine learning algorithms, which significantly improves the estimation accuracy. At the same time, through comprehensive analysis of multiple data such as meteorology, vegetation, and terrain, and prediction using machine learning models, efficient estimation of grassland yield is achieved, which is suitable for promotion and application. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a schematic diagram of a grass production remote sensing estimation system based on a machine learning algorithm of the present invention;

[0048] Figure 2 This is a flow chart of the method for remote sensing estimation of grassland yield based on machine learning algorithm of the present invention. DETAILED DESCRIPTION

[0049] The technical solution of the present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0050] Reference Figure 1 The embodiment of the present invention provides a grass yield remote sensing estimation system based on a machine learning algorithm, including:

[0051] Data acquisition module: This module is responsible for acquiring data from remote sensing satellite data sources, meteorological observation stations and terrain data platforms, including Normalized Difference Vegetation Index (NDVI), Enhanced Vegetation Index (EVI), temperature, precipitation, sunshine hours and terrain factors. These data are processed and used as input features of the model to build a grassland yield estimation model.

[0052] Data preprocessing module: After data collection, multi-source data needs to be preprocessed to ensure data quality and consistency. Data preprocessing includes data missing value filling, outlier detection, and normalization. For vegetation index data, the maximum value synthesis method is also used to reduce the interference of clouds on remote sensing data.

[0053] Feature selection module: In order to eliminate the impact of multicollinearity between different variables on the accuracy of model estimation, the system uses the Shapley Additive Explanations (SHAP) method to calculate the relative importance of each input factor, and then screen out the target independent variables that have a significant impact on the grass yield. Through feature selection, the complexity of the model is reduced and the overfitting problem is avoided.

[0054] Machine learning algorithm module: This module applies a variety of machine learning algorithms, including support vector machine (SVM), random forest (RF) and multivariate linear regression. Each algorithm is used to build a grassland yield estimation model, and the performance of different models is compared to select the optimal model.

[0055] Model training and validation module: This module uses historical data to train the machine learning model and evaluates the model by dividing the data into training set and validation set. Model training includes data input, algorithm selection and parameter optimization. After training, the model will output the prediction results of the validation set. 2 The accuracy of the prediction results is evaluated by using indicators such as value and root mean square error to select the optimal model.

[0056] Grassland yield estimation module: After the optimal model training is completed, this module uses new remote sensing data and meteorological data to estimate grassland yield. The system can estimate the grassland yield in the area based on the input characteristics of the target area provided by the user and generate the corresponding grassland yield spatial distribution map.

[0057] like Figure 2 As shown, an embodiment of the present invention provides a method for estimating grass yield by remote sensing based on a machine learning algorithm, including:

[0058] Obtain meteorological data, MODIS vegetation index data (NDVI, EVI), and digital elevation model (DEM) data in the observation area and perform data preprocessing;

[0059] Interpolate the acquired meteorological data and crop it to the specified range according to the provided shp file;

[0060] Perform maximum synthesis of vegetation index data (NDVI, EVI) and use shp files to crop the data;

[0061] Use DEM data to obtain the terrain features of the area and clip it to the boundaries defined by the shp file;

[0062] Based on the acquired meteorological data, vegetation index and DEM data, machine learning models such as random forest, SVM and decision tree were trained, and the features with greater impact on grassland yield were screened out by combining the SHAP method;

[0063] The selected features are used as input and the trained model is used to predict the grassland yield in the target year;

[0064] Generate a tif file of the spatial distribution of the prediction results, and perform statistics on grassland yields in different areas, and calculate the average and total values ​​within the area;

[0065] The deviation percentage between the predicted grassland yield and the yield of the previous year, five years and ten years is calculated, and the corresponding deviation percentage change tif file is generated.

[0066] After the data resampling is completed, preferably, the preprocessed data is clipped and processed, including the following steps:

[0067] Interpolation and clipping of meteorological data:

[0068] First, meteorological observation data (such as temperature, precipitation, sunshine, etc.) are read and interpolated. The griddata function in the scipy library is used to interpolate irregular observation point data into regular grid data to form a meteorological data grid with a specified resolution (such as 250 meters). Then, by using the geographic boundary information in the shapefile and combining the crop_to_shape function, the interpolated meteorological data is spatially cropped to ensure that the cropped data is consistent with the geographic boundary of the study area. The cropping process defines the spatial range of the data according to the boundaries in the shapefile file (such as administrative division boundaries), and finally generates a meteorological data grid consistent with the shapefile boundaries.

[0069] Processing and maximum value synthesis of MODIS vegetation index (NDVI and EVI):

[0070] Obtain the vegetation index (NDVI and EVI) data provided by MODIS and read it as raster data through the gdal.Open function. By using the maximum function in numpy, the NDVI and EVI data are synthesized at the maximum value to generate a new raster data representing the maximum value of these two types of vegetation indices. This step ensures that the comprehensive effect of the vegetation index is maximized to better reflect the vegetation growth. After that, using the boundary information in the shapefile file, the maximum value synthesized raster data is spatially clipped through the clipping function to retain only the vegetation index data in the target area. This process ensures that the vegetation data coincides with the boundary of the observation area, which facilitates subsequent analysis and processing.

[0071] Machine learning model training and feature screening:

[0072] A machine learning model is constructed using processed meteorological data, vegetation index data, and DEM data. Specifically, classic algorithms such as random forest, support vector machine (SVM), and decision tree are used for model training. The input data of the model includes various environmental factors (temperature, precipitation, sunshine, terrain, etc.), and the output is grassland yield data. In order to further improve the prediction accuracy of the model, the SHAP (Shapley Additive Explanations) method is introduced in the code. The contribution of each feature in the model is calculated by the SHAP value, and the features are screened. Through this method, several key features that have the greatest impact on grassland yield are screened to ensure the interpretability and usability of the model.

[0073] Grassland production forecast:

[0074] The selected important features are used as input, and the trained optimal model (such as random forest or SVM) is used to predict the grassland yield in the target year. The input data of the model are the meteorological data and other environmental factors for the next year, and the output is the predicted value of grassland yield. Through the prediction results, a spatially distributed tif file is generated. The tif file represents the spatial variation of grassland yield in the entire region. The data can be further used to analyze the spatial distribution characteristics of grassland productivity.

[0075] Zonal statistics calculations:

[0076] The zonal_stats function in the rasterstats library is used to perform statistical analysis on the prediction results in combination with the regional boundary information in the shapefile file. The specific steps include calculating the average and total grassland yield values ​​for each zone and saving these statistical results as Excel files. This step provides users with detailed yield information based on geographical regions, which facilitates more detailed decisions in subsequent grassland management.

[0077] Calculation of deviation percentage:

[0078] In order to analyze the interannual changes in grassland yield, the code implements the calculation of the deviation percentage. By comparing the predicted grassland yield with historical data (such as the previous year, the previous five years, and the previous ten years), the deviation percentage is calculated, which can quantify the change in grassland yield relative to the historical level. Each deviation percentage result will generate a corresponding tif file, which can intuitively understand the temporal and spatial changes and growth trends of grassland yield.

[0079] Example

[0080] In order to verify the effectiveness of the system of the present invention, we took the Qilian Mountains and Sanjiangyuan area of ​​Qinghai Province as the research objects, used 20 years of measured grassland production data and a variety of remote sensing and meteorological data, and constructed a variety of machine learning models for comparison. The experimental results show that the random forest model has the highest accuracy. In the Qilian Mountains area, its R 2 The value is 0.68, and the RMSE is 91.93 g / ㎡, which is better than the multiple linear regression model (R 2 =0.58, RMSE = 104.79 g / ㎡), decision tree model (R 2 =0.49, RMSE = 115.99 g / ㎡) and support vector machine model (R 2 =0.48, RMSE = 117.34 g / ㎡). In the Sanjiangyuan area, the random forest model also performed well, with R 2 The value is 0.75 and the RMSE is 92.55 g / ㎡, which is significantly better than other models.

[0081] The above is only a preferred specific implementation manner of the present invention, and the protection scope of the present invention is not limited thereto. Any simple change or equivalent replacement of the technical solution that can be obviously obtained by any technician familiar with the technical field within the technical scope disclosed in the present invention falls within the protection scope of the present invention.

Claims

1. A grass yield remote sensing estimation system based on machine learning algorithm, characterized by: It includes a data acquisition module, a data preprocessing module, a feature selection module, a machine learning algorithm module, a model training and verification module and a grassland grass yield estimation module, wherein the data acquisition module, the data preprocessing module, the feature selection module, the machine learning algorithm module, the model training and verification module and the grassland grass yield estimation module are connected in sequence; The data acquisition module is used to obtain data from remote sensing satellite data sources, meteorological observation stations and terrain data platforms. After being processed, the data are used as input features of the model to construct a grassland production estimation model; The data preprocessing module is used to preprocess multi-source data; The feature selection module is used to eliminate the influence of multicollinearity between different variables on the model estimation accuracy; The machine learning algorithm module applies a variety of machine learning algorithms, each algorithm is used to build a grass yield estimation model, and compare the performance of different models to select the optimal model; The model training and verification module uses historical data to train the machine learning model and evaluates the model by dividing the data into a training set and a verification set. The model training includes data input, algorithm selection and parameter optimization. After the training is completed, the model will output the prediction results for the verification set. 2 The accuracy of the prediction results is evaluated by the value and root mean square error indicators, so as to select the optimal model; The grassland yield estimation module is used to estimate the grassland yield using new remote sensing data and meteorological data after the optimal model training is completed. The system estimates the grassland yield in the area and generates a corresponding grassland yield spatial distribution map based on the input characteristics of the target area provided by the user.

2. The grass yield remote sensing estimation system based on machine learning algorithm according to claim 1 is characterized by: The data acquired by the data acquisition module include normalized difference vegetation index NDVI, enhanced vegetation index EVI, temperature, precipitation, sunshine hours and terrain factors.

3. The grass yield remote sensing estimation system based on machine learning algorithm according to claim 1 is characterized by: The data preprocessing in the data preprocessing module includes data missing value filling, outlier detection and normalization processing.

4. The grass yield remote sensing estimation system based on machine learning algorithm according to claim 1 is characterized by: The feature selection module uses the Shapley Additive Explanations method to calculate the relative importance of each input factor, and then screens out the target independent variables that have a significant impact on the grass yield.

5. The grass yield remote sensing estimation system based on machine learning algorithm according to claim 1 is characterized by: The machine learning algorithm module applies a variety of machine learning algorithms, including support vector machine SVM, random forest RF and multivariate linear regression.

6. A method for estimating grass yield using remote sensing based on a machine learning algorithm, characterized in that: The following steps are involved: Step 1: Obtain meteorological data, MODIS vegetation index data NDVI, EVI, and digital elevation model DEM data in the observation area and perform data preprocessing; Step 2: interpolate the acquired meteorological data and crop it to the specified range according to the provided shp file; Step 3: Synthesize the maximum values ​​of vegetation index data NDVI and EVI, and use shp file to clip the data; Step 4: Use DEM data to obtain the terrain features of the area and clip it to the boundaries defined by the shp file; Step 5: Based on the acquired meteorological data, vegetation index and DEM data, the random forest, SVM and decision tree machine learning models were trained, and the features with greater impact on grassland yield were screened out by combining the SHAP method; Step 6: Using the selected features as input, the trained model is used to predict the grassland yield in the target year; Step 7: Generate a spatial distribution tif file of the prediction results, and perform statistics on grassland yields in different areas, and calculate the average and total values ​​within the area; Step 8: Calculate the deviation percentage between the predicted grassland yield and the yield of the previous year, the previous five years and the previous ten years, and generate the corresponding deviation percentage change tif file.

7. The method for estimating grass yield by remote sensing based on machine learning algorithm according to claim 6, characterized in that: After the raw data collection in step 1 is completed, the data preprocessed in step 1 is trimmed and processed, including the following steps: Step 1) Interpolation and clipping of meteorological data: First, the meteorological observation data is read and interpolated. The griddata function in the scipy library is used to interpolate irregular observation point data into regular grid data to form a meteorological data grid with a specified resolution. Then, the interpolated meteorological data is spatially cropped using the geographic boundary information in the shapefile and the crop_to_shape function to ensure that the cropped data is consistent with the geographic boundary of the study area. The cropping process defines the spatial range of the data according to the boundary in the shapefile, and finally generates a meteorological data grid consistent with the shapefile boundary. Step 2), processing and maximum value synthesis of MODIS vegetation index NDVI and EVI: Obtain the vegetation index NDVI and EVI data provided by MODIS, and read them as raster data through the gdal.Open function; use the maximum function in numpy to synthesize the maximum value of the NDVI and EVI data to generate a new raster data representing the maximum value of these two types of vegetation indices; this step ensures that the comprehensive effect of the vegetation index is maximized to better reflect the vegetation growth; then, use the boundary information in the shapefile file to spatially clip the raster data after the maximum value synthesis through the clipping function to retain only the vegetation index data in the target area; this process ensures that the vegetation data coincides with the boundary of the observation area, which is convenient for subsequent analysis and processing; Step 3) Training and feature screening of machine learning models: A machine learning model is constructed using processed meteorological data, vegetation index data, and DEM data. Specifically, the model training uses random forest, support vector machine (SVM), and decision tree algorithms. The model input data includes various environmental factors, and the output is grassland yield data. The SHAP method is introduced in the code to calculate the contribution of each feature in the model through the SHAP value, and the features are screened. Through this method, several key features that have the greatest impact on grassland yield are screened to ensure the interpretability and usability of the model. Step 4) Grassland yield prediction: The selected important features are used as input, and the trained optimal model is used to predict the grassland yield in the target year. The input data of the model are the meteorological data and other environmental factors of the next year, and the output is the predicted value of grassland yield. Through the prediction results, a spatial distribution tif file is generated, which represents the spatial variation of grassland yield in the entire region. The data is further used to analyze the spatial distribution characteristics of grassland productivity. Step 5), regional statistical calculation: Use the zonal_stats function in the rasterstats library and the regional boundary information in the shapefile file to perform statistical analysis on the prediction results; Step 6) Calculation of deviation percentage: By comparing the predicted grassland production with historical data, a percentage deviation is calculated, which quantifies the magnitude of the change in grassland production relative to historical levels.

8. The method for estimating grass yield by remote sensing based on machine learning algorithm according to claim 7, characterized in that: In step 1), the meteorological observation data includes temperature, precipitation, and sunshine.

9. The method for estimating grass yield by remote sensing based on machine learning algorithm according to claim 7, characterized in that: In step 3), the environmental factors include temperature, precipitation, sunshine, and topography.

10. The method for estimating grass yield by remote sensing based on machine learning algorithm according to claim 7, characterized in that: In step 5), the specific steps include calculating the average and total values ​​of grassland yield in each area, and saving these statistical results as an Excel file.