Thyme biomass prediction method based on multispectral data and machine learning

Through the combination of UAV multispectral remote sensing technology and multiple machine learning models, a high-precision thyme biomass prediction model was established, solving the problems of low biomass estimation accuracy and low efficiency in traditional methods, and achieving accurate and efficient prediction of thyme biomass.

CN120182869AActive Publication Date: 2025-06-20INNER MONGOLIA UNIVERSITY

Patent Information

Application Number
CN202510266430.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-20
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

The prior art is difficult to accurately and efficiently estimate thyme biomass. The traditional methods have problems such as manual sampling, high destructiveness and low accuracy, and the existing vegetation index methods have failed to make full use of the information of multispectral data.

Method used

The low-altitude multispectral image data of the drone is used and combined with a variety of machine learning models. Through preprocessing, single-plant scale segmentation, vegetation index calculation and machine learning modeling, a high-precision thyme biomass prediction model is established.

Benefits of technology

Real-time and accurate prediction of large-area thyme biomass is achieved, prediction accuracy and stability are improved, the limitations of a single model are overcome, and a basis for grassland ecosystem management and resource evaluation are provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182869A_ABST
    Figure CN120182869A_ABST
Patent Text Reader

Abstract

The invention relates to a thyme biomass prediction method based on multispectral data and machine learning. The thyme biomass prediction method comprises the steps that multispectral image data of a research area is acquired and preprocessed; single plant scale segmentation is carried out, and a target extraction result of the single plant scale is acquired; calculating a vegetation index; and modeling the vegetation index by adopting a machine learning model to obtain a biomass prediction model, training the biomass prediction model, and automatically predicting the thyme biomass to obtain two prediction targets of dry weight and fresh weight. According to the method, the low-altitude multispectral image data of the unmanned aerial vehicle is combined with various machine learning models, a high-precision biomass prediction model is established, the spatial distribution characteristics of large-area thyme biomass are accurately obtained in real time, and a basis is provided for grassland ecological system management and resource evaluation; through prediction of 11 machine learning models, limitation of a single model is overcome, and stability and reliability of a prediction result are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of ecological monitoring, and in particular to a method for predicting the biomass of thyme based on multispectral data and machine learning. Background Art

[0002] Thyme is an aromatic plant of the genus Thymus in the Lamiaceae family, which is widely distributed in northern China and has important ecological and economic values. In recent years, relevant researchers have found that thyme essential oil contains various alkenes and terpenoid compounds, which have a positive effect on the treatment of diseases such as cancer and obesity. Therefore, thyme has gradually been developed and applied in medical and other fields. In addition, thyme has extremely strong environmental adaptability and can survive in a variety of habitats. Its developed adventitious roots and strong ductility enable it to play an important role in soil and water conservation and ecological restoration management. Therefore, the accurate estimation of thyme biomass provides an important reference for the evaluation of its medicinal value and the sustainable utilization of resources.

[0003] Traditional biomass measurement methods mainly rely on manual field sampling and weighing. This method has many disadvantages. For example, manual sampling is time-consuming and laborious, with low efficiency, and is destructive; the sampling points are limited, and it is difficult to obtain the spatial distribution information of a large area. In addition, manual measurement is easily affected by subjective factors, and the consistency and repeatability of the measurement results are poor. With the development of remote sensing technology, methods for estimating vegetation biomass based on remote sensing data have gradually been applied.

[0004] At present, scholars at home and abroad have carried out a large number of studies on remote sensing estimation of vegetation biomass. The research shows that the combination of unmanned aerial vehicle (UAV) remote sensing and machine learning provides a reliable technical path for vegetation biomass estimation. However, due to the small size of thyme individuals, the spatial resolution of traditional satellite remote sensing data is difficult to meet the fine monitoring requirements. In addition, most of the existing vegetation index methods are based on simple band combinations and fail to make full use of the rich information contained in multispectral data; in terms of biomass estimation models, traditional statistical regression methods rely too much on a single empirical model and are difficult to accurately describe the complex non-linear relationship between vegetation indices and biomass.

[0005] Therefore, how to use low-cost and high-efficiency UAV multispectral remote sensing technology, combined with advanced machine learning algorithms, to establish an accurate and reliable thyme biomass prediction model is a technical problem that urgently needs to be solved at present. Summary of the Invention

[0006] To solve the problems of low accuracy and efficiency in the estimation of thyme biomass in the existing technology, and the serious destructiveness of manual sampling, the purpose of the present invention is to provide a thyme biomass prediction method based on multi-spectral data and machine learning, which uses low-altitude multi-spectral image data of unmanned aerial vehicles combined with a variety of machine learning models to accurately obtain the spatial distribution characteristics of large-area thyme biomass in real time, improve the prediction accuracy, and enhance the stability and reliability of the prediction results.

[0007] To achieve the above purpose, the present invention adopts the following technical solutions: A thyme biomass prediction method based on multi-spectral data and machine learning, the method includes the following steps in sequence:

[0008] (1) Obtain multi-spectral image data of the research area, including red light band, near-infrared band and green light band;

[0009] (2) Preprocess the obtained multi-spectral images;

[0010] (3) Perform single-plant scale segmentation on the preprocessed images, and use a method combining mask and vector file to extract the target area, and obtain the target extraction result at the single-plant scale;

[0011] (4) Calculate the vegetation index according to the target extraction result at the single-plant scale;

[0012] (5) Use a machine learning model to model the vegetation index, obtain a biomass prediction model and train it, input the multi-spectral image to be predicted into the trained biomass prediction model, and automatically predict the thyme biomass to obtain two prediction targets of dry weight and fresh weight.

[0013] Step (1) specifically refers to: Using an unmanned aerial vehicle equipped with a multi-spectral camera, at a flight altitude of 15m, obtain the multi-spectral image data of the research area according to the parameters of a heading overlap rate of 80% and a side overlap rate of 70%.

[0014] Step (2) includes the following steps in sequence:

[0015] (2a) Perform radiometric calibration and atmospheric correction on the obtained multi-spectral images;

[0016] (2b) Set the NDVI threshold range to 0 to 1, and automatically segment the soil background by adjusting the threshold;

[0017] (2c) Generate a binary mask image based on the set NDVI threshold, mark the area where the NDVI value is greater than the NDVI threshold as 1, that is, the vegetation area, and the area less than the threshold as 0, that is, the background area;

[0018] (2d) Output the processing result after removing the background area, that is, the preprocessed image data.

[0019] Step (3) includes the following steps in sequence:

[0020] (3a) Import the preprocessed image data;

[0021] (3b) Import the vector boundary file containing the boundary information of the target area;

[0022] (3c) Extract the vegetation area;

[0023] (3d) Crop the extracted vegetation area according to the vector boundary file to obtain the target extraction result at the single-plant scale.

[0024] In step (4), the formulas for calculating the vegetation indices include:

[0025] NDVI = (NIR - Red) / (NIR + Red);

[0026] GNDVI = (NIR - Green) / (NIR + Green);

[0027] RVI = NIR / Red;

[0028] SAVI = ((NIR - Red) × 1.5) / (NIR + Red + 0.5);

[0029] OSAVI = ((NIR - Red) × 1.16) / (NIR + Red + 0.16));

[0030] In the formulas, NDVI is the normalized difference vegetation index, GNDVI is the green-band normalized difference vegetation index, SAVI is the soil-adjusted vegetation index, OSAVI is the optimized soil-adjusted vegetation index; NIR represents the reflectance value of the near-infrared band, Red represents the reflectance value of the red band, Green represents the reflectance value of the green band, and RVI is the ratio vegetation index.

[0031] Step (5) specifically refers to: standardizing all vegetation index features using a standard scaler; constructing a biomass prediction model including the following 11 machine learning models, the 11 machine learning models including lasso regression, ridge regression, improved partial least squares regression, linear regression, improved KNN regression, improved decision tree regression, improved neural network regression, improved random forest regression, improved gradient boosting regression, improved support vector regression, improved XGBoost regression; evaluating the performance of each machine learning model using the cross-validation method; recording the prediction results and evaluation metrics of each machine learning model.

[0032] The improved random forest regression refers to hyperparameter optimization through randomized search cross-validation: the number of trees is 50, the maximum depth is 5, the minimum number of samples for splitting is 2, and the minimum number of samples for leaf nodes is 1; the improved gradient boosting regression refers to hyperparameter optimization through RandomizedSearchCV: the number of trees is 50, the learning rate is 0.01, the maximum depth is 3, the minimum number of samples for splitting is 2, and the minimum number of samples for leaf nodes is 2.

[0033] The improved partial least squares regression means that the number of principal components is fixed at 2; the improved neural network regression refers to hyperparameter optimization through grid search cross-validation: the hidden layer structure is (50, 50), the activation function is set to relu, the solver is set to adam, and the learning rate is 0.001.

[0034] The improved support vector regression refers to hyperparameter optimization through grid search cross-validation: the kernel function type is set to linear, the regularization parameter is 1, and the gamma parameter is set to scale; the improved XGBoost regression means that the objective function is set to mean squared error.

[0035] The improved KNN regression means: hyperparameter optimization through grid search cross-validation: the number of neighbors is 5, and the weight scheme is set to distance; the improved decision tree regression refers to hyperparameter optimization through grid search cross-validation: the maximum depth is 10, the minimum number of samples for splitting is 5, and the minimum number of samples for leaf nodes is 2.

[0036] As can be seen from the above technical solutions, the beneficial effects of the present invention are as follows: First, compared with the prior art, the present invention uses low-altitude multi-spectral image data of unmanned aerial vehicles in combination with a variety of machine learning models to establish a high-precision biomass prediction model, and obtains the spatial distribution characteristics of large-area thyme biomass in real time and accurately, providing a basis for grassland ecosystem management and resource assessment; Second, the present invention applies multiple vegetation indices (NDVI, GNDVI, RVI, SAVI, OSAVI) to thyme biomass prediction, and through standardization processing and range constraint mechanism, highlights the key factors for biomass estimation, improves the prediction accuracy, and overcomes the limitations of a single model through the prediction of 11 machine learning models, enhancing the stability and reliability of the prediction results. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is the flowchart of the method of the present invention;

[0038] Figure 2 is the schematic diagram of the preprocessing interface;

[0039] Figure 3 is the schematic diagram of the single-plant segmentation interface;

[0040] Figure 4 It is a schematic diagram of the biomass prediction interface. Specific implementation manners

[0041] As Figure 1 shown, a thyme biomass prediction method based on multispectral data and machine learning, the method includes the following steps in sequence:

[0042] (1) Obtain multispectral image data of the research area, including red light band, near-infrared band and green light band;

[0043] (2) Preprocess the obtained multispectral images;

[0044] (3) Segment the preprocessed images at the single-plant scale, and use the method combining mask and vector file to extract the target area, and obtain the target extraction result at the single-plant scale;

[0045] (4) Calculate the vegetation index according to the target extraction result at the single-plant scale, the calculated vegetation indices form a data set, and divide the data set into a training set and a test set;

[0046] (5) Use a machine learning model to model the vegetation index, obtain a biomass prediction model and train it through the training set, input the multispectral image to be predicted into the trained biomass prediction model, and automatically predict the thyme biomass to obtain two prediction targets of dry weight and fresh weight.

[0047] Step (1) specifically refers to: using a drone equipped with a multispectral camera, at a flight altitude of 15m, and obtaining multispectral image data of the research area according to the parameters of a heading overlap rate of 80% and a side overlap rate of 70%.

[0048] Step (2) includes the following steps in sequence:

[0049] (2a) Perform radiometric calibration and atmospheric correction on the obtained multispectral images;

[0050] (2b) Set the NDVI threshold range to 0 to 1, and automatically segment the soil background by adjusting the threshold;

[0051] (2c) Generate a binary mask image based on the set NDVI threshold, mark the area where the NDVI value is greater than the NDVI threshold as 1, that is, the vegetation area, and mark the area less than the threshold as 0, that is, the background area;

[0052] (2d) Output the processing result after removing the background area, that is, the preprocessed image data.

[0053] Step (3) includes the following steps in sequence:

[0054] (3a) Import the preprocessed image data;

[0055] (3b) Import a vector boundary file containing the boundary information of the target area;

[0056] (3c) Perform an overlay operation on the preprocessed image data and the binary mask image to extract the vegetation area;

[0057] (3d) Crop the extracted vegetation area according to the vector boundary file to obtain the target extraction result at the single-plant scale.

[0058] In step (4), the formulas for calculating the vegetation indices include:

[0059] NDVI = (NIR - Red) / (NIR + Red);

[0060] GNDVI = (NIR - Green) / (NIR + Green);

[0061] RVI = NIR / Red;

[0062] SAVI = ((NIR - Red) × 1.5) / (NIR + Red + 0.5);

[0063] OSAVI = ((NIR - Red) × 1.16) / (NIR + Red + 0.16));

[0064] In the formulas, NDVI is the normalized difference vegetation index, GNDVI is the green-band normalized difference vegetation index, SAVI is the soil-adjusted vegetation index, and OSAVI is the optimized soil-adjusted vegetation index; NIR represents the reflectance value of the near-infrared band, Red represents the reflectance value of the red band, Green represents the reflectance value of the green band, and RVI is the ratio vegetation index.

[0065] Step (5) specifically refers to: performing standardization processing on all vegetation index features using a standard scaler; constructing a biomass prediction model including the following 11 machine learning models, and the 11 machine learning models include lasso regression, ridge regression, improved partial least squares regression, linear regression, improved KNN regression, improved decision tree regression, improved neural network regression, improved random forest regression, improved gradient boosting regression, improved support vector regression, and improved XGBoost regression; using the cross-validation method to evaluate the performance of each machine learning model; and recording the prediction results and evaluation metrics of each machine learning model.

[0066] Lasso regression is used to handle the situation of multicollinearity among features; Ridge regression is used to prevent overfitting through L2 regularization; Partial least squares regression is applicable when the number of features is more than the number of samples; Linear regression is used to establish a linear relationship between features and the target variable; KNN regression makes predictions based on neighboring samples; Decision tree regression performs feature partitioning and prediction through a tree structure; Neural network regression uses a multi-layer perceptron for non-linear modeling; Random forest regression integrates the prediction results of multiple decision trees; Gradient boosting regression improves the prediction performance through iteration; Support vector regression uses a kernel function to handle non-linear problems; XGBoost regression adopts a boosting algorithm optimized by second-order derivatives.

[0067] The improved random forest regression refers to hyperparameter optimization through RandomizedSearchCV: the number of trees is 50, the maximum depth is 5, the minimum number of samples for splitting is 2, and the minimum number of samples for leaf nodes is 1; The improved gradient boosting regression refers to hyperparameter optimization through RandomizedSearchCV: the number of trees is 50, the learning rate is 0.01, the maximum depth is 3, the minimum number of samples for splitting is 2, and the minimum number of samples for leaf nodes is 2.

[0068] The improved partial least squares regression means that the number of principal components is fixed at 2; The improved neural network regression refers to hyperparameter optimization through GridSearchCV: the hidden layer structure is (50, 50), the activation function is set to relu, the solver is set to adam, and the learning rate is 0.001.

[0069] The improved support vector regression refers to hyperparameter optimization through GridSearchCV: the kernel function type is set to linear, the regularization parameter is 1, and the gamma parameter is set to scale; The improved XGBoost regression means that the objective function is set to mean squared error.

[0070] The improved KNN regression means: hyperparameter optimization through GridSearchCV: the number of neighbors is 5, and the weight scheme is set to distance; The improved decision tree regression refers to hyperparameter optimization through GridSearchCV: the maximum depth is 10, the minimum number of samples for splitting is 5, and the minimum number of samples for leaf nodes is 2. Figure 2 It shows the soil segmentation interface based on NDVI, including the input parameter setting area, threshold adjustment area, result display area, and operation button area.

[0071] As Figure 3 shown, the single-plant segmentation interface provides an operation process of "image upload → mask upload → SHP import → single-plant extraction", and users can conveniently complete the extraction process from the original image to the single-plant target.

[0072] Using the NDVI segmentation result as a mask, the original multispectral image is masked and extracted to obtain the pure vegetation area image; import the pre-made vector boundary file (in shp format), ensure the use of the WGS1984 coordinate system, and crop based on the vector boundary to obtain the target image at the single-plant scale.

[0073] The dataset is divided according to the following rules: measure the biomass (fresh weight and dry weight) of each test sample as the label data, extract the features of each vegetation index at the corresponding position as the input features, and then randomly divide them into a training set and a test set in a ratio of 7:3.

[0074] Table 1 Comparison table of fresh weight prediction accuracies of each model

[0075] Model <![CDATA[Training set R 2 > <![CDATA[Test set R 2 > RMSE of training set RMSE of test set Linear regression 0,72 0,73 29,57 24,15 Ridge regression 0,72 0,72 29,66 24,50 Lasso regression 0,72 0,71 29,73 24,83 Random forest 0,85 0,67 21,95 26,72 Gradient boosting 0,96 0,48 10,90 33,44 Partial least squares 0,72 0,72 29,81 24,47 Neural network 0,73 0,69 29,12 25,75 Support vector 0,72 0,72 29,75 24,51 XGBoost 1,00 0,38 0,01 36,39 KNN 0,67 0,63 32,09 28,21 Decision tree 0,80 0,63 25,33 27,99

[0076] Table 2 Comparison table of dry weight prediction accuracies of each model

[0077]

[0078]

[0079] As shown in Table 1 and Table 2, the prediction accuracies of different models on the training set and test set of fresh weight and dry weight are listed.

[0080] As Figure 4 shown, the present invention provides an intuitive prediction result display interface, which includes the following contents:

[0081] The band data input area is used to select the red light, near-infrared, and green light band image files;

[0082] The vegetation index display area calculates and displays the values of various vegetation indices;

[0083] The prediction result display area displays the prediction results of each model.

[0084] The present invention uses multispectral remote sensing data combined with a variety of machine learning models to predict the biomass of thyme. By calculating a variety of vegetation indices (NDVI, GNDVI, RVI, SAVI, and OSAVI) as feature inputs, 11 machine learning models including linear regression, ridge regression, Lasso regression, random forest, gradient boosting, partial least squares regression, neural network regression, support vector regression, XGBoost regression, K-nearest neighbor regression, and decision tree regression are constructed. The present invention provides an intuitive prediction result display interface. Users only need to input the reflectance data of the red light, near-infrared, and green light bands to obtain the biomass prediction result. Through experimental verification, the present invention can accurately predict the biomass of thyme, providing important technical support for the evaluation and management of grassland ecosystems.

Claims

1. A thyme biomass prediction method based on multispectral data and machine learning, characterized in that: The method comprises the following steps in order: (1) Obtain multispectral image data of the study area, including red light band, near infrared band and green light band; (2) Preprocessing the acquired multispectral images; (3) Segment the preprocessed image at the scale of a single plant, extract the target area using a method combining mask and vector file, and obtain the target extraction result at the scale of a single plant; (4) Calculate vegetation index based on the target extraction results at the individual plant scale; (5) The vegetation index was modeled using a machine learning model to obtain a biomass prediction model and train it. The multispectral image to be predicted was input into the trained biomass prediction model to automatically predict the biomass of thyme and obtain the two prediction targets of dry weight and fresh weight.

2. The thyme biomass prediction method based on multispectral data and machine learning according to claim 1, characterized in that: Step (1) specifically refers to: using an unmanned aerial vehicle equipped with a multispectral camera to obtain multispectral image data of the study area at a flight altitude of 15 m according to the parameters of a heading overlap rate of 80% and a lateral overlap rate of 70%.

3. The thyme biomass prediction method based on multispectral data and machine learning according to claim 1, characterized in that: Step (2) comprises the following steps in order: (2a) Perform radiometric calibration and atmospheric correction on the acquired multispectral images; (2b) Setting the NDVI threshold range to 0 to 1, the soil background is automatically segmented by adjusting the threshold; (2c) Generate a binary mask image based on the set NDVI threshold, mark the area with NDVI value greater than the NDVI threshold as 1, i.e., vegetation area, and mark the area with NDVI value less than the threshold as 0, i.e., background area; (2d) Output the processing result after removing the background area, that is, the pre-processed image data.

4. The thyme biomass prediction method based on multispectral data and machine learning according to claim 1, characterized in that: Step (3) comprises the following steps in order: (3a) importing pre-processed image data; (3b) importing a vector boundary file containing target area boundary information; (3c) Extracting vegetation areas; (3d) The extracted vegetation area is cropped according to the vector boundary file to obtain the target extraction result at the single plant scale.

5. The thyme biomass prediction method based on multispectral data and machine learning according to claim 1, characterized in that: In step (4), the formula for calculating the vegetation index includes: NDVI = (NIR-Red) / (NIR+Red); GNDVI=(NIR-Green) / (NIR+Green); RVI=NIR / Red; SAVI=((NIR-Red)×1,5) / (NIR+Red+0,5); OSAVI=((NIR-Red)×1,16) / (NIR+Red+0,16)); In the formula, NDVI is the normalized difference vegetation index, GNDVI is the green band normalized difference vegetation index, SAVI is the soil adjusted vegetation index, OSAVI is the optimized soil adjusted vegetation index, NIR represents the reflectance value of the near infrared band, Red represents the reflectance value of the red light band, Green represents the reflectance value of the green light band, and RVI is the ratio vegetation index.

6. The thyme biomass prediction method based on multispectral data and machine learning according to claim 1, characterized in that: Step (5) specifically refers to: standardizing all vegetation index features using a standardized scaler; constructing a biomass prediction model comprising the following 11 machine learning models, wherein the 11 machine learning models include Lasso regression, ridge regression, improved partial least squares regression, linear regression, improved KNN regression, improved decision tree regression, improved neural network regression, improved random forest regression, improved gradient boosting regression, improved support vector regression, and improved XGBoost regression; using a cross-validation method to evaluate the performance of each machine learning model; and recording the prediction results and evaluation indicators of each machine learning model.

7. The thyme biomass prediction method based on multispectral data and machine learning according to claim 6, characterized in that: The improved random forest regression refers to hyperparameter optimization through randomized search cross validation: the number of trees is 50, the maximum depth is 5, the minimum number of split samples is 2, and the minimum number of leaf node samples is 1; the improved gradient boosting regression refers to hyperparameter optimization through RandomizedSearchCV: the number of trees is 50, the learning rate is 0.01, the maximum depth is 3, the minimum number of split samples is 2, and the minimum number of leaf node samples is 2.

8. The thyme biomass prediction method based on multispectral data and machine learning according to claim 6, characterized in that: The improved partial least squares regression refers to setting the number of principal components to 2; the improved neural network regression refers to performing hyperparameter optimization through grid search cross validation: the hidden layer structure is (50,50), the activation function is set to relu, the solver is set to adam, and the learning rate is 0.

001.

9. The thyme biomass prediction method based on multispectral data and machine learning according to claim 6, characterized in that: Improved support vector regression refers to hyperparameter optimization through grid search cross validation: the kernel function type is set to linear, the regularization parameter is 1, and the gamma parameter is set to scale; improved XGBoost regression refers to setting the objective function to mean square error.

10. The thyme biomass prediction method based on multispectral data and machine learning according to claim 6, characterized in that: Improved KNN regression refers to: hyperparameter optimization through grid search cross validation: the number of neighbors is 5, and the weight scheme is set to distance; improved decision tree regression refers to hyperparameter optimization through grid search cross validation: the maximum depth is 10, the minimum number of split samples is 5, and the minimum number of leaf node samples is 2.

Citation Information

Patent Citations

  • Cultured porphyra biomass measurement method based on spectral remote-sensing images

    CN110398465A

  • Multi-vegetation-index rice yield estimation method based on machine learning algorithm

    CN111241912A

  • Vegetation classification and biomass inversion method based on remote sensing data

    CN114445719A

  • Corn aboveground biomass estimation method and system

    CN116187478A

  • Wheat biomass monitoring method and system based on multi-modal data and phenological information

    CN119027811A

Cited By

  • Pennisetum alopecuroides cold resistance evaluation method based on unmanned aerial vehicle multispectral image

    CN120992519A