Tobacco starch content prediction method based on map fusion
Through drones, the hyperspectral and image data of tobacco fields is collected, combined with integrated learning algorithms, an efficient tobacco starch content prediction model is built, which solves the high cost and low accuracy problems of traditional detection methods, and realizes high-precision tobacco starch content prediction, supporting precise agricultural management.
Patent Information
- Application Number
- CN202510162824.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-06
AI Technical Summary
Traditional tobacco starch content detection methods are time-consuming, cost-effective and low-precision, making them difficult to meet the needs of large-scale field operations, and existing detection methods based on remote sensing data are also difficult to provide sufficient accuracy.
UAVs are used to collect hyperspectral and image data from tobacco fields, and combined with integrated learning algorithms, an efficient tobacco starch content prediction model is built, and the prediction accuracy is achieved through data preprocessing, feature extraction and machine learning and deep learning model training.
It improves the accuracy of tobacco starch content prediction, reduces production costs, enhances the robustness and prediction accuracy of the model, and provides accurate management support for tobacco production.
Smart Images

Figure CN120107731A_ABST
Abstract
Description
Technical Field
[0002] The present invention relates to precision agricultural management technology, specifically a tobacco starch content prediction method based on graph fusion, which is a tobacco starch content prediction method based on remote sensing technology to obtain graph data combined with an integrated learning algorithm. The method is suitable for precision agricultural management of tobacco planting, helping farmers to accurately predict tobacco starch content through remote sensing data, optimize production management, and improve yield and quality. Background Art
[0003] Starch is a key chemical component in tobacco, which directly affects important indicators such as its quality, burning characteristics and taste. Traditional methods for detecting tobacco starch content mainly rely on chemical analysis, which is not only time-consuming and costly, but also requires a high level of manual experience and is difficult to adapt to large-scale field operations. In recent years, with the development of drone remote sensing technology, it has become possible to obtain the growth status of field crops through hyperspectral imaging and image data, but how to accurately predict the starch content of tobacco is still a difficult problem in agricultural precision management.
[0004] Existing crop detection methods based on remote sensing data are usually limited to simple spectral data analysis or image feature extraction. For tobacco, a special crop, a single data source often cannot provide sufficient accuracy. Therefore, there is an urgent need for a tobacco starch content prediction method that combines multiple data sources (such as hyperspectral data and image data) and uses efficient machine learning and deep learning methods for integrated analysis to improve prediction accuracy and reduce costs. Summary of the invention
[0005] The present invention aims to provide a method for predicting tobacco starch content based on graph fusion, so as to solve the problems of high cost, long cycle and low precision of traditional chemical analysis methods. By integrating multi-source data and integrated learning methods, an efficient prediction model is constructed, which can accurately predict tobacco starch content and provide support for precision agricultural management.
[0006] The objective of the present invention is achieved through the following technical solutions: A method for predicting tobacco starch content based on graph fusion mainly comprises the following steps: Step 1: Use drones to collect hyperspectral and image data of tobacco canopies in tobacco fields, and obtain the starch content of upper leaves; the details are as follows: Step 1.1, the data collection period is based on the key growth period after the flue-cured tobacco topping, including: in Taining County, Fujian Province, data collection was carried out during the flue-cured tobacco topping period (111 days after transplanting), the lower leaf maturity period (125 days after transplanting), and the middle leaf maturity period (132 days after transplanting); in Jianyang District, Fujian Province, the data collection time is the topping period (107 days after transplanting), the lower leaf maturity period (126 days after transplanting), and the middle leaf maturity period (139 days after transplanting); in Wuyishan City, Fujian Province, the data collection time is the topping period (70 days after transplanting), the lower leaf maturity period (87 days after transplanting), and the middle leaf maturity period (100 days after transplanting).
[0007] Step 1.2, use drones to collect hyperspectral data and visible light data. The drone used to obtain hyperspectral and image data is DJI M350RTK, equipped with Gaiasky-mini3-VN hyperspectral camera, with a spectral range of 400~1000 nm, a spectral resolution of 5.5 nm, a hovering built-in push-broom, a pixel count of 1024*448, and 224 spectral channels. The waypoint overlap and route overlap are both 60%, and the flight altitude is 40 m. The visible light camera is Zenmuse P1, with a heading overlap and a lateral overlap of 80%, and a flight altitude of 20 m. Data acquisition was carried out between 10:00 and 14:00 in clear and cloudless weather.
[0008] Step 1.3, starch content determination After collecting the UAV hyperspectral image data, three representative tobacco plants were selected from each plot for destructive sampling to obtain fresh tobacco leaf samples from the upper leaves of flue-cured tobacco (the third leaf from the top leaf). The samples were frozen after destemming, embedded in dry ice and transported to the laboratory for freeze-drying. The starch content was determined according to YC / T 216-2013 "Continuous Flow Method for Determination of Starch in Tobacco and Tobacco Products".
[0009] Step 2: perform data preprocessing, including processing of hyperspectral data and segmentation of image data; the details are as follows: Step 2.1: Preprocess the hyperspectral image data by using SpecView software for lens calibration, reflectance calibration, and atmospheric correction, and splice the image data using PhotoScan and HiRegistrator software. Then, extract the region of interest (ROI) of the spliced image using ENVI 5.6 software, and extract the average spectral reflectance of the canopy tobacco leaves in each plot as the spectral reflectance of the upper leaves of the flue-cured tobacco in that plot.
[0010] Step 2.2: Visible light images: Use DJI Zhitu software to stitch the images captured by the visible light camera to generate a complete image of the experimental field. Labelme software was used to annotate the 72 images to construct a training set. Segmentation: Use the U2-Net network model to perform image segmentation.
[0011] Step 2.3, in order to address the problem of insufficient data, the present invention introduces the synthetic minority over-sampling technique (SMOTE) as a data enhancement method to improve the training effect of the model.
[0012] Step 3: extract the features of hyperspectral data and image data and perform feature fusion; the details are as follows: Step 3.1, extract the color features and texture features of the segmented image, including four statistical features of mean, standard deviation, skewness, third-order moment and peak of three commonly used color spaces such as RGB, HSV and YCbCr. In addition, seven color features widely used in crop monitoring were selected. For texture features, this study selected three types of texture analysis methods suitable for high-density green leafy plant images: gray-level co-occurrence matrix features (GLCM), local binary patterns (LBP) and Gabor filters. A total of 43 color features and 48 texture features were extracted by the above method, totaling 91 field tobacco plant image features. These features provide rich characterization information for subsequent starch content prediction.
[0013] Step 3.2, the hyperspectral screening characteristic bands are screened using PCA, SiPLS, SPA, CARS, and Random-Frog methods, and the collaborative interval partial least squares (SiPLS) method is preferred. According to the optimal results of the number of bands screened, it is shown that selecting 30 bands is the best.
[0014] Step 4: Use machine learning (such as support vector regression (SVR)) and deep learning (such as GRU) to train and predict the data; the details are as follows: Step 4.1, the present invention uses 9 machine learning models, namely: support vector regression (SVR), partial least squares regression (PLSR), Lasso regression, ridge regression (Ridge), kernel ridge regression (KRR), K nearest neighbor regression (KNN), random forest regression (RF), gradient boosting regression (GBR) and XGBoost regression (XGB). At the same time, the optimal parameters of each model are tuned using hyperparameter grid search (GridSearchCV) to further optimize the regression prediction performance. Support vector regression (SVR) is preferred.
[0015] In step 4.2, the present invention selects 6 deep learning models, including multi-layer perceptron (MLP), convolutional neural network (CNN), recurrent neural network (RNN), long short-term memory network (LSTM), gated recurrent unit (GRU), and attention mechanism (Transformer). The training process of each model adopts the early stopping method, and the training is terminated when the model loss value no longer decreases within 50 consecutive epochs. The gated recurrent unit (GRU) is preferred.
[0016] Step 5: Integrate the prediction results of multiple optimal models through stacking technology to obtain the final starch content prediction value.
[0017] The aforementioned machine learning and deep learning have shown relatively excellent performance. Therefore, the present invention proposes an optimal combination scheme based on the above experimental results, selecting support vector regression (SVR) and gated recurrent unit (GRU) as the basic learner (i.e., the original learner), using multi-layer perceptron (MLP) as the meta-learner, and stacking different learning methods through the stacking generalization method to form an integrated learning model.
[0018] The beneficial effects of the present invention are as follows: 1. The present invention improves the accuracy of starch content prediction by combining hyperspectral data and image data, avoiding the limitations of traditional methods that rely on manual experience and chemical analysis.
[0019] 2. Adopt integrated learning and deep learning models to make full use of information from multi-source data, effectively improving the robustness and prediction accuracy of the model.
[0020] 3. This method can be widely used in the precise management of tobacco production, reduce production costs, and improve the quality control level of tobacco, and has strong application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is the technical roadmap of the present invention; Figure 2 The image segmentation result diagram in step 2 of the present invention; Figure 3 It is a function diagram of the training result in step 4 of the present invention; Figure 4 This is a scatter plot of the prediction results of step 5 of the present invention; Figure 5 This is a linear comparison diagram of the prediction results in step 5 of the present invention. DETAILED DESCRIPTION
[0022] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0023] The following is a detailed description of the tobacco starch content prediction method based on graph fusion provided by the present invention in conjunction with the accompanying drawings (see Figure 1 ): First, the UAV is used to collect hyperspectral data and image data as well as starch content data at fixed points in the tobacco field at regular intervals. In the data preprocessing process, SiPLS technology is used to reduce the dimension of the hyperspectral data, remove redundant information, and segment the image data. Figure 2 As shown in the figure, after segmentation, the color features and texture features are extracted. After preprocessing, the data is input into the machine learning and deep learning models for training and optimization.
[0024] Among them, the support vector regression (SVR) model performed better in machine learning models, with an R² of 0.88 and an RMSE of 4.21; and in deep learning models, the GRU model performed best, with an R² of 0.94 and an RMSE of 2.03. In order to further improve the prediction performance of the model, this paper introduced stacking generalization (Stacking) and data enhancement technology (SMOTE), which enhances the generalization ability of the model by integrating the SVR and GRU models and using the multi-layer perceptron (MLP) as a meta-learner. The research results show that after the integrated learning method is combined with the screened hyperspectral and image feature data, the R² of the model reaches 0.97 and the RMSE drops to 1.50, which significantly improves the accuracy of tobacco starch content prediction. The results are as follows: Figure 3 , Figure 4 , Figure 5 Through the above method, the model can accurately predict the starch content of tobacco and provide strong decision support for tobacco production.
[0025] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention are equivalent replacement methods and are included in the protection scope of the present invention.
Claims
1. A method for predicting tobacco starch content based on graph fusion, characterized in that: The following steps are involved: Step 1: Use drones to collect hyperspectral and image data of tobacco canopies in the field, and collect fresh tobacco leaf samples from the upper part to obtain starch content; Step 2, data preprocessing, including hyperspectral data processing and image data segmentation processing; Step 3, extracting the features of hyperspectral data and image data, and performing feature fusion; Step 4: Use machine learning models and deep learning models to train and predict data; Step 5: Integrate the prediction results of multiple optimal models through stacking technology to obtain the final starch content prediction value.
2. The method according to claim 1, characterized in that Step 1 is as follows: Step 1.1, the collection period is based on the key growth period after the flue-cured tobacco is toppled, including four periods: topping period, lower leaf maturity period, middle leaf maturity period, and upper leaf maturity period; Step 1.2: Use drones to collect hyperspectral data and visible light data. Data acquisition is performed between 10:00 and 14:00 in clear and cloudless weather. Step 1.3, starch content determination: The starch content was determined by freeze-drying fresh tobacco leaf samples of the upper leaves of flue-cured tobacco obtained in the field and referring to YC / T216-2013 "Continuous flow method for determination of starch in tobacco and tobacco products".
3. The method according to claim 1, characterized in that Step 2 is as follows: Step 2.1, spectral data calibration and splicing were performed using SpecView software, PhotoScan and HiRegistrator software for the hyperspectral image data; then, the average spectral reflectance of the canopy tobacco leaves in each plot was extracted using ENVI 5.6 software as the spectral reflectance of the upper leaves of the flue-cured tobacco in the plot; Step 2.2, visible light image segmentation uses the U2-Net network model to perform image segmentation; In step 2.3, the synthetic minority over-sampling technique (SMOTE) is introduced as a data enhancement method to improve the training effect of the model.
4. The method according to claim 1, characterized in that: Step 3 is as follows: Step 3.1, extract the color features and texture features of the segmented image, including four statistical features of mean, standard deviation, skewness, third-order moment and peak value of three commonly used color spaces such as RGB, HSV and YCbCr. In addition, seven color features widely used in crop monitoring were selected; for texture features, three types of texture analysis methods suitable for high-density green leafy plant images were selected: gray-level co-occurrence matrix features (GLCM), local binary patterns (LBP) and Gabor filters; a total of 43 color features and 48 texture features were extracted through the above methods, totaling 91 field tobacco plant image features; Step 3.2, the hyperspectral screening characteristic bands are screened using PCA, SiPLS, SPA, CARS, and Random-Frog methods, and the collaborative interval partial least squares (SiPLS) method is preferred. According to the optimal results of the number of bands screened, it is shown that selecting 30 bands is the best.
5. The method according to claim 1, characterized in that In step 4, 9 machine learning models and 6 deep learning models are used to train and predict the data, and the best model is selected after training; The nine machine learning models are: Support Vector Regression (SVR), Partial Least Squares Regression (PLSR), Lasso Regression, Ridge Regression (Ridge), Kernel Ridge Regression (KRR), K-Nearest Neighbor Regression (KNN), Random Forest Regression (RF), Gradient Boosting Regression (GBR) and XGBoost Regression (XGB). At the same time, the hyperparameter grid search (GridSearchCV) is used to tune the best parameters of each model to further optimize the regression prediction performance. The six deep learning models include multi-layer perceptron (MLP), convolutional neural network (CNN), recurrent neural network (RNN), long short-term memory network (LSTM), gated recurrent unit (GRU), and attention mechanism (Transformer); the training process of each model adopts the early stopping method (Early Stopping) and the training is terminated when the model loss value no longer decreases within 50 consecutive epochs.
6. The method according to claim 1, characterized in that In step 5, based on the excellent performance of the aforementioned machine learning and deep learning, an optimal combination scheme based on the above experimental results is proposed. Support vector regression (SVR) and gated recurrent unit (GRU) are selected as the original learners, and multi-layer perceptron (MLP) is used as the meta-learner. Different learning methods are stacked together through the stacking method to form an integrated learning model.
7. The method according to claim 5, characterized in that Support vector regression (SVR) is selected as the optimal model among the 9 machine learning models.
8. The method according to claim 5, characterized in that The gated recurrent unit (GRU) is the preferred deep learning model among the 6 models.