Random forest model-based tobacco leaf yield estimation method and device
Tobacco leaf yield prediction is solved by integrating random forest models with multiple data sources, and the low efficiency and high cost problems of traditional manual production estimation methods are achieved, and high-precision and low-cost tobacco leaf yield prediction and scientific decision-making support are achieved.
Patent Information
- Application Number
- CN202510314087.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-08
AI Technical Summary
Traditional manual production estimation methods have high subjectivity, low efficiency and high cost problems in tobacco leaf yield prediction, especially in large-scale agricultural production, which is difficult to achieve efficient and accurate yield prediction.
The tobacco leaf production estimation method based on the random forest model is adopted, and the tobacco farmer tobacco field binding data, tobacco farmer acquisition data, field vector data and field remote sensing data are integrated. Yield prediction is performed through the random forest regression model, and data processing and model training are carried out in combination with vegetation index, enhanced vegetation index and leaf area index.
It realizes high-precision and low-cost tobacco field yield prediction, reduces artificial errors, improves the accuracy and efficiency of yield prediction, supports scientific decision-making, adapts to different environments and planting models, and has the ability to update dynamically.
Smart Images

Figure CN120278544A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of agricultural yield estimation, and particularly relates to a tobacco leaf yield estimation method and device based on a random forest model. Background Art
[0002] Currently, tobacco leaf yield measurement still relies on traditional manual estimation methods, which have significant limitations, especially in large-scale agricultural production. Manual yield estimation usually requires agricultural technicians to conduct on-site investigations on the growth of tobacco leaves in each field block. By collecting data (such as the growth status, coverage, and number of leaves of tobacco leaves), and then combining experience to predict the yield. However, this method highly depends on the experience and judgment of tobacco technicians, resulting in a large degree of subjectivity in the yield measurement results, and the accuracy is often limited by the experience of the operator. Only experienced tobacco technicians can make relatively accurate estimations under different climatic conditions and growth environments. Therefore, there may be large fluctuations in the yield measurement results among different tobacco technicians.
[0003] In addition, manual yield estimation is a labor-intensive task, usually requiring a large amount of human resources for field investigations. To obtain relatively accurate yield data, it is usually necessary to dispatch multiple technicians to the fields for data collection, analysis, and summary. This not only takes time and effort but also increases the labor cost. Especially in large-scale planting areas, the cycle of manual yield estimation is long, and timely yield measurement feedback cannot be achieved. Therefore, the manual method has low efficiency in yield measurement for large-scale field blocks and cannot quickly and accurately obtain the yield data of the entire planting area.
[0004] Meanwhile, the yield of tobacco leaves is affected by various environmental factors, including climate, soil quality, precipitation, etc. Traditional manual yield estimation methods usually cannot fully consider these complex environmental variables. This makes the accuracy of manual estimation results vary greatly in different regions and different growth environments, resulting in local inaccuracies in the yield measurement results. With the expansion of agricultural production scale, the limitations of the manual yield estimation method become more obvious when covering large areas of field blocks, and it cannot meet the requirements of large-scale, efficient, and real-time yield measurement.
[0005] In the process of modern agricultural development, with the gradual reduction of the labor force and the increase of labor costs, traditional manual yield estimation methods are facing severe challenges. Especially for large agricultural cooperatives and enterprises, relying on manual yield estimation for large-scale yield prediction has become increasingly infeasible. Therefore, how to improve the efficiency and accuracy of yield measurement has become an urgent problem to be solved. Based on modern technical means such as remote sensing technology and satellite images, through data collection, analysis, and modeling to achieve efficient and accurate yield measurement, a solution for a tobacco leaf yield estimation method based on a random forest model is provided. With the help of these technologies, efficient and accurate yield prediction can be achieved in large-scale field blocks, greatly reducing manual intervention and improving the accuracy and efficiency of yield measurement. Summary of the Invention
[0006] Based on the problems mentioned in the above background art, the present invention provides a method for estimating tobacco leaf yield based on a random forest model.
[0007] The technical solution adopted by the present invention is as follows: A method and device for estimating tobacco leaf yield based on a random forest model.
[0008] In the first aspect, the present invention provides a method for estimating tobacco leaf yield based on a random forest model, including the following steps: S1: Data acquisition: including data on the binding of tobacco farmers to their tobacco fields, tobacco farmer purchase data, vector data of the fields, and remote sensing data of the fields; S2: Data preprocessing: splicing the tobacco farmer purchase data, remote sensing data of the fields through the data on the binding of tobacco farmers to their tobacco fields to form a data set, and dividing the data set into a training set and a test set; S3: Constructing a random forest model: training using the training set and testing the model using the test set to obtain the prediction result of the tobacco leaf yield; S4: Model evaluation: verifying the accuracy of the random forest model through cross-validation and evaluation metrics.
[0009] Further, the acquisition of the remote sensing data of the fields in S1 is specifically as follows: S11: By performing regional superposition of the vector data of the fields and the remote sensing spectral data, extracting the remote sensing data of the corresponding regions of the fields, where the remote sensing data of the fields specifically includes: vegetation index, enhanced vegetation index, and leaf area index data; S12: Performing standardization processing on the remote sensing data such as the vegetation index, enhanced vegetation index, and leaf area index: Making it suitable for the training of the regression model.
[0010] Further, the data preprocessing in S2 is specifically as follows: S21: Splicing the data on the binding of tobacco farmers to their tobacco fields and the tobacco farmer purchase data through the tobacco farmer code and the business year; S22: Extracting the remote sensing data of the fields through the contour information of the fields planted by the tobacco farmers obtained from the data on the binding of tobacco farmers to their tobacco fields; S23: Splicing the obtained remote sensing data of the fields and the tobacco farmer purchase data through the data id.
[0011] Further, the ratio of the training set to the test set is 8:2, and the random seed is set to 42 to ensure the reproducibility of the results.
[0012] Further, the specific construction of the random forest model in step S3 includes: S31: Defining the hyperparameter space, including: the number of trees: the number of decision trees in the random forest, generating a sequence with a step size of 50 from 50 to 500; the minimum number of samples for splitting: the minimum number of samples required to split an internal node, set to [2, 5, 10]; the minimum number of samples in a leaf: the minimum number of samples required for a leaf node, set to [1, 2, 4]; the maximum number of features: the maximum number of features considered when looking for the best split, set to ['auto','sqrt', 'log2'].
[0013] Further, in S32: Use RandomizedSearchCV for hyperparameter optimization, including: setting the number of cross-validation folds to 5; setting the number of iterations to 50; setting the evaluation metric to negative mean squared error; and setting the number of parallel processes to 4.
[0014] Further, in S33: Run random search on the training set to find the best combination of hyperparameters; S34: Output the best combination of hyperparameters found through random search as the construction parameters of the random forest model.
[0015] In a second aspect, the present invention also discloses a tobacco leaf yield estimation device based on a random forest model, including a data acquisition module: used to acquire data on the binding of tobacco farmers to their tobacco fields, tobacco farmer purchase data, field vector data of digital tobacco fields, and field remote sensing data; a data processing module: used to splice the tobacco farmer purchase data, field remote sensing data, and data on the binding of tobacco farmers to their tobacco fields to form a data set, and divide the data set into a training set and a test set; a prediction module: construct a random forest model, train it using the training set, and test the model using the test set to obtain the prediction result of the tobacco leaf yield; an evaluation module: verify the accuracy of the random forest model through cross-validation and evaluation metrics.
[0016] In a third aspect, the present invention also discloses an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the above method.
[0017] In a fourth aspect, the present invention also discloses a computer-readable storage medium, on which a computer program is stored, and the computer program is used to cause a computer to execute the above method.
[0018] The beneficial effects of the present invention:
[0019] By integrating data on the binding of tobacco farmers to their tobacco fields, tobacco farmer purchase data, field vector data of digital tobacco fields, and vegetation index data obtained by remote sensing, and using a regression model to accurately predict the field yield of tobacco leaves, the following remarkable advantages are achieved:
[0020] 1. High-precision yield prediction:
[0021] By combining field remote sensing data (such as vegetation index, enhanced vegetation index, leaf area index) with tobacco farmer purchase data, accurate field yield prediction (accuracy is about 93.315%) can be achieved through a random forest model. This prediction result can reflect the growth status of each field, thereby avoiding errors caused by manual yield measurement and improving the accuracy of prediction.
[0022] 2. Large-scale and efficient field yield measurement:
[0023] Different from the traditional method of estimating field yields relying on manual experience, this technology can efficiently predict large-scale fields through field remote sensing data and existing tobacco farmer purchase data. By obtaining the remote sensing data of the planted field area, the yield can be predicted. Especially in large-scale tobacco fields or fields that are difficult to directly access, this technology provides a highly operable and automated solution, significantly reducing the time cost and human resource consumption of manual yield measurement.
[0024] 3. Reducing manual intervention and minimizing human errors:
[0025] Traditional field yield prediction relies on experienced experts or tobacco farmers for manual estimation. The process is not only time-consuming but also easily affected by the personal experience of the operators, prone to errors. By introducing the automated analysis of field remote sensing data and random forest models, this technology can objectively and quickly process data and make predictions, greatly reducing the errors caused by manual intervention and improving the reliability of yield prediction.
[0026] 4. Data-driven scientific decision support:
[0027] This technology can provide scientific and accurate yield predictions based on the prediction data, providing decision support for tobacco farmers and agricultural managers. These prediction data can help tobacco farmers reasonably plan planting schemes, optimize purchase plans, and take appropriate management measures at different growth stages, improving agricultural production efficiency.
[0028] 5. Dynamic update and continuous optimization:
[0029] With the continuous acquisition of new field remote sensing data and tobacco farmer purchase data, this technology supports the regular update and optimization of the model. By regularly training the model and adjusting parameters, the technology can self-adapt to different environmental changes and production conditions, ensuring the timeliness and accuracy of prediction results. This feature enables the technology to be applied to tobacco leaf yield prediction in the long term and can continuously improve the prediction accuracy with the accumulation of data volume.
[0030] 6. Multi-dimensional data integration for comprehensive assessment of field health status:
[0031] This technology is not limited to using a single data source but integrates various data such as tobacco farmer purchase data, remote sensing vegetation index, and leaf area index. This multi-dimensional data integration method can more comprehensively reflect the growth status of the field. Through the complementarity of different data sources, the technology can provide more comprehensive and accurate yield predictions at different times and in different environments.
[0032] 7. Strong scalability and adaptability:
[0033] This technology can not only be applied to tobacco fields, but also adjust the prediction of the yield of other crop fields according to needs. By modifying and expanding the data input model and adjusting the regression analysis method, it can easily adapt to different agricultural planting patterns, regions and climate conditions. Therefore, this technology has good scalability and can meet the needs of different agricultural production. Description of the Drawings
[0034] The present invention can be further illustrated by non-limiting embodiments given in the accompanying drawings;
[0035] Figure 1 is a flow chart of the present invention;
[0036] Figure 2 is a block diagram of the tobacco leaf yield estimation device of the present invention;
[0037] Figure 3 is a schematic structural diagram of a computer system of an electronic device suitable for implementing the present invention. Detailed Embodiments
[0038] As Figure 1 shown, a tobacco leaf yield estimation method based on a random forest model includes the following steps:
[0039] S1: Data acquisition: including tobacco farmer-tobacco field binding data, tobacco farmer purchase data, field vector data and field remote sensing data;
[0040] Among them, the tobacco farmer-tobacco field binding data: collect the binding information between tobacco farmers and fields, which will help locate each field and be associated with other data sources (such as tobacco farmer purchase data, field remote sensing data, etc.).
[0041] Specifically as follows: When the annual tobacco leaf production work is carried out, the relationship between tobacco farmers and tobacco fields will be bound during the contract signing process, and the tobacco technicians will operate on a dedicated platform. This data includes the binding relationship between tobacco farmers and tobacco fields, as well as information such as the actual planting area of tobacco farmers and the longitude and latitude of tobacco fields. The data structure is as follows:
[0042] Table 1. Structure of the annual tobacco farmer-tobacco field binding relationship table
[0043]
[0044]
[0045] The purchase data of tobacco farmers: collect the purchase data of tobacco farmers, including the purchase volume of the upper, middle and lower parts of tobacco leaves. These data will be used to calculate the average yield per mu of the fields planted by tobacco farmers, as the prediction target data for training the model and used for regression analysis.
[0046] Specifically as follows: The tobacco farmer acquisition data is sourced from the Tobacco ADB, which includes the tobacco farmer ID, acquisition part, and acquisition quantity. The data structure is as follows:
[0047] Table 2. Tobacco Farmer Acquisition Data Structure Table
[0048] Indicator English Name Indicator Chinese Name FARMER_CD Farmer ID LEAF_POSITION_CD Tobacco Leaf Position ID LEAF_POSITION_NAME Tobacco Leaf Position Name BUY_WGHT Purchase Quantity BUSINESS_YEAR Data Year
[0049] Field vector data acquisition: Used to obtain field contour information, extract the geometric shapes (boundary coordinates and area) of each field from the field vector data, and determine the specific location of the field. This information can help identify the spatial distribution and positional relationship of the fields, providing a basis for subsequent remote sensing data analysis.
[0050] Specifically as follows: The contour data of the tobacco fields comes from the vector layer data of the existing digital tobacco fields, which includes data on more than 310,000 fields within the scope of Chongqing. This data includes the longitude and latitude, area, and contour vector data of the fields.
[0051] Remote sensing data acquisition: Obtain multi-temporal vegetation index, enhanced vegetation index, and leaf area index data for the tobacco-growing fields during the main field stage from early May to the end of August each year.
[0052] Vegetation Index (NDVI): Obtain the vegetation index (Normalized Difference Vegetation Index, NDVI) of each field through remote sensing technology. This index can reflect the growth status of crops, especially the green coverage during the growing season.
[0053]
[0054] Among them:
[0055] NIR is the reflectance of the near-infrared band (usually in the wavelength range of 0.76–0.90μm).
[0056] RED is the reflectance of the red light band (usually in the wavelength range of 0.63–0.69μm).
[0057] Enhanced Vegetation Index (EVI): Obtain the Enhanced Vegetation Index (EVI). This index corrects the effects of the atmosphere, soil, and vegetation based on NDVI and can provide a more accurate assessment of crop growth in more complex environments.
[0058]
[0059] Among them:
[0060] G is the gain factor, usually with a value of 2.5.
[0061] C1 and C2 are the coefficients of the red and blue light channels, which are 6 and 7.5 respectively.
[0062] L is the background light reflection value, usually taking the value of 10000.
[0063] NIR, RED, and BLUE are the reflectance of the near-infrared, red, and blue light bands respectively.
[0064] Leaf Area Index (LAI): Obtain the Leaf Area Index (LAI), which reflects the leaf area density of the crop and is an important indicator to measure the photosynthesis ability of the crop.
[0065]
[0066] Where:
[0067] Aleaf represents the total area of the plant leaves, usually in square meters.
[0068] Aground represents the total area of the ground, usually also in square meters.
[0069] Data acquisition is as follows: Through the GEE platform, overlay the vector data of the tobacco fields and the remote sensing spectral data in the area, and extract the remote sensing data corresponding to the fields.
[0070] S2: Data preprocessing: Splice the tobacco farmer acquisition data, the field remote sensing data, and the data bound by the tobacco farmer and the tobacco field to form a data set, and divide the data set into a training set and a test set;
[0071] Splice the tobacco farmer acquisition data and the field remote sensing data through the data bound by the tobacco farmer and the tobacco field to ensure that the actual yield data of each field can be spliced with the corresponding field contour information and remote sensing data.
[0072] Specifically as follows: Since the acquisition of multiple data sources is involved, the data needs to be spliced before the modeling data analysis. First, splice the data bound by the tobacco farmer and the tobacco field and the tobacco farmer acquisition data through the tobacco farmer code and the business year. Obtain the field contour information planted by the tobacco farmer through the data bound by the tobacco farmer and the tobacco field, and extract the remote sensing feature data thereof. Finally, splice all the obtained data through the data id.
[0073] Table 3. Data structure table
[0074]
[0075]
[0076] Standardize remote sensing data such as vegetation index, enhanced vegetation index, and leaf area index so that it is suitable for the training of regression models. The data standardization formula is as follows:
[0077] S3: Build a random forest model: Use the training set for training and use the test set to test the model to obtain the prediction results of tobacco leaf yield;
[0078] Details of the construction of the random forest regression model:
[0079] Use the Python programming language and rely on libraries such as numpy, pandas, and sklearn for data processing and model construction.
[0080] Use the train_test_split function to divide the dataset into a training set and a test set, where the test set accounts for 20% of the total data and the training set accounts for 80%. Set the random seed to 42 to ensure the repeatability of the results.
[0081] Initialize a random forest regression model RandomForestRegressor and set the random seed to 42 to ensure the reproducibility of the results.
[0082] Define the hyperparameter space of the random forest model, including the number of trees (n_estimators), the minimum number of samples for splitting (min_samples_split), the minimum number of samples for leaves (min_samples_leaf), and the maximum number of features (max_features).
[0083] 1. n_estimators
[0084] Definition: The number of decision trees in the random forest.
[0085] Function: More trees usually improve the performance of the model, but also increase the computational cost; too few trees may lead to underfitting of the model, and too many trees may lead to waste of computing resources.
[0086] Range of values:
[0087] In this example, use np.arange to generate a sequence from 50 to 500 with a step of 50:
[0088] 2. min_samples_split
[0089] Definition: The minimum number of samples required to split an internal node.
[0090] Function: Controls the splitting condition of the decision tree. If the number of samples at a node is less than this value, the splitting will not continue; a larger value can prevent overfitting but may lead to underfitting.
[0091] Range of values:
[0092] In this example, it is set to [2, 5, 10].
[0093] 3. min_samples_leaf
[0094] Definition: The minimum number of samples required for a leaf node.
[0095] Function: Controls the minimum number of samples in a leaf node. If the number of samples in a leaf node after splitting is less than this value, the splitting will not occur; a larger value can prevent overfitting but may lead to underfitting.
[0096] Range of values:
[0097] In this example, it is set to [1, 2, 4].
[0098] 4. max_features
[0099] Definition: The maximum number of features to consider when looking for the best split.
[0100] Function: Controls the number of features considered during the splitting of each decision tree; a smaller value can reduce the variance of the model, prevent overfitting, but may increase the bias.
[0101] Range of values:
[0102] In this example, it is set to ['auto','sqrt', 'log2']:
[0103] 'auto': Considers all features.
[0104] 'sqrt': Considers the square root of the total number of features.
[0105] 'log2': Considers the logarithm of the total number of features.
[0106] Use RandomizedSearchCV for hyperparameter optimization, specifically:
[0107] Set the number of cross - validation folds to 5. The specific operation of 5 - fold cross - validation is to randomly divide the dataset into 5 parts (referred to as "folds"). Each time, 4 of them are used as the training set, and the remaining 1 part is used as the validation set. This process is repeated 5 times, each time choosing a different validation set, and finally, the average of the 5 evaluation results is taken as the performance metric of the model.
[0108] The number of iterations is 50. The number of iterations refers to the number of parameter combinations randomly sampled from the hyperparameter space during the RandomizedSearchCV process. In the present invention, n_iter = 50 is set, indicating that RandomizedSearch will randomly sample 50 different hyperparameter combinations from the defined hyperparameter space (param_dist) for training and evaluation.
[0109] The evaluation metric is neg_mean_squared_error. The evaluation metric is used to measure the performance of the model on the validation set. In the present invention, neg_mean_squared_error is selected as the evaluation metric. Mean Squared Error (MSE) is a commonly used performance metric in regression problems, representing the average of the squares of the differences between the predicted values and the true values. The smaller the MSE, the better the prediction performance of the model. Negative Mean Squared Error is the negative form of MSE. Since RandomizedSearchCV defaults to maximizing the evaluation metric, using negative MSE can indirectly achieve the goal of minimizing MSE.
[0110] The number of parallel processes is 4. Parallel processing refers to using a multi-core CPU to simultaneously process multiple tasks during model training and cross-validation, thereby accelerating the calculation. n_jobs = 4 means using 4 CPU cores to process tasks in parallel. For example, in 5-fold cross-validation, the training and validation for each fold can be assigned to different cores and carried out simultaneously.
[0111] Run RandomizedSearch on the training set to find the best hyperparameter combination.
[0112] Output the best hyperparameter combination found through RandomizedSearch as the construction parameters for the Random Forest regression model. In the present invention, the final hyperparameter combination is:
[0113] n_estimators: 100
[0114] min_samples_split: 5
[0115] min_samples_leaf: 1
[0116] max_features:'sqrt'.
[0117] S4: Model evaluation: Verify the accuracy of the Random Forest model through cross-validation and the evaluation metric.
[0118] This model splits the dataset into a training set and a validation set in the ratio of 8:2. Among them, there are more than 63,000 data in the training set and more than 15,000 data in the validation set. Through cross-validation and evaluation metrics (such as R 2 , mean squared error, etc.), the regression model is verified to evaluate the accuracy and generalization ability of the model.
[0119] 2. Evaluation Metrics
[0120] To comprehensively evaluate the prediction performance of the model, the following common regression model evaluation metrics are used in this study:
[0121] Accuracy: To intuitively measure the prediction accuracy of the model, we use the Mean Absolute Percentage Error (MAPE), and calculate the accuracy of the model through the following formula:
[0122] Accuracy = 100×(1 - MAPE)
[0123] The calculation formula of MAPE is:
[0124]
[0125] Among them, yi is the actual value, is the predicted value, and n is the number of samples. Through 100×(1 - MAPE), we obtain the accuracy of the model, and the higher the value, the smaller the prediction error of the model.
[0126] Coefficient of Determination (R2): The coefficient of determination is an important indicator to evaluate the goodness of fit of the regression model, indicating the ability of the model to explain the data. Its calculation formula is:
[0127]
[0128] Among them, is the mean of the actual values. The closer the R2 value is to 1, the better the fitting effect of the model on the data and the stronger the prediction ability.
[0129] Mean Absolute Error (MAE): The mean absolute error measures the average of the absolute errors between the predicted values and the true values. The formula is:
[0130]
[0131] Among them, yi is the actual value, is the predicted value, n is the number of samples, and the smaller the MAE, the closer the prediction result of the model is to the true value.
[0132] Root Mean Square Error (RMSE): The root mean square error is the standard deviation of the errors, reflecting the magnitude of the model prediction errors. Its calculation formula is:
[0133]
[0134] Among them, yi is the actual value, is the predicted value, n is the number of samples. The smaller the RMSE value, the smaller the prediction error of the model, and the better the prediction effect of the model.
[0135] Model application: Using the trained regression model, predict new fields or fields without collected yield data. Input the vegetation index data of the new field (such as NDVI, EVI, LAI) to obtain the corresponding predicted average yield per mu.
[0136] Result visualization and analysis: Present the prediction results in a visual way, such as maps, charts, etc., so that relevant personnel such as tobacco farmers and agricultural managers can make decisions and manage. More accurate acquisition, scheduling and production plans can be formulated based on the predicted yield data.
[0137] Model optimization: Evaluate the performance of the model according to the comparison between the actual acquisition data and the prediction results, and further adjust the model parameters to improve the accuracy of the model. With the acquisition of new acquisition data and remote sensing data, regularly update the regression model to maintain the accuracy and timeliness of the model.
[0138] Please refer to Figure 2 , which is a block diagram of a tobacco leaf yield estimation device based on a random forest model shown in an exemplary embodiment of this application. As Figure 2 shown, in an exemplary embodiment, the tobacco leaf yield estimation device based on a random forest model at least includes a data acquisition module 110, a data processing module 120, a prediction module 130, and an evaluation module 140, which are introduced in detail as follows:
[0139] Data acquisition module 110: Used to acquire tobacco farmer tobacco field binding data, tobacco farmer acquisition data, field vector data, and field remote sensing data;
[0140] Data processing module 120: Used to splice the tobacco farmer acquisition data, field remote sensing data, and form a data set through the tobacco farmer tobacco field binding data, and divide the data set into a training set and a test set;
[0141] Prediction module 130: Construct a random forest model, train it using the training set, and test the model using the test set to obtain the prediction result of the tobacco leaf yield;
[0142] Evaluation module 140: Verify the accuracy of the random forest model through cross-validation and evaluation metrics.
[0143] It should be noted that the tobacco leaf yield estimation device based on the random forest model provided in the above embodiments and the tobacco leaf yield estimation method based on the random forest model provided in the above embodiments belong to the same concept. The content of the operations performed by each module has been described in detail in the method embodiments and will not be repeated here.
[0144] Figure 2 The structure diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown. It should be noted that Figure 2 The computer system 200 of the shown electronic device is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0145] As Figure 3 shown, the computer system 200 includes a central processing unit (CPU) 201, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 202 or the program loaded from the storage section 208 into the random access memory (RAM) 203, such as executing the method in the above embodiments. In the RAM 203, various programs and data required for system operation are also stored. The CPU 201, ROM 202, and RAM 203 are connected to each other via a bus 204. The input / output (I / O) interface 205 is also connected to the bus 204.
[0146] The following components are connected to the I / O interface 205: an input section 202 including a keyboard, a mouse, etc.; an output section 207 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 208 including a hard disk, etc.; and a communication section 209 including a network interface card such as a LAN (Local Area NetworK) card, a modem, etc. The communication section 209 performs communication processing via a network such as the Internet. A drive 210 is also connected to the I / O interface 209 as required. A removable medium 211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 210 as required so that the computer program read from it can be installed into the storage section 208 as required.
[0147] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 209, and / or installed from the removable medium 211. When the computer program is executed by the central processing unit (CPU) 201, various functions defined in the system of the present application are executed.
[0148] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor of a computer, the computer is caused to execute the method for configuring a rule engine for early warning as described above. The computer-readable storage medium may be included in the electronic device described in the above embodiment, or may exist separately without being assembled into the electronic device.
[0149] It should be noted that the computer-readable medium shown in the embodiments of the present application may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable computer program is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0150] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0151] The units involved in the embodiments described in the present application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not, in some cases, constitute a limitation on the unit itself.
[0152] The above embodiments are only illustrative of the principles and effects of the present application and are not intended to limit the present application. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present application. Therefore, all equivalent modifications or changes made by those of ordinary skill in the art in the technical field without departing from the spirit and technical idea disclosed in the present application should still be covered by the claims of the present application.
Claims
1. A tobacco leaf yield estimation method based on a random forest model, characterized in that: It includes the following steps: S1: Data acquisition: including data binding of tobacco farmers and their tobacco fields, tobacco farmer purchase data, field vector data, and field remote sensing data; S2: Data preprocessing: splicing the tobacco farmer purchase data and the field remote sensing data through the data binding of tobacco farmers and their tobacco fields to form a data set, and dividing the data set into a training set and a test set; S3: Construct a random forest model: use the training set to train the random forest model, and use the test set to test the model to obtain the prediction result of the tobacco leaf yield; S4: Model evaluation: verify the accuracy of the random forest model through cross-validation and evaluation metrics.
2. The tobacco leaf yield estimation method based on a random forest model according to claim 1, wherein: The acquisition of the field remote sensing data in S1 is specifically as follows: S11: By performing regional overlay of the field vector data and the remote sensing spectral data, extract the remote sensing data of the corresponding area of the field, where the field remote sensing data specifically includes: vegetation index, enhanced vegetation index, and leaf area index data; S12: Perform standardization processing on the vegetation index, enhanced vegetation index, and leaf area index data respectively to obtain the standardized vegetation index, enhanced vegetation index, and leaf area index data.
3. The tobacco leaf yield estimation method based on the random forest model according to claim 1, characterized in that: The data preprocessing in S2 is specifically as follows: S21: Splice the data binding of tobacco farmers and their tobacco fields and the tobacco farmer purchase data through the tobacco farmer code and the business year; S22: Extract the field remote sensing data through the field contour information of the tobacco fields planted by the tobacco farmers obtained from the data binding of tobacco farmers and their tobacco fields; S23: Splice the obtained field remote sensing data and the tobacco farmer purchase data through the data id.
4. The tobacco leaf yield estimation method based on a random forest model according to claim 1, wherein: The ratio of the training set to the test set is 8:2, and the random seed is set to 42 to ensure the reproducibility of the results.
5. The tobacco leaf yield estimation method based on a random forest model according to claim 1, characterized in that: The specific construction of the random forest model in step S3 includes: S31: Define the hyperparameter space, including: Number of trees: the number of decision trees in the random forest, generating a sequence with a step size of 50 from 50 to 500; Minimum number of samples for splitting: the minimum number of samples required to split an internal node, set to [2, 5, 10]; Minimum number of samples in a leaf: the minimum number of samples required for a leaf node, set to [1, 2, 4]; Maximum number of features: the maximum number of features considered when looking for the best split, set to ['auto','sqrt', 'log2'].
6. A tobacco leaf yield estimation method based on a random forest model according to claim 5, characterized in that: In S3, the hyperparameters are optimized through random search, including: Using cross-validation to evaluate the model performance; Performing iterative search to determine the best hyperparameter combination; Using the negative mean squared error as the evaluation metric.
7. A tobacco leaf yield estimation method based on a random forest model according to claim 6, characterized in that: Run random search on the training set to find the best hyperparameter combination; output the best hyperparameter combination found through random search as the construction parameters of the random forest model.
8. A tobacco leaf yield estimation device based on a random forest model, characterized in that: It includes: Data acquisition module: used to acquire data binding of tobacco farmers and their tobacco fields, tobacco farmer purchase data, field vector data, and field remote sensing data; Data processing module: used to splice the tobacco farmer purchase data and the field remote sensing data through the data binding of tobacco farmers and their tobacco fields to form a data set, and divide the data set into a training set and a test set; Prediction module: Build a random forest model, train it using the training set, and test the model using the test set to obtain the prediction result of the tobacco leaf yield; Evaluation module: Verify the accuracy of the random forest model through cross-validation and evaluation metrics.
9. An electronic device, characterized in that: The electronic device includes at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: Stored thereon is a computer program, and the computer program is used to cause a computer to execute the method according to any one of claims 1-7.