A method and device for predicting crop yield
By determining the disturbed area in the crop yield prediction model and adding sample points, combined with the meteorological adjustment coefficient, the problem of large-area regional prediction in the prior art ignores local interference, and improves the prediction accuracy.
Patent Information
- Application Number
- CN202410754900.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-12
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-06-12
AI Technical Summary
When dealing with large areas, the existing crop yield prediction model ignores local areas affected by topography, meteorology and other factors, resulting in low yield prediction accuracy.
By determining the disturbed area in the area to be predicted based on remote sensing images, adding multiple enhancement sample points, building a historical production estimation model, and making predictions based on meteorological adjustment coefficients.
Improve the accuracy of crop yield prediction, especially in areas disturbed by natural factors, and can more accurately output crop yield results.
Smart Images

Figure CN118570645B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of remote sensing technology, and in particular to a method and device for predicting crop yield. Background Art
[0002] With the development of satellite remote sensing technology, remote sensing technology has been applied to crop growth, crop pests and diseases, crop planting area extraction, etc., and can also predict crop yields. Crop yield prediction is of great significance to the formulation of national economic plans, food policies, and ensuring national food security.
[0003] Existing technologies usually estimate and predict future crop yields based on measured yield data of previous years, meteorological data, remote sensing images, etc. Measured yield data refers to the actual statistical crop yield data. The predicted yield is often an overall prediction for crops in a large area, such as predicting the yield of a certain crop in a certain administrative city. In a large area, the yield of a small part of the area is easily affected by factors such as terrain and weather. For example, low-lying areas are prone to waterlogging, causing crop disasters; areas with relatively poor soil quality will have lower yields. The current crop prediction model ignores the consideration of this part of the area, thus affecting the accuracy of crop yield prediction. Summary of the invention
[0004] In view of the above problems, the present application proposes a crop yield prediction method and device, the main purpose of which is to improve the accuracy of crop yield prediction.
[0005] In order to achieve the above objectives, this application mainly provides the following technical solutions:
[0006] In a first aspect, the present application provides a crop yield prediction method, comprising:
[0007] Determine the disturbed area in the area to be predicted based on the remote sensing image of the area to be predicted in the current year;
[0008] determining a plurality of enhanced sample points in the disturbed area;
[0009] Using the enhanced sample points and the historical production measured data of the area to be predicted, a historical production estimation model is constructed;
[0010] The predicted crop yield results are output based on the historical yield estimation model, the remote sensing images of the current year, and the meteorological adjustment coefficient automatically constructed based on meteorological data.
[0011] In a second aspect, the present application provides a crop yield prediction device, the device comprising:
[0012] A first determining unit is used to determine the disturbed area in the area to be predicted based on the remote sensing image of the area to be predicted in the current year;
[0013] A second determining unit, configured to determine a plurality of enhanced sample points in the disturbed area;
[0014] A construction unit, used to construct a historical production estimation model using the enhanced sample and the historical production measured data of the area to be predicted;
[0015] The output unit is used to output the predicted crop yield results based on the historical yield estimation model, the remote sensing image of the year, and the meteorological adjustment coefficient automatically constructed based on meteorological data.
[0016] On the other hand, the present application also provides a storage medium, which is used to store a computer program, wherein when the computer program is running, it controls the device where the storage medium is located to execute the crop yield prediction method of the first aspect mentioned above.
[0017] On the other hand, the present application also provides an electronic device, comprising at least one processor, and at least one memory and a bus connected to the processor; wherein the processor and the memory communicate with each other through the bus; and the processor is used to call program instructions in the memory to execute the crop yield prediction method of the first aspect as described above.
[0018] By means of the above technical solution, a crop yield prediction method and device provided by the present application can improve the accuracy of crop yield prediction. The present application obtains the remote sensing image of the current year for the area to be predicted, and determines the disturbed area in the area to be predicted based on the remote sensing image of the current year. In order to solve the problem of insufficient accuracy of the yield prediction model due to insufficient samples in the disturbed area, the present application determines multiple enhanced sample points after determining the disturbed area, which is equivalent to adding samples of the disturbed area to the present application, and constructing a historical yield estimation model using existing samples, i.e., historical yield measured data, and newly added samples. Compared with the existing historical yield estimation model obtained by only using historical yield measured data as sample training, the present application can output the crop yield prediction results more accurately.
[0019] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present application. Also, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:
[0021] Figure 1 A flow chart of a crop yield prediction method proposed in an embodiment of the present application is shown;
[0022] Figure 2 A flowchart of another crop yield prediction method proposed in an embodiment of the present application is shown;
[0023] Figure 3 A flowchart of another crop yield prediction method proposed in an embodiment of the present application is shown;
[0024] Figure 4 The NDVI distribution map of soybean area in Nehe City proposed in the embodiment of the present application is shown;
[0025] Figure 5a The distribution map of measured soybean production data in Nehe City in 2022 is shown;
[0026] Figure 5b The sample intensification area map of soybean production in Nehe City is shown;
[0027] Figure 6 The distribution map of soybean production intensification sample points in Nehe City in 2022 is shown;
[0028] Figure 7 A schematic diagram of a crop yield prediction device proposed in an embodiment of the present application is shown;
[0029] Figure 8 A schematic diagram of another crop yield prediction device proposed in an embodiment of the present application is shown;
[0030] Fig. 9 A schematic structural diagram of an electronic device proposed in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0031] The exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.
[0032] At present, the crop yield prediction model is constructed using historical yield measured data, remote sensing images and other data. Due to the large prediction area, the historical yield measured data as sample data is prone to insufficient samples and uneven distribution. However, too few or missing samples in the disturbed area will cause a decrease in the overall prediction accuracy.
[0033] To this end, an embodiment of the present application provides a method for predicting crop yields, which improves the accuracy of crop yield prediction by strengthening samples in areas affected by nature and increasing the number of sample points in this area.
[0034] The specific implementation steps of the embodiment of the present application are as follows: Figure 1 As shown, including:
[0035] 101. Based on the remote sensing images of the area to be predicted in the current year, determine the disturbed area in the area to be predicted.
[0036] The disturbed area refers to the area where crop yields are reduced or damaged due to natural factors such as precipitation and soil quality. Studies have shown that the crop yields in the disturbed area are often lower than those in other areas in the same year.
[0037] In this step, the remote sensing images of the current year of the area to be predicted for crops are obtained, and the remote sensing images of the current year are preprocessed to obtain the preprocessed remote sensing images of the current year. The vegetation index of each block is calculated based on the preprocessed remote sensing images of the current year, and the disturbed area is determined by judging the size of the vegetation index.
[0038] Exemplarily, based on the preprocessed remote sensing images of the year, the vegetation index of each plot is determined, and the plots with a vegetation index less than a first threshold are determined to be disturbed plots; then the area to be predicted is divided into several unit areas, and when the density of disturbed plots in a unit area is greater than a second threshold, the unit area is determined to be a disturbed area.
[0039] Exemplarily, based on the field survey data of the area to be predicted and the remote sensing images of the year, the Normalized Difference Vegetation Index (NDVI) value is used to determine whether each plot is a disturbed area. If the NDVI of the plot is less than the disturbance threshold, the plot is considered to be a disturbed area. Otherwise, it is considered to be a normal plot.
[0040] 102. Determine a plurality of enhanced sample points in the disturbed area.
[0041] Among them, the enhanced sample point is a sample point that is added or strengthened relative to the existing sample point. The sample point is the point corresponding to the coordinate position of the sample in the coordinate system. The existing samples in this application are historical production measured data. The existing sample points are the points corresponding to the coordinate positions of each historical production measured data in the coordinate system.
[0042] In this step, it can be determined whether the existing sample points in each disturbed area are insufficient or missing. For disturbed areas with insufficient or missing sample points, they can be supplemented by the measured yield data of previous years or the yield forecast data of the same year.
[0043] Exemplarily, the actual production data of the previous year for the area to be predicted is obtained, and the distribution of the actual production data of the previous year for each disturbed area is determined. If the actual production data of the previous year for the disturbed area is insufficient or missing, the enhanced sample points are determined based on the production forecast data of the previous year for the disturbed area. If the production forecast data of the previous year cannot be obtained, the actual production data of the disturbed area in previous years are used to determine the enhanced sample points.
[0044] 103. Use the historical production measurement data of enhanced sample points and the area to be predicted to construct a historical production estimation model.
[0045] The historical yield measured data refers to the actual yield data of a certain year that has been counted. The historical yield estimation model refers to the crop yield estimation model trained using historical yield data.
[0046] In this step, the samples corresponding to the enhanced sample points and the historical production measured data constitute the training sample data, and the original production estimation model is trained using the training sample data and the corresponding remote sensing images to obtain the historical production estimation model.
[0047] Exemplarily, based on the longitude and latitude coordinates of each enhanced sample point, the yield forecast data of the previous year is extracted as an enhanced sample, and the enhanced sample and the measured yield data of the previous year are used as training samples to train and obtain the yield estimation model of the previous year.
[0048] Exemplarily, it is determined whether there is yield forecast data from the previous year corresponding to each enhanced sample point. If not, the actual yield data of the year closest to the previous year is extracted as the enhanced sample based on the latitude and longitude coordinates of each enhanced sample point. The enhanced sample and the actual yield data of the previous year are used to train the yield estimation model for the previous year.
[0049] It should be noted that the historical production estimation model can be obtained by training with historical production measured data and historical remote sensing images, or by training with historical production measured data, historical production forecast data and historical remote sensing images. That is, the historical production estimation model establishes the correlation between historical production measured data and its corresponding historical remote sensing images, or the historical production estimation model establishes the correlation between historical production measured data, historical production forecast data and their respective corresponding historical remote sensing images.
[0050] 104. Output the predicted crop yield results based on the historical yield estimation model, the remote sensing images of the current year, and the meteorological adjustment coefficient automatically constructed based on meteorological data.
[0051] In this step, the characteristic data in the remote sensing image of the current year can be extracted, and the characteristic data can be input into the historical yield estimation model to obtain the estimated yield data of the crop in the current year. Meteorological data is obtained, and the meteorological adjustment coefficient is automatically constructed according to the meteorological data. The estimated yield data of the current year and the meteorological adjustment coefficient are used to predict the predicted yield of the crop. Among them, the meteorological adjustment coefficient is used to correct the estimated yield data of the crop in the current year so as to obtain a highly accurate predicted yield result of the crop.
[0052] This application predicts the current year's crop yield through the previous year's yield estimation model, which has a strong timeliness and provides more accurate decision-making support for the formulation of agricultural policies.
[0053] In order to explain the crop yield prediction method proposed by the present invention in more detail, the present application proposes another embodiment of the crop yield prediction method. The specific implementation steps of the embodiment of the present invention are as follows: Figure 2 As shown, including:
[0054] 201. Determine phenological periods and process multivariate data.
[0055] Among them, multivariate data include remote sensing images, meteorological data, historical production measured data, etc.
[0056] In this step, the phenological period of crops in the area to be predicted is determined according to the growth law of crops and climate characteristics. Combined with field surveys and agricultural observation data, multivariate data such as remote sensing images, meteorological data, and historical yield measured data for the specified time period of the target phenological period are obtained and preprocessed. The remote sensing images and meteorological data can be data of the same phase of the previous year and the current year, and the historical yield measured data can be the yield measured data of the previous year.
[0057] 202. Construct a historical production estimation model based on enhanced sample points.
[0058] In this step, sample enhancement is achieved by generating enhanced sample points in order to build a more accurate historical production estimation model. The specific implementation method of building a historical production estimation model is as follows:
[0059] Step 1: Determine the disturbed area in the area to be predicted.
[0060] In one achievable method, the interference affected area is confirmed by:
[0061] Step 1a: Calculate the Normalized Difference Vegetation Index (NDVI) of each block in the predicted area based on the remote sensing image of the year. If the NDVI is less than the interference threshold, execute step 1b; if the NDVI is greater than or equal to the interference threshold, execute step 1c.
[0062] Step 1b: determine that the land parcel is a disturbed area, and determine the boundary vector of the land parcel to determine the boundary of the disturbed area.
[0063] Step 1c: Determine that the plot is a normal area.
[0064] NDVI is highly sensitive to green plants. The embodiment of the present application uses NDVI to determine the disturbed area, which can accurately distinguish the disturbed area from the normal area.
[0065] Step 2: Determine multiple enhanced sample points in the disturbed area.
[0066] In one achievable method, the enhanced sample point confirmation method is:
[0067] Step 2a: traverse each disturbed area to determine whether the disturbed area contains historical production measured data.
[0068] In this step, each disturbed area is judged one by one to determine whether the disturbed area contains the actual measured data of production in the previous year. If so, step 2b is executed; if not, step 2c is executed.
[0069] Step 2b: Determine the disturbed area as the normal area of the sample.
[0070] Step 2c: Determine the disturbed area as the area to be strengthened.
[0071] The area to be strengthened is a disturbed area where additional strengthening sample points are required.
[0072] Step 2d: Generate random sample points in the area to be strengthened using a random point generation algorithm.
[0073] In this step, a random point creation method can be used to randomly place a specified number of points in the area to be strengthened in a range window corresponding to the area to be strengthened to obtain random sample points.
[0074] In one possible implementation, random sample points are generated as follows:
[0075] Create a random number data stream through a random number generator and a seed, create a polygon in the area to be strengthened, and use a standard polygon partitioning algorithm to divide the polygon according to triangles of different sizes. Specifically, first place the first point in the first polygon, randomly select a triangle in the entire polygon, and use the two sides of the triangle as the basis axis X and Y for placing random points; randomly select a point on the basis axis, and transform the next unused value in the random number data stream into a uniform distribution. The two random values obtained are used as the first random sample points, and the above steps are repeated until the specified number of random sample points is reached.
[0076] Step 2e: According to the number of historical production measured data and the principle of random point proportion, select enhanced sample points from the random sample points.
[0077] In this step, the number of historical production measured data is first counted, and then the number of enhanced sample points is determined based on the random point proportion principle and the number of historical production measured data, so as to select a certain number of enhanced sample points from the random sample points.
[0078] Step 3: Use the historical production measurement data of enhanced sample points and the area to be predicted to build a historical production estimation model.
[0079] In one achievable approach, the historical production estimation model is constructed as follows:
[0080] Step 3a: Determine the longitude and latitude coordinates of each enhanced sample point, and extract the historical production forecast data corresponding to the longitude and latitude coordinates.
[0081] Among them, the historical production forecast data is the forecast data for the same year as the historical production measured data.
[0082] In this step, the longitude and latitude coordinates of each enhanced sample point are obtained, and the predicted data of the production in the previous year at the longitude and latitude coordinates are extracted based on the longitude and latitude coordinates. That is, in the absence of the measured data of the production in the previous year, the predicted data is used instead of the measured data.
[0083] Step 3b: Use historical production forecast data and historical production measured data as sample data.
[0084] It should be noted that the sample data in this step refers to historical production data. When actually training the historical production estimation model, in addition to the historical production data, historical remote sensing images are also needed to train the model.
[0085] Step 3c: Use sample data and target vegetation index to construct a historical yield estimation model based on the target vegetation index.
[0086] The target vegetation index is one or more vegetation indices selected from a plurality of vegetation indices.
[0087] The historical yield estimation model constructed in the embodiment of the present application is a historical yield estimation model based on vegetation index. The present application can select a target vegetation index from a plurality of vegetation indexes according to actual scenarios. The historical yield estimation model is constructed using sample data and target vegetation index.
[0088] In one achievable method, the target vegetation index may be generated as follows:
[0089] According to the historical remote sensing images of the area to be predicted, various historical vegetation indices are calculated; sensitivity analysis is performed on various historical vegetation indices, and the N historical vegetation indices with the highest sensitivity are screened out from all historical vegetation indices; and the N historical vegetation indices are determined as target vegetation indices.
[0090] Among them, the historical vegetation index corresponds to the historical remote sensing image. For example, the remote sensing image of 2022 can calculate the vegetation index of 2022. When screening the historical vegetation index, the sensitivity of each historical vegetation index will be sorted from high to low, and the top N historical vegetation indices will be used as the target vegetation index, where N is a positive integer.
[0091] This application uses a highly sensitive vegetation index to construct a historical yield estimation model, which improves the accuracy of the historical yield estimation model. In addition, the embodiment of this application overcomes the problem of small sample data and uneven distribution of existing crop yield estimation models, performs sample enhancement processing on disturbed areas, and increases sample points through a random point generation algorithm to improve the accuracy of the yield estimation model, thereby improving the accuracy of crop yield prediction and providing scientific support for agricultural assessment and farmland management.
[0092] 203. Generate the estimated production data for the current year based on the historical production estimation model and the remote sensing images of the current year.
[0093] Among them, the estimated production data for the current year is the output data of the historical production estimation model.
[0094] In this step, input data of the historical yield estimation model is obtained based on the remote sensing image of the current year, and the input data is input into the historical yield estimation model to obtain the estimated yield data of the crop for the current year.
[0095] 204. Automatically construct a meteorological adjustment coefficient based on meteorological data.
[0096] Since most of the existing crop prediction models are built for specific regions, the applicability of existing crop prediction models in other regions is poor, which is not conducive to the promotion and large-scale application of the models. In order to improve the applicability of crop prediction models in different regions, this application constructs a meteorological adjustment coefficient, and realizes accurate prediction of crop yields in different regions by automatically constructing the meteorological adjustment coefficient.
[0097] The meteorological adjustment coefficient of the present application is constructed based on the weight of each meteorological factor. In order to accurately and reasonably assign weights to each meteorological factor, the present invention adopts a random forest algorithm to perform feature screening and measure the importance of each meteorological factor. The meteorological factors are sorted and the feature importance scores are output according to the measurement results. The normalized feature importance scores are used as the impact weights of the meteorological factors.
[0098] In one achievable manner, the specific construction method of the meteorological adjustment coefficient is as follows:
[0099] Step 1: Use the random forest algorithm to evaluate feature importance.
[0100] In this step, the Gini index (Gini) can be used as an evaluation index to measure the impact of each meteorological factor on yield. Variable importance measures (VIM) are represented by VIM, and Gini is represented by GI. Assuming there are J features X1, X2, X3, ... XJ, I decision trees, and C categories, calculate the Gini index score of each feature Xj
[0101] The calculation formula of the Gini index of the i-th tree node q is:
[0102]
[0103] Among them, C means there are C categories, c means the cth category, p qc Indicates the proportion of category C in node q.
[0104] Feature X h The importance of node q in the i-th tree, that is, the change in the Gini index before and after node q is:
[0105]
[0106] in, and They represent the Gini index of the two new nodes (l, r) after the node q branches. If the node where feature Xj appears in decision tree i is set Q, then the importance of Xj in the i-th tree is:
[0107]
[0108] Assume there are a total of I trees, then:
[0109]
[0110] Finally, the obtained results are normalized:
[0111]
[0112] Step 2: Use the normalized feature importance score as the meteorological factor weight.
[0113] Among them, the feature importance score is the Gini index score. The meteorological factor weight is used to indicate the impact of meteorological factors on crop yield.
[0114] Step 3: Obtain the meteorological adjustment coefficient according to the weight of each meteorological factor.
[0115] The inventors have found through research that among all meteorological factors, temperature, total precipitation and solar radiation have a greater impact on crop yields. This application selects the above three meteorological factors to establish a meteorological adjustment coefficient model.
[0116] In this step, the average temperature, total precipitation and solar radiation of the previous year and the current year in the area to be strengthened are compared to obtain the difference ratio of each meteorological factor, and then the meteorological regulation coefficient is obtained according to each difference ratio and its corresponding meteorological factor weight. This application realizes the automatic construction of the meteorological regulation system model through the calculation formula of the meteorological regulation coefficient (Coefficient of Meteorological Regulation Index, CMRI).
[0117]
[0118] Among them, tmp is the annual mean temperature, tpre is the total precipitation, ssrd is the solar radiation, α, β, and γ are the characteristic importance scores of annual mean temperature, total precipitation, and solar radiation to crop yield, respectively.
[0119] The embodiment of the present application comprehensively considers the important influence of geographical location, climate factors, etc. on crop yield, evaluates the influence weights of temperature, precipitation and solar radiation on crop yield through the random forest method, determines the proportion of meteorological factors through feature importance scoring, makes the meteorological adjustment coefficient more reasonable, and realizes accurate yield prediction of crops.
[0120] 205. Generate crop yield forecast results based on the estimated yield data for the current year and the meteorological adjustment coefficient.
[0121] In this step, the crop yield forecast result is generated based on the estimated yield data of the current year output by the historical yield estimation model and the meteorological adjustment coefficient output by the meteorological adjustment system model. That is, this application is based on the historical yield estimation model and the automatically constructed meteorological adjustment coefficient model to obtain the crop yield forecast model of the current year, and the yield forecast of the current year is completed by the forecast model without actual measured data.
[0122] Yield 当年_G=Yield_Modle 上一年 ×Picture 当年 (7)
[0123] Yield 当年_Y =Yield 当年_G ×CMRI (8)
[0124] Among them, Yield 当年_G Yield_Modle is the estimated yield result for the current year. 上一年 For the historical production estimation model, Picture 当年 is the remote sensing image of the year, CMRI is the meteorological adjustment coefficient, Yield 当年_Y The expected production results for the current year.
[0125] In addition, after generating the predicted crop yield results, this application can also verify the accuracy based on the actual yield data of the year and the predicted crop yield results:
[0126]
[0127] Among them, δ is the prediction accuracy, N is the predicted value, and L is the actual value.
[0128] In order to explain the above embodiment in more detail, this application takes the soybean yield remote sensing prediction of Nehe City, Heilongjiang Province as an example to explain the specific implementation method in detail:
[0129] like Figure 3 As shown, when predicting crop yields, preliminary data preparation is required, that is, multivariate data is obtained in advance. Multivariate data includes remote sensing data, yield data, meteorological data, and auxiliary data. The above multivariate data are data for 2022 and 2023. After obtaining the above data, the following steps are performed specifically.
[0130] Step 1: Determine the phenological period and process the multivariate data.
[0131] Step 1a: Determine the phenological period of soybean in Nehe City.
[0132] According to agricultural meteorological observation data and field investigation, the growing period of soybeans in Nehe City is divided into the following table:
[0133] Table 1 Phenological characteristics of soybean in Nehe City
[0134]
[0135] Step 1b: Process multivariate data.
[0136] 1) Acquisition and preprocessing of remote sensing images.
[0137] The best time for estimating soybean yield in Nehe City was determined to be July, August, and September. The remote sensing images without cloud cover in this time period were searched and downloaded. The remote sensing images finally selected are shown in the following table:
[0138] Table 2 Remote sensing image data
[0139]
[0140] The downloaded Sentinel-2 remote sensing images are L2A-level product data. The data has completed radiation correction, geometric correction, and atmospheric correction. The downloaded data only needs to be resampled to 10m and the coordinate system converted.
[0141] 2) Acquisition and preprocessing of meteorological data.
[0142] It is determined that the meteorological factors that affect soybean growth are temperature, precipitation and solar radiation. According to the time phase of remote sensing image data, the monthly temperature, precipitation and solar radiation data of ERA5 in August 2022 and 2023 are downloaded, and the meteorological data are converted and spatially interpolated to obtain meteorological data with a spatial resolution of 10m×10m to match the remote sensing image.
[0143] Step 2: Construct a historical production estimation model based on sample enhancement.
[0144] Step 2a: Determine the disturbed area.
[0145] Based on the remote sensing images of Nehe City in 2023, the NDVI value was calculated using ENVI software, and the disturbed area was determined by the NDVI threshold. In the present invention, NDVI = 0.6 is selected as the threshold of the interference area. When NDVI < 0.6, the area is identified as a disturbed area; when NDVI > 0.6, the area is considered to be a normal area. Figure 3 This is the NDVI change map of the soybean distribution area in Nehe City in 2023. As can be seen from the figure, the red part is the soybean-disaster area, which is distributed in the southwest and northeast of Nehe City.
[0146] Step 2b: Sample strengthening of historical production measured data.
[0147] This application performs sample enhancement on the measured data of the disturbed area, that is, adds enhanced sample points. This application ensures the accuracy of model construction by adding sample volume to improve its precision.
[0148] like Figure 5a As shown, there are 24 measured data on soybean yield in Nehe City in 2022, distributed in 9 plots throughout the city. Combined with the above threshold judgment of the disturbed area, it is found that there is measured data distribution in the northeastern disturbed area of Nehe City, and there is no measured data in the southwestern disturbed area. Therefore, the embodiment of this application mainly carries out sample strengthening for the southwest region.
[0149] like Figure 5b As shown in the figure, based on the remote sensing image of Nehe City, ArcGIS was used to select the interference area and form a shp mask file to facilitate the subsequent random sample generation.
[0150] The embodiment of the present application is based on the shp mask file, creates random samples in the area through a random point generation algorithm, and determines the proportion of new sample points based on the random point proportion principle according to the area of the disturbed area, the original measured production data, etc. Figure 6 As shown, the present invention adds a total of 10 enhanced sample points in the interfered area.
[0151] The longitude and latitude coordinates of the enhanced sample points are used to extract the predicted soybean yield in Nehe City in 2022 as the newly added measured soybean yield data. Combined with the existing 24 measured yield data, the integration of the measured soybean yield data in Nehe City in 2022 is completed. There are a total of 34 sample points in this model construction.
[0152] Step 2c: Construct a historical soybean yield estimation model for Nehe City.
[0153] The present invention is based on the preprocessed 2022 remote sensing image, and completes the calculation of the vegetation index based on the reflectivity value as shown in Table 3 below.
[0154] Table 3 Calculation of vegetation index
[0155]
[0156] Through sensitivity analysis, the three most sensitive vegetation indices in Table 3 were screened out, and a historical soybean yield estimation model for Nehe City in 2022 based on vegetation indices was constructed.
[0157] Yield_Modle 2022 =419.971×PSRI-49.0902×GNDVI+82.2164×SIPI+125.9 14(10)
[0159] Step 3: Automatically construct the meteorological adjustment coefficient.
[0160] The embodiment of the present application is based on the preprocessed meteorological data, and first calculates the ratio difference of temperature, precipitation and solar radiation in Nehe City in 2022 and 2023.
[0161]
[0162] Secondly, the feature importance was used to evaluate the impact weights of the three meteorological factors on crop yields. The obtained feature importance values were normalized and used as the adjustment coefficients of temperature, precipitation and solar radiation, respectively. The results showed that the adjustment coefficients of temperature, precipitation and solar radiation on soybean yield in Nehe City in August were 0.5, 0.3 and 0.2, respectively.
[0163]
[0164] Among them, VIM is the feature importance score, Gini is the Gini index, tmp is temperature, tpre is precipitation, and ssrd is solar radiation.
[0165] The above results are calculated through the meteorological adjustment coefficient formula, and the feature importance score is introduced into the formula to complete the construction of the meteorological adjustment model of Nehe City. The final calculation result is used as the meteorological adjustment coefficient CMRI of Nehe City in 2023.
[0166] CMRI 2023 =tmp×0.5+tpre×0.3+ssrd×0.2 (17)
[0167] Step 4: Construct a crop yield prediction model based on sample enhancement.
[0168] The present invention obtains the historical yield estimation model Yield_Model 2022 Applied to the remote sensing images of Nehe City in 2023, the soybean yield of Nehe City in 2023 was estimated based on the historical yield estimation model, and then the meteorological adjustment coefficient CMRI of Nehe City in 2023 was introduced. 2023, Constructing a soybean yield forecasting model for Nehe City in 2023 2023 The final result is the predicted soybean production in Nehe City in 2023.
[0169] Yield 2023_G =Yield_Modle 2022 ×Picture 2023 (18)
[0170] Yield 2023_Y =Yield 2023_G ×CMRI 2023 (19)
[0171] Among them, Yield 2023_G The soybean production estimate for 2023 is as follows: 2023_Y For the 2023 soybean production forecast, CMRI 2023 is the meteorological adjustment coefficient in 2023, Yield_Modle 2022It is the historical production estimation model for 2022. Picture 2023 It is the remote sensing image of 2023.
[0172] Step 5: Model accuracy verification.
[0173] In order to effectively verify and evaluate the application method, the accuracy of the predicted yield results was verified based on the measured soybean yield data of Nehe City in 2023. The calculation formula is as follows:
[0174]
[0175] Among them, δ is the prediction accuracy, N is the predicted value, and L is the actual value.
[0176] The verification results are shown in Table 4:
[0177] Table 4 Model accuracy verification results
[0178]
[0179] The average accuracy is 91.74%, which meets the requirements of practical applications.
[0180] Furthermore, as a response to the above Figure 1-2 The implementation of the method embodiment shown in the figure, the embodiment of the present application provides a crop yield prediction device, which is used to improve the accuracy of crop yield prediction. The embodiment of the device corresponds to the aforementioned method embodiment. For ease of reading, this embodiment will not repeat the details of the aforementioned method embodiment one by one, but it should be clear that the device in this embodiment can correspond to all the contents of the aforementioned method embodiment. Specifically, Figure 7 As shown, the device comprises:
[0181] A first determining unit 31 is used to determine the disturbed area in the area to be predicted based on the remote sensing image of the area to be predicted in that year;
[0182] A second determining unit 32, configured to determine a plurality of enhanced sample points of the interfered area obtained by the first determining unit 31;
[0183] A construction unit 33 is used to construct a historical production estimation model by using the enhanced samples obtained by the second determination unit 32 and the historical production measured data of the area to be predicted;
[0184] The output unit 34 is used to output the predicted crop yield results based on the historical yield estimation model constructed by the construction unit 33, the remote sensing image of the year, and the meteorological adjustment coefficient automatically constructed based on meteorological data.
[0185] Further, such as Figure 8 As shown, the first determining unit 31 includes:
[0186] The calculation module 311 is used to calculate the normalized difference vegetation index NDVI of each block in the area to be predicted based on the remote sensing image of the year;
[0187] A first determination module 312 is configured to determine that the land parcel is a disturbed area if the NDVI obtained by the calculation module 311 is less than a disturbed threshold;
[0188] The second determination module 313 is used to determine that the land parcel is a normal area if the NDVI obtained by the calculation module 311 is greater than or equal to the interference threshold.
[0189] Further, such as Figure 8 As shown, the second determining unit 32 includes:
[0190] The judgment module 321 is used to traverse each disturbed area and judge whether the disturbed area contains historical production measured data;
[0191] The determination module 322 is used to determine that the disturbed area is an area to be strengthened if the determination module 321 determines that it does not contain;
[0192] A generating module 323, used to generate random sample points in the area to be enhanced determined by the determining module 322 using a random point generating algorithm;
[0193] The screening module 324 is used to screen out enhanced sample points from the random sample points generated by the generating module 323 according to the quantity of the historical production measured data and the principle of random point proportion.
[0194] Further, such as Figure 8 As shown, the construction unit 33 includes:
[0195] An extraction module 331 is used to determine the longitude and latitude coordinates of each enhanced sample point and extract the historical production prediction data corresponding to the longitude and latitude coordinates;
[0196] A sample generation module 332 is used to use the historical production forecast data obtained by the extraction module 331 and the historical production measured data as sample data;
[0197] The construction module 333 is used to construct a historical yield estimation model based on the target vegetation index using the sample data and the target vegetation index obtained by the sample generation module 332.
[0198] Furthermore, the construction unit 33 is specifically used for:
[0199] Calculating various historical vegetation indices based on the historical remote sensing images of the area to be predicted;
[0200] Conduct sensitivity analysis on each historical vegetation index, and select the N historical vegetation indices with the highest sensitivity from all historical vegetation indices, where N is a positive integer;
[0201] The N historical vegetation indices are determined as the target vegetation indices.
[0202] Further, such as Figure 8 As shown, the meteorological data includes the meteorological data of the current year and the meteorological data of the previous year, and the output unit 34 includes:
[0203] A generating module 341 is used to generate the estimated production data for the current year based on the historical production estimation model and the remote sensing images of the current year;
[0204] A construction module 342, for automatically constructing the meteorological adjustment coefficient according to the meteorological data of the current year and the meteorological data of the previous year;
[0205] The generation module 343 is used to generate the predicted yield result of the crop by using the estimated yield data for the current year generated by the generation module 341 and the meteorological adjustment coefficient obtained by the construction module 342.
[0206] Furthermore, the construction module 342 is specifically used for:
[0207] Extracting the meteorological factor data of each meteorological factor of the current year from the meteorological data of the current year, and extracting the meteorological factor data of each meteorological factor of the previous year from the meteorological data of the previous year, wherein the meteorological factors include: temperature, precipitation, and solar radiation;
[0208] For each meteorological factor, the ratio of the meteorological factor data of the current year to the meteorological factor data of the previous year is calculated to obtain the difference ratio of the meteorological factor;
[0209] The Gini index based on the random forest algorithm evaluates the importance of each meteorological factor and obtains the meteorological factor weight, which is used to indicate the degree of influence of the meteorological factor on the crop yield.
[0210] The meteorological adjustment coefficient is obtained according to each difference ratio and its corresponding meteorological factor weight.
[0211] Furthermore, the embodiment of the present application also provides a processor, the processor is used to run a program, wherein the program executes the above Figure 1-3 The crop yield prediction method described in .
[0212] Furthermore, an embodiment of the present application further provides a storage medium, wherein the storage medium is used to store a computer program, wherein when the computer program is running, the device where the storage medium is located is controlled to execute the above Figure 1-3 The crop yield prediction method described in .
[0213] Furthermore, the embodiment of the present application provides an electronic device 4, as shown in FIG5, the device includes at least one processor 41, and at least one memory 42 and a bus 43 connected to the processor 41; wherein the processor 41 and the memory 42 communicate with each other through the bus 43; the processor 41 is used to call the program instructions in the memory 42 to execute the above-mentioned crop yield prediction method. The device in this article can be a server, a PC, a PAD, a mobile phone, etc.
[0214] Furthermore, the present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing a program initialized with the steps of the crop yield prediction method as described above.
[0215] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0216] It is understandable that the related features in the above methods and devices can be referenced to each other. In addition, the "first", "second" and the like in the above embodiments are used to distinguish the embodiments, but do not represent the advantages and disadvantages of the embodiments.
[0217] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0218] The algorithm and display provided herein are not inherently related to any particular computer, virtual system or other device. Various general purpose systems can also be used together with the teachings based on this. According to the above description, it is obvious that the structure required for constructing such systems. In addition, the application is not directed to any specific programming language either. It should be understood that various programming languages can be utilized to realize the content of the application described herein, and the description of the specific language above is for the purpose of disclosing the best mode of implementation of the application.
[0219] In addition, the memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0220] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0221] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0222] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0223] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0224] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0225] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0226] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0227] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0228] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0229] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. A method for predicting crop yield, characterized in that: The method comprises: Determine the disturbed area in the area to be predicted based on the remote sensing image of the area to be predicted in the current year; When the disturbed area does not contain the historical production measured data, a random point generation algorithm is used to generate random sample points, and a plurality of enhanced sample points in the disturbed area are determined from the random sample points, wherein the random sample points are generated in the following manner: a random number data stream is created through a random number generator and a seed, and a polygon is created in the area to be enhanced; a standard polygon partitioning algorithm is used to divide the polygon according to triangles of different sizes; a triangle in the polygon is randomly selected, and two sides of the triangle are used as the basis axes X and Y for placing random points; a point is randomly selected on the basis axis, and the next unused value in the selected random number data stream is transformed into a uniformly distributed value to obtain a new random point; the first random sample point on the basis axes X and Y is determined based on the two random points obtained, and the above steps are repeated until a specified number of random sample points are reached, wherein the area to be enhanced is a disturbed area that does not contain the historical production measured data; Using the enhanced sample points and the historical yield measured data of the area to be predicted, a historical yield estimation model is constructed, including: determining the longitude and latitude coordinates of each enhanced sample point, and extracting the historical yield prediction data corresponding to the longitude and latitude coordinates; using the historical yield prediction data and the historical yield measured data as sample data; using the sample data and the target vegetation index to construct a historical yield estimation model based on the target vegetation index; The predicted crop yield results are output based on the historical yield estimation model, the remote sensing images of the current year, and the meteorological adjustment coefficient automatically constructed based on meteorological data.
2. The method according to claim 1, characterized in that Based on the remote sensing image of the area to be predicted in the current year, the disturbed area in the area to be predicted is determined, including: Calculate the Normalized Difference Vegetation Index (NDVI) of each block in the area to be predicted based on the remote sensing images of the year; If the NDVI is less than the disturbance threshold, the plot is determined to be a disturbed area; If the NDVI is greater than or equal to the interference threshold, the plot is determined to be a normal area.
3. The method according to claim 1, characterized in that Determining a plurality of enhanced sample points in the disturbed area includes: Traverse each disturbed area to determine whether the disturbed area contains historical production measured data; If not, the disturbed area is determined to be an area to be strengthened; Generate random sample points in the area to be strengthened using a random point generation algorithm; According to the quantity of the historical production actual measurement data and the principle of random point proportion, enhanced sample points are selected from the random sample points.
4. The method according to claim 3, characterized in that The target vegetation index is generated as follows: Calculating various historical vegetation indices based on the historical remote sensing images of the area to be predicted; Conduct sensitivity analysis on each historical vegetation index, and select the N historical vegetation indices with the highest sensitivity from all historical vegetation indices, where N is a positive integer; The N historical vegetation indices are determined as the target vegetation indices.
5. The method according to claim 1, characterized in that The meteorological data includes the meteorological data of the current year and the meteorological data of the previous year. Based on the historical yield estimation model, the remote sensing image of the current year, and the meteorological adjustment coefficient automatically constructed based on the meteorological data, the predicted yield results of crops are output, including: Generate the estimated production data for the current year based on the historical production estimation model and the remote sensing images of the current year; Automatically construct the meteorological adjustment coefficient according to the meteorological data of the current year and the meteorological data of the previous year; The crop yield forecast result is generated by using the estimated yield data for the current year and the meteorological adjustment coefficient.
6. The method according to claim 5, characterized in that The meteorological adjustment coefficient is automatically constructed based on the meteorological data of the current year and the meteorological data of the previous year, including: Extracting the meteorological factor data of each meteorological factor of the current year from the meteorological data of the current year, and extracting the meteorological factor data of each meteorological factor of the previous year from the meteorological data of the previous year, wherein the meteorological factors include: temperature, precipitation, and solar radiation; For each meteorological factor, calculate the ratio of the meteorological factor data of the current year to the meteorological factor data of the previous year to obtain the difference ratio of the meteorological factor; The Gini index based on the random forest algorithm evaluates the importance of each meteorological factor and obtains the meteorological factor weight, which is used to indicate the degree of influence of the meteorological factor on the crop yield. The meteorological adjustment coefficient is obtained according to each difference ratio and its corresponding meteorological factor weight.
7. A crop yield prediction device, characterized in that: The device comprises: A first determining unit is used to determine the disturbed area in the area to be predicted based on the remote sensing image of the area to be predicted in the current year; The second determination unit is used to generate random sample points using a random point generation algorithm when the disturbed area does not contain historical production measured data, and determine multiple enhanced sample points of the disturbed area from the random sample points, wherein the random sample points are generated in the following manner: a random number data stream is created by a random number generator and a seed, and a polygon is created in the area to be strengthened; a standard polygon partitioning algorithm is used to divide the polygon according to triangles of different sizes; a triangle in the polygon is randomly selected, and two sides of the triangle are used as the basis axes X and Y for placing random points; a point is randomly selected on the basis axis, and the next unused value in the selected random number data stream is transformed into a uniformly distributed value to obtain a new random point; a first random sample point on the basis axes X and Y is determined based on the two random points obtained, and the above steps are repeated until a specified number of random sample points are reached, wherein the area to be strengthened is a disturbed area that does not contain historical production measured data; A construction unit is used to construct a historical yield estimation model using the enhanced sample and the historical yield measured data of the area to be predicted, including: determining the longitude and latitude coordinates of each enhanced sample point, and extracting the historical yield prediction data corresponding to the longitude and latitude coordinates; using the historical yield prediction data and the historical yield measured data as sample data; using the sample data and the target vegetation index to construct a historical yield estimation model based on the target vegetation index; The output unit is used to output the predicted crop yield results based on the historical yield estimation model, the remote sensing image of the year, and the meteorological adjustment coefficient automatically constructed based on meteorological data.
8. A storage medium, characterized in that: The storage medium is used to store a computer program, wherein when the computer program is running, it controls the device where the storage medium is located to execute the crop yield prediction method according to any one of claims 1 to 6.
9. An electronic device, characterized in that: The device includes at least one processor, and at least one memory and a bus connected to the processor; wherein the processor and the memory communicate with each other through the bus; the processor is used to call program instructions in the memory to execute the crop yield prediction method as described in any one of claims 1-6.
Citation Information
Patent Citations
Crop yield estimation method and device, storage medium and electronic equipment
CN115545311A
Meteorological-combined remote sensing prediction method and device for crop yield in northwest region
CN116187525A
Crop yield prediction method and device, electronic equipment and storage medium
CN116757338A