A soybean identification method based on a standard curve

By constructing a standard curve and using a random forest classifier, the problem of low soybean identification accuracy in remote sensing technology was solved, and high-precision and automated soybean planting area monitoring was achieved.

CN116824384BActive Publication Date: 2025-10-17HANGZHOU NORMAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310197975.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-03
Publication Date
2025-10-17
Estimated Expiration
2043-03-03

AI Technical Summary

Technical Problem

Existing remote sensing technology has problems in monitoring soybean planting area, such as low recognition accuracy, significant impact of cloud and rain pollution, inability to reflect crop changes, and difficulty in obtaining samples.

Method used

A standard curve-based method was used to process satellite data, construct time series normalized phenological index characteristics, extract sample point standard curves, calculate phenological parameter characteristic images, and use random forest classifiers for soybean identification.

Benefits of technology

It achieves high accuracy and automation in soybean identification, has certain predictive functions, and can adapt to and improve identification accuracy in different climatic zones.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116824384B_ABST
    Figure CN116824384B_ABST
Patent Text Reader

Abstract

The application relates to a soybean identification method based on a standard curve, which comprises the following steps: S1, satellite data processing; S2, construction of a time series normalized vegetation index feature: calculating a vegetation index to obtain a vegetation index image, and using a time series to construct a time series vegetation index feature for the vegetation index image; S3, construction of a sample point standard curve: forming a sample point standard curve according to sample point time series vegetation index features, positions and meteorological data factors; S4, sample point standard curve parameter information extraction and vegetation feature calculation: extracting main vegetation information in the sample point standard curve; calculating and generating a vegetation parameter feature image for other pixels; S5, crop variety identification and precision evaluation: inputting sample information into a classifier to obtain a sample classifier, inputting other pixel data into the sample classifier for identification, and using a classification confusion matrix to calculate identification precision. The application has high automation and high precision.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a soybean identification method based on a standard curve, and mainly applies to evaluation and monitoring of a soybean planting area. BACKGROUND

[0002] Soybean is a high-protein food, a main raw material for livestock feed, and an important source of edible oil, and plays an important role in world food production. The traditional agricultural survey method for estimating the soybean planting area is relatively time-consuming and labor-intensive, is easily affected by subjective factors, and the result data cannot provide specific spatial distribution information. Remote sensing technology can more timely, efficiently and objectively realize large-scale crop planting area monitoring, and is low in cost. There are four types of methods for identifying soybean by using mainstream remote sensing technology at present.

[0003] (1) Soybean is identified by using spectral features, that is, each band of satellite data is used as a classification feature to identify soybean, and some people find that the red edge band and near-infrared band of satellite data such as Sentinel-2 data can well distinguish soybean and corn. This method uses a single time phase image, and when the research area is large (above the provincial scale), the Sentinel-2 single time phase image missing due to pollution such as cloud and rain cannot completely cover the entire research area, so that subsequent research cannot be carried out. The single time phase image also cannot reflect the change process of crops, especially crops with similar reflectivity values in the same period are easily confused, leading to the inability to identify, and due to the noise of the original reflectivity value of the satellite sensor, lens anomaly and the like, errors may be caused.

[0004] (2) Soybean is identified by using vegetation index features, that is, each band of satellite data is used to calculate a vegetation index to identify soybean, such as enhanced vegetation index EVI, normalized difference vegetation index NDVI and the like. This method uses a single time phase image, and when the research area is large (above the provincial scale), the Sentinel-2 single time phase image missing due to pollution such as cloud and rain cannot completely cover the entire research area, so that subsequent research cannot be carried out. The single time phase image also cannot reflect the change process of crops, especially crops with similar reflectivity values in the same period are easily confused, leading to the inability to identify.

[0005] (3) Soybean is identified by using time series features, such as using time series satellite data spectral bands, time series vegetation indexes and the like, time series methods have a percentage number calculation to synthesize a feature image, that is, a plurality of continuous satellite image data within a period of time are processed into a time series image data, which can show the characteristics changing with time. This method only processes satellite data and vegetation indexes, and the growth characteristics of soybean cannot be visualized and analyzed, the key growth period for distinguishing soybean cannot be found, and the method does not have interpretability, cannot solve the problem of difficult sample acquisition, has few related parameters, and is low in soybean identification accuracy.

[0006] (4) Using phenological characteristics to identify soybeans, the key parameters are obtained by processing time series vegetation index. This method only uses a few key phenological points, which is not universal, cannot solve the problem of sample acquisition difficulty, cannot adapt to different climate zones, cannot identify soybeans in areas without samples, uses fewer technical parameters, and has low soybean identification accuracy. SUMMARY

[0007] The technical problem solved by the present application is to overcome the above-mentioned deficiencies in the prior art, and to provide a standard curve-based soybean identification method that can accurately identify soybeans and has a certain prediction function.

[0008] The technical solution adopted by the present application to solve the above technical problem is: a standard curve-based soybean identification method, characterized by comprising the following steps

[0009] Step S1, satellite data processing:

[0010] The satellite data processing step includes obtaining all relevant satellite data of soybeans in the study area during the growth period, and performing cloud control and cloud removal processing on the satellite data.

[0011] Step S2, constructing time series normalized phenological index characteristics:

[0012] The processed satellite data is used to calculate the vegetation index of all pixels in the study area to obtain a vegetation index image of all pixels. The time series method is used to process the vegetation index image of each pixel to construct the time series vegetation index characteristics of the pixel. The basic unit of this step is a pixel of satellite data, and the calculation processing method is pixel by pixel.

[0013] Step S3, constructing sample point standard curve:

[0014] There are two methods to construct the sample point standard curve. One is to directly use the sample point standard curve of the previous year or the average value of the sample point over the years. The other method is: obtain the location of the typical representative sample point (referred to as sample point) of crops including soybeans, use the sample point time series vegetation index characteristics obtained in step S2 to calculate the sample point time series curve, and use the sample point time series curve and add meteorological data factors to construct the sample point standard curve.

[0015] Step S4, sample point standard curve parameter information extraction and phenological characteristics calculation:

[0016] The sample point standard curve is extracted from the soybean sample point key growth period information to obtain the main sample point information (sample information) ; the soybean sample point key growth period information (9 main phenological feature parameters) of all pixels (referred to as other pixels) in the study area is calculated, and the phenological parameter feature image is synthesized.

[0017] Step S5, crop variety identification and precision evaluation:

[0018] The sample information is input into the classifier to obtain a sample classifier, the phenological parameter feature image of all pixels except the sample points is input into the sample classifier for identification and output of the identification result, and the identification accuracy is calculated using the classification confusion matrix.

[0019] The cloud amount control of the satellite data in step S1 means that only (screening) images with a total cloud amount less than 30% are used; the cloud removal processing means removing outliers, for example, using NA to assign values to pixels in clouds, snow and shadows, to achieve the effect of removing outliers.

[0020] In step S2, the vegetation index used is the normalized phenological vegetation index, which is not sensitive to soil and the reflectivity of the short-wave infrared band included in the normalized phenological vegetation index can capture the change in tree crown water content, so that the normalized phenological vegetation index has good effect in soybean identification, and the normalized phenological vegetation index calculation formula is as follows:

[0021]

[0022] In the formula, NIR is the near-infrared band reflectivity, RED is the red band reflectivity, and SWIR is the short-wave infrared band reflectivity.

[0023] In step S2, the time series method is harmonic fitting, and the linear harmonic model uses a series of sine and cosine waves to fit complex signals, which has been widely used in fitting of vegetation index time series, and the linear harmonic model calculation formula is as follows:

[0024]

[0025] Where f(t) is the fitted vegetation index value at time t; a is a constant term; b is the coefficient of the first-order term; M is the number of harmonic combinations, in this application, M is 2, i.e. a two-term linear harmonic model; C and D are the coefficients of the cosine function and the sine function; ω is the reciprocal of the number of days in a year (1 / 365); t is DOY (a day in a year).

[0026] In step S3, the time series vegetation index features obtained in step S2 are used to calculate the time series curve of the crop representative sample point, i.e., the sample point position is corresponded to the sample point time series vegetation index features and unified together, the change of the time series vegetation index pixel of the position is visualized to form a sample point time series curve, the horizontal coordinate is Julian day DOY (day), and the vertical coordinate is the value of normalized phenology vegetation index.

[0027] In step S3, the sample point time series curve is used to construct a sample point standard curve by adding meteorological data factors, i.e., the meteorological data is added to the sample point time series curve to obtain the standard curve of the pixel (which may be soybean, may be corn or other crops), and the formula of the sample point standard curve is:

[0028]

[0029] wherein, Part is a harmonic fitting formula, Temp is a temperature variable, hour is a sunshine duration variable, rain is a rainfall variable, and c0, c1, and c2 are meteorological weight factors.

[0030] In step S4, the soybean sample point key growth period information in the sample point standard curve is extracted to obtain sample information, and the soybean sample point key growth period information adopts main phenology point parameter information (9 parameters), and the 9 parameters are respectively the Julian day DOY values of three nodes in the soybean growth period, the Julian day DOY values of three nodes in the senescence period, and the lengths corresponding to the three soybean growth stages.

[0031] The main phenology point (time point) is determined by the following formula:

[0032] u=(V min -V max )×p+V min

[0033] wherein, u is a reference threshold value, used to determine the soybean sample point start of season (SOS) and the soybean sample point end of season (EOS), V min , and V maxThe minimum and maximum values in the time series of the soybean sample point, respectively, p is the amplitude ratio, p has three values, the first value corresponds to the growth start node of the soybean sample point, the soybean sample point dormancy start node, the second value corresponds to the soybean sample point growth mid-node, the soybean sample point senescence mid-node, and the third value corresponds to the soybean sample point maturity node, and the soybean sample point senescence start node. In the present application, p is 0.15, 0.5 and 0.90, respectively, so 6 main phenological point information is obtained, which is the Julian day DOY value of the three nodes of the soybean sample point growth period, and the Julian day DOY value of the three nodes of the senescence period.

[0034] In the present application, the soybean sample point phenological start time (SOS) and the soybean sample point phenological end time (EOS) are respectively designated as the first day and the last day exceeding the threshold value u, and u is a dynamic value which depends on the annual amplitude of the soybean sample point time series.

[0035] SOS = first (f (x) > u)

[0036] EOS = last (f (x) > u)

[0037] Wherein, f(t) is the formula of the soybean sample point standard curve in step S3, first() is the Julian day DOY corresponding to the first value of the soybean sample point exceeding the threshold value, and last() is the Julian day DOY corresponding to the last value of the soybean sample point exceeding the threshold value.

[0038] And calculate the length length of the soybean sample point phenological start time (SOS) and the soybean sample point phenological end time (EOS) when p is 0.15, 0.5 and 0.90, respectively, and the length calculation formula is:

[0039] length = EOS-SOS

[0040] Wherein, EOS is the Julian day DOY of the soybean sample point phenological start, SOS is the Julian day DOY of the soybean sample point phenological end, and length is the length of the soybean sample point growth stage.

[0041] In step S4, the method for calculating the key growth period information of all pixels in the study area range except the sample points and synthesizing the phenological parameter feature image is: calculating the phenological parameter feature image of all pixels in the study area range except the sample points at the main phenological point of the soybean sample point with the closest growth condition, and a total of 9 feature images are taken as the phenological parameter features.

[0042] In the step S5, when the phenological parameter feature image of a certain pixel is input into the sample classifier for identification, the sample point data with the closest growth condition is used for comparison and identification (for example, first compared with the soybean sample point with the closest growth condition, if the similarity between the pixel data and the soybean sample point data meets the standard, the pixel data is marked as soybean, otherwise, if it needs to be compared with the soybean reference, if it does not need to be compared with the soybean reference, it is directly marked as other), the classification result is output, the soybean is extracted, and the accuracy is evaluated using the classification confusion matrix, and the evaluation indexes are overall accuracy (OA), plot accuracy (PA), and user accuracy (UA).

[0043] Compared with the prior art, the present application has the following advantages and effects: convenient to use, high degree of automation, and accurate identification of soybeans. BRIEF DESCRIPTION OF DRAWINGS

[0044] Figure 1 is a flowchart of an embodiment of the present application.

[0045] Figure 2 is a comparison diagram of the soybean 2020 standard curve (prediction curve) and the actual 2020 time curve.

[0046] Figure 3 is a comparison diagram of the soybean normalized phenological vegetation index disaster curve and the soybean standard curve.

[0047] Figure 4 is a diagram showing the design principle of the soybean normalized phenological vegetation index and key phenological points. DETAILED DESCRIPTION

[0048] The present application will be further described in detail below in conjunction with the drawings and through embodiments, and the following embodiments are an explanation of the present application, but the present application is not limited to the following embodiments.

[0049] Referring to Figures 1-4 , the present application is mainly applicable to the evaluation and monitoring of soybean planting area. By fully utilizing resource satellite remote sensing data and combining the soybean identification method of the present application, it is determined whether the agricultural product planting type of a certain block (the image of the block is composed of a large number of pixels, such as the pixel spatial resolution of the Sentinel 2 image is 10 m, each pixel is 10 m*10 m in size, and each pixel has relatively independent parameters, and the calculation involved in the present application is pixel-by-pixel calculation, i.e., each pixel constituting the image is calculated separately) is soybean, the size and distribution of the soybean planting area, and the like, thereby providing basic information for scientific prediction and monitoring of the soybean planting area, type, growth condition, yield size, and pest situation. The soybean identification in Heilongjiang Province will be described below as an example.

[0050] The soybean sowing area and yield of Heilongjiang Province account for 48.90% and 46.95% of the national soybean sowing area and yield, respectively, and it is a major soybean production area in China. Therefore, soybean in Heilongjiang Province is representative as an object of the soybean identification method.

[0051] Step S1, satellite data processing;

[0052] The satellite data is Sentine1-2 MSI (MultiSpectral Instrument, Level-2A) data from the Google Earth Engine cloud platform. The data is a surface reflectance product that has been atmospherically corrected. No preprocessing step is required. Cloud cover less than 30% is screened out using a cloud cover screening function for all available images from 2019 to 2020. Abnormal values are removed using a cloud removal function specific to the Sentine1-2 satellite data. Finally, a total of 7026 satellite images are obtained, which completely cover the entire soybean growing season in Heilongjiang Province.

[0053] Step S2, constructing a time series normalized vegetation index (referred to as a time series normalized vegetation index);

[0054] First, the normalized vegetation index is calculated using the processed satellite data to synthesize a vegetation index image. The normalized vegetation index calculation formula is as follows:

[0055]

[0056] In the formula, NIR is the near-infrared band reflectance, RED is the red band reflectance, and SWIR is the short-wave infrared band reflectance.

[0057] Then, the time series normalized vegetation index (i.e., the normalized vegetation index with time series characteristics) is calculated using a harmonic fitting method. The harmonic fitting formula is as follows:

[0058]

[0059] In the formula, f(t) is the fitted vegetation index value at time t; a is a constant term; b is the coefficient of the first-order term; M is the number of harmonic combinations, which is 2 in this application, i.e., a two-term linear harmonic model; C and D are the coefficients of the cosine function and the sine function; ω is the reciprocal of the number of days in a year (1 / 365); and t is DOY (day of the year).

[0060] Step S3, constructing a sample point standard curve:

[0061] First, typical representative sample points of soybeans and crops as soybean reference (crops growing in the same season as soybeans, such as corn and / or rice, soybean reference can be but not necessarily) are obtained, and each sample point of each crop is composed of several pixels and only one crop is planted. The field sample point data used in this example is mainly 315 soybean sample points, 498 corn samples, and 362 other crop sample points in Heilongjiang Province in 2019, and 1000 soybean sample points in 2020. After obtaining the complete data of a certain growth season of the sample point in this application, it can be used for many years without repeated sampling, or it can be adjusted to the data of the last year or the average data of several years; In special cases (such as newly built large soybean fields or a certain sample point is changed for other purposes), the sample point can also be adjusted accordingly. Then, the time series curve of the sample point is calculated using the S2 step time series normalized vegetation index feature and the typical representative sample point position of soybeans and other crops (here, the same as the soybean reference, the same below) on the GEE cloud platform, and the ui.Chart.image.doySeriesByRegion() function is used for drawing. Finally, the time series curve is used to construct the sample point standard curve by adding meteorological data factors, and the meteorological data uses the temperature and precipitation data in the ERA5 reanalysis data (ERA5-Land Hourly-ECMWF Climate Reanalysis). The precipitation data of each day is accumulated for 24 hours to obtain the daily precipitation, and the average value of the temperature data of each day is taken to obtain the daily temperature. According to the solar elevation angle, azimuth angle and atmospheric transmittance, the sunshine duration is calculated, and the temperature, precipitation and sunshine duration are added to the time series processing harmonic fitting to obtain the sample point standard curve (referred to as standard curve, including soybean sample point standard curve, corn sample point standard curve, and other crops can also be included).

[0062] In addition, in order to prove the robustness of the obtained standard curve, the curve is compared and verified from the aspects of time scale and disaster area. First, the verification in time scale, taking 2019 as the reference year and 2020 as the target year, using the normalized vegetation index fitting time series curve in 2019, and then adding climate factors using ERA5 meteorological data in 2020, finally obtaining the predicted standard curve of 2020. Then, the normalized vegetation index fitting time series curve of the actual 2020 is used, which is recorded as the actual 2020 curve. The comparison chart of the standard curve and the actual curve is as follows Figure 2As shown, the two curves are very close in trend, and in comparison, the actual curve has a certain up and down fluctuation trend in some time periods. The figure shows that the standard curve can achieve better prediction effect by adding the climate data of the target year, and the actual median fitting curve of the target year is very similar. Then the disaster area is verified, and the standard curve of 2019 flood disaster meteorological factor is compared with the actual curve of 2020, and the comparison chart is as follows Figure 3 As shown, the two are basically consistent in trend, and the standard curve is higher than the disaster curve in value due to the influence of large change meteorological data, but the two are very consistent in the rising and falling stages of the curve, that is, they are very close in the key soybean growth stage, which shows that the standard curve method has a certain robustness.

[0063] The specific formula of the standard curve is as follows:

[0064]

[0065] Wherein, Part of the harmonic fitting formula, Temp is the temperature variable, hour is the sunshine duration variable, rain is the rainfall variable, c0, c1, c2 are meteorological weight factors.

[0066] The sunshine duration calculation formula is as follows:

[0067] h = ∑ (τ i > τ min ) x Δt i

[0068] Wherein, ∑ represents the summation of the value at each time; τ i represents the atmospheric transmittance at the i th time; τ min represents the threshold atmospheric transmittance, that is, the minimum transmittance at which the sun can be considered to be “directly” incident on the ground, usually 0.1-0.3; Δt i represents the time interval at the i th time.

[0069] Step S4, sample point standard curve parameter information extraction and phenology feature calculation; the key growth period information in the sample point standard curve (6 main phenology point information of soybean, growth period length, middle growth period length and mature period length are taken in this embodiment) is extracted to obtain the main phenology information of soybean or soybean reference (referred to as sample information or sample) of sample point; the key growth period information of all pixels in the research area range except the sample point is calculated, and the phenology parameter feature image is synthesized.

[0070] The method for extracting the key growth period information in the sample point standard curve to obtain the soybean or soybean reference crop information of the sample point is as follows: the key growth period main phenological point parameter information in the sample point standard curve is extracted as a sample, the main phenological points are the Julian day DOY values of the three nodes of the soybean growth period, the Julian day DOY values of the three nodes of the soybean senescence period, a total of 6 main phenological points, and the lengths of the corresponding three growth stages of soybean, a total of 9 parameters. The soybean sample point data (soybean sample) is used as the basis for classifying a certain pixel as soybean in the next step, and all soybean reference (for example, corn) sample point data (soybean reference sample) is used as the basis for classifying a certain pixel as a soybean reference (for example, corn) in the next step. The crop that is neither similar to the soybean sample nor similar to the soybean reference sample is classified as other, according to whether the similarity meets the standard.

[0071] The formula for determining the main phenological points is:

[0072] u=(V mim -V max )×p+V min

[0073] wherein u is a reference threshold value for determining the soybean sample point phenological start time (SOS, such as the growth start time shown in Figure 4 ) and the phenological end time (EOS, such as the dormancy start time in Figure 4 ), V min and V max are the minimum value and the maximum value in the time sequence of the soybean sample point (the two values are the minimum value and the maximum value in the time sequence of the pixel, which are obtained by using the.min() and.max() functions in the Google Earth Engine (Google Earth Engine) cloud platform), and p is the amplitude ratio, p is taken as 0.15 (corresponding to the growth start node of the soybean sample point, the dormancy start node of the soybean sample point), 0.5 (corresponding to the middle growth node of the soybean sample point, the middle senescence node of the soybean sample point) and 0.90 (corresponding to the mature node of the soybean sample point, the senescence start node of the soybean sample point) in the present application, see Figure 4 , so 6 main phenological point information is obtained, which is the Julian day DOY value of the three nodes of the growth period and the Julian day DOY value of the three nodes of the senescence period.

[0074] In the present application, the soybean sample point phenological start time (SOS) and the soybean sample point phenological end time (EOS) are respectively designated as the first day and the last day of the threshold value u, and u is a dynamic value which depends on the annual amplitude of the time sequence of the soybean sample point.

[0075] SOS=first(f(x)>u)

[0076] EOS = last(F(x) > u)

[0077] where f(t) is the formula of the soybean sample point standard curve in step S3, first() is the Julian day DOY corresponding to the first value of the soybean sample point exceeding the threshold, and last() is the Julian day DOY corresponding to the last value of the soybean sample point exceeding the threshold.

[0078] and the length length of the soybean sample point phenological beginning time (SOS) and the soybean sample point phenological end time (EOS) when p is 0.15, 0.5 and 0.90 respectively, the length calculation formula is:

[0079] length = EOS - SOS

[0080] where EOS is the Julian day DOY of the soybean sample point phenological beginning, SOS is the Julian day DOY of the soybean sample point phenological end, and length is the length of the soybean sample point growth stage time.

[0081] The method for calculating the key growth period information of all pixels in the study area and synthesizing the phenological parameter characteristic image is as follows: the method of this step is the same as the sample point parameter information extraction method in step S4, but the application object is different, that is, the Julian day DOY values of the three nodes of the sample point closest to the growth condition (usually the closest sample point is used, on the one hand, it is convenient to calculate, on the other hand, the soil and meteorological conditions are relatively close, or a soybean sample point and a soybean reference sample point can be set in a township or a block, and all pixels in the township or the block are compared with the data of the two sample points) are calculated, the Julian day DOY values of the three nodes of the senescence period are calculated, and the length of the corresponding three soybean sample points growth stage is calculated. All (9) parameters of each pixel are synthesized into a characteristic image, that is, a phenological parameter characteristic image of 9 parameters, which reflects the complete crop information of the pixel. When it is input into the sample classifier and compared with the crop information of the closest soybean sample point and the crop information of the soybean reference sample point, the pixel planted with soybean or soybean reference or other plants can be identified according to the principle of the highest similarity.

[0082] Step S5, soybean extraction (identification) and accuracy evaluation: A random forest classifier is used, the number of trees is 500, and other parameters use the default values. The sample of step S4 is input into the random forest classifier for training to obtain a random forest classifier (sample classifier) ​​with all sample information. The main phenological characteristics of soybeans in the key growth period of all pixels are input into the sample classifier for identification. During identification, the sample point with the closest growth conditions is used for comparison. If the similarity between the data of a certain pixel and the data of the soybean sample point with the closest growth conditions meets the standard (for example, 90%), it is identified as soybean. If the similarity between the data of a certain pixel and the data of the corn sample point with the closest growth conditions meets the standard (for example, 90%), it is identified as corn. The identification results are marked on the soybean planting distribution map of Heilongjiang Province in that year by color (in this example, only soybeans are marked as green, and other plants are marked as khaki. If necessary, corn or other crops can also be marked). At the same time, the random forest classification confusion matrix is ​​calculated to obtain four indicators: overall accuracy (OA), mapping accuracy (PA), user accuracy (UA), and F1-Score. The random forest classification confusion matrix is ​​defined as follows:

[0083]

[0084] In formula (6), n represents the number of categories, m ij The classification confusion matrix represents the number of pixels that actually belong to class i but are classified into class j in the classification result map. The values ​​of the diagonal elements are the number of pixels correctly classified in each category. Therefore, the larger the value of the diagonal elements in the classification confusion matrix, the more pixels are correctly classified and the more reliable the classification result. The classification confusion matrix is ​​used to calculate the overall accuracy (OA), mapping accuracy (PA), user accuracy (UA) and F1-Score of this application. The calculation formulas are:

[0085]

[0086]

[0087]

[0088]

[0089] Among them, m i+ is the row sum in the classification confusion matrix, m +i is the column sum in the classification confusion matrix.

[0090] The overall precision of soybeans in Heilongjiang Province in 2020 extracted according to the above method is 86.95%, the user precision is 90.91%, the mapping precision is 86.14%, the F1-Score is 0.8846, and compared with the statistical data, the area precision reaches 95.94%, and the recognition effect is better in the field of soybean recognition research.

[0091] The application provides a soybean recognition method based on a standard curve. The method focuses on the vegetation index time sequence curve itself, analyzes and summarizes the curve characteristic change rule and difference of soybeans and corns through visualizing the vegetation index time sequence curves of the soybeans and the corns, finds the key growth stage time sequence curve characteristics of the soybeans, improves the soybean recognition precision, and adds a meteorological weight factor, so that the standard curve has stronger robustness, has the advantages of predicting the standard curve of other years, solving the problem of sample acquisition difficulty, and having higher fault tolerance in disaster areas.

Claims

1. A soybean identification method based on a standard curve, characterized by The following steps are included Step S1: Satellite data processing: The satellite data processing step acquires all relevant satellite data of soybean growth period in the study area and performs cloud control and cloud removal on the data; Step S2: Constructing time series normalized phenological index features: The processed satellite data are used to calculate the vegetation index of all pixels in the study area to obtain the vegetation index image of all pixels. For each pixel, the vegetation index image of the pixel is processed using the time series method of the pixel to construct the time series vegetation index feature of the pixel. Step S3, constructing a sample point standard curve; The construction of the sample point standard curve includes: obtaining the locations of typical representative sample points of crops including soybeans, calculating the sample point time series curve using the sample point time series vegetation index characteristics obtained in step S2 and the sample point locations, and constructing the sample point standard curve using the sample point time series curve and adding meteorological data factors; The formula of the sample point standard curve is: in, The first part is the harmonic fitting formula, where Temp is the temperature variable, hour is the sunshine duration variable, rain is the rainfall variable, and c0, c1, and c2 are meteorological weight factors. The key growth period information of soybean sample points adopts the parameter information of main phenological points, which are determined by the following formula: u=(V min -V max )×p+V min Among them, u is the reference threshold, which is used to determine the start time and end time of soybean phenology at the sample point, V min and V max are the minimum and maximum values ​​in the time series of the soybean sample point, respectively. p is the amplitude ratio. p has three values. The first value corresponds to the growth start node and the dormancy start node of the soybean sample point. The second value corresponds to the mid-growth node and the mid-senescence node of the soybean sample point. The third value corresponds to the maturity node and the senescence start node of the soybean sample point. Step S4: Extracting standard curve parameter information of sample points and calculating phenological characteristics: Extract the key growth period information of the soybean sample point from the sample point standard curve to obtain the main phenological information of the sample point; calculate the key growth period information of the soybean sample point for all pixels except the sample point within the study area to synthesize the phenological parameter characteristic image; Step S5: Crop variety identification and accuracy assessment: The sample information is input into the classifier to obtain the sample classifier. The phenological parameter characteristic images of all pixels other than the sample point are input into the sample classifier for recognition and the recognition results are output. The recognition accuracy is calculated using the classification confusion matrix.

2. The soybean identification method based on a standard curve according to claim 1, wherein: The vegetation index in step S2 is the normalized phenological vegetation index, and the calculation formula of the normalized phenological vegetation index is as follows: Where NIR is the reflectance of the near-infrared band, RED is the reflectance of the red band, and SWIR is the reflectance of the short-wave infrared band.

3. The soybean identification method based on the standard curve according to claim 1, characterized in that: The time series method in step S2 adopts a linear harmonic model, and the linear harmonic model calculation formula is as follows: Where f(t) is the fitted vegetation index value at time t; a is a constant term; b is the coefficient of the first-order term; M is the number of harmonic combinations. In this application, M is 2, that is, a binomial linear harmonic model; C and D are the coefficients of the cosine and sine functions; ω is 1 / 365; and t is the day of the year.

4. The soybean identification method based on a standard curve according to claim 1, wherein: When the phenological parameter characteristic images of all pixels other than the sample point are input into the sample classifier for identification as described in step 5, the phenological parameter characteristic images of all pixels other than the sample point are compared and identified with the sample point data with the closest growth conditions in the sample classifier.

5. The soybean identification method based on the standard curve according to claim 1 is characterized in that: The sample classifier adopts the random forest classifier, and the classification confusion matrix adopts the random forest classification confusion matrix. The random forest classification confusion matrix is ​​defined as follows: In formula (6), n represents the number of categories, m ij It represents the number of pixels that actually belong to class i but are classified into class j in the classification result map. The value of the diagonal element is the number of pixels that are correctly classified in each category. The classification confusion matrix is ​​used to calculate the overall accuracy (OA), mapping accuracy (PA), user accuracy (UA) and F1-Score. The calculation formulas are: Among them, m i+ is the row sum in the classification confusion matrix, m +i is the column sum in the classification confusion matrix.

6. The soybean identification method based on the standard curve according to claim 1, characterized in that: The three values ​​of the amplitude ratio p are 0.15, 0.5 and 0.90 respectively.

Citation Information

Patent Citations

  • Method and system for determining crop planting area based on phenological characteristics and storage medium

    CN114219847A

  • Soybean planting area identification method based on Sentinel-1, 2 image and soybean planting area measurement and calculation method based on Sentinel-1, 2 image

    CN115063610A