Peanut planting distribution extraction method and system based on phenological characteristics

By constructing phenological spectral index characteristics and using a random forest classifier, the problem of difficulty in accurately obtaining peanut planting distribution in the prior art is solved, and efficient and accurate peanut planting distribution extraction is achieved.

CN119992184APending Publication Date: 2025-05-13SHENYANG LIGONG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510063523.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

It is difficult to accurately obtain peanut planting distribution in the prior art, especially in the unknown distribution areas of peanuts and soybeans, and the processing cost of multi-spectral and hyperspectral remote sensing image data is high.

Method used

The peanut planting distribution extraction method based on phenological characteristics is constructed, and the timing image set is classified using a random forest classifier to obtain the peanut planting distribution.

Benefits of technology

It improves the accuracy and efficiency of peanut planting distribution extraction, reduces the difficulty of field sample point collection, and reduces the calculation cost of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992184A_ABST
    Figure CN119992184A_ABST
Patent Text Reader

Abstract

The invention discloses a peanut planting distribution extraction method and system based on phenological characteristics, and belongs to the technical field of crop phenotype acquisition, and the method comprises the steps: generating a time sequence image set according to the obtained Sentinel-2 image data in a research region; establishing phenological spectral index characteristics for representing crop density and yellowness, adding all characteristic variables to each time sequence image in the time sequence image set to obtain a characteristic variable time sequence image set, obtaining to-be-divided sample data containing characteristic variable information, and performing division and crop type marking on sample points of the to-be-divided sample data to obtain a to-be-divided sample image set; taking image sample data sets of other crops except rice, corn, peanut and soybean as a training set, and training a classifier; and obtaining a feature variable time sequence image set of the to-be-extracted peanut planting area, classifying the feature variable time sequence image set through the classifier, and obtaining the distribution condition of the peanut samples in the research area. According to the method, the precision of peanut planting distribution condition extraction can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of crop phenotype acquisition, and more specifically to a method and system for extracting peanut planting distribution based on phenological characteristics. Background Art

[0002] Peanuts and soybeans are the main oil crops grown in Northeast my country, but there are few studies on the distribution of peanuts in Northeast China. Most studies only focus on three crops: corn, rice and soybeans. The main way to obtain the distribution of peanut planting is still the official statistics bureau-led field data collection and statistics at the county and city level. However, this method is not only time-consuming and labor-intensive, but can only obtain the planting area of ​​peanuts, without more detailed planting distribution data.

[0003] Optical remote sensing satellite images have the characteristics of wide coverage, large amount of information and good timeliness. Therefore, the use of optical remote sensing images to carry out crop information extraction research has become the main means of crop classification and identification. Studies have shown that corn, soybeans and peanuts have similar growth and maturity periods, and the spectral differences caused by phenological differences are the key features for accurately distinguishing these crops. Existing studies usually use multispectral or hyperspectral remote sensing image data to construct indicative feature sets to distinguish crop categories, but the data volume of large-scene multispectral and hyperspectral remote sensing images is extremely large, and high computational costs are required for processing. In addition, in supervised classification research, it is also necessary to obtain pure pixel samples of crops as evenly as possible in the study area, resulting in the fact that the study area of ​​most studies is often limited to townships or counties.

[0004] Google Earth Engine (GEE) is a new generation of remote sensing big data computing and analysis cloud platform developed by Google. It can quickly access massive multi-source remote sensing data and call huge computing resources for analysis. The GEE platform can easily and quickly call remote sensing images from multispectral / hyperspectral optical satellites such as Landsat, Sentinel, and MODIS, and can also upload other multi-source remote sensing data for calculation, which has great advantages over traditional analysis methods. The GEE platform can solve the difficulties in calculating large-scale remote sensing data. Research related to existing technologies has shown the feasibility and effectiveness of crop classification based on remote sensing images under the GEE platform.

[0005] However, these studies based on the GEE platform also require the collection of a large number of field sample points. For the problem of peanut planting distribution extraction, it is still very difficult to obtain a large number of real sample points when the exact distribution areas of peanuts and soybeans are unknown. Summary of the invention

[0006] In response to the problems existing in the above-mentioned fields, the present invention proposes a method and system for extracting peanut planting distribution based on phenological characteristics. The phenological difference factors of peanuts and soybeans in the late growth period are used to construct a new phenological characteristic combination index, sample point division and crop type marking are performed, a training set is constructed, and a random forest classifier is trained. The trained random forest classifier is used to classify the feature variable time series image set, so as to accurately obtain the planting distribution of peanuts.

[0007] In order to solve the above technical problems, the present invention discloses a method for extracting peanut planting distribution based on phenological characteristics, comprising the following steps:

[0008] Construct image sample datasets of crops other than rice, corn, peanuts and soybeans;

[0009] Generate a time series image set based on the acquired Sentinel-2 image data in the study area; construct a phenological spectral index feature for characterizing crop lushness and brownness by analyzing the NDVI and EVI values ​​of crop leaf growth at different stages; and add all characteristic variables to each time series image in the time series image set by determining the number of characteristic variables based on the phenological spectral index feature to obtain a characteristic variable time series image set.

[0010] According to the feature variable time series image set, the sample data to be divided containing the feature variable information is obtained, and the sample points are divided and the crop types are marked; the samples with crop type marks after division and the image sample data sets of other crops are used as training sets to train the random forest classifier;

[0011] The characteristic variable time series image set of the peanut planting area to be extracted is obtained, and the characteristic variable time series image set is classified by the trained random forest classifier to obtain the distribution of peanut planting areas in the study area.

[0012] Preferably, the constructing of the image sample dataset of crops other than rice, corn, peanut and soybean comprises the following steps:

[0013] Obtain a 30m-level distribution dataset of multiple crops in a preset time period in the study area and determine the initial sample points of the crops; based on the initial sample points of the crops, generate farmland discrimination areas by screening out key images that can cover the maturity period of the crops; based on the generated farmland discrimination areas, determine the sample datasets of other crops except rice, corn, peanuts and soybeans;

[0014] The acquisition of other crop sample point datasets includes two stages:

[0015] In the first stage, several sample points are randomly generated within the study area, and the generated sample points are compared with the generated farmland discrimination area. When the current sample point is located in the discrimination area, and the current sample point is at a pixel with a pixel value of 0 in the discrimination area, that is, the current sample point falls in the non-farmland area, then the current sample point is marked as other, and so on, all other sample points falling in the non-farmland area are obtained. The areas represented by the other sample points obtained in the first stage include built-up areas, water areas, woodlands and bare land;

[0016] The 30m-level distribution dataset of various crops in the preset time period in the study area is taken as dataset A;

[0017] In the second stage, when the current sample point is located in the farmland discrimination area, but the current sample point is at a pixel whose pixel value in the discrimination area is not 0, that is, the current sample point falls in the farmland area; continue to compare the current sample point with the pixel feature category at the corresponding position in data set A. When the current sample point is also classified as other categories in data set A, a buffer zone with a radius of 100m is established with the current sample point as the center, and the pixel proportion of other categories in the buffer zone is counted in data set A. When the proportion exceeds 90%, the current sample point is marked as other, and so on, to obtain all other sample points falling in the farmland area. The other sample points obtained in the second stage mainly represent crops other than rice, corn, peanuts and soybeans.

[0018] Preferably, generating the farmland identification area specifically includes:

[0019] Import Dynamic World feature coverage data into the GEE platform. The feature coverage type of Dynamic World feature coverage data is determined by the value of the image label band.

[0020] By binarizing each image, the pixels with a label band value of 4 are retained as farmland pixels, and the remaining pixel values ​​are set to 0;

[0021] By retaining all pixels with non-zero pixel values ​​at the same position in the binary images, a binary image is synthesized as the farmland discrimination area.

[0022] Preferably, the construction of the phenological spectral index features for characterizing crop lushness and brownness specifically includes:

[0023] The density and yellowness of vegetation leaves will significantly affect the vegetation index;

[0024] When the crop leaves grow most vigorously, the NDVI and EVI values ​​reach their peak values. After entering the scorching yellow stage, the NDVI and EVI values ​​will drop significantly. Based on this phenological feature, two new normalized indices NorNDVI and Nor EVI , to demarcate the peanut and soybean sample points;

[0025] Nor NDVI and Nor EVI The calculation formulas are:

[0026]

[0027] In the formula, NDVI may and EVI may are the NDVI and EVI values ​​of the images in May in the time series image set. NDVI sep and EVI sep They are the NDVI and EVI values ​​of the September images in the time series image set.

[0028] Preferably, the determining the number of characteristic variables specifically includes:

[0029] From the acquired Sentinel-2 images, three original spectral bands, B6, B11 and B12, were selected as original feature variables;

[0030] Seven vegetation indices, NDVI, EVI, LSWI, NDSVI, NDTI, GCVI and REP, were selected as vegetation characteristic variables;

[0031] Calculate the NDVI value and EVI value of each time-series image in the time-series image set respectively to obtain the NDVI map and EVI map, and use the gray-level co-occurrence matrix to generate the texture feature quantity of the NDVI map and the EVI map;

[0032] In the entire time series image set, the 10th, 25th, 50th, 75th, 90th and 25-75th percentiles of each pixel are selected as image percentile feature variables;

[0033] Harmonic regression analysis was performed on the curves of the three vegetation indices NDVI, EVI and LSWI in the time series image collection to fit the periodic changes in the time series data. The periodic components in the data were estimated by linear regression, and the harmonic analysis results, i.e. the parameters obtained by fitting, were used as periodic characteristic variables.

[0034] The number of characteristic variables is determined based on the original characteristic variables, vegetation characteristic variables, texture characteristic quantities, image percentile characteristic variables, periodic characteristic variables and constructed phenological spectral index characteristics.

[0035] Preferably, the acquisition of samples with crop type labels after the division specifically includes:

[0036] Obtain the 20m-level soybean planting distribution dataset in Northeast China in 2021, use it as dataset B, and obtain the initial sample points of soybeans;

[0037] Obtain the 20m-level peanut planting distribution dataset in Northeast China in 2018, which is generated using Sentinel-2 satellite remote sensing images with an accuracy of 20m. It covers the three provinces in Northeast my country and has two types of land features: peanuts and other crops. The dataset is represented by dataset C, and the initial sample points of peanuts are obtained.

[0038] According to the initial sample points of soybeans and peanuts obtained, the K-means clustering algorithm is used on the GEE platform to cluster the sample data sets B and C to be divided, which contain characteristic variable information, and assign different crop type labels to the sample points. The crop types are correspondingly labeled as soybeans and peanuts, and are used as samples with crop type labels after division.

[0039] Preferably, obtaining the distribution of peanut planting areas within the study area specifically includes:

[0040] The obtained samples with crop type labels after division and the constructed image sample datasets of other crops except rice, corn, peanuts and soybeans are used as training samples and input into the random forest classifier. The feature variable time series image set is classified by the random forest classifier to extract the distribution of peanut samples and soybean planting areas in dataset B and dataset C respectively;

[0041] The distribution of peanut samples and soybean planting areas in the extracted datasets B and C is compared with the statistical data of the main producing counties. By analyzing the comparison results, when there are differences in the comparison results, the peanut and soybean sample points are divided again; when there are no differences in the comparison results, the division is terminated to obtain the trained random forest classifier;

[0042] The acquired characteristic variable time series image set of the peanut planting area to be extracted is input into the trained random forest classifier to output the distribution of peanut planting areas in the study area.

[0043] Preferably, the generation of the time-series image set specifically includes:

[0044] Based on the acquired Sentinel-2 image data in the study area, the L2A-level remote sensing images of the Sentinel-2 satellite covering the study area and the time period were screened, and the QA band was used for cloud removal. The full-band values ​​of each pixel point of the cloud-free image were processed by median synthesis to obtain cloud-free high-quality synthetic images.

[0045] Using crop image data throughout the entire growth period with a three-month window size, linear interpolation processing is performed on cloud-free high-quality synthetic images to obtain images covering the entire study area without data missing, forming a time series image set.

[0046] Preferably, a peanut planting distribution extraction system based on phenological characteristics is also included, comprising:

[0047] Other crop sample acquisition module, used to construct image sample datasets of other crops except rice, corn, peanuts and soybeans;

[0048] The characteristic variable time series image set acquisition module is used to generate a time series image set based on the acquired Sentinel-2 image data in the study area; by analyzing the NDVI and EVI values ​​of crop leaf growth at different periods, the phenological spectral index characteristics used to characterize the density and brownness of crops are constructed; according to the phenological spectral index characteristics, by determining the number of characteristic variables, all characteristic variables are added to each time series image in the time series image set to obtain the characteristic variable time series image set;

[0049] The classifier training module is used to obtain the sample data to be divided containing the feature variable information according to the feature variable time series image set, and divide the sample points and mark the crop types; the samples with crop type marks after division and the image sample data sets of other crops are used as training sets to train the random forest classifier;

[0050] The peanut planting area distribution extraction module is used to obtain the characteristic variable time series image set of the peanut planting area to be extracted, classify the characteristic variable time series image set through the trained random forest classifier, and obtain the distribution of peanut planting areas in the study area.

[0051] Compared with the prior art, the present invention has the following beneficial effects:

[0052] The peanut planting distribution extraction method based on phenological characteristics proposed in the present invention constructs a sample data set of other crops except rice, corn, peanuts and soybeans; the phenological difference factors of soybeans and peanuts in the late growth period are used to construct the phenological spectral index characteristics used to characterize the crop density and brownness, determine the characteristic variable time series image set, and use the phenological characteristics as key features, so that it plays an obvious role in obtaining sample points and crop classification, and can effectively improve the classification accuracy. The sample data to be divided containing characteristic variable information is obtained, and its sample points are divided and crop types are marked. The samples with crop type marks after division and other crop sample data sets are used as training sets to train the classifier. The trained classifier is used to classify the characteristic variable time series image set of the peanut planting area to be extracted, so as to accurately obtain the distribution of peanut planting areas in the study area. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 This is a flow chart of the peanut planting distribution extraction method based on phenological characteristics proposed by the present invention;

[0054] Figure 2 A process for generating a time series image set provided by an embodiment of the present invention;

[0055] Figure 3 Dynamic World classification results provided by the embodiment of the present invention;

[0056] Figure 4 The vegetation index calculation formula proposed in the embodiment of the present invention;

[0057] Figure 5 A schematic diagram of the distribution of initial sample points provided by an embodiment of the present invention;

[0058] Figure 6 The process of dividing the original soybean sample points and classifying them provided in the embodiment of the present invention;

[0059] Figure 7 The final sample point distribution provided by the embodiment of the present invention;

[0060] Figure 8 NDVI mean curves of corn, soybean and peanut sample points provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0061] The following will be combined with the attached embodiment of the present invention Figure 1-Figure 8 , the technical solutions in the embodiments of the present invention are clearly and completely described. It should be understood that the terms described in the present invention are only used to describe specific implementation methods and are not used to limit the present invention.

[0062] like Figure 1 As shown, the present invention proposes a peanut planting distribution extraction method based on phenological characteristics, comprising the following steps:

[0063] S1: Construct image sample datasets of crops other than rice, corn, peanuts and soybeans;

[0064] S2: Generate a time series image set based on the acquired Sentinel-2 image data in the study area; construct a phenological spectral index feature for characterizing crop lushness and brownness by analyzing the NDVI and EVI values ​​of crop leaf growth at different stages; according to the phenological spectral index feature, add all characteristic variables to each time series image in the time series image set by determining the number of characteristic variables, and obtain a characteristic variable time series image set;

[0065] S3: According to the feature variable time series image set, the sample data to be divided containing the feature variable information is obtained, and the sample points are divided and the crop types are marked; the samples with crop type marks after division and the image sample data sets of other crops are used as training sets to train the random forest classifier;

[0066] S4: Obtain a time series image set of characteristic variables of the peanut planting area to be extracted, classify the time series image set of characteristic variables through the trained random forest classifier, and obtain the distribution of peanut planting areas in the study area.

[0067] Specifically, in step S1, constructing an image sample dataset of crops other than rice, corn, peanuts and soybeans includes the following steps:

[0068] Obtain a 30m-level distribution dataset of multiple crops in a preset time period in the study area and determine the initial sample points of the crops; based on the initial sample points of the crops, generate farmland discrimination areas by screening out key images that can cover the maturity period of the crops; based on the generated farmland discrimination areas, determine the sample datasets of other crops except rice, corn, peanuts and soybeans;

[0069] The acquisition of other crop sample point datasets includes two stages:

[0070] In the first stage, several sample points are randomly generated within the study area, and the generated sample points are compared with the generated farmland discrimination area. When the current sample point is located in the discrimination area, and the current sample point is at a pixel with a pixel value of 0 in the discrimination area, that is, the current sample point falls in the non-farmland area, then the current sample point is marked as other, and so on, all other sample points falling in the non-farmland area are obtained. The areas represented by the other sample points obtained in the first stage include built-up areas, water areas, woodlands and bare land;

[0071] The 30m-level distribution dataset of various crops in the preset time period in the study area is taken as dataset A;

[0072] In the second stage, when the current sample point is located in the farmland discrimination area, but the current sample point is at a pixel whose pixel value in the discrimination area is not 0, that is, the current sample point falls in the farmland area; continue to compare the current sample point with the pixel feature category at the corresponding position in data set A. When the current sample point is also classified as other categories in data set A, a buffer zone with a radius of 100m is established with the current sample point as the center, and the pixel proportion of other categories in the buffer zone is counted in data set A. When the proportion exceeds 90%, the current sample point is marked as other, and so on, to obtain all other sample points falling in the farmland area. The other sample points obtained in the second stage mainly represent crops other than rice, corn, peanuts and soybeans.

[0073] Among them, generating farmland discrimination area includes:

[0074] Import Dynamic World feature coverage data into the GEE platform. The feature coverage type of Dynamic World feature coverage data is determined by the value of the image label band.

[0075] By binarizing each image, the pixels with a label band value of 4 are retained as farmland pixels, and the remaining pixel values ​​are set to 0;

[0076] By retaining all pixels with non-zero pixel values ​​at the same position in the binary images, a binary image is synthesized as the farmland discrimination area.

[0077] In step S2, the generation of the time series image set specifically includes:

[0078] Based on the acquired Sentinel-2 image data in the study area, the L2A-level remote sensing images of the Sentinel-2 satellite covering the study area and the time period were screened, and the QA band was used for cloud removal. The full-band values ​​of each pixel point of the cloud-free image were processed by median synthesis to obtain cloud-free high-quality synthetic images.

[0079] Using crop image data throughout the entire growth period with a three-month window size, linear interpolation processing is performed on cloud-free high-quality synthetic images to obtain images covering the entire study area without data missing, forming a time series image set.

[0080] Construct phenological spectral index features for characterizing crop density and brownness, including:

[0081] The density and yellowness of vegetation leaves will significantly affect the vegetation index;

[0082] When the crop leaves grow most vigorously, the NDVI and EVI values ​​reach their peak values. After entering the scorching yellow stage, the NDVI and EVI values ​​will drop significantly. Based on this phenological feature, two new normalized indices Nor NDVI and Nor EVI , to demarcate the peanut and soybean sample points;

[0083] Nor NDVI and Nor EVI The calculation formulas are:

[0084]

[0085] In the formula, NDVI may and EVI may are the NDVI and EVI values ​​of the images in May in the time series image set. NDVI sep and EVI sep They are the NDVI and EVI values ​​of the September images in the time series image set.

[0086] Determine the number of feature variables, including:

[0087] From the acquired Sentinel-2 images, three original spectral bands, B6, B11 and B12, were selected as original feature variables;

[0088] Seven vegetation indices, NDVI, EVI, LSWI, NDSVI, NDTI, GCVI and REP, were selected as vegetation characteristic variables;

[0089] Calculate the NDVI value and EVI value of each time-series image in the time-series image set respectively to obtain the NDVI map and EVI map, and use the gray-level co-occurrence matrix to generate the texture feature quantity of the NDVI map and the EVI map;

[0090] In the entire time series image set, the 10th, 25th, 50th, 75th, 90th and 25-75th percentiles of each pixel are selected as image percentile feature variables;

[0091] Harmonic regression analysis was performed on the curves of the three vegetation indices NDVI, EVI and LSWI in the time series image collection to fit the periodic changes in the time series data. The periodic components in the data were estimated by linear regression, and the harmonic analysis results, i.e. the parameters obtained by fitting, were used as periodic characteristic variables.

[0092] The number of characteristic variables is determined based on the original characteristic variables, vegetation characteristic variables, texture characteristic quantities, image percentile characteristic variables, periodic characteristic variables and constructed phenological spectral index characteristics.

[0093] In step S3, samples with crop type labels are obtained after classification, including:

[0094] Obtain the 20m-level soybean planting distribution dataset in Northeast China in 2021, use it as dataset B, and obtain the initial sample points of soybeans;

[0095] Obtain the 20m-level peanut planting distribution dataset in Northeast China in 2018, which is generated using Sentinel-2 satellite remote sensing images with an accuracy of 20m. It covers the three provinces in Northeast my country and has two types of land features: peanuts and other crops. The dataset is represented by dataset C, and the initial sample points of peanuts are obtained.

[0096] According to the initial sample points of soybeans and peanuts obtained, the K-means clustering algorithm is used on the GEE platform to cluster the sample data sets B and C to be divided, which contain characteristic variable information, and assign different crop type labels to the sample points. The crop types are correspondingly labeled as soybeans and peanuts, and are used as samples with crop type labels after division.

[0097] Obtain the distribution of peanut planting areas in the study area, including:

[0098] The obtained samples with crop type labels after division and the constructed image sample datasets of other crops except rice, corn, peanuts and soybeans are used as training samples and input into the random forest classifier. The feature variable time series image set is classified by the random forest classifier to extract the distribution of peanut samples and soybean planting areas in dataset B and dataset C respectively;

[0099] The distribution of peanut samples and soybean planting areas in the extracted datasets B and C is compared with the statistical data of the main producing counties. By analyzing the comparison results, when there are differences in the comparison results, the peanut and soybean sample points are divided again; when there are no differences in the comparison results, the division is terminated to obtain the trained random forest classifier;

[0100] The acquired characteristic variable time series image set of the peanut planting area to be extracted is input into the trained random forest classifier to output the distribution of peanut planting areas in the study area.

[0101] The present invention also proposes a peanut planting distribution extraction system based on phenological characteristics, comprising:

[0102] Other crop sample acquisition module, used to construct image sample datasets of other crops except rice, corn, peanuts and soybeans;

[0103] The characteristic variable time series image set acquisition module is used to generate a time series image set based on the acquired Sentinel-2 image data in the study area; by analyzing the NDVI and EVI values ​​of crop leaf growth at different periods, the phenological spectral index characteristics used to characterize the density and brownness of crops are constructed; according to the phenological spectral index characteristics, by determining the number of characteristic variables, all characteristic variables are added to each time series image in the time series image set to obtain the characteristic variable time series image set;

[0104] The classifier training module is used to obtain the sample data to be divided containing the feature variable information according to the feature variable time series image set, and divide the sample points and mark the crop types; the samples with crop type marks after division and the image sample data sets of other crops are used as training sets to train the random forest classifier;

[0105] The peanut planting area distribution extraction module is used to obtain the characteristic variable time series image set of the peanut planting area to be extracted, classify the characteristic variable time series image set through the trained random forest classifier, and obtain the distribution of peanut planting areas in the study area.

[0106] The method proposed by the present invention can effectively improve the extraction accuracy of peanut planting distribution conditions.

[0107] Example

[0108] The embodiment provided by the present invention selects a certain city administrative area as the research area, and the administrative area covered is located in the central and northern part of a certain province, with flat terrain, fertile land, and a climate suitable for agricultural development. The main crops in the research area include rice, corn, soybeans and peanuts, and the total planting area of ​​the four crops accounts for more than 80% of all crops in the city.

[0109] Data preprocessing

[0110] Sentinel-2 is a high-resolution multispectral imaging satellite. Its main payload is a multispectral imager with a total of 13 bands, and a spatial resolution of 10m for visible light, 20m for near infrared, and 60m for short-wave infrared. The present invention screens the L2A-level remote sensing images of the Sentinel-2 satellite covering the study area from January 1, 2022 to January 1, 2023, and uses the QA band for cloud removal. The full-band values ​​of each pixel point of the cloud-free image within each month are subjected to median synthesis processing, and a total of 12 cloud-free high-quality synthetic images are obtained.

[0111] In order to prevent the loss of data of individual pixel points, the image data of the crop growth period (April to October) was used to perform linear interpolation processing with a window size of three months to obtain 7 high-quality images covering the entire study area without data loss, forming a time series image set. The process of generating a time series image set is as follows: Figure 2shown.

[0112] Dynamic World is a digital product launched by Google in 2022. It uses a deep learning model to determine the type of land cover in near real time based on remote sensing images from the Sentinel-2 satellite and gives a classification result. Its advantages are sufficient data volume and excellent timeliness of the data given in near real time. At the same time, due to the excessive pursuit of near real-time performance, the classification results of the same land cover area may be different in different time periods, that is, misclassification may occur. Figure 3 The following is the Dynamic World classification result of Sentinel-2 image data in May 2022. A large number of rice fields around the urban area of ​​a certain city were mistakenly classified as water areas. Figure 3 In addition, there are a lot of areas with missing data, such as Figure 3 Shown in black area.

[0113] The present invention obtains the 30m-level three-crop planting distribution dataset in Northeast China in 2021, which is generated by Landsat 8 satellite remote sensing images with an accuracy of 30m, covering the three provinces in Northeast my country. The ground object types are rice, corn, soybeans and others (including buildings, water bodies and other crops, etc.). The dataset is represented by dataset A.

[0114] The present invention also obtains the 20m-level soybean planting distribution dataset in Northeast China in 2021 using Sentinel-2 satellite remote sensing images, with an accuracy of 20m, covering the three provinces in Northeast my country, and the types of land features are soybeans and others (including buildings, water bodies and other crops, etc.). The dataset is represented by dataset B.

[0115] At the same time, the present invention obtains the 20m-level peanut planting distribution dataset in Northeast China in 2018 using Sentinel-2 satellite remote sensing images, with an accuracy of 20m, covering the three provinces in Northeast my country, and the ground object types are peanuts and others (including buildings, water bodies and other crops, etc.). The dataset is represented by dataset C.

[0116] In the present invention, under the GEE platform, dataset A and dataset B are used to obtain initial sample points. In addition, these two datasets and the 20m-level peanut planting distribution dataset in Northeast China in 2018 will be used together for comparative analysis of crop area extraction results.

[0117] The present invention uses the official annual statistical yearbook data of a certain city to assist in verifying the classification results of crops. The yearbook provides the sowing area data of major crops including peanuts in each county within the administrative area of ​​a certain city.

[0118] Table 1 Comparison of data sets with statistical yearbook data

[0119]

[0120] As shown in Table 1, the crop area data provided by the Statistical Yearbook, Dataset A, and Dataset B of a city in 2021 are given. As shown in the data in Table 1, the crop areas in the dataset are significantly larger than those in the Statistical Yearbook, and the data results in the yearbook are very different, indicating that there is a phenomenon of crop misclassification.

[0121] This paper uses the 2022 Sentinel-2 image data to generate a time series image set, and uses dataset A and dataset B, combined with the Dynamic World data classification results, to generate crop sample points. Using three original spectral features, seven vegetation index features, three texture features, and two phenological spectral index features, the time series image set is classified using the K-means clustering algorithm and the random forest algorithm on the GEE platform.

[0122] Through the visualization of classification results, it is shown that a large number of river beaches are misclassified into a variety of crops, mainly soybeans. Considering that the rain and heat in a certain city are synchronized, the arrival of the flood season will cause the water area including the river floodplain to expand to the maximum. At this time, the crops enter the growth and maturity period with the most suitable temperature, and the spectral characteristics of the crops are also most significant. Therefore, selecting images during the crop maturity period can effectively reduce the possibility of misclassification of river beaches and improve the reliability of subsequent sample point selection. To this end, the present invention imports Dynamic World land feature coverage data from cloud storage in GEE to screen key images that can cover the crop maturity period. After many experiments, the Dynamic World land feature coverage data from June 1, 2022 to August 30, 2022 was finally selected.

[0123] The type of ground feature coverage in Dynamic World data is determined by the value of the label band of the image. In order to obtain a more accurate sample discrimination area, each image is binarized, and the pixels with a label band pixel value of 4 are retained as farmland pixels, and the remaining pixel values ​​are set to 0. Then the NonZero command is used to retain all pixels with non-0 pixel values ​​at the same position in all binary images, and a binary image is synthesized as the farmland discrimination area.

[0124] From the Sentinel-2 image data, three original spectral bands, B6, B11 and B12, are selected as original feature variables. These three spectral bands cover the red edge band and the short-wave infrared band that is more sensitive to leaf water content. The calculation method is as follows Figure 4As shown, PGreen, PRed, PRed1, PRed2, PRed3, PNIR, PSWIR1, and PSWIR2 are the surface reflectance values ​​of the B3, B4, B5, B6, B7, B8, B11, and B12 bands of the Sentinel-2 remote sensing image, respectively.

[0125] In addition, seven commonly used vegetation indices were selected as characteristic variables. NDVI is the most commonly used vegetation index, which uses the characteristics of vegetation reflecting more near-infrared light and absorbing more red light to distinguish vegetation from other landforms. EVI inherits the advantages of NDVI and also improves its problems such as saturation of high vegetation areas, incomplete correction of atmospheric effects, and soil background. LSWI is based on the principle that water bodies strongly absorb short-wavelength infrared light, and is highly sensitive to vegetation water content and water bodies. NDSVI and NDTI are sensitive to chlorotic vegetation and soil, and can effectively distinguish vegetation with different phenology when processing time series data. GCVI can effectively reflect chlorophyll content and is also highly sensitive to vegetation. REP is a precise red edge index designed specifically for Sentinel-2 remote sensing images, which can more effectively reflect the chlorophyll content of vegetation.

[0126] The NDVI and EVI values ​​were calculated for each time series image to obtain the NDVI and EVI maps. Then, the gray-level co-occurrence matrix was used to generate the texture feature quantities of the NDVI and EVI maps, and three features, namely, the second moment of the autocorrelation angle (ASM), the correlation of the autocorrelation matrix (CORR), and the entropy of the autocorrelation matrix (ENT), were selected and added to each time series image according to the order of magnitude and importance of different texture feature variables. Compared with the spectral features that reflect the spectral information of pixels, the texture features can effectively describe the spatial distribution information between local pixels.

[0127] For the entire time series image set, the 10th, 25th, 50th, 75th, 90th and 25-75th percentiles of each pixel are selected as feature variables. In the time series image set, these features are used to describe the distribution of pixel values ​​and extract their range of variation.

[0128] The calculation formula of the periodic characteristic variable is:

[0129] VIt=Constant+asin3sin(3πt)+acos3cos(3πt)

[0130] +asin6sin(6πt)+acos6cos(6πt)

[0131] Where t is the time point in the time series image set, VI t is the value of the vegetation index used that changes over time; constant, asi n3 、aco s3 、asin6 and aco s6 are the parameters obtained after fitting, which serve as the overall periodic characteristic variables.

[0132] Finally, a total of 189 feature variables were obtained, and their categories and quantities are shown in Table 2.

[0133] Table 2 Characteristic variables

[0134]

[0135]

[0136] The four major crops planted in a city are corn, rice, peanuts and soybeans. However, most of the current research on the distribution of crop planting is limited to rice, corn and soybeans, and there is little research on the specific planting distribution of peanuts. In order to make full use of the existing data set results to improve the classification accuracy, the present invention obtains rice, corn, peanuts, soybeans and other five types of ground feature sample points for training random forest classifiers.

[0137] Other land features include crops other than rice, corn, peanuts and soybeans, and non-agricultural land features (such as buildings, water areas, etc.).

[0138] The acquisition of other ground sample points includes two stages:

[0139] In the first stage, several sample points are randomly generated within the administrative area of ​​a city and compared with the generated farmland discrimination map. If the current sample point is located at a pixel with a pixel value of 0 in the discrimination map, that is, it falls in the non-farmland area, the sample point is marked as other, and a total of 273 other sample points are obtained. These other sample points mainly represent built-up areas, water areas, woodlands, and bare land.

[0140] In the second stage, if the current sample point is located at a pixel with a non-zero pixel value in the farmland discrimination map, that is, it falls in the farmland area, it will continue to be compared with the pixel object category at the corresponding position in data set A. If it is also classified as other categories in data set A, a buffer zone with a radius of 100m will be established with the sample point as the center, and the proportion of pixels of other categories in the buffer zone will be counted in data set A. If the proportion exceeds 90%, the sample point will be marked as other, and a total of 97 other sample points will be obtained. These other sample points mainly represent crops other than the four major crops of rice, corn, peanuts and soybeans.

[0141] Corn is the most widely planted crop in a city, and its sample points are highly reliable and have high classification accuracy. Rice also usually has high classification accuracy due to its unique spectral characteristics during its growth period.

[0142] The strategy for obtaining rice and corn sample points is similar to the strategy for generating other sample points in the second stage. For sample points falling in the farmland discrimination area, if they are also classified as rice (corn) in dataset A, a buffer with a radius of 100m is established with the sample point as the center, and the proportion of rice (corn) pixels in the buffer is counted in dataset A. If the proportion exceeds 90%, the sample point is marked as rice (corn). Finally, a total of 307 rice sample points and 310 corn sample points were obtained.

[0143] Initial soybean sample point acquisition

[0144] Compared with corn and rice, soybeans are more likely to be misclassified and missed. For example, the visualized classification results of the simulated dataset A show that there are large areas of soybean planting in L3 and L4 of the city. According to the statistical yearbook data, L3 and L4 are the main peanut producing areas, and the peanut planting area far exceeds the soybean planting area, as shown in Table 3. Therefore, it can be inferred that a large number of peanut planting areas in dataset A are misclassified as soybeans.

[0145] Table 3 Statistical Yearbook Data of Major Soybean and Peanut Producing Areas in 2021

[0146] County Soybean / hectare Peanuts / hectare L1 City 2905 3041 L2 County 2865 10022 L3 County 1688 11780 L4 Zone 571 5718

[0147] Compared with data set A, although data set B also has misclassification in the study area, the data deviation is smaller. In order to improve the reliability of sample data, the present invention uses data set B to generate original soybean sample points, and selects L1 and L2 with the largest soybean planting area as random sample point generation areas, and then adopts the same strategy as that of generating rice and corn sample points to obtain a total of 294 sample points. These sample points include soybean sample points and peanut sample points that are misclassified as soybeans.

[0148] After the above sample points are obtained, the sample point distribution of rice, corn, soybean and other landform categories is obtained as follows Figure 5 shown.

[0149] Selected soybean sample points and peanut sample points

[0150] The present invention uses the K-means algorithm in the GEE platform to divide the 294 initial sample points into peanut sample points and soybean sample points, and then uses the random forest classifier to synchronously complete the peanut planting distribution mapping. The specific process is as follows Figure 6 shown.

[0151] In the GEE platform, all feature variables are added to each time series image to obtain a feature variable time series image set. Sample data to be divided containing feature variable information is obtained from this image set, and clustering is performed using the K-means algorithm. Different crop type labels (peanuts or soybeans) are assigned to the sample points.

[0152] The random forest classifier was trained using the divided soybean and peanut samples as well as the previously obtained rice, corn and other category samples, and the feature variable time series image set was classified. The classification results of peanuts and soybeans were compared with the statistical data of the main producing counties. Based on the difference in the results, it was determined whether the peanut and soybean sample points needed to be divided again or the division needed to be terminated.

[0153] The initial number of clusters is set to 2. The number of clusters is adjusted according to the visualization results of sample points in different groups after each division and the field data. The wrong sample points are eliminated based on the results of classifying the time series image set by the classifier trained by the sample points obtained this time.

[0154] Precision analysis and data comparison

[0155] The present invention adopts the confusion matrix method, takes producer's accuracy (PA) and user's accuracy (UA) as reference indicators to evaluate the performance of the classifier, and uses 30% test samples to generate the confusion matrix. The results obtained by the present invention are compared with the official statistical yearbook data county by county, and the determination coefficient R is used. 2 The root mean square error (RMSE) is used as a comparison indicator.

[0156] Coefficient of determination R 2 It is used to measure the degree of fit of the model to the observed data. The closer its value is to 1, the better the model fits the data. The calculation formula is as follows:

[0157]

[0158] Among them, SS res is the residual sum of squares, that is, the sum of squares of the differences between the model predictions and the actual observed values; SS tot is the total sum of squares, that is, the sum of the squares of the differences between the actual observations and the mean of the observations.

[0159] RMSE is used to measure the difference between the model prediction value and the actual observation value. It is the square root of the average value of the sum of squares of the residuals. The smaller the RMSE, the higher the accuracy of the model prediction; its calculation formula is as follows:

[0160]

[0161] Where n is the number of samples, y i is the actual observed value, is the model prediction value.

[0162] The original soybean sample points were divided 6 times in total. After removing the abnormal sample points, 51 soybean sample points, 55 peanut sample points, and 64 sample points classified as other sample points were finally obtained. The distribution of the sample points finally obtained is as follows: Figure 7 shown.

[0163] In order to verify the credibility and effectiveness of the peanut sample points, three crop sample points with similar spectral characteristics during the growth period, corn, soybean and peanut, were selected. The monthly NDVI mean values ​​of different crop sample points were calculated based on the time series image set for comparison. The results are shown in Figure 2. Figure 8 As shown. Although the monthly NDVI means of the three crops are similar and their changing trends are basically the same, the NDVI value of peanuts is slightly lower than that of soybeans and corn in the key growth period (150-220 days), while the NDVI value of soybeans decreases more than peanuts and corn in the later period (220-270 days). This result is consistent with the results of existing studies, that is, among the three crops with similar growth cycles, peanut leaves grow slowly and reach prosperity in August, while soybean leaves begin to turn yellow earlier in August, which verifies the phenological characteristics constructed based on this phenological factor. NDVI and Nor EVI The feasibility and effectiveness of peanut sample point division and classification extraction.

[0164] PA and UA are obtained based on the confusion matrix obtained from 30% of the samples, and the results are shown in Table 4.

[0165] Table 4 Crop extraction accuracy of the method of the present invention

[0166] category PA / % UA / % peanut 41.18 87.50 Rice 84.78 90.70 corn 69.77 55.05 Soybean 71.43 50.00 other 72.00 69.77

[0167] For the distribution of peanuts, the classification results give a higher user accuracy, but the producer accuracy is lower. The reason is that without collecting sample data on the spot, the present invention selects sample points from existing data sets, and the existing data sets have problems such as misclassification and omission. Although the official prior information on crop planting distribution areas is referenced, there is still the possibility that individual sample points are incorrectly marked. UA is the accuracy that users are concerned about when using classification results. Without considering the reliability of sample points, low PA and high UA mean that the classifier has fewer misclassifications in a certain category, but there is still a problem of omission. This problem often occurs when the classifier has a relatively strong ability to recognize a certain category, but a weak ability to recognize other categories. The higher UA of the method of the present invention indicates that the classifier has a higher accuracy. Combined with the method of obtaining peanut sample points, it shows that the classifier designed by the present invention still has great practical significance when it is difficult to obtain field sample data.

[0168] In order to verify the classification performance of crops in county-level regions, the statistical yearbook data of a certain city in 2022 were selected for comparison with the extraction results of the present invention. The results are shown in Tables 5 and 6.

[0169] Table 5: Fitting of the results of the present invention calculated based on the statistical yearbook of a certain city

[0170]

[0171]

[0172] Table 6 Comparison of crop area extracted by the present invention with the statistical yearbook data of a certain city

[0173]

[0174] S is the corresponding crop sowing area in the statistical yearbook. Since RMSE is not a dimensionless indicator, the ratio of RMSE to crop sowing area S is used as an additional comparison indicator. 2 The results show that the classification results of rice, corn and peanuts by the method proposed in the present invention have a high degree of fit with the statistical yearbook data. Among the four crops, the RMSE value of peanuts is the smallest, and the RMSE / S value is at the same order of magnitude as that of rice and corn, which fully demonstrates the excellent extraction effect of the method of the present invention on peanut planting areas. Combined with Table 6, it can be seen that the crop planting area at the county level is consistent with the statistical yearbook data, among which the peanut planting area is very close to the statistical data.

[0175] In order to ensure the reliability of the sample points, the present invention obtains the original soybean sample points from L1 city and L2 county, where soybean misclassification is less common. Therefore, the peanut sample points divided from the original soybean sample points are also taken from L1 city and L2 county. Although these small and not completely regionally representative peanut sample points are used to train the classifier, the simulation still gives classification results that are highly consistent with the real data in the two major peanut producing counties of L3 County and L4 District, and effectively corrects the error of classifying a large number of peanut planting areas as soybean planting areas in these two counties in data set A. This also shows that the sample point division method proposed in the present invention can obtain accurate crop sample points with good classification feature indications. In addition, due to the use of the phenological characteristics Nor constructed based on phenological factors NDVI and Nor EVI ,The comparison results of peanut planting areas at other county and district levels also show a high accuracy, as shown in Table 6.

[0176] In summary, thanks to the reliability and effectiveness of the sample point division method and classification features, the crop distribution extraction method proposed in the present invention shows excellent classification performance.

[0177] In order to further verify the validity of the peanut crop distribution extraction results, the present invention selects an existing data set to compare with the annual statistical yearbook data of the city. The method proposed in the present invention mainly extracts peanut crop distribution, and can also extract three crops: rice, corn and soybean. At present, there are few peanut crop distribution data sets covering the study area, only data set C, so data set A and data set B are also selected to compare the distribution extraction of rice, corn and soybean crops, with R 2 , RMSE and RMSE / S as evaluation criteria, the crop classification results involved in the comparison were compared with the statistical yearbook data of the corresponding years at the county level for fitting accuracy, and the results are shown in Table 7. As can be seen from Table 7, the peanut classification results of the present invention are better than the results of data set C under the three evaluation indicators, and the classification results of the other three crops are also better than the results of data set A and data set B.

[0178] Table 7 Comparison results with existing datasets

[0179]

[0180] In order to obtain a more intuitive comparison effect, the classification results of the present invention and the classification results of data set A are visualized and compared with the 10m-level Sentinel-2 high-resolution remote sensing image (red, green, and blue three-band synthesis). The comparison results show that the peanut planting area extracted by the present invention has clear boundaries, and the distribution and shape of the plots are highly consistent with the Sentinel-2 high-resolution image.

[0181] As an important oil crop, peanut is the third largest crop in a certain province and a certain city, but there are few studies on the distribution of peanut planting in the study area of ​​the city. Based on the GEE platform, the present invention makes full use of multi-source data including Sentinel-2 original remote sensing images, Dynamical World land cover data, and the distribution data of three major crops and soybeans in Northeast China in 2021. Three original spectra, seven vegetation indices, and 189 characteristic variables such as new combination indexes constructed according to phenological factors are used to characterize crops. The K-means clustering algorithm and random forest classifier are used to extract the distribution of peanut planting in a certain city as the study area. The final classification results obtained were compared and verified at the county level with the existing data sets and official statistical data. The results show that the method of the present invention obtains the main crop distribution data of the study area of ​​the city with the highest practicality and accuracy. The results obtained by the present invention can fill the gap in the high-precision distribution extraction of peanuts in the study area, providing important support for the development of precision agriculture.

[0182] In the case of difficulty in collecting crop sample data on the spot, the present invention generates crop sample points using the crop distribution data set disclosed by previous studies, which not only improves the efficiency of obtaining crop sample points, but also ensures the reliability and validity of sample data, and has important practical significance. In view of the large number of misclassifications in the data set (the river beach is misclassified as a soybean planting area, and the main peanut producing counties are very different from the official statistical yearbook data), the Dynamic World land feature coverage data of the crop maturity period is used to generate farmland discrimination areas, reducing the possibility of selecting sample points in non-farmland areas, which can effectively solve the problem of misclassification of river beach.

[0183] In multispectral remote sensing images, the spectral characteristics of soybeans and peanuts are slightly different, and crop misclassification is prone to occur. The present invention uses the phenological difference factors of soybeans and peanuts in the late growth period to construct a new phenological characteristic combination index, and uses this feature to divide sample points and classify crops. The experimental results show that phenological characteristics, as key features, play an obvious role in obtaining sample points and crop classification, and can effectively improve classification accuracy.

[0184] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

[0185] In addition, unless otherwise specified, all technical and scientific terms used in the present invention have the same meanings as commonly understood by those skilled in the art to which the present invention belongs. All documents mentioned in this specification are incorporated by reference to disclose and describe the methods related to the documents. In the event of any conflict with any incorporated document, the content of this specification shall prevail.

Claims

1. A method for extracting peanut planting distribution based on phenological characteristics, characterized in that: The following steps are involved: Construct image sample datasets of crops other than rice, corn, peanuts and soybeans; Based on the acquired Sentinel-2 image data in the study area, a time series image set was generated; by analyzing the NDVI and EVI values ​​of crop leaf growth at different stages, a phenological spectral index feature was constructed to characterize the density and yellowness of crops; According to the characteristics of the phenological spectral index, by determining the number of characteristic variables, all characteristic variables are added to each time series image in the time series image set to obtain a characteristic variable time series image set; According to the feature variable time series image set, the sample data to be divided containing the feature variable information is obtained, and the sample points are divided and the crop types are marked; the samples with crop type marks after division and the image sample data sets of other crops are used as training sets to train the random forest classifier; The characteristic variable time series image set of the peanut planting area to be extracted is obtained, and the characteristic variable time series image set is classified by the trained random forest classifier to obtain the distribution of peanut planting areas in the study area.

2. The peanut planting distribution extraction method based on phenological characteristics according to claim 1, characterized in that: The method of constructing an image sample dataset of other crops except rice, corn, peanut and soybean comprises the following steps: Obtain a 30m-level distribution dataset of multiple crops in a preset time period in the study area and determine the initial sample points of the crops; based on the initial sample points of the crops, generate farmland discrimination areas by screening out key images that can cover the maturity period of the crops; based on the generated farmland discrimination areas, determine the sample datasets of other crops except rice, corn, peanuts and soybeans; The acquisition of other crop sample point datasets includes two stages: In the first stage, several sample points are randomly generated within the study area, and the generated sample points are compared with the generated farmland discrimination area. When the current sample point is located in the discrimination area, and the current sample point is at a pixel with a pixel value of 0 in the discrimination area, that is, the current sample point falls in the non-farmland area, then the current sample point is marked as other, and so on, all other sample points falling in the non-farmland area are obtained. The areas represented by the other sample points obtained in the first stage include built-up areas, water areas, woodlands and bare land; The 30m-level distribution dataset of various crops in the preset time period in the study area is taken as dataset A; In the second stage, when the current sample point is located in the farmland discrimination area, but the current sample point is at a pixel whose pixel value in the discrimination area is not 0, that is, the current sample point falls in the farmland area; continue to compare the current sample point with the pixel feature category at the corresponding position in data set A. When the current sample point is also classified as other categories in data set A, a buffer zone with a radius of 100m is established with the current sample point as the center, and the pixel proportion of other categories in the buffer zone is counted in data set A. When the proportion exceeds 90%, the current sample point is marked as other, and so on, to obtain all other sample points falling in the farmland area. The other sample points obtained in the second stage mainly represent crops other than rice, corn, peanuts and soybeans.

3. The peanut planting distribution extraction method based on phenological characteristics according to claim 2 is characterized in that: The generating of the farmland identification area specifically includes: Import Dynamic World feature coverage data into the GEE platform. The feature coverage type of Dynamic World feature coverage data is determined by the value of the image label band. By binarizing each image, the pixels with a label band value of 4 are retained as farmland pixels, and the remaining pixel values ​​are set to 0; By retaining all pixels with non-zero pixel values ​​at the same position in the binary images, a binary image is synthesized as the farmland discrimination area.

4. The peanut planting distribution extraction method based on phenological characteristics according to claim 3 is characterized in that: The phenological spectral index characteristics for characterizing crop lushness and brownness specifically include: The density and yellowness of vegetation leaves will significantly affect the vegetation index; When the crop leaves grow most vigorously, the NDVI and EVI values ​​reach their peak values. After entering the scorching yellow stage, the NDVI and EVI values ​​will drop significantly. Based on this phenological feature, two new normalized indices Nor NDVI and Nor EVI , to demarcate the peanut and soybean sample points; Nor NDVI and Nor EVI The calculation formulas are: In the formula, NDVI may and EVI may are the NDVI and EVI values ​​of the images in May in the time series image set. NDVI sep and EVI sep They are the NDVI and EVI values ​​of the September images in the time series image set.

5. The peanut planting distribution extraction method based on phenological characteristics according to claim 4 is characterized in that: The determining of the number of characteristic variables specifically includes: From the acquired Sentinel-2 images, three original spectral bands, B6, B11 and B12, were selected as original feature variables; Seven vegetation indices, NDVI, EVI, LSWI, NDSVI, NDTI, GCVI and REP, were selected as vegetation characteristic variables; Calculate the NDVI value and EVI value of each time-series image in the time-series image set respectively to obtain the NDVI map and EVI map, and use the gray-level co-occurrence matrix to generate the texture feature quantity of the NDVI map and the EVI map; In the entire time series image set, the 10th, 25th, 50th, 75th, 90th and 25-75th percentiles of each pixel are selected as image percentile feature variables; Harmonic regression analysis was performed on the curves of the three vegetation indices NDVI, EVI and LSWI in the time series image collection to fit the periodic changes in the time series data. The periodic components in the data were estimated by linear regression, and the harmonic analysis results, i.e. the parameters obtained by fitting, were used as periodic characteristic variables. The number of characteristic variables is determined based on the original characteristic variables, vegetation characteristic variables, texture characteristic quantities, image percentile characteristic variables, periodic characteristic variables and constructed phenological spectral index characteristics.

6. The peanut planting distribution extraction method based on phenological characteristics according to claim 5, characterized in that: The acquisition of samples with crop type labels after the division specifically includes: Obtain the 20m-level soybean planting distribution dataset in Northeast China in 2021, use it as dataset B, and obtain the initial sample points of soybeans; Obtain the 20m-level peanut planting distribution dataset in Northeast China in 2018, which is generated using Sentinel-2 satellite remote sensing images with an accuracy of 20m. It covers the three provinces in Northeast my country and has two types of land features: peanuts and other crops. The dataset is represented by dataset C, and the initial sample points of peanuts are obtained. According to the initial sample points of soybeans and peanuts obtained, the K-means clustering algorithm is used on the GEE platform to cluster the sample data sets B and C to be divided, which contain characteristic variable information, and assign different crop type labels to the sample points. The crop types are correspondingly labeled as soybeans and peanuts, and are used as samples with crop type labels after division.

7. The peanut planting distribution extraction method based on phenological characteristics according to claim 6, characterized in that: The obtaining of the distribution of peanut planting areas in the study area specifically includes: The obtained samples with crop type labels after division and the constructed image sample datasets of other crops except rice, corn, peanuts and soybeans are used as training samples and input into the random forest classifier. The feature variable time series image set is classified by the random forest classifier to extract the distribution of peanut samples and soybean planting areas in dataset B and dataset C respectively; The distribution of peanut samples and soybean planting areas in the extracted datasets B and C is compared with the statistical data of the main producing counties. By analyzing the comparison results, when there are differences in the comparison results, the peanut and soybean sample points are divided again; when there are no differences in the comparison results, the division is terminated to obtain the trained random forest classifier; The acquired characteristic variable time series image set of the peanut planting area to be extracted is input into the trained random forest classifier to output the distribution of peanut planting areas in the study area.

8. The peanut planting distribution extraction method based on phenological characteristics according to claim 7 is characterized in that: The generation of the time series image set specifically includes: Based on the acquired Sentinel-2 image data in the study area, the L2A-level remote sensing images of the Sentinel-2 satellite covering the study area and the time period were screened, and the QA band was used for cloud removal. The full-band values ​​of each pixel point of the cloud-free image were processed by median synthesis to obtain cloud-free high-quality synthetic images. Using crop image data throughout the entire growth period with a three-month window size, linear interpolation processing is performed on cloud-free high-quality synthetic images to obtain images covering the entire study area without data missing, forming a time series image set.

9. A peanut planting distribution extraction system based on phenological characteristics, characterized in that: include: Other crop sample acquisition module, used to construct image sample datasets of other crops except rice, corn, peanuts and soybeans; The characteristic variable time series image set acquisition module is used to generate a time series image set based on the acquired Sentinel-2 image data in the study area; by analyzing the NDVI and EVI values ​​of crop leaf growth at different periods, the phenological spectral index characteristics used to characterize crop lushness and brownness are constructed; According to the characteristics of the phenological spectral index, by determining the number of characteristic variables, all characteristic variables are added to each time series image in the time series image set to obtain a characteristic variable time series image set; The classifier training module is used to obtain the sample data to be divided containing the feature variable information according to the feature variable time series image set, and divide the sample points and mark the crop types; the samples with crop type marks after division and the image sample data sets of other crops are used as training sets to train the random forest classifier; The peanut planting area distribution extraction module is used to obtain the characteristic variable time series image set of the peanut planting area to be extracted, classify the characteristic variable time series image set through the trained random forest classifier, and obtain the distribution of peanut planting areas in the study area.

Citation Information

Cited By

  • Crop planting structure extraction method based on hierarchical extraction and multi-feature integration

    CN120708048A