Peanut extraction method based on time sequence characteristics
The timing characteristics of crops are extracted and the characteristics are selected through Sentinel-2 time series data. Combined with the random forest classification method, the problem of difficulty in distinguishing different crops in a single-phase remote sensing image is solved, and effective extraction of peanuts and accurate identification of planting structures is achieved.
Patent Information
- Application Number
- CN202510214068.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to effectively distinguish different crop planting areas on single-stage remote sensing images, especially when the crop phenology period is small, and it is difficult to effectively distinguish different crops based on the timing characteristics of NDVI alone.
The time series data of Sentinel-2 are used to extract the time series spectral reflectivity, vegetation index, red edge index and texture feature vectors of peanuts and other typical crops. Diagnostic feature extraction is achieved through feature optimization, and effective extraction of peanuts is achieved in combination with random forest classification methods.
By extracting multiple timing characteristics and optimizing characteristics, the differences between different crops are maximized, effective extraction of peanuts and accurate identification of planting structure distribution are achieved.
Smart Images

Figure CN120147882A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of satellite image feature extraction, and specifically to a peanut extraction method based on temporal features. Background Art
[0002] The extraction of crop planting structures usually uses single-phase images, mainly including visual interpretation, unsupervised and supervised classification methods based on image statistical classification, and various integration methods based on image semantic texture information. However, different crops often exhibit similar spectral characteristics on the same-phase remote sensing images. Therefore, it is difficult to effectively distinguish the planting areas of different crops on single-phase remote sensing images. In recent years, the crop recognition method based on phenological features has added the concept of key phenological periods, using the phase change law of time-series remote sensing data and comprehensively considering the key phenological period features of crops for crop type recognition. It can effectively solve the problems encountered by the method based on single-phase image spectral features, making crop type recognition more targeted and avoiding the blind use of satellite remote sensing data. In the current research on crop classification based on temporal features, most are based on NDVI time-series data to achieve the extraction of crop planting structures by utilizing the differences in phenological periods among crops. However, when the differences in crop phenological periods are small, it is often difficult to effectively distinguish different crops only based on the temporal features of NDVI. Summary of the Invention
[0003] The purpose of the present invention is to address the above deficiencies by providing a peanut extraction method based on temporal features. By taking peanuts as the research object, using Sentinel-2 time-series data, extracting the temporal features such as spectra, vegetation indices, red-edge indices, and textures of peanuts and other typical crops, and then realizing the extraction of diagnostic features through feature optimization to maximize the differences between different crops. On this basis, the effective extraction of peanuts is achieved by combining the random forest classification method.
[0004] To achieve the above purpose, the present invention adopts the following technical solutions:
[0005] A peanut extraction method based on temporal features, comprising the following steps:
[0006] Step 1: Obtain long-temporal surface reflectance data;
[0007] Obtain Sentinel-2 time-series image data covering the peanut phenological period. After the preprocessing processes of radiometric calibration, atmospheric correction, and geometric correction, convert the DN value of the image into surface reflectance information to obtain long-temporal surface reflectance data;
[0008] Step 2: Combine the surface reflectance data obtained in Step 1 to extract the temporal spectral reflectance, vegetation index, red-edge index, and texture feature vectors of peanuts and other typical crops;
[0009] Step 3: For the temporal feature vectors extracted in Step 2, calculate the Euclidean distance between the feature vectors of peanuts and other crops, and perform normalization processing on it to evaluate the distinguishability of this feature between peanuts and other crops, and achieve feature optimization;
[0010] Step 4: Based on the feature optimization results, such as red-edge position index, contrast, chlorophyll sensitivity index, entropy, ratio vegetation index, and correlation, construct a temporal feature sample set;
[0011] Step 5: Obtain the temporal feature sample set D, and each sample element includes the red-edge position index REPi, contrast Contrasti, chlorophyll sensitivity index MTCIi, entropy Entropyi, ratio vegetation index RVIi, correlation Correlationi, and yi crop type;
[0012] Step 6: Based on the random forest machine learning classification method, train the samples to construct a random forest classifier model;
[0013] Step 7: Realize peanut extraction.
[0014] Further, in Step 1, the radiometric calibration is a processing process of converting the digital quantization value (DN) of the image into a physical quantity (radiance);
[0015] The atmospheric correction is a process of converting the radiance of the image into the true surface reflectance of the ground object;
[0016] The geometric correction is to select a certain number of ground control points on the image, and use the DEM data within the image range to correct the tilt and projection error of the image at the same time, and give the image real plane and elevation information.
[0017] Further, in Step 2, the extraction of temporal spectral reflectance: extract 4 spectral reflectance features of blue-band reflectance, green-band reflectance, red-band reflectance, and near-infrared band reflectance;
[0018] Extraction of vegetation indices: According to Formula 1 - Formula 3, extract 3 vegetation index features of normalized vegetation index, enhanced vegetation index, and ratio vegetation index:
[0019]
[0020] Among them, NDVI is the normalized vegetation index, EVI is the enhanced vegetation index; RVI is the ratio vegetation index, ρ NIR is the near-infrared band reflectance, ρ RED is the red-band reflectance.
[0021] Further, red-edge index extraction: According to Formulas 4 - 12, extract nine red-edge index features including red-edge position index, red-edge normalized difference vegetation index, plant senescence reflectance index, red-edge chlorophyll index, new inverted red-edge chlorophyll index, modified chlorophyll absorption ratio index, modified simple ratio vegetation index, chlorophyll-sensitive index, and normalized red-edge index:
[0022]
[0023]
[0024] Among them, REP is the red-edge position index, NDVIRE is the red-edge normalized difference vegetation index, PSRI is the plant senescence reflectance index, CIRE is the red-edge chlorophyll index, IRECI is the new inverted red-edge chlorophyll index, MCARI is the modified chlorophyll absorption ratio index, MSRRE is the modified simple ratio vegetation index, MTCI is the chlorophyll-sensitive index, NERE is the normalized red-edge index, and B3 - B8 are the band numbers of Sentinel-2 data.
[0025] Further, texture feature vector extraction: According to Formulas 13 - 17, extract energy, entropy, contrast, homogeneity, and correlation:
[0026]
[0027] Among them, i and j are the gray levels of the image, quant k is the set total number of gray levels, and p ij is the element of the gray-level co-occurrence matrix p corresponding to gray levels i and j. Mean is the average value of matrix p, and Variance is the variance of matrix p; Energy is energy, Entropy is entropy, Contrast is contrast, Homogeneity is homogeneity, and Correlation is correlation.
[0028] Further, in Step 3, calculate the Euclidean distance between the feature vectors of peanuts and other crops according to Formula 18:
[0029]
[0030] Among them, xi represents the eigenvalue of peanuts at the i-th time phase; yi represents the eigenvalue of other crops at the i-th time phase; d(x,y)k represents the distance between the k-th type of feature peanuts and other crops. The closer this value is to 1, the smaller the difference between the time-series feature vectors of the two crops, and vice versa, the closer it is to 0, the greater the time-series difference between the two types of crops.
[0031] Further, in step five, REPi, Contrasti, MTCIi, Entropyi, RVIi, and Correlationi each contain time series of t periods. Therefore, a total of 6t features are included in one sample data:
[0032]
[0033] Further, in step six, the construction of the random forest classifier model includes the following steps:
[0034] First, use the resampling technique to draw about 2 / 3 of the subsamples from the samples to construct a training sample set Dt;
[0035] Second, based on the training set Dt, create a CART decision tree. During the creation process, the splitting rule for each node is to first randomly select m features from all the features, and then select the optimal splitting point from these m features to make the left and right subtree divisions;
[0036] Third, repeat the first and second steps N times to create N decision trees to form a random forest classifier.
[0037] Further, in step seven, first input the time-series feature image data into the constructed random forest classifier, and use the voting method for classification prediction. The final classification result of each pixel point is determined by the mode of the output categories of each decision tree. Finally, extract the pixel points belonging to the peanut category to form a peanut planting structure distribution map.
[0038] The beneficial effects of the present invention are:
[0039] In practical applications, by obtaining long-term surface reflectance data, extracting the time-series spectral reflectance, vegetation indices, red-edge indices, and texture feature vectors of peanuts and other typical crops, calculating the Euclidean distance between the time-series feature vectors of peanuts and other crops through the time-series feature vectors, and normalizing it to evaluate the distinguishability of this feature between peanuts and other crops, realizing feature optimization. Based on the feature optimization results, such as the red-edge position index, contrast, chlorophyll sensitivity index, entropy, ratio vegetation index, and correlation, construct a time-series feature sample set. By obtaining the time-series feature sample set D, each sample element contains the optimized feature vector and crop category information. Based on the random forest machine learning classification method, train the samples to construct a random forest classifier model to achieve peanut extraction; the present invention takes peanuts as the research object, uses Sentinel-2 time-series data, extracts the spectral, vegetation index, red-edge index, texture, etc. time-series features of peanuts and other typical crops, and then realizes the extraction of diagnostic features through feature optimization, maximizing the differences between different crops. On this basis, combined with the random forest classification method, the effective extraction of peanuts is realized. Description of the Drawings
[0040] Figure 1 It is a flowchart of the method steps of the present invention. Detailed Description of the Invention
[0041] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments in the following description are only examples, and other obvious variations can be conceived by those skilled in the art.
[0042] As Figure 1 shown, a peanut extraction method based on temporal characteristics includes the following steps:
[0043] Step 1: Obtain long-term surface reflectance data;
[0044] Obtain Sentinel-2 time series image data covering the peanut phenological period. After preprocessing processes such as radiometric calibration, atmospheric correction, and geometric correction, convert the DN value of the image into surface reflectance information to obtain long-term surface reflectance data;
[0045] Step 2: Combine the surface reflectance data obtained in Step 1 to extract the temporal spectral reflectance, vegetation index, red edge index, and texture feature vectors of peanuts and other typical crops;
[0046] Step 3: For the temporal feature vectors extracted in Step 2, calculate the Euclidean distance between the peanut and other crop feature vectors and normalize it to evaluate the distinguishability of this feature between peanuts and other crops and achieve feature optimization;
[0047] Step 4: After feature optimization, use the red edge position index, contrast, chlorophyll sensitivity index, entropy, ratio vegetation index, and correlation as feature vectors to construct a temporal feature sample set;
[0048] Step 5: Obtain the temporal feature sample set D, and each sample element includes the red edge position index REPi, contrast Contrasti, chlorophyll sensitivity index MTCIi, entropy Entropyi, ratio vegetation index RVIi, correlation Correlationi, and yi crop type;
[0049] Step 6: Based on the random forest machine learning classification method, train the samples to construct a random forest classifier model;
[0050] Step 7: Realize peanut extraction.
[0051] As Figure 1As shown, in Step 1, the radiometric calibration is a process of converting the digital quantization value (DN) of an image into a physical quantity (radiance); in this embodiment, by converting the digital quantization value (DN) of the image into a physical quantity (radiance), the error of the sensor itself is eliminated;
[0052] The atmospheric correction is a process of converting the radiance of an image into the true surface reflectance of the ground object; in this embodiment, by converting the radiance of the image into the true surface reflectance of the ground object, the radiation error caused by the effects of atmospheric reflection, absorption, scattering, etc. is eliminated.
[0053] The geometric correction is to select a certain number of ground control points on the image and use the DEM data within the image range to correct the tilt and projection difference of the image simultaneously, and endow the image with real plane and elevation information.
[0054] As Figure 1 shown, in Step 2, the extraction of temporal spectral reflectance: extract four spectral reflectance features of the blue band reflectance (central wavelength 0.49μm), green band reflectance (central wavelength 0.56μm), red band reflectance (central wavelength 0.665μm), and near-infrared band reflectance (central wavelength 0.842μm);
[0055] Extraction of vegetation indices: According to Formula 1 - Formula 3, extract three vegetation index features of the normalized difference vegetation index, enhanced vegetation index, and ratio vegetation index:
[0056]
[0057] where NDVI is the normalized difference vegetation index, EVI is the enhanced vegetation index; RVI is the ratio vegetation index, ρ NIR is the near-infrared band reflectance, ρ RED is the red band reflectance.
[0058] As Figure 1 shown, the extraction of red-edge indices: According to Formula 4 - Formula 12, extract nine red-edge index features of the red-edge position index, red-edge normalized difference vegetation index, plant senescence reflectance index, red-edge chlorophyll index, new inverted red-edge chlorophyll index, improved chlorophyll absorption index, modified ratio vegetation index, chlorophyll sensitivity index, and normalized red-edge index:
[0059]
[0060] Among them, REP is the red-edge position index, NDVIRE is the red-edge normalized vegetation index, PSRI is the plant senescence reflectance index, CIRE is the red-edge chlorophyll index, IRECI is the new inverted red-edge chlorophyll index, MCARI is the modified chlorophyll absorption index, MSRRE is the modified ratio vegetation index, MTCI is the chlorophyll sensitivity index, NERE is the normalized red-edge index, and B3 - B8 are the band numbers of Sentinel-2 data.
[0061] As Figure 1 shown, texture feature vector extraction: According to Formulas 13 - 17, extract energy, entropy, contrast, homogeneity, and correlation:
[0062]
[0063]
[0064] Among them, i and j are the gray-level series of the image, quant k is the set total number of gray levels, p ij is the element of the gray-level co-occurrence matrix p corresponding to gray levels i and j, Mean is the average value of matrix p, Variance is the variance of matrix p; Energy is energy, Entropy is entropy, Contrast is contrast, Homogeneity is homogeneity, and Correlation is correlation.
[0065] As Figure 1 shown, in Step 3, calculate the Euclidean distance between the feature vectors of peanuts and other crops according to Formula 18:
[0066]
[0067] Among them, xi represents the eigenvalue of peanuts at the i-th time phase; yi represents the eigenvalue of other crops at the i-th time phase; d(x,y)k represents the distance between the k-th type of characteristic peanuts and other crops. The closer this value is to 1, the smaller the difference between the time-series feature vectors of the two crops, and vice versa, the closer it is to 0, the greater the time-series difference between the two types of crops.
[0068] As Figure 1 shown, in Step 5, REPi, Contrasti, MTCIi, Entropyi, RVIi, and Correlationi each contain a time series of t periods. Therefore, a total of 6t features are included in one sample data:
[0069]
[0070] As Figure 1 shown, in Step 6, the construction of the random forest classifier model includes the following steps:
[0071] In the first step, a reset sampling technique is used to extract approximately two-thirds of the sub-samples from the sample to construct a training sample set Dt;
[0072] In the second step, based on the training set Dt, a CART decision tree is created. During the creation process, the splitting rule for each node is to randomly select m features from all the features first, and then select the optimal splitting point from these m features to make the division of the left and right sub-trees;
[0073] In the third step, the first and second steps are repeated N times to create N decision trees to form a random forest classifier.
[0074] As Figure 1 shown, in step seven, the time-series feature image data (time-series red-edge position index data, contrast data, chlorophyll sensitivity index data, entropy data, ratio vegetation index data, and correlation data) is first input into the constructed random forest classifier, and voting is used for classification prediction. The final classification result of each pixel point is determined by the mode of the output classes of each decision tree. Finally, the pixel points belonging to the peanut category are extracted to form a peanut planting structure distribution map.
[0075] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Those skilled in the art of the present invention can make various modifications or supplements to the described specific embodiments or use similar methods for substitution, but will not deviate from the scope defined by the spirit of the present invention.
Claims
1. A peanut extraction method based on time series characteristics, characterized in that: The following steps are involved: Step 1: Obtain long-term surface reflectivity data; Acquire Sentinel-2 time series image data covering the flower biological phase, and after the preprocessing of radiometric calibration, atmospheric correction and geometric correction, convert the DN value of the image into surface reflectance information to obtain long-term surface reflectance data; Step 2: Combine the surface reflectance data obtained in step 1 to extract the time series spectral reflectance, vegetation index, red edge index and texture feature vector of peanuts and other typical crops; Step 3: for the temporal texture feature vector extracted in step 2, the Euclidean distance between the feature vectors of peanuts and other crops is calculated, and normalized to evaluate the distinguishability of the feature between peanuts and other crops, so as to achieve feature optimization; Step 4: Based on the feature optimization results obtained in step 3, a time series feature sample set including the crop growth cycle is constructed; Step 5: Obtain a time series feature sample set, where each sample element includes an optimized feature vector and crop category information; Step 6: Train the samples based on the random forest machine learning classification method and build a random forest classifier model; Step seven: realize peanut extraction.
2. The peanut extraction method based on time series characteristics according to claim 1, characterized in that: In step 1, the radiation calibration is a process of converting the digital quantization value (DN) of the image into a physical quantity (radiance); The atmospheric correction is a process of converting the radiance of the image into the real surface reflectivity of the ground object; The geometric correction is to select a certain number of ground control points on the image and use the DEM data within the image range to simultaneously perform tilt correction and projection error correction on the image, thereby giving the image real plane and elevation information.
3. The peanut extraction method based on time series characteristics according to claim 2, characterized in that: In step 2, time series spectral reflectance extraction: extracting four spectral reflectance features: blue band reflectance, green band reflectance, red band reflectance and near infrared band reflectance; Vegetation index extraction: According to formula 1-formula 3, three vegetation index features, namely normalized vegetation index, enhanced vegetation index and ratio vegetation index, are extracted: Among them, NDVI is the normalized vegetation index, EVI is the enhanced vegetation index; RVI is the ratio vegetation index, ρ NIR is the reflectivity in the near infrared band, ρ RED is the red band reflectivity.
4. The peanut extraction method based on time series characteristics according to claim 3 is characterized in that: Red edge index extraction: According to formula 4-formula 12, nine red edge index features are extracted, including red edge position index, red edge normalized vegetation index, plant senescence reflectance index, red edge chlorophyll index, new inverted red edge chlorophyll index, improved chlorophyll absorption index, modified ratio vegetation index, chlorophyll sensitivity index and normalized red edge index: Among them, REP is the red edge position index, NDVIRE is the red edge normalized vegetation index, PSRI is the plant senescence reflectance index, CIRE is the red edge chlorophyll index, IRECI is the new inverted red edge chlorophyll index, MCARI is the improved chlorophyll absorption index, MSRRE is the modified ratio vegetation index, MTCI is the chlorophyll sensitivity index, NERE is the normalized red edge index, and B3-B8 are the band numbers of Sentinel-2 data.
5. The peanut extraction method based on time series characteristics according to claim 4 is characterized in that: Texture feature vector extraction: According to Formula 13-Formula 17, energy, entropy, contrast, homogeneity and correlation are extracted: Among them, i and j are the gray levels of the image, quant k is the total number of gray levels set, p ij is the element of the gray-level co-occurrence matrix p corresponding to gray levels i and j, Mean is the average value of matrix p, Variance is the variance of matrix p, Energy is energy, Entropy is entropy, Contrast is contrast, Homogeneity is homogeneity, and Correlation is correlation.
6. The peanut extraction method based on time series characteristics according to claim 1, characterized in that: In step 3, the Euclidean distance between the feature vectors of peanuts and other crops is calculated according to formula 18: Among them, xi represents the characteristic value of peanuts in the i-th phase; yi represents the characteristic value of other crops in the i-th phase; d(x,y)k represents the distance between the k-th characteristic peanuts and other crops. The closer the value is to 1, the smaller the difference between the time series characteristic vectors of the two crops. Conversely, the closer it is to 0, the greater the time series difference between the two types of crops.
7. The peanut extraction method based on time series characteristics according to claim 1, characterized in that: In step 5, the sample elements in the time series feature sample set D include the optimized feature vectors and crop category information, such as red edge position index REPi, contrast Contrasti, chlorophyll sensitivity index MTCIi, entropy Entropyi, ratio vegetation index RVIi, correlation Correlationi and yi crop type. Each feature contains a time series of t periods, so a sample data includes 6t features in total:
8. The peanut extraction method based on time series characteristics according to claim 1, characterized in that: In step 6, the construction of the random forest classifier model The following steps are involved: In the first step, about 2 / 3 of the subsamples are extracted from the sample using the reset sampling technique to construct a training sample set Dt; The second step is to create a CART decision tree based on the training set Dt. During the creation process, the splitting rule of each node is to first randomly select m features from all features, and then select the optimal splitting point from these m features to divide the left and right subtrees; In the third step, repeat the first and second steps N times to create N decision trees to form a random forest classifier.
9. The peanut extraction method based on time series characteristics according to claim 1, characterized in that: In step seven, the time series feature image data is first input into the constructed random forest classifier, and classification prediction is performed by voting. The final classification result of each pixel is determined by the majority of the output categories of each decision tree. Finally, the pixel points belonging to the peanut category are extracted to form a peanut planting structure distribution map.