Large-area crop identification method and system based on multi-source spatiotemporal distance fusion features
Through the multi-source spatiotemporal distance fusion feature method, using the BGSI and JM distance optimization features, combined with the weaving month probability random forest classifier, the problems of high accuracy and cost in large-scale crop identification are solved, and efficient automatic crop identification is achieved.
Patent Information
- Application Number
- CN202411762121.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-03
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-12-03
AI Technical Summary
In large areas, feature fusion and classifier design of multi-source remote sensing data are difficult to effectively identify crops, especially since crop characteristics at different growth stages vary greatly and differences between classes are small, resulting in high classification accuracy and computational cost, making it difficult to achieve automated and accurate identification.
A multi-source spatiotemporal distance fusion feature method is adopted. By collecting multi-source image data of the crop growth cycle, feature indexes are extracted, and a feature combination variable dataset is constructed. The spatiotemporal bidirectional feature separation index (BGSI) and Jeffries-Matusita (JM) distance index are used to select features. Combined with the distance feature DF, a probabilistic random forest classifier for weaving month is constructed for crop identification.
It improves the accuracy and generalization ability of crop identification in large areas, reduces the complexity and computational cost of the classification model, and improves the accuracy and efficiency of identification.
Smart Images

Figure CN119904741B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multi-source remote sensing feature fusion and crop identification application, and in particular relates to a large-area crop identification method based on multi-source spatiotemporal distance fusion features. Background Art
[0002] Reliable crop type classification and spatial distribution provide a scientific basis for food security and sustainable agricultural development. The essence of automatic and accurate crop type identification lies in the spatiotemporal expression of different crop types and the complex geographic distance relationships. Currently, research is underway on automated sample feature extraction and selection over large areas to achieve precise mapping of different crops. The development of multi-source, medium- to high-resolution remote sensing data has promoted refined spatiotemporal monitoring of agriculture, but has overlooked the presence of feature distance differences during the complex crop sowing process. Especially within large regional spaces, multi-image source feature expression, spatial distance feature fusion, and optimized classification algorithms play a key role in determining the accuracy and timeliness of crop classification over large areas. Therefore, the multi-source feature extraction and fusion of optical, radar, and SRTM elevation data face numerous challenges in crop classification. In crop identification using multi-source remote sensing image feature fusion, multi-source feature fusion provides the classifier with richer input information, thereby improving crop classification accuracy. In addition, since different crops have different characteristics and morphologies at different growth stages, and large differences within crop classes and small differences between classes easily increase the complexity and computational cost of the classifier, it is necessary to enhance the generalization ability of the classification model for crop identification in large areas. Therefore, optimizing feature selection fusion and optimizing classifier design have become urgent issues to be solved in order to improve crop identification accuracy and automation. To this end, the present invention proposes a large-area crop identification method based on multi-source spatiotemporal distance fusion features. Summary of the Invention
[0003] To address the problem of automated and accurate crop identification over large areas, this paper proposes a large-scale crop identification method that integrates multi-source spatiotemporal distance features. This method achieves higher accuracy in identifying different crops and improves crop identification capabilities over large areas.
[0004] The above-mentioned object of the present invention can be achieved by the following technical solution: a large-area crop identification method based on multi-source spatiotemporal distance fusion features, specifically comprising the following steps:
[0005] Step S1, collecting available multi-source image data during the crop growth cycle, extracting characteristic indexes, and constructing a feature combination variable dataset;
[0006] Step S2: Feature optimization is performed based on the spatiotemporal bidirectional feature separation index (BGSI). The importance and separation of different crop variables are calculated using the JM distance index. When the JM value is greater than a certain threshold, step S3 is entered.
[0007] Step S3: Express the preferred spatiotemporal features of different crops in a large area, and combine the preferred features with the distance features DF to construct an enhanced dataset that integrates the preferred features and the distance features.
[0008] Step S4, constructing a random forest classifier for the probability of weaving the network month;
[0009] Step S5: input the enhanced data set into the weaving month probability random forest classifier for classification and recognition.
[0010] Furthermore, the multi-source available image data in step S1 includes synthetic aperture radar data Sentinel1, optical image data Sentinel2, and SRTM elevation data, and the characteristic index includes an optical characteristic index, a radar characteristic index, and an SRTM characteristic index. The characteristic combination variable data set includes: a characteristic index combination corresponding to Sentinel2+Sentinel1+SRTM; a characteristic index combination corresponding to Sentinel2+SRTM; and a characteristic index combination corresponding to Sentinel1+SRTM.
[0011] Furthermore, the optical characteristic index includes:
[0012] (1) Normalized Difference Vegetation Index:
[0013]
[0014] In the above formula, ρ Nir is the near-infrared band, ρ Red is the red band;
[0015] (2) Renormalized vegetation index:
[0016] REP=705+35*(0.5*(ρ red3 +ρ red )-ρ red1 ) / (ρ red2 -ρ red1 )) / 1000 (2)
[0017] In the above formula, ρ red3 is the red light reflectance, ρ red is the near-infrared light reflectivity, ρ red1 is the reflectivity of the boundary between red light and near-infrared light, which is very sensitive to the chlorophyll concentration of vegetation. red2is the reflectivity of another key point on the red edge, which is used to measure the slope change of the red edge;
[0018] (3) Canopy reflectance index:
[0019]
[0020] In the above formula, ρ Red1 is the red edge light reflectivity, ρ Nir is the near-infrared light reflectance;
[0021] (4) Improved soil adjustment vegetation index 2:
[0022]
[0023] In the above formula, ρ Nir is the near-infrared band, ρ Red is the red band;
[0024] (5) Land surface water index:
[0025]
[0026] In the above formula, ρ Nir is the near-infrared band, ρ swir1 It is the short-wave infrared band;
[0027] (6) Normalized difference tillage index:
[0028]
[0029] In the above formula, ρ swir1 is the near-infrared band, ρ Swir2 It is the red edge band;
[0030] (7) Normalized difference snow index:
[0031]
[0032] In the above formula, ρ Green is the green band, ρ Swir It is the shortwave infrared band.
[0033] Furthermore, radar signature indices include:
[0034] (8) Radar polarization characteristics and indices:
[0035] sum=VV+VH (8)
[0036] (9) Radar polarization characteristic difference index:
[0037] subtract=VV-VH (9)
[0038] (10) Radar surface water index:
[0039]
[0040] (11) Radar reflectivity index:
[0041]
[0042] In the above formula, VV means that the radar wave is vertically polarized during both transmission and reception, that is, it is transmitted vertically upward and received in the vertical direction; VH means that the radar wave is vertically polarized during transmission and horizontally polarized during reception, that is, it is transmitted in the vertical direction and received in the horizontal direction.
[0043] Furthermore, the SRTM characteristic index is calculated as follows:
[0044]
[0045] Where aspect is the slope, It is the rate of change of elevation Z in the X direction, that is, the horizontal gradient, and is the elevation difference between adjacent pixels in the X direction on the horizontal distance; It is the rate of change of elevation Z in the Y direction, that is, the vertical gradient, which is the elevation difference between adjacent pixels in the vertical direction; Z is the elevation value of the terrain; X and Y are the horizontal and vertical coordinate axes in the image, expressed in plane coordinates on the ground; atan2 is the inverse tangent function, which is used to calculate the slope angle.
[0046] Furthermore, in step S2, the intra-class and inter-class distances between different crops are first calculated based on the intra-class and inter-class differences, thereby generating feature separability index features. The calculation formula of the separability index GSI is as follows:
[0047]
[0048] Among them, Inter-class Distance represents the distance between classes, and Intra-class Distance represents the distance within a class; let the eigenvalue be x t , the category label is y t , time is t, for each category C i , calculate the mean μ of all sample eigenvalues in this category i ;
[0049] Calculate the sum of the squared distances of all samples in the category to the mean, which is the intra-class variance. The formula is:
[0050]
[0051] The total intra-class distance is the sum of the intra-class distances of all classes:
[0052] Intra-class Distance=∑ i Intra-class Distance i (15)
[0053] Calculate the feature mean μ of all categories i The mean μ of each category is calculated, and the square distance from the mean of each category to the total mean is multiplied by the number of samples in that category. The formula is:
[0054] Inter-class Distance i =|C i |(μ i -μ) 2 (16)
[0055] The total inter-class distance is the sum of the inter-class distances of all classes:
[0056] Inter-class Distance=∑ i Inter-class Distance i (17)
[0057] Then, the GSI is subjected to double row and column iterative normalization using the Sinkhorn-Knopp algorithm. By iteratively adjusting the rows and columns of the matrix and setting the convergence threshold, the algorithm is considered to have converged when the difference between the sum of all rows and columns is close to 1 and is less than the convergence threshold ∈. The feature separability index BGSI that can represent different crops is calculated.
[0058] Furthermore, JM distance is an indicator that measures the separability between different categories in classification problems. The JM distance J between categories a and b is ab It can be expressed as:
[0059]
[0060] Among them, m a and m b are the feature means of categories a and b, respectively, l a and l b is the corresponding characteristic variance, B ab represents Bhattacharyya distance; J ab Represents the separability measure between sample sets a and b.
[0061] Furthermore, the specific implementation of step S3 is as follows:
[0062] 1) Standardized latitude and longitude data, assuming the original latitude and longitude data is (lat i ,lon i), the standardization can be done using the following formula:
[0063]
[0064] Among them, μ lat and σ lat are the mean and standard deviation of the latitude data, μ lon and σ lon are the mean and standard deviation of the longitude data respectively;
[0065] 2) Calculate the distance between adjacent points, fixed points, or construct a distance matrix as needed; assuming there are n sample points, the Euclidean distance between sample points can be calculated using the following formula:
[0066]
[0067] For each sample point i, find the sample point j that is closest to it:
[0068]
[0069] 3) Use interpolation method to interpolate unknown positions to obtain more eigenvalues; using Kriging interpolation method, first need to establish covariance function C (h) , where h is the distance between sample points, Kriging interpolation estimates The formula is:
[0070]
[0071] in, is the known eigenvalue of sample point i, λ i is the interpolation weight, obtained by solving the following linear equations:
[0072]
[0073] Here, C(d ij ) is the covariance between sample points i and j, C(d i0 ) is the unknown point x0 and the known point x i The covariance of the distance between them, μ is the Lagrange multiplier;
[0074] 4) Merge the distance feature fusion interpolation result with the original 12 features in step 1 to construct an enhanced feature set; 1i , X 2i ,...,X 12i ] and the standardized longitude and latitude data Distance feature And the interpolation results [K 1i , K 2i , ..., Kni ] to obtain the enhanced distance feature set:
[0075]
[0076] Furthermore, in step S4, the weaving semi-monthly probabilistic random forest classifier is based on the grid, and its input is the spatiotemporal characteristics of the samples in the grid and the training samples from 8 adjacent grid tiles. The grid division is based on the standard grid scheme within a large area, and the study area is divided into 4*4 fixed-size grid units. For each sample point in the grid, the classifier votes on the classification results every few days during the crop growing season and selects the category with the most occurrences as the final category.
[0077] Furthermore, in step S5, the recognition results of the classifier are evaluated and optimized using a cross-validation technique. Several measurement indicators are calculated during the evaluation and optimization, including confusion matrix, producer accuracy (PA), user accuracy (UA), overall accuracy (OA), Kappa coefficient and F1-score, etc., to evaluate the mapping accuracy.
[0078] Among them, the F1 score calculation formula is defined as follows:
[0079]
[0080] The present invention also provides a large-area crop identification system based on multi-source spatiotemporal distance fusion features, including a processor and a memory, the memory being used to store program instructions, and the processor being used to call the stored instructions in the memory to execute the large-area crop identification method based on multi-source spatiotemporal distance fusion features as described in the above technical solution.
[0081] Compared with the prior art, the present invention has the following beneficial technical effects:
[0082] (1) The present invention uses a dataset based on a combination of spectral index features, radar index features, and SRTM features. The bidirectional spatiotemporal feature separation index (BGSI) is constructed to determine the optimal feature combination over a large area. The Jeffries-Matusita (JM) distance, an indicator of the separability between different crop features, is used to further determine the importance and separability of the features, thereby ensuring the improvement of the classification accuracy of different crops.
[0083] (2) The present invention further introduces distance features, which work in conjunction with preferred features to promote the expression of spatiotemporal features between different types of crops and to associate complex geographical distances of crops in large areas. By coupling the preferred features of multi-source time series images with the spatial distance features of large areas, it is beneficial to the promotion and application of different crop identification in large areas.
[0084] (3) The web-based monthly probability random forest (WMRF) classification model designed in this paper performs better than the traditional random forest classifier model in crop classification and identification. Under the condition of complex spatiotemporal heterogeneity of large-scale crops, it verifies the application of multi-source remote sensing features in synergistic spatiotemporal distance features in the identification of different crops in a large area. The fusion of multi-source spatiotemporal distance features has a certain effect on reducing the complexity of the classification model and improving the operating efficiency. The synergistic spatiotemporal distance features highlight the importance of multi-source feature fusion in the classification of complex heterogeneity of large-scale crops, providing a theoretical basis and support for the classification of crops in large areas.
[0085] In summary, the present invention extracts and optimally combines the features of multi-source remote sensing images, constructs a spatiotemporal bidirectional feature separation index (BGSI), reduces the time for processing feature data, can further determine the importance and separability of preferred features between different crops, and constructs a collaborative spatiotemporal distance feature dataset of preferred features to promote the complex geographical distance association of crops in large areas. At the same time, a weaving month probability random forest (WMRF) classification model is designed to fully utilize the spatiotemporal distance characteristics of multi-source remote sensing images, thereby improving the accuracy and generalization ability of large-scale crop classification and recognition models. BRIEF DESCRIPTION OF THE DRAWINGS
[0086] Figure 1 This is a flow chart of a large-area crop identification method based on multi-source spatiotemporal distance fusion features provided by an embodiment of the present invention;
[0087] Figure 2 This is the distribution map of field sample collection for multiple crops in the Hetao Plain in the study area in 2023;
[0088] Figure 3 It is the average landing point accuracy and overall accuracy based on the BGSI method integrating multi-source remote sensing feature combination (Sentinel-2, Sentinel-1 and SRTM feature selection);
[0089] Figure 4 It is a comparison of the classification accuracy of the Weaving Month Probabilistic Random Forest (WMRF) classification model under different feature fusion datasets;
[0090] Figure 5 It is a large-scale crop mapping based on multi-source spatiotemporal distance fusion features and a local crop visualization map based on different feature fusion datasets;
[0091] Figure 6 This is the F1 score result of crop classification under the fusion of multi-source spatiotemporal distance features of different classifiers. DETAILED DESCRIPTION
[0092] In order to facilitate those skilled in the art to understand and implement the present invention, Figure 1The following is a flow chart of the technical method. The technical solution of the present invention is further elaborated in detail in combination with the drawings in the embodiments. It should be understood that the embodiments described here are only used to illustrate and explain part of the embodiments of the present invention and are not used to limit the present invention.
[0093] The embodiment of the present invention provides a large-area crop identification method based on multi-source spatiotemporal distance fusion features, which utilizes multi-source remote sensing image feature fusion to improve the timeliness and accuracy of crop identification in a large area. Figure 1 The method provided in the embodiment of the present invention is shown in the flowchart, which specifically includes the following:
[0094] Step 1: Data acquisition and preprocessing: Collect available image data from multiple sources during the crop growth cycle to extract feature variables and construct a feature combination variable dataset;
[0095] The crop sample dataset for the 2023 field survey was divided into 80% of the data as a training set and 20% of the data as a test set. The different sample areas were pre-processed with the time series Sentinel-1 and Sentinel-2 high-resolution remote sensing images and SRTM elevation remote sensing image data of the Hetao Plain from May to October 2023. Figure 2 This is the distribution map of field sample collection for multiple crops in the Hetao Plain in the study area in 2023;
[0096] Step 2: Multi-source data variable acquisition and feature dataset construction: Multi-source image data (time-series optical, time-series SAR, and SRTM data) are used to extract feature index variables and construct a dataset of important combination features. The combined feature dataset includes: the feature index combination corresponding to Sentinel2+Sentinel1+SRTM; the feature index combination corresponding to Sentinel2+SRTM; and the feature index combination corresponding to Sentinel1+SRTM.
[0097] Step 3: Optimization based on spatiotemporal crop phenology features: Feature extraction of corresponding images of different crops during their growth period is performed, and feature optimization is performed based on the spatiotemporal bidirectional feature separation index (BGSI). On this basis, the importance and separation of different crop variables are calculated using the JM distance indicator.
[0098] Step 4. Construct an enhanced data set under distance feature fusion: Use the preferred feature collaborative distance feature (DF) to construct an enhanced data set under different feature fusions. The feature data set includes: feature index combination corresponding to Sentinel2+Sentinel1+SRTM+spatial distance feature (distance feature fusion DFF); feature index combination corresponding to Sentinel2+Sentinel1+SRTM (multi-source feature fusion FF); feature index combination corresponding to Sentinel2+SRTM; feature index combination corresponding to Sentinel1+SRTM; feature index combination corresponding to Sentinel1+SRTM;
[0099] Step 5: Design and optimize the classifier: Use the Web Monthly Probability Random Forest (WMRF) classifier to integrate the recognition probabilities of each month during the crop growth period to achieve recognition of different crops;
[0100] Step 6: Perform Web Monthly Probabilistic Random Forest (WMRF) classification and recognition on the multi-source fusion feature combination dataset. Verify the recognition performance by comparing the results of different classifiers, generate classification results, and evaluate the classification and recognition accuracy to achieve accurate identification of different crops in a large area and crop mapping applications.
[0101] Specifically, feature acquisition is performed by preprocessing the field sample set using multi-source remote sensing images to obtain sample sets corresponding to different feature indices and extracting multi-source feature indices. The specific implementation of step 1 in this embodiment includes the following: The optical feature index is calculated as follows:
[0102] (1) Normalized Difference Vegetation Index:
[0103]
[0104] In the above formula, ρ Nir is the near-infrared band, ρ Red is the red band;
[0105] (2) Renormalized vegetation index:
[0106] REP=705+35*(0.5*(ρ red3 +ρ red )-ρ red1 ) / (ρ red2 -ρ red1 )) / 1000 (2)
[0107] In the above formula, ρ red3 is the red light reflectance, ρ red is the near-infrared light reflectivity, ρ red1 is the reflectivity of the boundary between red light and near-infrared light, which is very sensitive to the chlorophyll concentration of vegetation. red2 is the reflectivity of another key point on the red edge, which is used to measure the slope change of the red edge;
[0108] (3) Canopy reflectance index:
[0109]
[0110] In the above formula, ρ Red1 is the red edge light reflectivity, ρ Nir is the near-infrared light reflectance;
[0111] (4) Improved soil adjustment vegetation index 2:
[0112]
[0113] In the above formula, ρ Nir is the near-infrared band, ρ Red is the red band;
[0114] (5) Land surface water index:
[0115]
[0116] In the above formula, ρ Nir is the near-infrared band, ρ Swir1 It is the short-wave infrared band;
[0117] (6) Normalized difference tillage index:
[0118]
[0119] In the above formula, ρ Swir1 is the near-infrared band, ρ Swir2 It is the red edge band;
[0120] (7) Normalized difference snow index:
[0121]
[0122] In the above formula, ρ Green is the green band, ρ Swir It is the short-wave infrared band;
[0123] The radar signature index is calculated as follows:
[0124] (8) Radar polarization characteristics and indices:
[0125] sum=VV+VH (8)
[0126] In the above formula, VV means that the radar wave is vertically polarized when transmitting and receiving (i.e., it is transmitted vertically upward and received vertically), and VH means that the radar wave is vertically polarized when transmitting and horizontally polarized when receiving (i.e., it is transmitted vertically and received horizontally).
[0127] (9) Radar polarization characteristic difference index:
[0128] subtract=VV-VH (9)
[0129] (10) Radar surface water index:
[0130]
[0131] (11) Radar reflectivity index:
[0132]
[0133] The SRTM characteristic index is calculated as follows:
[0134] (12) Slope:
[0135]
[0136] In the above formula It is the rate of change of elevation (Z) in the X direction (i.e., the horizontal gradient). It is usually the elevation difference in the X direction between adjacent pixels in the horizontal distance; It is the rate of change of elevation (Z) in the Y direction (i.e., the vertical gradient), usually the elevation difference between adjacent pixels in the vertical direction; Z is the elevation value of the terrain; X and Y are the horizontal and vertical coordinate axes in the image, usually expressed in plane coordinates on the ground; atan2 is the inverse tangent function, which is used to calculate the slope angle.
[0137] Step 3: The calculation of the bidirectional spatiotemporal feature separation index (BGSI) and the JM distance index is implemented as follows:
[0138] The purpose of feature optimization is to select key features with higher discriminative power for crop classification, thereby improving classification accuracy and model efficiency. For the bidirectional temporal and spatial feature separation index (BGSI) feature optimization, the intra-class and inter-class distances between different crops are calculated based on the intra-class and inter-class differences, thereby generating a feature separability index (GSI) to quantify the discriminative power of the features.
[0139] Among them, the calculation formula of feature separability index (GSI) is as follows
[0140]
[0141] Let the eigenvalue be x t , the category label is y t , time is t, for each category C i , i is used to identify each sample and calculate the mean μ of all sample feature values in the category i .
[0142] Calculate the sum of the squared distances of all samples in the category to the mean (intra-class variance), the formula is:
[0143]
[0144] Among them C i Represents the sample set of the i-th category. Specifically, t usually represents the index or identifier of a sample, and C i is the set of all samples belonging to category i. So, t refers to the time series data, belonging to category C i a certain moment or a certain sample.
[0145] The total intra-class distance is the sum of the intra-class distances of all categories:
[0146] Intra-class Distance=∑ i Intra-class Distance i (15)
[0147] Calculate the feature mean μ of all categories i The mean μ of each category is calculated, and the square distance from the mean of each category to the total mean is multiplied by the number of samples in that category. The formula is
[0148] Inter-class Distance i =|C i |(μ i -μ) 2 (16)
[0149] The total inter-class distance is the sum of the inter-class distances of all classes:
[0150] Inter-class Distance=∑ i Inter-class Distance i (17)
[0151] The Sinkhorn-Knopp algorithm performs dual iterative row-column normalization on the GSIs of different characteristics of a specific crop at the same time stage and the GSIs of the same characteristics of a specific crop at different time stages to further calculate the importance of spatiotemporal features to the expression of different crops. The Sinkhorn-Knopp algorithm iteratively adjusts the rows and columns of the matrix and sets a convergence threshold. The algorithm is considered converged when the difference between the sum of all rows and columns is close to 1 and is less than the convergence threshold ∈. Therefore, the GSI is combined with the Sinkhorn-Knopp algorithm to calculate the feature separability index (BGSI) that represents different crops. The resulting BGSI matrix is further subjected to feature optimization. Based on the spatial dimension of the matrix, features suitable for the specific conditions of the eastern and western regions can be selected to improve the accuracy and robustness of the classification model. The temporal dimension of the matrix can also be used to identify key features for different crop growth stages.
[0152] Jeffries-Matusita (JM) distance is an indicator that measures the separability between different categories in classification problems. Calculating the Jeffries-Matusita (JM) distance can effectively evaluate the separability of different categories in the feature space, thereby solving the problem of quantifying the importance and separability of different crop characteristic variables. For the calculation of the importance and separability of crop characteristic variables in a large area, JM distance is a classic indicator to measure the separability between different categories. It combines the influence of the mean difference and covariance between categories and reflects the degree of overlap of the category distribution. The size of the JM distance has the following meanings: the larger the distance, the more separable the category; the smaller the distance, the more difficult it is to separate the category. JM distance J of different categories a and b ab It can be expressed as:
[0153]
[0154] Among them, m a and m b are the feature means of categories a and b, respectively, l a and l b is the corresponding characteristic variance, B ab stands for Bhattacharyya distance. J ab Represents the separability measure between sample sets a and b. A JM value greater than 1.4 indicates strong separability between different classes. When strong separability between samples under the feature is verified, an enhanced feature set is constructed.
[0155] Figure 3The mean drop accuracy (MDA) and overall accuracy results are calculated for a dataset using a combination of multi-source remote sensing features (Sentinel-2, Sentinel-1, and SRTM feature selection) using the BGSI method. Mean Decrease Accuracy (MDA) is a method used to assess feature importance, particularly in random forest models. MDA measures the importance of specific features to overall model performance. This method helps identify and select the most influential features, thereby improving model efficiency and accuracy.
[0156] In step 4, the distance matrix of neighboring points and fixed point features is constructed by using the latitude and longitude standardized data; an enhanced feature set is constructed by combining the distance feature fusion Kriging interpolation result with the original 12 features. The specific implementation method is as follows:
[0157] The optimized feature and the collaborative distance feature (DF) will be used to construct an enhanced dataset under different feature fusion;
[0158] By using the latitude and longitude standardized data to construct the neighboring point distance and fixed point feature distance matrix;
[0159] Standardized latitude and longitude data, assuming the original latitude and longitude data is (lat i ,lon i ), the standardization can be done using the following formula:
[0160]
[0161] Among them, μ lat and σ lat are the mean and standard deviation of the latitude data, μ lon and σ lon are the mean and standard deviation of the longitude data respectively.
[0162] The Kriging interpolation method will be used to interpolate the unknown locations to obtain more eigenvalues, and then obtain distance feature fusion; assuming there are n sample points, the Euclidean distance between sample points can be calculated using the following formula:
[0163]
[0164] For each sample point i, find the sample point j that is closest to it:
[0165]
[0166] Establish the covariance function C (h) , where h is the distance between sample points, Kriging interpolation estimates The formula is:
[0167]
[0168] Among them, z (xi) is the known eigenvalue of sample point i, λ i The interpolation weights can be obtained by solving the following linear equations:
[0169]
[0170] Among them, C(d ij ) is the covariance between sample points i and j, C(d i0 ) is the unknown point x0 and the known point x i The covariance of the distance between them, μ is the Lagrange multiplier;
[0171]
[0172] In order to meet the crop growth differences caused by spatial distance in a large area, an enhanced feature set is constructed by merging the distance feature fusion interpolation results with the original 12 features to more comprehensively express the crop characteristics and their regional differences.
[0173] The original 12 features [X 1i , X 2i ,...,X 12i ] and the standardized longitude and latitude data Distance feature And the interpolation results [K 1i , K 2i , ..., K ni ] to obtain the enhanced distance feature set:
[0174]
[0175] The specific implementation of step 5 is as follows:
[0176] A Weaving Web Semi-Monthly Probabilistic Random Forest (WMRF) classification model was optimized and designed based on a grid. Its inputs were the spatiotemporal characteristics of samples within the grid and training samples from eight adjacent grid tiles. The gridding was based on a standard grid scheme used for large regions, dividing the study area into fixed-size 4x4 grid cells. For each sample point in the grid, the model voted on the classification results every 15 days during the crop growing season, selecting the category with the highest number of occurrences as the final category. A consensus voting mechanism was used to improve the accuracy of the classification results and conduct precision evaluation.
[0177] Among them, the model was optimized and cross-validated, and several accuracy evaluation indicators were calculated, including the use of confusion matrix, producer accuracy (PA), user accuracy (UA), overall accuracy (OA), Kappa coefficient and F1-score to evaluate the mapping accuracy.
[0178] Evaluation of classification mapping accuracy The F1 score calculation formula is defined as follows:
[0179]
[0180] Step 6 is implemented as follows:
[0181] By combining datasets with multi-source feature fusion, the capabilities of the Web Month Probabilistic Random Forest (WMRF) classifier were validated. The crop recognition performance using different multi-source feature fusion datasets was verified by comparing the results with those of different classifiers. This enabled accurate identification of different crops over a large area and its application in crop mapping. To evaluate the impact of the WMRF classifier on the ability to classify and identify different crops, the accuracy was compared with that of the RF classification model.
[0182] In order to evaluate the classification and recognition ability of the Weaving Month Probabilistic Random Forest (WMRF) model, the recognition results of different crops under the enhanced dataset containing distance feature fusion constructed in step 4 and the multi-source feature combination dataset are compared. Figure 4 This is a comparison of crop classification accuracy of the Web Month Probabilistic Random Forest (WMRF) classification model under different multi-source feature fusion datasets. Figure 5 It is a large-scale crop mapping under multi-source spatiotemporal distance fusion features under the WMRF classifier and a local crop visualization map under different feature fusion data sets.
[0183] Then, we compare the results of different classifiers to verify the recognition performance and compare the classification accuracy of the three classification models:
[0184] (1) Traditional random forest (RF) model;
[0185] (2) Optimize the partitioned network random forest method (FFRF) model;
[0186] (3) Web Month Probabilistic Random Forest (WMRF) model.
[0187] Experiments comparing the three classifiers were conducted using the same multi-source spatiotemporal distance feature fusion and optimal parameters. RF used all sample points within a large region, FFRF used sample points corresponding to each region of a 4x4 block network, and WMRF used sample points from a 3x3 sliding window combined with adjacent windows in a 4x4 block network. Table 1 compares the crop classification results generated by the three classifiers.
[0188] Table 1 Performance comparison of WMRF classification model and other random forest classifier models
[0189]
[0190] In order to further evaluate the classification and recognition accuracy of different classifiers, the effectiveness of the WMRF classifier in large-area crop mapping applications was demonstrated. Figure 6 This is the F1 score result of crop classification under three different classifiers under the fusion of multi-source spatiotemporal distance features.
[0191] An embodiment of the present invention also provides a large-area crop identification system based on multi-source spatiotemporal distance fusion features, including a processor and a memory, the memory being used to store program instructions, and the processor being used to call the stored instructions in the memory to execute the large-area crop identification method based on multi-source spatiotemporal distance fusion features as described in the above technical solution.
[0192] The above-mentioned specific embodiments are merely embodiments of the preferred fusion of corresponding characteristics of several crops of the present invention, and are not intended to limit the present invention. Based on the technical solutions of the present invention and the relevant inspiration of the above-mentioned embodiments, equivalent alternative improvements and combinations made by those skilled in the art to the above-mentioned specific embodiments should all be included in the scope of protection of the present invention.
Claims
1. A large-area crop recognition method based on multi-source spatiotemporal distance fusion features, characterized by: The following steps are involved: Step S1, collecting available multi-source image data during the crop growth cycle, extracting characteristic indexes, and constructing a feature combination variable dataset; Step S2: Feature optimization is performed based on the spatiotemporal bidirectional feature separation index (BGSI). The importance and separation of different crop variables are calculated using the JM distance index. When the JM value is greater than a certain threshold, step S3 is entered. In step S2, the intra-class and inter-class distances between different crops are first calculated based on the intra-class and inter-class differences, thereby generating the feature separability index feature. Then, the GSI is subjected to double row and column iterative normalization processing using the Sinkhorn-Knopp algorithm. By iteratively adjusting the rows and columns of the matrix and setting the convergence threshold, the algorithm is considered to have converged when the sum of all rows and columns is close to 1 and the difference is less than the convergence threshold ∈. The feature separability index BGSI that can represent different crops is calculated. Step S3: Express the preferred spatiotemporal features of different crops in a large area, and combine the preferred features with the distance features DF to construct an enhanced dataset that integrates the preferred features and the distance features. The specific implementation of step S3 is as follows: 1) Standardized latitude and longitude data; 2) Calculate the distance between adjacent points, fixed points or construct a distance matrix as needed; 3) Use interpolation method to interpolate unknown positions to obtain more eigenvalues; using Kriging interpolation method, first need to establish covariance function C (h) , where h is the distance between sample points, Kriging interpolation estimates The formula is: in, is the known eigenvalue of sample point i, λ i is the interpolation weight, obtained by solving the following linear equations: Here, C(d ii ) is the covariance between sample points i and j, C(d i0 ) is the unknown point x0 and the known point x i The covariance of the distance between them, μ is the Lagrange multiplier; 4) Merge the distance feature fusion interpolation result with the original 12 features in step 1 to construct an enhanced feature set; 1i , X 2i ,...,X 12i ] and the standardized longitude and latitude data Distance feature And the interpolation results [K 1i , K 2i , ..., K ni ] to obtain the enhanced distance feature set: ; Step S4, constructing a random forest classifier for the probability of weaving the network month; Step S5: input the enhanced data set into the weaving month probability random forest classifier for classification and recognition.
2. The large-area crop identification method based on multi-source spatiotemporal distance fusion features according to claim 1, characterized in that: The multi-source available image data in step S1 includes synthetic aperture radar data Sentinel1, optical image data Sentinel2, and SRTM elevation data. The characteristic index includes the optical characteristic index, the radar characteristic index, and the SRTM characteristic index. The characteristic combination variable data set includes: the characteristic index combination corresponding to Sentinel2+Sentinel1+SRTM; the characteristic index combination corresponding to Sentinel2+SRTM; and the characteristic index combination corresponding to Sentinel1+SRTM.
3. The large-area crop identification method based on multi-source spatiotemporal distance fusion features according to claim 2 is characterized by: Optical characteristic indices include: (1) Normalized Difference Vegetation Index: In the above formula, ρ Nir is the reflectivity in the near-infrared band, ρ Red is the red band reflectivity; (2) Renormalized vegetation index: REP=705+35*(0.5*(ρ red3 +r red )-r red1 ) / (ρ red2 -r red1 )) / 1000 (2) In the above formula, ρ red3 is the red light reflectance, ρ red is the near-infrared light reflectivity, ρ red1 is the reflectivity of the boundary between red light and near-infrared light, which is very sensitive to the chlorophyll concentration of vegetation. red2 is the reflectivity of another key point on the red edge, which is used to measure the slope change of the red edge; (3) Canopy reflectance index: In the above formula, ρ Red1 is the red edge light reflectivity, ρ Nid is the reflectivity in the near-infrared band; (4) Improved soil adjustment vegetation index 2: In the above formula, ρ Nir is the reflectivity in the near-infrared band, ρ Red is the red band reflectivity; (5) Land surface water index: In the above formula, ρ Nir is the reflectivity in the near-infrared band, ρ Swir1 It is the short-wave infrared band; (6) Normalized difference tillage index: In the above formula, ρ Swir1 is the near-infrared band, ρ Swir2 It is the red edge band; (7) Normalized difference snow index: In the above formula, ρ Green is the green band, ρ Swir It is the shortwave infrared band.
4. The large-area crop identification method based on multi-source spatiotemporal distance fusion features according to claim 2, characterized in that: Radar signature indices include: (8) Radar polarization characteristics and indices: sum=VV+VH (8) (9) Radar polarization characteristic difference index: subtract=VV-VH (9) (10) Radar surface water index: (11) Radar reflectivity index: In the above formula, VV means that the radar wave is vertically polarized when transmitting and receiving, that is, it is transmitted vertically upward and received in the vertical direction; VH means that the radar wave is vertically polarized when transmitting and horizontally polarized when receiving, that is, it is transmitted in the vertical direction and received in the horizontal direction. The SRTM characteristic index is calculated as follows: Where aspect is the slope, It is the rate of change of elevation Z in the X direction, that is, the horizontal gradient, and is the elevation difference between adjacent pixels in the X direction on the horizontal distance; It is the rate of change of elevation Z in the Y direction, that is, the vertical gradient, which is the elevation difference between adjacent pixels in the vertical direction; Z is the elevation value of the terrain; X and Y are the horizontal and vertical coordinate axes in the image, expressed in plane coordinates on the ground; atan2 is the inverse tangent function, which is used to calculate the slope angle.
5. The large-area crop identification method based on multi-source spatiotemporal distance fusion features according to claim 1 is characterized in that: In step S2, the separability index GSI is calculated as follows: Among them, Inter-class Distance represents the distance between classes, and Intra-class Distance represents the distance within a class; let the eigenvalue be x t , the category label is y t , time is t, for each category C i , calculate the mean μ of all sample eigenvalues in this category i ; Calculate the sum of the squared distances of all samples in the category to the mean, which is the intra-class variance. The formula is: The total intra-class distance is the sum of the intra-class distances of all classes: Intra-class Distance=∑ i Intra-class Distance i (15) Calculate the feature mean μ of all categories i The mean μ of each category is calculated, and the square distance from the mean of each category to the total mean is multiplied by the number of samples in that category. The formula is: Inter-class Distance i =|C i |(m i -m) 2 (16) The total inter-class distance is the sum of the inter-class distances of all classes: Inter-class Distance=∑ i Inter-class Distance i (17) 。 6. The large-area crop identification method based on multi-source spatiotemporal distance fusion features according to claim 1, characterized in that: JM distance is an indicator that measures the separability between different categories in classification problems. The JM distance J between categories a and b is ab Expressed as: Among them, m a and m b are the feature means of categories a and b, respectively, l a and l b is the corresponding characteristic variance, B ab represents Bhattacharyya distance; J ab Represents the separability measure between sample sets a and b.
7. The large-area crop identification method based on multi-source spatiotemporal distance fusion features according to claim 1, characterized in that: The calculation method of standardized latitude and longitude data is: Assuming the original latitude and longitude data is (lat i ,lon i ), the standardization process uses the following formula: Among them, μ lat and σ lat are the mean and standard deviation of the latitude data, μ lon and σ lon are the mean and standard deviation of the longitude data respectively; The calculation method for calculating the distance between neighboring points, fixed point distance or constructing the distance matrix is as follows: Assuming there are n sample points, the Euclidean distance between the sample points is calculated using the following formula: For each sample point i, find the sample point j that is closest to it:
8. The large-area crop identification method based on multi-source spatiotemporal distance fusion features according to claim 1, characterized in that: In step S4, the weaving semi-monthly probabilistic random forest classifier is based on the grid, and its input is the spatiotemporal characteristics of the samples in the grid and the training samples from 8 adjacent grid tiles. The grid division is based on the standard grid scheme in a large area, and the study area is divided into 4*4 fixed-size grid units. For each sample point in the grid, the classifier votes on the classification results every few days during the crop growing season and selects the category with the most occurrences as the final category.
9. The large-area crop identification method based on multi-source spatiotemporal distance fusion features according to claim 1, characterized in that: In step S5, the recognition results of the classifier are evaluated and optimized using a cross-validation technique. Several measurement indicators are calculated during the evaluation and optimization, including confusion matrix, producer accuracy (PA), user accuracy (UA), overall accuracy (OA), Kappa coefficient and F1-score, etc., to evaluate the mapping accuracy. Among them, the F1 score calculation formula is defined as follows:
10. A large-area crop identification system based on multi-source spatiotemporal distance fusion features, characterized by: The method comprises a processor and a memory, wherein the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the large-area crop identification method based on multi-source spatiotemporal distance fusion features as described in any one of claims 1 to 9.