A method and system for quantitatively expressing crop spatial heterogeneity
Through correlation analysis and factor analysis, the landscape index is screened, combined with the K-means clustering method, the problem of difficult to quantify the spatial heterogeneity of crops in complex areas is solved, and high-precision crop distribution description and classification are achieved.
Patent Information
- Application Number
- CN202310425522.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-17
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2043-04-17
AI Technical Summary
The prior art is difficult to identify and quantify crop spatial heterogeneity in complex areas with high accuracy, and the high correlation of landscape index results in redundant data, making it impossible to effectively describe crop distribution status.
The initial landscape index was screened by correlation analysis and factor analysis methods, the farmland landscape was partitioned through K-means clustering method, and the crop spatial heterogeneity quantification expression system was established, and the crop distribution status was described using common factor vectors and cluster analysis.
It improves the quantitative accuracy and robustness of crop distribution, reduces the redundancy of landscape index, realizes quantitative description of farmland landscape, and improves the accuracy of remote sensing recognition and classification.
Smart Images

Figure CN116469008B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing image processing of farmland landscape areas, and in particular relates to a method and system for quantitatively expressing crop spatial heterogeneity. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] The spatial distribution of crop types is crucial for crop growth monitoring, yield estimation, and disaster assessment. Accurately describing crop spatial distribution facilitates the classification of crops at different growing stages, from local to global scales, and facilitates more targeted analysis of the mechanisms by which crop distribution influences crop identification and area estimation using remote sensing.
[0004] In recent years, scholars at home and abroad have utilized various key technologies to conduct research on precise mapping of crop spatial distribution, achieving significant progress in the theory, methods, and applications of remote sensing crop identification. Current research suggests that in large-scale crop-producing areas with regular crop distribution and concentrated distribution of a single crop type, high-precision crop mapping can be achieved using only single-temporal imagery. However, in my country, regions with complex crop spatial heterogeneity, characterized by fragmented farmland plots, complex crop planting structures, and diverse planting patterns, remote sensing imagery poses challenges for high-precision crop identification. Previous studies have found that indicators of farmland landscape heterogeneity within the monitoring area, such as crop fragmentation, plot shape, and clustering, significantly impact the accuracy of remote sensing crop identification and classification algorithms. Therefore, quantifying regions with complex crop spatial heterogeneity and identifying crops based on their distribution characteristics is a key area of current research.
[0005] The inventors found that existing research on the quantitative expression of crop spatial heterogeneity mainly involves two problems:
[0006] First, how to quantify crop spatial heterogeneity? On the one hand, domestic and international scholars often use qualitative terms to describe crop spatial distribution from the perspective of agricultural landscapes, such as smallholder agriculture, heterogeneous landscapes, and fragmented areas, which lack theoretical basis. Others use single indicators such as average cultivated land patch area to measure spatial heterogeneity, but this fails to provide a comprehensive description of crop distribution. On the other hand, existing studies treat regions with complex crop spatial heterogeneity as a whole, without zoning crop distribution, and ignoring the complex crop types caused by internal heterogeneity. As a result, fragmented cultivated land areas are difficult to identify with high precision.
[0007] Second, how can landscape ecology knowledge be integrated into the expression of crop spatial heterogeneity? From a landscape ecology perspective, landscape pattern is generally used to describe landscape patches of varying shapes, sizes, numbers, and spatial arrangements. Landscape indices are indicators for landscape pattern analysis and important parameters reflecting its pattern, representing a major innovation in landscape ecology research methods. Incorporating landscape indices into the quantification of crop spatial heterogeneity can fully capture crop spatial distribution. However, many landscape indices are highly correlated, leading to data redundancy. Summary of the Invention
[0008] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a method and system for quantitatively expressing crop spatial heterogeneity. Targeting agricultural landscapes with complex and diverse crop spatial distributions, the method provides a theoretical basis and guidance for crop mapping and the selection of spatial spectral features.
[0009] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:
[0010] A first aspect of the present invention provides a method for quantitatively expressing crop spatial heterogeneity.
[0011] A method for quantitatively expressing crop spatial heterogeneity comprises the following steps:
[0012] Obtain remote sensing images of the study area, preprocess the remote sensing images, and obtain a sample data set;
[0013] Select the initial landscape indicators and calculate the initial landscape index of the study area based on the sample data set;
[0014] The correlation analysis of the initial landscape indices was performed, and the initial landscape indices with strong correlation reflecting the same spatial characteristics were eliminated to obtain the optimal landscape indices;
[0015] A factor analysis model including a common factor vector was established, and factor analysis was performed on the optimal landscape index to estimate the common factor score of the common factor vector. Based on the common factor score, the contribution rate of the optimal landscape index to each type of common factor was determined. The optimal landscape index was screened to obtain the optimal landscape index associated with describing the importance of each type of common factor. The optimal landscape index associated with describing the importance of each type of common factor was used as the crop spatial heterogeneity index.
[0016] The K-means clustering method was used to divide the sample data set into plot categories, and the common factor score of each plot category was calculated. According to the common factor score of each plot category, the crop spatial heterogeneity index of each common factor was used to describe the ecological significance represented by the common factor, and then the crop distribution status and classification results of each plot category were described.
[0017] Preferably, the initial landscape indicators include:
[0018] Area_perimeter_density index: percentage of landscape type area, average patch area, landscape element patch density, edge density;
[0019] Shape indices: landscape shape index, area-weighted average shape index, area-weighted average patch fractal dimension;
[0020] Spread index: average minimum distance, landscape separation, aggregation index, spread index;
[0021] Diversity index: fragrant diversity index, uniformity index.
[0022] Preferably, Spearman correlation analysis is used for the initial landscape index to reflect the degree of correlation between the indices, and the rank correlation coefficient is calculated for the data of the initial landscape index at the patch type and landscape level respectively. The closer the absolute value of the rank correlation coefficient is to 1, the stronger the correlation is.
[0023] Preferably, a factor analysis model including a common factor vector is established, factor analysis is performed on the optimal landscape index, and the common factor score of the common factor vector is estimated, specifically including:
[0024] S104-1: Conduct correlation analysis on the optimal landscape indicators to test whether the optimal landscape indicators are suitable for factor analysis;
[0025] S104-2: Establish a factor analysis model, standardize the sample data set, and obtain the correlation matrix of the standardized sample data set;
[0026] S104-3: Solve the factor loading matrix of the factor analysis model based on principal component analysis and determine the common factor variables;
[0027] S104-4: Rotate the factor loading matrix using the varimax method;
[0028] S104-5: Estimate the factor scores of the common factor vector based on the rotated factor loading matrix, judge the importance of the optimal landscape indicator based on the contribution rate of the optimal landscape indicator to each type of common factor, and screen the optimal landscape indicator to finally obtain the indicator that describes the spatial heterogeneity of crops.
[0029] Preferably, the factor analysis model is:
[0030] X=A×F+E
[0031] Among them, X is the original variable vector, A is the common factor loading matrix, F is the common factor vector, and E is the impact of the common factor on the variance of the data.
[0032] Preferably, the score of a single common factor is obtained by using a regression method, and the score function is:
[0033] F j =c j1 X1+…+c jp X p ;
[0034] Among them, c jp Factor F j In the scalar X p The score on ; j = 1, 2,…m.
[0035] Preferably, the K-means clustering method is used to divide the sample data set into categories, specifically including:
[0036] Randomly select k sample data as the initial cluster center;
[0037] Calculate the distance between each data object to be clustered and the initial cluster center, assign the data object to be clustered to the cluster with the smallest distance for classification;
[0038] After classification, recalculate the centroid of each class and update the values of k cluster centers;
[0039] For the updated k cluster centers, S104 - 2 and S104 - 3 are repeatedly executed. When the set conditions are met, the iteration ends and the classification is completed.
[0040] A second aspect of the present invention provides a system for quantitatively expressing crop spatial heterogeneity.
[0041] A quantitative expression system for crop spatial heterogeneity, comprising:
[0042] The dataset acquisition module is configured to: acquire remote sensing images of the study area, pre-process the remote sensing images, and obtain a sample dataset;
[0043] The initial landscape index calculation module is configured to: select initial landscape indicators and calculate the initial landscape index of the study area based on the sample data set;
[0044] The optimal landscape index selection module is configured to: perform correlation analysis on the initial landscape indexes, eliminate the initial landscape indexes with strong correlation reflecting the same spatial characteristics, and obtain the optimal landscape index;
[0045] The crop spatial heterogeneity index acquisition module is configured to: establish a factor analysis model including a common factor vector, perform factor analysis on the optimal landscape index, estimate the common factor score of the common factor vector, determine the contribution rate of the optimal landscape index to each type of common factor based on the common factor score, screen the optimal landscape index, obtain the optimal landscape index associated with describing the importance of each type of common factor, and use the optimal landscape index associated with describing the importance of each type of common factor as the crop spatial heterogeneity index;
[0046] The description module is configured to: use the K-means clustering method to divide the sample data set into plot categories, calculate the common factor score of each plot category, and based on the common factor score of each plot category, describe the ecological significance represented by the common factor through the crop spatial heterogeneity index of each common factor, and then describe the crop distribution status and classification results of each plot category.
[0047] A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps of the method for quantitatively expressing crop spatial heterogeneity as described in the first aspect of the present invention.
[0048] The fourth aspect of the present invention provides an electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor. When the processor executes the program, the steps in the method for quantitatively expressing crop spatial heterogeneity as described in the first aspect of the present invention are implemented.
[0049] One or more of the above technical solutions have the following beneficial effects:
[0050] The present invention provides a method and system for expressing crop spatial heterogeneity. It creatively uses landscape indices to quantitatively evaluate the spatial heterogeneity of crop distribution. A series of initial landscape indices are selected and screened using methods such as correlation analysis and factor analysis to determine quantitative descriptive indicators for farmland landscapes. The system also reduces the correlation and redundancy between indices, effectively avoiding information duplication and improving the robustness of quantitative results. This system achieves a quantitative description of the distribution of farmland landscapes. The research results provide a path for establishing a unified crop heterogeneity evaluation framework and offer guidance and suggestions for improving effective large-scale crop mapping.
[0051] The present invention provides a method and system for quantitatively expressing crop spatial heterogeneity. By performing dimensionality reduction and analytical evaluation on multiple variables through factor analysis, the method can compress the original data while retaining as much information as possible, making the significance of the factors clearer. By performing orthogonal rotation on the factor loading matrix, the common factors can be used to classify all indicator variables, explore the potential factors of the problem, and reasonably summarize the analysis results.
[0052] The present invention provides a method and system for expressing crop spatial heterogeneity, which realizes the partitioning of landscape units through the K-means clustering method, obtains the optimal number of classifications through multiple experiments, and uses a hierarchical cluster analysis method to perform ecological partitioning of landscape units. The research results can be used to classify and discuss crop heterogeneous areas, which helps to improve the accuracy of remote sensing identification and crop classification.
[0053] Advantages of additional aspects of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0055] Figure 1 This is a flow chart of the method according to the first embodiment of the present invention.
[0056] Figure 2 3 is a geographical location distribution data diagram of the study area used in the embodiment of the present invention.
[0057] Figure 3 This is the result of Spearman correlation analysis.
[0058] Figure 4 This is the classification result diagram of configuration heterogeneity.
[0059] Figure 5 This is a system structure diagram of embodiment 2 of the present invention. DETAILED DESCRIPTION
[0060] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0061] It should be noted that the terms used herein are for describing particular embodiments only and are not intended to limit the exemplary embodiments according to the present invention.
[0062] In the absence of conflict, the embodiments of the present invention and the features thereof may be combined with each other.
[0063] Glossary:
[0064] K-means clustering method: K-means clustering algorithm (k-means clustering algorithm), is an iterative clustering analysis algorithm.
[0065] Spearman's correlation coefficient: In statistics, the Spearman rank correlation coefficient, also known as the Spearman correlation coefficient, is named after Charles Edward Spearman. Often denoted by the Greek letter ρ, it is a nonparametric measure of the dependence between two variables. It uses a monotonic equation to evaluate the correlation between two statistical variables. If there are no duplicate values in the data and the two variables are perfectly monotonically correlated, the Spearman correlation coefficient is +1 or -1.
[0066] SNAP: The full name is Stanford Network Analysis Project, which is a very powerful open source tool provided by Stanford University.
[0067] Sen2Cor Processor: Sen2Cor is the processor that generates and formats Sentinel-2Level 2A products. It performs atmospheric, terrain, and cirrus corrections on Atmosphere Level 1C input data. Sen2Cor creates bottom-of-atmosphere (BOA) and optionally terrain- and cirrus-corrected reflectance images. Additionally, it generates quality metrics such as aerosol optical depth, water vapor, scene classification maps, and cloud and snow probability.
[0068] Fragstats: Fragstats software is a professional landscape pattern index calculation tool that helps users more conveniently analyze and control the process of environmental variables.
[0069] Example 1
[0070] This embodiment discloses a method for quantitatively expressing crop spatial heterogeneity.
[0071] like Figure 1 As shown, a method for quantitatively expressing crop spatial heterogeneity includes the following steps:
[0072] Obtain remote sensing images of the study area, preprocess the remote sensing images, and obtain a sample data set;
[0073] Select the initial landscape indicators and calculate the initial landscape index of the study area based on the sample data set;
[0074] The correlation analysis of the initial landscape indices was performed, and the initial landscape indices with strong correlation reflecting the same spatial characteristics were eliminated to obtain the optimal landscape indices;
[0075] A factor analysis model including a common factor vector was established, and factor analysis was performed on the optimal landscape index to estimate the common factor score of the common factor vector. Based on the common factor score, the contribution rate of the optimal landscape index to each type of common factor was determined. The optimal landscape index was screened to obtain the optimal landscape index associated with describing the importance of each type of common factor. The optimal landscape index associated with describing the importance of each type of common factor was used as the crop spatial heterogeneity index.
[0076] The K-means clustering method was used to divide the sample data set into plot categories, and the common factor score of each plot category was calculated. According to the common factor score of each plot category, the crop spatial heterogeneity index of each common factor was used to describe the ecological significance represented by the common factor, and then the crop distribution status and classification results of each plot category were described.
[0077] Furthermore, the step S101: acquiring remote sensing images and related data of the area to be measured to obtain a sample data set of the study area specifically includes:
[0078] S101-1: Determine the scope of the study area;
[0079] S101-2: Acquire remote sensing images of the area and perform preprocessing;
[0080] S101-3: Perform manual visual interpretation based on ground-based measured data to supplement the study area grid.
[0081] Furthermore, the S101-1: determining the scope of the study area specifically includes:
[0082] In order to describe crop spatial heterogeneity more accurately and make it more representative, the study area was divided into a 5 km × 5 km grid, and only areas with cultivated land area greater than 30% were retained for research. Figure 2 shown.
[0083] This study also collected 10-meter crop classification data for Heilongjiang Province in 2019. The overall accuracy of this data was 0.87, with user accuracies of 0.94 for rice, 0.84 for corn, and 0.84 for soybeans. Because the cultivated areas of rice, corn, and soybeans in Heilongjiang Province are large and the classification results are good, this study selected 5 km × 5 km grids in Heilongjiang Province where the cultivated areas of the three major crops, rice, corn, and soybeans, accounted for more than 30%, resulting in a total of 8,362 grids for further research.
[0084] Further, S102-2: remote sensing images of the area are acquired and preprocessed to initially generate a sample set, including:
[0085] For example, during the 2019 crop growing season, 15 Sentinel-2 images with cloud cover less than 10% were collected, using nine bands from each image as input. The image data was preprocessed, including atmospheric correction, format conversion, and image cropping based on the collected farmland data map.
[0086] In the SNAP environment, the top-of-atmosphere (TOA) data were first converted to bottom-of-atmosphere (BOC) reflectance using the Sen2Cor processor. The BOC reflectance was then resampled to a 10-meter spatial resolution using bilinear interpolation. Finally, the images covering the study area were stitched and cropped according to the grid boundaries to obtain a complete remote sensing image of the study area.
[0087] Furthermore, the S101-3: performing manual visual interpretation based on ground-measured data to supplement the study area grid includes:
[0088] The field survey data used in this paper was collected from July to September 2019 within the study area using a handheld GPS device with a positioning accuracy of 5 meters to collect ground feature samples and control point information. During the survey, detailed information was recorded for representative crops such as rice, corn, and soybeans in the study area. The total number of corn, rice, and soybean samples collected across the study area was 2,088, 1,818, and 1,223, respectively.
[0089] Based on the field survey data, grids with less than 100 samples were supplemented by manual visual interpretation with reference to Google Earth.
[0090] Further, S102: preliminary selection of landscape indices and calculation of the landscape indices of the study area using Fragstats; including:
[0091] Landscape indices were selected at both the patch type level and the landscape level by comprehensively considering the landscape spatial structure, spatial characteristics, landscape diversity and the physical meaning of the indicators;
[0092] Landscape indices were calculated using Fragstats software using crop type data of the study area.
[0093] Among them, the 13 landscape indicators selected at the patch type level and the landscape level include: area_perimeter_density index, shape index, spread index, and diversity index.
[0094] Landscape indices are indicators for landscape pattern analysis and can also be used to describe the spatial distribution patterns of different crops. Currently, landscape indices are generally analyzed at three levels: individual patches, patch types composed of several individual patches, and landscapes composed of several patch types. However, due to the large number of patches in the study area and the high computational complexity, patch-level analysis is not used in practical applications. Only indices at the patch type and landscape levels are used.
[0095] Due to the strong correlation between some landscape indicators, this paper comprehensively understands the ecological significance of each indicator and the focus of the landscape structure it reflects, while ensuring the scientificity, comprehensiveness and accessibility of the selected index. According to the characteristics of crop distribution in the study area, a series of indicators in four categories, totaling 13, are selected that have the greatest significance for characterizing farmland landscapes and reflect the most important characteristics of crop distribution.
[0096] Among them, the 13 landscape indicators selected at the patch type level and the landscape level include: area_perimeter_density index, shape index, spread index, and diversity index.
[0097] The area_perimeter_density index includes: percentage of landscape type area (PLAND), average patch area (AREA_MN), landscape element patch density (PD), and edge density (ED).
[0098] Shape indices include landscape shape index (LSI), area-weighted mean shape index (AWMSI), and area-weighted mean patch fractal dimension (FRAC_AM).
[0099] The sprawl index includes: mean nearest distance (MNN / ENN_MN), landscape separation (SPLIT), aggregation index (AI), and sprawl index (CONTAG).
[0100] The diversity indices include the Shannon Diversity Index (SHD I) and the Evenness Index (SHE I).
[0101] The calculation and description of the 13 landscape indices selected in this invention are shown in the following table.
[0102] Table 1 Initial landscape index selected by the present invention
[0103]
[0104]
[0105]
[0106]
[0107] Furthermore, the S103: performing correlation analysis on the initial landscape index, eliminating indicators with strong correlation reflecting the same spatial characteristics, and obtaining the optimal landscape index includes:
[0108] In this invention, Spearman correlation analysis is mainly used to reflect the correlation between the initial landscape indices reflecting different characteristics of crop space. The correlation coefficients of various landscape index data are analyzed and calculated at the patch type and landscape level. The calculation formula of the rank correlation coefficient is as follows:
[0109]
[0110] Among them, n is the number of level pairs of variables x and y, that is, the sample content; x and y are the initial landscape indicators. i is the difference in rank between the same pair (i=1,2,...,n), and r is between -1 and 1.
[0111] A null hypothesis test was performed to determine whether the overall correlation coefficient was 0, and a two-tailed t-test (with significance levels of 0.05 and 0.1, respectively) was used to test the significance of the correlation coefficient.
[0112] When P > 0.05, we considered that there was no correlation between the variables;
[0113] When P ≤ 0.05, there is a correlation between the variables. In this case, the closer the value of |r| is to 1, the stronger the correlation.
[0114] The p-value is the probability that the null hypothesis is true.
[0115] The correlation between different landscape indicators is obtained by the method described in the invention, such as Figure 3 As shown. According to the results of correlation calculation, at the patch and landscape levels, the Spearman rank correlation coefficients between different landscape indices vary greatly. The indices with higher correlation are eliminated, and at least one landscape index of each type is retained. The average value of the sum of the absolute values of the correlation coefficients of landscape separation SPLIT, area-weighted average patch fractal dimension FRAC_AM, uniformity index SHEI, average patch area AREA_MN, and average nearest distance ENN_MN is 0.43, which can be considered to be relatively independent. In summary, the present invention selects 5 optimal landscape indices (SPLI T, FRAC_AM, SHEI, AREA_MN and ENN_MN) for subsequent analysis.
[0116] Furthermore, the S104: factor analysis is performed on the optimal landscape index to screen out indicators describing crop spatial heterogeneity, including:
[0117] S104-1: Conduct correlation analysis on several existing indices;
[0118] S104-2: Establish a factor analysis model and calculate the standardized correlation matrix R;
[0119] S104-3: Solve the factor loading matrix using the principal component analysis method based on the principal component analysis model to determine the common factor variables;
[0120] S104-4: Use the varimax method to rotate the loading matrix to make the contribution rate of each factor to each principal component clearer;
[0121] S103-5: Estimate the factor scores of the common factor vector based on the rotated factor loading matrix. Determine the importance of each common factor based on the contribution rate of the relatively independent landscape index to each type of common factor, and screen them to ultimately obtain an index indicator describing the farmland landscape.
[0122] Furthermore, the step S104-1: determining whether the original variables are suitable for factor analysis includes:
[0123] The basic logic of factor analysis is to utilize the concept of dimensionality reduction and, based on the inherent connections between variables, aggregate numerous complex variables into a small number of independent, common factors that reflect the primary information of the original variables. This requires strong correlations between the original variables. Therefore, when applying factor analysis, it is necessary to first conduct a correlation analysis on the variables to test whether the original variables are suitable for factor analysis.
[0124] The correlation analysis method described above was used to calculate and analyze the correlation of the five selected optimal landscape indices at the patch type level. The process will not be repeated here.
[0125] Furthermore, the step S104-2: establishing a factor analysis model to obtain a standardized matrix R specifically includes:
[0126] The mathematical form of the factor analysis model is as follows:
[0127] X = A × F + E;
[0128] Among them, X is the original variable vector, A is the common factor loading matrix, F is the common factor vector, and E is the impact of the common factor on the variance of the data, which can be ignored.
[0129] The sample correlation coefficient matrix R established after standardization of the original data is as follows:
[0130]
[0131] Elements of
[0132] Where, ρ ii =1,ρij =ρ ji ,ρ ij is the correlation coefficient of the i-th indicator, for example, ρ11 is the correlation coefficient between the area indicator and the lsi indicator, ρ12 is the correlation coefficient between the area indicator and the enn indicator, etc.; a ki ,a kj It is a standardized landscape index.
[0133] Furthermore, the step S104-3: solving the common factor loading matrix based on the principal component analysis method of the principal component analysis model to determine the common factor variables specifically includes:
[0134] Solve the eigenvalues and eigenvectors of a matrix;
[0135] The factor loading matrix was solved based on principal component analysis.
[0136] The eigenvalues and eigenvectors of the matrix are:
[0137] Construct the equation:
[0138] f(λ)=|λI-R|;
[0139] Let f(λ) = 0 and solve the equation;
[0140] Among them, λ that makes the equation have non-zero solutions is the eigenvalue of R, I is the unit matrix, and the n-dimensional non-zero column vector composed of eigenvalue λ is the eigenvector of the matrix.
[0141] The factor loading matrix obtained by principal component analysis is:
[0142] The core of factor analysis is to construct factor variables and combine the original variables into a few factors. According to the principal component theory, the first m principal components with eigenvalues greater than 1 are selected. The eigenvalues obtained above are expressed as λ1≥λ2≥…≥λ p , the corresponding orthogonalized eigenvectors are represented as e1,e2,…,e p , then the estimation of the loading matrix of the principal component factor analysis of matrix R is:
[0143]
[0144] After obtaining the factor loading matrix, the factor analysis model can be expressed as:
[0145]
[0146] Among them, F1, F2, …, F m are common factors, ε1,ε2,…,ε p It is a special factor and can be ignored in actual analysis.
[0147] Further, step S104-4: Rotate the load matrix using the method of maximum variance, including:
[0148] Since the present invention needs to understand the specific meaning of the common factors and further explain their practical significance, but the load matrix of the factors is too comprehensive, resulting in the unclear meaning of the common factors, it is necessary to perform coordinate transformation by rotating the load matrix, so as to simplify the common factors, reduce the comprehensiveness of the matrix, make the contribution rate of each factor to each principal component more clear, and facilitate the analysis of its meaning.
[0149] Let Q be an m-order orthogonal matrix, and perform an orthogonal transformation B = AQ, then
[0150] BB T = AA T ;
[0151] where B is the new factor load matrix after rotation.
[0152] Further, S103-5: Estimate the factor scores of the common factor vectors according to the rotated factor load matrix, judge the importance of each type of common factor according to the contribution rate of the relatively independent landscape index, and screen them, and finally obtain the index indicators describing the farmland landscape. Including:
[0153] For the model X = AF + E, if the influence of the special factor E is not considered, X = AF can be obtained, but the matrix A is pX m order. In the model, we require m <= p. Usually, the number of factors is much smaller than the number of variables, that is, m < p. Therefore, the load matrix is irreversible and the estimate of F cannot be directly obtained. So the present invention selects the regression method to obtain the scores of the common factors.
[0154] The score function of a single factor obtained by using the regression method is:
[0155] F j = c j1 X1 + … + c jp X p , (j = 1, 2, …, m);
[0156] where c jp is the score of the factor F j on the scalar X p , Xp is the correlation size of the landscape index. For example, X1 is the correlation size of the obtained factor 1 and area (the contribution rate of area to factor 1), and X2 is the value of the contribution rate of ENN to factor 1.
[0157] The final formula of the score function is:
[0158] F = XR -1B′;
[0159] Where R is the correlation matrix, R = X′X.
[0160] The factor analysis results calculated by the above equation are shown in Table 2.
[0161] Table 2 Factor analysis results of the optimal landscape index subset
[0162]
[0163] Factor scores reflect the contribution of variables to the common factors. The relationship between landscape indices and principal component factors allows for interpretation of the ecological significance of each common factor from a landscape ecology perspective. The results showed that the two factors accounted for 85.37% of the variance (Table 2). Clearly, among the landscape indices, AREA_MN, ENN_MN, and SHD I scored higher on Factor 1, indicating that they represent Factor 1 well. SPLI T and FACE_AM had higher loadings on Factor 2 and can represent Factor 2. Factor 1 describes the average size of objects, the distribution of the same crop type within the cultivated land, and the number of species, while Factor 2 determines the spatial distribution and geometry of the cultivated land.
[0164] Furthermore, S105: based on the common factor scores, the sample data is divided into several representative categories using the K-means clustering method, and the crop distribution status and classification results are described through the ecological significance represented by the common factors. This includes:
[0165] S105-1: Classify the plots using the K-means clustering method, and then select the landscape area with the closest grid number through multiple experiments;
[0166] S105-2: Describe the distribution status of each type of crop based on the importance value and correlation of the common factors to each type.
[0167] Among them, the plots were classified by the K-means clustering classification method, and then the landscape area with the closest grid number was selected through multiple experiments, including:
[0168] From the perspective of landscape ecology, landscape pattern is generally defined as homogeneous patches with similar shapes, sizes, numbers, and spatial combinations. The study area was first divided into 8,362 plots for subsequent study zoning.
[0169] Randomly select k sample data as the initial cluster centers {μ1,μ2,…,μ k};
[0170] For each remaining data object to be clustered, assign the sample data to the cluster with the smallest distance according to the following formula; calculate the distance from the training data set to each cluster center, and assign it to the class with the nearest centroid by comparing the size, and iterate n times.
[0171]
[0172] Among them, c is the initial cluster center, sample data x is the comprehensive score of the factor analysis of the plot sample point, c is the set initial cluster center, and d is the distance between the data of the plot sample data and the initial cluster center.
[0173] After classification, the class mean method is used to recalculate the centroid of each class in each iteration according to the following formula, and update the values of the k cluster centers:
[0174]
[0175] For all k cluster centers, S104-2 and S104-3 are repeatedly executed. If the updated classification result remains unchanged and the position point changes very little, it is considered that a stable state has been reached, the iteration ends and the algorithm is stopped to obtain the clustering result. Otherwise, the iteration continues.
[0176] Furthermore, based on the importance and correlation of the common factors for each type, the distribution of each type of crop is described, including:
[0177] Table 3 Cluster centers of landscape heterogeneity partitions and average scores of landscape indicators
[0178]
[0179] Landscape heterogeneity includes compositional heterogeneity and configurational heterogeneity. Based on K-Means cluster analysis and the physical meaning of the five selected landscape indicators, regions with higher Factor-1 scores were designated A1, while regions with lower Factor-1 scores were designated A2. Regions with higher Factor-2 scores were designated B1, while B2 had lower Factor-2 scores. Region A1 exhibited larger average patch area, discrete distribution of homogeneous patches, and fewer crop varieties, while the opposite was true for Region A2. Region B1 was characterized by discrete distribution of landscape patches and simple geometry, while the opposite was true for Region B2. Ultimately, four landscape regions were identified (A1B1, A2B2, A1B2, and A2B1). Taking Region A1B1, which had higher Factor-1 and Factor-2 scores, as an example, its AREA_MN, ENN_MN, and SHDI scores were 73.43, 11.53, and 0.66, respectively. AREA_MN represents landscape fragmentation, ENN_MN measures the spatial pattern of the landscape, and SHDI indicates the number and proportion of landscape elements. Therefore, the average patch area in this region is large, the distribution of patches of the same type is discrete, the crop types are relatively small, the distribution of landscape patches is discrete, and the geometric shape is simple.
[0180] Figure 4 (a), (b), (c), and (d) represent representative landscape patches of the four types of regions derived using cluster analysis. Region a represents the first type of region, and the colors above it represent the distribution of different crops on that patch. This is analogous to the other three types of patches. The arrows from top to bottom indicate increasing configurational heterogeneity (increasing number of crop types) from region (a) to (c), and increasing component heterogeneity from region (c) to (d) and from (a) to (b).
[0181] In this study, we used the K-Means cluster analysis method to divide the 8,362 plots in the study area into landscape zones with similar objects to determine the best features for crop classification. For the number of landscape zones, we repeated the experiment five times and selected the one with the closest number of grids. The crop heterogeneity within each zone was described as compositional heterogeneity and configurational heterogeneity ( Figure 4 ). The former refers to the increase in the number of crop types, and the latter refers to the increase in the spatial dispersion of crops.
[0182] Example 2
[0183] This embodiment discloses a quantitative expression system for crop spatial heterogeneity.
[0184] like Figure 5 As shown, a crop spatial heterogeneity quantitative expression system includes:
[0185] The dataset acquisition module is configured to: acquire remote sensing images of the study area, pre-process the remote sensing images, and obtain a sample dataset;
[0186] The initial landscape index calculation module is configured to: select initial landscape indicators and calculate the initial landscape index of the study area based on the sample data set;
[0187] The optimal landscape index selection module is configured to: perform correlation analysis on the initial landscape indexes, eliminate the initial landscape indexes with strong correlation reflecting the same spatial characteristics, and obtain the optimal landscape index;
[0188] The crop spatial heterogeneity index acquisition module is configured to: establish a factor analysis model including a common factor vector, perform factor analysis on the optimal landscape index, estimate the common factor score of the common factor vector, determine the contribution rate of the optimal landscape index to each type of common factor based on the common factor score, screen the optimal landscape index, obtain the optimal landscape index associated with describing the importance of each type of common factor, and use the optimal landscape index associated with describing the importance of each type of common factor as the crop spatial heterogeneity index;
[0189] The description module is configured to: use the K-means clustering method to divide the sample data set into plot categories, calculate the common factor score of each plot category, and based on the common factor score of each plot category, describe the ecological significance represented by the common factor through the crop spatial heterogeneity index of each common factor, and then describe the crop distribution status and classification results of each plot category.
[0190] Example 3
[0191] The purpose of this embodiment is to provide a computer-readable storage medium.
[0192] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the method for quantitatively expressing crop spatial heterogeneity as described in Example 1 of the present disclosure.
[0193] Example 4
[0194] The purpose of this embodiment is to provide an electronic device.
[0195] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor. When the processor executes the program, the steps in the method for quantitatively expressing crop spatial heterogeneity as described in Example 1 of the present disclosure are implemented.
[0196] The steps involved in the apparatuses of Examples 2, 3, and 4 above correspond to those of Method Example 1. For detailed implementations, please refer to the relevant description of Example 1. The term "computer-readable storage medium" should be understood to mean a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and causing the processor to perform any method of the present invention.
[0197] Those skilled in the art will appreciate that the modules or steps of the present invention described above can be implemented using a general-purpose computer device. Alternatively, they can be implemented using program code executable by a computing device, which can then be stored in a storage device and executed by the computing device. Alternatively, they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.
[0198] Although the above describes the specific embodiments of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without any creative work are still within the scope of protection of the present invention.
Claims
1. A method for quantitatively expressing crop spatial heterogeneity, characterized in that: The following steps are involved: Obtain remote sensing images of the study area, preprocess the remote sensing images, and obtain a sample data set; Select the initial landscape indicators and calculate the initial landscape index of the study area based on the sample data set; The correlation analysis of the initial landscape indices was performed, and the initial landscape indices with strong correlation reflecting the same spatial characteristics were eliminated to obtain the optimal landscape indices; A factor analysis model including a common factor vector was established, and factor analysis was performed on the optimal landscape index to estimate the common factor score of the common factor vector. Based on the common factor score, the contribution rate of the optimal landscape index to each type of common factor was determined. The optimal landscape index was screened to obtain the optimal landscape index associated with describing the importance of each type of common factor. The optimal landscape index associated with describing the importance of each type of common factor was used as the crop spatial heterogeneity index. The K-means clustering method was used to divide the sample data set into plot categories, and the common factor score of each plot category was calculated. According to the common factor score of each plot category, the crop spatial heterogeneity index of each common factor was used to describe the ecological significance represented by the common factor, and then the crop distribution status and classification results of each plot category were described.
2. The method for quantitatively expressing crop spatial heterogeneity according to claim 1, wherein: Initial landscape indicators include: Area_perimeter_density index: percentage of landscape type area, average patch area, landscape element patch density, edge density; Shape indices: landscape shape index, area-weighted average shape index, area-weighted average patch fractal dimension; Spread index: average minimum distance, landscape separation, aggregation index, spread index; Diversity index: fragrant diversity index, uniformity index.
3. The method for quantitatively expressing crop spatial heterogeneity according to claim 1, wherein: Spearman correlation analysis was used to analyze the correlation between the initial landscape indices. The rank correlation coefficients were calculated for the initial landscape index data at the patch type and landscape levels. The closer the absolute value of the rank correlation coefficient was to 1, the stronger the correlation was.
4. The method for quantitatively expressing crop spatial heterogeneity according to claim 1, wherein: A factor analysis model containing a common factor vector is established, and factor analysis is performed on the optimal landscape index to estimate the common factor score of the common factor vector, including: S104-1: Conduct correlation analysis on the optimal landscape indicators to test whether the optimal landscape indicators are suitable for factor analysis; S104-2: Establish a factor analysis model, standardize the sample data set, and obtain the correlation matrix of the standardized sample data set; S104-3: Solve the factor loading matrix of the factor analysis model based on principal component analysis and determine the common factor variables; S104-4: Rotate the factor loading matrix using the varimax method; S104-5: Estimate the factor scores of the common factor vector based on the rotated factor loading matrix, judge the importance of the optimal landscape indicator based on the contribution rate of the optimal landscape indicator to each type of common factor, and screen the optimal landscape indicator to finally obtain the indicator that describes the spatial heterogeneity of crops.
5. The method for quantitatively expressing crop spatial heterogeneity according to claim 4, wherein: The factor analysis model is: X=A×F+E Among them, X is the original variable vector, A is the common factor loading matrix, F is the common factor vector, and E is the impact of the common factor on the variance of the data.
6. The method for quantitatively expressing crop spatial heterogeneity according to claim 4, wherein: The regression method is used to obtain the score of a single common factor. The score function is: F j =C j1 X1+…+C jp X p ; Among them, c jp Factor F j In the scalar X p The score on ; j = 1, 2,…m.
7. The method for quantitatively expressing crop spatial heterogeneity according to claim 4, wherein: The K-means clustering method is used to divide the sample data set into categories, including: Randomly select k sample data as the initial cluster center; Calculate the distance between each data object to be clustered and the initial cluster center, assign the data object to be clustered to the cluster with the smallest distance for classification; After classification, recalculate the centroid of each class and update the values of k cluster centers; For the updated k cluster centers, S104 - 2 and S104 - 3 are repeatedly executed. When the set conditions are met, the iteration ends and the classification is completed.
8. A quantitative expression system for crop spatial heterogeneity, characterized by: include: The dataset acquisition module is configured to: acquire remote sensing images of the study area, pre-process the remote sensing images, and obtain a sample dataset; The initial landscape index calculation module is configured to: select initial landscape indicators and calculate the initial landscape index of the study area based on the sample data set; The optimal landscape index selection module is configured to: perform correlation analysis on the initial landscape indexes, eliminate the initial landscape indexes with strong correlation reflecting the same spatial characteristics, and obtain the optimal landscape index; The crop spatial heterogeneity index acquisition module is configured to: establish a factor analysis model including a common factor vector, perform factor analysis on the optimal landscape index, estimate the common factor score of the common factor vector, determine the contribution rate of the optimal landscape index to each type of common factor based on the common factor score, screen the optimal landscape index, obtain the optimal landscape index associated with describing the importance of each type of common factor, and use the optimal landscape index associated with describing the importance of each type of common factor as the crop spatial heterogeneity index; The description module is configured to: use the K-means clustering method to divide the sample data set into plot categories, calculate the common factor score of each plot category, and based on the common factor score of each plot category, describe the ecological significance represented by the common factor through the crop spatial heterogeneity index of each common factor, and then describe the crop distribution status and classification results of each plot category.
9. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method for quantitatively expressing crop spatial heterogeneity as described in any one of claims 1 to 7 are implemented.
10. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method for quantitatively expressing crop spatial heterogeneity according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Crop classification feature optimization method and system for heterogeneous farmland landscape area
CN114550008A
Use of microbiome and metobolome clusters to evaluate skin health
WO2022097069A1