Method for rapidly evaluating growth suitability of large-area characteristic crops

By combining traditional multi-layer weighted overlay method and species distribution model, we can identify the natural environment gradient and explore the response intervals of key environmental factors, and solve the problem that traditional methods can quickly and effectively evaluate the growth suitability of characteristic crops on a large scale, achieving efficient and objective growth suitability evaluation.

CN119940692APending Publication Date: 2025-05-06INST OF SOIL SCI CHINESE ACAD OF SCI
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411787153.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Traditional plant growth suitability evaluation methods are difficult to quickly and effectively evaluate the growth suitability of characteristic crops on a large scale, especially under the influence of factors such as spatial scale effect, regional protection and cost.

Method used

Combining the traditional "bottom-up" multi-layer weighted superposition plant growth suitability evaluation method and "top-down" species distribution model, through systematic point survey, sampling testing, species distribution model training, cluster analysis and multi-layer weighted superposition, we identify the natural environment gradient and explore the response interval of key environmental factors to achieve growth suitability evaluation of characteristic crops.

Benefits of technology

This method improves the objectivity and operability of the growth suitability evaluation of characteristic crops, can quickly and at low cost, and overcomes the problem of difficulty in applying traditional methods on a large scale.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940692A_ABST
    Figure CN119940692A_ABST
Patent Text Reader

Abstract

The invention discloses a method for rapidly evaluating the growth suitability of large-area characteristic crops, and the method comprises the following steps: S1, obtaining the growth distribution point information and key quality data of a multi-production area of the characteristic crops, collecting the environment variables of a research area, and constructing a database; s2, constructing a species distribution model based on the constructed database and a maximum entropy principle; s3, training a species distribution model, and calculating the distribution probability of characteristic crops; collecting a plurality of classification thresholds, and selecting an optimal threshold to obtain a potential distribution area of characteristic crops; s4, obtaining quality classification and origin environment classification results; s5, taking the origin environment classification result as a natural environment gradient, analyzing a response mode of the key quality of the characteristic crops to each environment factor, and establishing suitability classification of the key environment factors on a large scale; and S6, constructing a characteristic crop growth suitability evaluation model. According to the method, the growth suitability evaluation of large-area characteristic agricultural products can be quickly realized at low cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of ecological environment protection, and in particular to a method for rapidly evaluating the growth suitability of characteristic crops in a large region. Background Art

[0002] Plant growth suitability evaluation is a method to assess whether a specific area is suitable for the growth of a certain plant. It comprehensively considers multiple environmental factors, such as climate, soil, water source and topography, to determine the growth potential of plants in the area. Traditionally, suitability evaluation is achieved through the superposition of multiple layers. It is a "bottom-up" process. Its basic process includes: clarifying the scope of the study area and determining the evaluation unit, selecting evaluation indicators, determining the indicator classification and indicator weight, calculating the comprehensive score, and suitability classification and zoning. Carrying out traditional plant growth suitability evaluation is a time-consuming and costly system engineering. It is necessary to set up control variable experiments and conduct long-term field experiments to determine the evaluation indicators, divide the suitability levels of indicators, etc., in order to establish a scientific and reliable suitability evaluation system. Therefore, traditional growth suitability evaluation is widely used in major food crops, such as rice, corn, wheat, etc., and important economic crops such as cotton, tobacco, tea, etc., but is less used in the growth suitability evaluation of relatively niche specialty crops, especially lacking large-scale research work.

[0003] Specialty crops refer to crops that are grown in a specific area and whose products have unique qualities. Specialty crops have the characteristics of "authenticity", such as "Wuchang rice", "Zhongning wolfberry", and "Chifeng astragalus". They are formed under the comprehensive factors of specific climate, regional soil environment, water quality, etc., and their ecological niche is relatively narrow, which limits their widespread planting and keeps their high economic value. However, in order to achieve efficient introduction and expansion, large-scale growth suitability zoning is needed as a guide. However, due to the comprehensive factors such as spatial scale effect, regional protection, and cost, it is difficult to achieve large-scale suitability zoning of thousands of specialty crops by relying solely on the traditional "bottom-up" suitability evaluation method. In addition, the suitability evaluation of traditional methods lacks an overall vision. From the existing research, the current only studies on the suitability of specialty crops based on local areas mostly lack corresponding reference materials in the evaluation process such as determination of the study area, selection of evaluation indicators, and division of suitable intervals. There are also problems of strong subjectivity and incomplete evaluation results.

[0004] The species distribution model provides a "top-down" overall perspective. It is based on the ecological niche theory and establishes a model to predict the potential distribution of species in uninvestigated areas by analyzing the relationship between known species distribution data and environmental variables (such as climate, soil type, vegetation cover, etc.). Species distribution models are widely used in the fields of wildlife protection, economic forest production area development, and traditional Chinese medicine.

[0005] Species distribution model is a kind of method. In the past 30 years, a large number of species distribution models have emerged based on different theories and calculation methods, such as the main domain analysis model (DOMAIN), the niche factor analysis model (ENFA), the Mahalanobis distance model (MD), the maximum entropy model (Maxent), the generalized linear model (GLM), the generalized additive model (GAM), the classification and regression tree model (CART), the boosted regression tree model (BRT), the genetic algorithm model (GARP), and the artificial neural network (ANN). Among them, the Maxent model is the most widely used model in recent years. It is based on the maximum entropy principle and estimates the potential distribution probability of species under given environmental conditions by maximizing entropy. A large number of studies have shown that the Maxent model has strong applicability and can obtain good prediction results when the sample data is insufficient.

[0006] Compared with the traditional plant growth suitability evaluation method, the advantages of species distribution model are: first, it provides a "top-down" large-scale perspective, which can quickly identify the potential growth area of ​​plants; second, it is data-driven, can integrate multi-source data, and the number of indicators is not limited. At the same time, it is highly objective and requires less professional level; third, in addition to traditional field verification, species distribution model can be verified and uncertainty evaluated through independent verification, cross-validation and other methods, and the credibility of prediction results is high. Of course, species distribution model also has disadvantages: first, there are many types of models, the applicability of different models is different, and there is uncertainty in model selection; second, it is highly data-dependent, and the research results are greatly affected by data accuracy; third, the suitability differences within the species distribution points are blurred. In order to deal with this shortcoming, some studies use hierarchical modeling to reduce the model bias caused by the suitability differences of known species distribution points, but on the whole, species distribution model is only suitable for predicting the potential distribution area of ​​species, but it is difficult to evaluate growth suitability.

[0007] In traditional methods, in order to determine the suitability interval of plants to a certain environmental factor, it is often achieved by controlling variables, that is, controlling other environmental factors to be in suitable conditions, setting up gradient tests for a single environmental factor, and studying the response of plants to this environmental factor. Although this method has high accuracy, it is extremely costly and difficult to apply to large-scale suitability studies of specialty crops. Under natural conditions, almost all environmental factors are not independent. Therefore, through clustering and other methods, the growth environment of a certain plant is classified and zoned, and a natural environmental gradient is constructed. Further research on the response of plant growth to environmental factors under different environmental gradients is an important means to solve the problem of the suitable interval of evaluation factors in the large-scale suitability evaluation of specialty crops.

[0008] Different from traditional bulk crops, specialty crops are small-scale and fine production methods that pursue economic benefits, requiring a rapid suitability evaluation process on a large scale. However, from the current research, there are technical difficulties, which can be summarized as follows:

[0009] (1) Although the traditional plant growth suitability evaluation method based on "bottom-up" multi-layer weighted superposition is mature and accurate, it has obvious shortcomings in application: first, it requires long-term observation and monitoring, and the research cycle is long; second, the workload increases exponentially with the increase in the number of evaluation indicators, which limits the number of indicators that can be considered; in addition, the scale effect is obvious, and research in one region is difficult to extend to other regions. In addition, the reference data is limited, making it difficult to independently carry out large-scale evaluation of the growth suitability of specialty crops.

[0010] (2) Although species distribution models provide a large-scale perspective, can quickly identify the potential distribution of species on a large scale, and increase the objectivity and verifiability of research, their disadvantages are: uncertainty in model selection; high data dependence and high data accuracy requirements; most importantly, they ignore the suitability differences within the known species distribution points used for model construction. Therefore, species distribution models are more suitable for predicting the potential distribution area of ​​plants rather than evaluating the growth suitability of plants.

[0011] (3) Combining the traditional "bottom-up" multi-layer weighted superposition suitability evaluation method and the "top-down" species distribution model, it is possible to quickly carry out large-scale regional characteristic crop growth suitability evaluation. However, due to the scale effect, the suitability interval of environmental factors at a large scale is difficult to be directly derived from the existing small-scale research data. There is still a lack of reference materials for the division of suitability intervals of key environmental factors that affect the growth of characteristic crops at a large scale. It is necessary to establish the response interval of key environmental factors based on environmental gradients from scratch. Summary of the invention

[0012] Purpose of the invention: The purpose of the present invention is to provide a method for quickly evaluating the growth suitability of specialty crops in a large area, combining the traditional "bottom-up" multi-layer weighted superposition plant growth suitability evaluation method with the "top-down" species distribution model, and based on cluster analysis, correlation analysis, regression analysis and other methods, identifying the spatial natural environment gradient, and exploring the response range of key environmental factors affecting the growth of specialty crops, so as to achieve objectivity in the evaluation of the growth suitability of specialty crops and strong operability.

[0013] Technical solution: A method for rapidly evaluating the growth suitability of specialty crops in a large area, comprising the following steps:

[0014] S1, through systematic survey and sampling test, obtain the growth distribution point information and key quality data of characteristic crops in high-yield areas, collect environmental variables covering the study area, and build a database;

[0015] S2, construct species distribution model based on the constructed database and maximum entropy principle;

[0016] S3, training species distribution model, calculating the distribution probability of characteristic crops; collecting multiple classification thresholds, classifying the distribution probability into two categories, and selecting the optimal threshold to obtain the potential distribution area of ​​characteristic crops;

[0017] S4, cluster analysis is used to cluster the growth environments of characteristic crops at multiple recording points to obtain quality classification and origin environment classification results;

[0018] S5, using the results of the origin environment classification as the natural environment gradient, analyzing the response pattern of the key qualities of characteristic crops to each environmental factor based on the natural environment gradient, and establishing the suitability classification of key environmental factors on a large scale;

[0019] S6, using the multi-layer weighted overlay method, constructs a growth suitability evaluation model for characteristic crops and obtains large-scale suitability evaluation results.

[0020] Further, the implementation steps for building a database are as follows:

[0021] S11, complete the systematic sampling by collecting the growth and distribution information of characteristic crops;

[0022] S12, based on the distribution map, carry out actual sampling work and record the actual distribution points of characteristic crops {P1, P2, P3, ..., P n}, where P n Indicate the distribution point of the nth species, and conduct surveys and sampling tests to obtain product quality and soil properties at the planting site;

[0023] S13, collects multiple environmental variables covering the study area, including climate, topography, soil, and organisms;

[0024] S14, preprocessing the collected species distribution point data:

[0025] The species distribution points were exported to the same coordinate system, the geographic coordinates were uniformly converted to GCS_WGS_1984, and the projection coordinates were uniformly converted to WGS_1984_Albers;

[0026] Set a minimum geographic distance threshold D to filter out other points whose geographic location is less than the threshold D; set the spatial resolution of the environmental variable grid to r, and the following conditions must be met:

[0027] Then the minimum geographic distance threshold D is expressed as: Where a is an arbitrary constant greater than 1;

[0028] By screening the points, we ensure that the Euclidean distance d ≥ D between all the recorded points, and finally obtain n species distribution points site: site = (long, lati), which is an n × 2 vector, where long represents longitude and lati represents latitude;

[0029] S15, preprocessing and standardizing the environmental variables;

[0030] S16, using GIS software to extract the raster values ​​of the environmental variable layer corresponding to the species distribution points. The data set is described as data = (site, z1, z2, ..., zm), which is an n × (m + 2) vector, where m represents the number of environmental layers.

[0031] Further, in step S11, the growth distribution information of characteristic crops is collected through multiple channels to complete the system point sampling, and the implementation steps are as follows:

[0032] S111, obtaining growth and distribution information of characteristic crops as reference information for site placement through literature search and database access;

[0033] S112, using satellite image maps to observe the image features of the characteristic crops at the reference point, and based on this, continue to search for the planting distribution of the characteristic crops around the reference point, and preliminarily obtain the precise distribution map set of characteristic crops {S1, S2, S3, ..., S n}, where S n represents the nth image spot;

[0034] S113, collect high-resolution remote sensing images of the study area, use distribution patch sets, combine remote sensing images, and identify similar image patches through supervised learning methods as characteristic crop planting points as actual point distribution references;

[0035] S114, confirm the total number of points according to the distribution range of characteristic crops and the complexity of the terrain in the distribution area, and complete the sampling of points based on the planned point density of the terrain and soil types in different areas.

[0036] Furthermore, the implementation steps for constructing a species distribution model are as follows:

[0037] S21, based on the collected environmental variables, define the characteristic function f i (x), represents the value of any grid x on the i-th environmental variable; then for any feature, there is an expected value constraint E p [f i (x)]:

[0038] E p [f i (x)]=Σ x∈X f i (x)P(x)=c i

[0039] Among them, c i is the expected value of all grids of the i-th environmental variable; P(x) represents the probability distribution of any grid value in the study area;

[0040] For any grid point, X represents the set of all grid points in the study area;

[0041] S22, construct the Lagrangian function and solve the probability distribution with maximum entropy under given constraints:

[0042]

[0043] Where L represents the Lagrangian function; -∑ x∈X P(x)logP(x) is the expression of entropy; λ0(∑ x∈X P(x)-1) is to ensure that the sum of the probabilities P(x) is equal to 1, and λ0 represents the Lagrange multiplier, which is used to adjust the weight of this item;

[0044] Contains m environmental variable constraints, λ i Indicates that each constraint corresponds to a Lagrange multiplier;

[0045] Take the partial derivative of each P(x) and make it equal to 0:

[0046]

[0047] Solving for this yields:

[0048]

[0049] Since the sum of all probabilities is equal to 1, for x∈X, the formula for calculating the sum of probabilities is as follows:

[0050]

[0051] definition but

[0052] After finishing, we get:

[0053]

[0054] As the species distribution probability, then λ iis regarded as the weight parameter of the i-th environmental variable, and Z(λ) is the normalization constant;

[0055] S23, perform parameter estimation and find the model parameter λ i , by maximizing the log-likelihood function, the predicted probability distribution l(λ) of the crop distribution model is made closest to the distribution of the actual observed data:

[0056]

[0057] In the formula, n represents the number of species distribution points, x (j) represents the set of environmental conditions at the jth species distribution point; substituting the probability expression P(x) into the above formula, we get:

[0058]

[0059] S24, establish the gradient function, for each λ i Find the partial derivative:

[0060]

[0061] In the formula, E p [f i ] indicates f i The expected value predicted by the model under the probability distribution P(x|λ) with given parameter λ; For each observed x (j) f i The sum of the values ​​of nE p [f i ] represents the comprehensive expected value under the model prediction;

[0062] S25, using the gradient ascent method to iterate and update the parameter λ i ,make Close to 0 to maximize l(λ) and find an optimal set of λ i , so that the predicted probability distribution of the model fits the observed data best, the expression is as follows:

[0063]

[0064] In the formula, α is the learning rate, which determines the magnitude of each update step.

[0065] Furthermore, in step S3, the steps for obtaining the potential distribution area of ​​characteristic crops are as follows:

[0066] S31, dividing the database into a training set and a validation set, the training set is used to train the constructed crop distribution model, and the validation set is used to test the prediction performance of the crop distribution model;

[0067] S32, running the crop distribution model and outputting species distribution probability; selecting the optimal threshold t, converting the species distribution probability into the potential distribution area of ​​characteristic crops, and converting the distribution probability into the presence or absence of the species to achieve binary classification;

[0068] S33, select threshold t based on sensitivity and specificity;

[0069] S34, using a threshold calculation method to obtain a set of classification thresholds {T1, T2, T3, ..., T n}, each classification threshold T i , calculate and obtain a potential distribution area R i , then we get the potential distribution area set {R1, R2, R3, ..., R n};

[0070] S35, combining literature query, network query and field investigation, verifies the actual boundaries of the potential distribution areas of characteristic crops, determines the optimal classification threshold, and obtains the most accurate potential distribution areas of characteristic crops.

[0071] Further, in step S4, cluster analysis is used to cluster the key qualities and growth environments of the characteristic crops at multiple recording points, and the implementation steps are as follows:

[0072] S41, conduct Pearson correlation analysis on multiple quality indicators of specialty crops to understand the relationship between quality indicators:

[0073]

[0074] In the formula, qa i ,qb i are the values ​​corresponding to the two quality indicators qa and qb at any recording point; f represents the number of recording points where the quality of the characteristic crop is obtained, qa mean ,qb mean represents the mean of the two quality indicators;

[0075] S42, principal component analysis is performed on the growth environment factors of characteristic crops. The specific steps are as follows:

[0076] S421, for an f×m matrix consisting of f points and m environmental variables, perform data standardization to convert each column of quality indicator data into a mean equal to 0 and a standard deviation of 1;

[0077] S422, calculate the covariance matrix, perform eigenvalue decomposition on the covariance matrix, obtain m principal components, select the first e principal components whose cumulative variance contribution rate reaches the set threshold, and record the set as {PC1, PC2, PC3, ..., PC e};

[0078] S423, using the eigenvector corresponding to the selected principal component, converting the original data into a principal component space;

[0079] S43, clustering the e principal components using K-means clustering to obtain clustering results of the growth environment of characteristic crops, and counting the quality characteristics of characteristic crops in different clusters;

[0080] S44, based on the clustering results of the production environment, construct the natural environment gradient of the growth environment of the characteristic crops.

[0081] Further, in step S5, the steps for establishing the suitability classification of key environmental factors on a large scale are as follows:

[0082] S51, based on the data of the recording points in the entire study area, the Pearson correlation analysis method was used to calculate the correlation coefficient between each growth environment factor of the characteristic crops and their quality index, and the correlation coefficient between the growth environment factors of the characteristic crops; find the environmental factor pairs with a correlation greater than 0.8, retain the factor with a higher correlation with the quality index, and delete the other one;

[0083] S52, for each production environment classification, use Pearson correlation analysis to establish the correlation between the key quality indicators of specialty crops and multiple environmental factors; combine the correlation coefficient and scatter plot to compare the response of the key quality indicators of specialty crops to various environmental factors under different environmental gradients, and explore the response pattern of specialty crops to each environmental factor in a large area; based on the response pattern, find the optimal threshold and divide the suitability of specialty crops to each environmental factor into grades.

[0084] Further, in step S6, the implementation steps of constructing a characteristic crop growth suitability evaluation model are as follows:

[0085] S61, the potential distribution area of ​​characteristic crops was taken as the study area;

[0086] S62, using stepwise linear regression or random forest, identify environmental factors that have a significant impact on the growth of specialty crops as evaluation indicators, and establish the weights of the evaluation indicators based on standardized regression coefficients or variable importance;

[0087] S63, unify the dimensions of the suitability levels of each evaluation index, and perform spatial overlay operations on the evaluation indexes of multiple layers in ArcGIS software to obtain a comprehensive evaluation score;

[0088] S64, using the natural breakpoint method, divides the comprehensive evaluation score into b levels as the growth suitability level of the characteristic agricultural crop, and finally completes the growth suitability evaluation of the characteristic agricultural products in the study area.

[0089] Compared with the prior art, the present invention has the following significant effects:

[0090] 1. The present invention systematically arranges the growth and distribution information of characteristic crops, conducts field investigation and sampling tests, and has stronger spatial representativeness than the historical data obtained by predecessors based on specimen libraries. It is not easy to cause model deviation when training species distribution models. At the same time, the key quality indicators of corresponding characteristic crops are obtained based on the recording points, which ensures the point correspondence between quality and environment, and the data accuracy is high;

[0091] 2. The present invention uses the species distribution model to divide the potential distribution area of ​​characteristic crops, and uses this as the research area to carry out further growth suitability evaluation, which is more objective than the traditional use of administrative regions as research areas;

[0092] 3. The present invention constructs a natural environment gradient and directly analyzes the response relationship between the key qualities of characteristic crops and environmental factors. It is more objective and more operable than the traditional method that highly relies on expert experience in terms of suitability classification and weight determination of evaluation factors;

[0093] 4. Combining the "top-down" species distribution model with the "bottom-up" traditional suitability evaluation method provides an effective way to quickly and low-costly evaluate the growth suitability of specialty agricultural products in large regions. BRIEF DESCRIPTION OF THE DRAWINGS

[0094] Figure 1 It is the overall flow chart of the present invention;

[0095] Figure 2 It is a flow chart for sampling points of characteristic crop systems;

[0096] Figure 3 It is a schematic diagram of the minimum distance threshold between adjacent sampling points. If point a is to be retained, points b and d need to be deleted;

[0097] Figure 4 It is a flow chart of the species distribution model (maximum entropy model) to identify the potential distribution areas of plants;

[0098] Figure 5 It is a basic flow chart for principal component analysis of the quality of specialty agricultural products and the environment of production areas;

[0099] Figure 6 It is a schematic diagram of the suitability classification process of environmental factors that affect the quality of specialty crops;

[0100] Figure 7a It is a schematic diagram of the receiver characteristic curve (ROC curve), which reflects the accuracy of the maximum entropy model. The selection area of ​​the classification threshold is in the upper left corner;

[0101] Figure 7b It is a schematic diagram of the omission rate curve, which reflects the proportion of the actual record points that the model fails to predict. When the threshold is low, the omission rate is low, and the corresponding potential distribution area of ​​wolfberry is large; the opposite is true when the threshold is high;

[0102] Figure 8 This is a partial schematic diagram of the potential distribution area of ​​wolfberry based on two binary classification thresholds. After query and verification, it was found that the low threshold division result is more consistent with the potential distribution of wolfberry;

[0103] Fig. 9 This is a partial schematic diagram of the wolfberry growth suitability zoning based on the national scale. DETAILED DESCRIPTION

[0104] The present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0105] like Figure 1 The figure shows a flow chart of a method for rapidly evaluating the growth adaptability of characteristic crops in a large region, which specifically includes the following steps:

[0106] Step 1: Through systematic site surveys and sampling tests, obtain the growth distribution point information and key quality data of the specialty crop production areas, and collect various environmental variables such as climate, topography, soil, etc. covering the whole country to build a database. The specific implementation steps are as follows:

[0107] (1.a) Collect the growth and distribution information of characteristic crops through multiple channels, and complete the systematic sampling work based on this information. The process is as follows Figure 2 .

[0108] (1.a.1) Obtain the growth and distribution information of specialty crops as reference for site placement through literature search, database access (such as China Digital Herbarium, Global Biodiversity Information Database), data published on the websites of government departments at all levels, and news reports.

[0109] (1.a.2) Using satellite image maps to observe the image features of the characteristic crop at the reference point, and based on this, continue to search for the planting distribution of the characteristic crop around the reference point, and preliminarily obtain the precise distribution map set of the characteristic crop {S1, S2, S3, ..., S n}, where S n Represents the nth image spot.

[0110] (1.a.3) Collect high-resolution remote sensing images of the study area, use distribution patch sets, combine remote sensing images, and use supervised learning to identify similar image patches, i.e., characteristic crop planting sites, which can be used as a reference for actual site layout.

[0111] (1.a.4) Based on the distribution range of specialty crops, the complexity of the terrain in the distribution area and other factors, combined with objective conditions, confirm the total number of distribution points, and plan the distribution point density based on the terrain, soil type, etc. of different areas to complete the distribution point work.

[0112] (1.b) Based on the distribution map, carry out actual sampling work and record the actual distribution points of characteristic crops {P1, P2, P3, ..., P n}, where P n It indicates the distribution point of the nth species, and conducts surveys and sampling tests to obtain product quality and soil properties at the planting site.

[0113] (1.c) Collect a variety of environmental variables covering the study area, including raster and vector formats. Common environmental variables include climate, topography, soil, and biology. Climate variables include multi-year average rainfall, accumulated temperature, average temperature, sunshine hours, diurnal temperature difference, drought index, etc. Large-scale climate variables are easy to obtain, and many data platforms can apply for climate variables with a resolution of 1km; terrain variables are mostly calculated and obtained by digital elevation model (DEM), such as slope, aspect, plane curvature, profile curvature, terrain moisture index, etc. SRTM provides 30m and 90m resolution DEM data covering the world; soil variables include soil type, total nitrogen, total phosphorus, total potassium, cation exchange capacity (CEC), particle composition, bulk density, etc. Soil data is difficult to obtain. In recent years, based on the national soil system survey, researchers have constructed a high-precision national soil information grid with a resolution of 90m in China, which can be obtained by applying for relevant platforms; biological data mainly include various vegetation indices extracted from multi-period remote sensing images, such as normalized vegetation index and leaf area index.

[0114] (1.d) Preprocess the collected species distribution point data to exclude species distribution points with close geographical locations, ensure that there is sufficient spatial isolation between distribution points, and avoid overfitting of the model.

[0115] (1.d.1) Export species distribution points to the same coordinate system, convert geographic coordinates to GCS_WGS_1984, and convert projection coordinates to WGS_1984_Albers.

[0116] (1.d.2) Set a minimum geographical distance threshold D, filter out other points whose geographical location is less than the threshold D, and keep only one ( Figure 3 ). The specific method is to calculate the Euclidean distance d between two adjacent sampling points. Assume that in the two-dimensional space, there are sample points p1(x1, y1) and p2(x2, y2).

[0117]

[0118] Assuming that the spatial resolution of the environmental variable grid is r, to avoid duplicate points in the same grid, the following conditions need to be met:

[0119]

[0120] Then the minimum geographic distance threshold D can be expressed as:

[0121]

[0122] Among them, a is an arbitrary constant greater than 1, which can be adjusted according to factors such as species type, terrain complexity, etc. Through point screening, the Euclidean distance d ≥ D between all record points is ensured, and n species distribution points are finally obtained; its longitude and latitude are described as site = (long, lati), which is an n × 2 vector, where long represents longitude and lati represents latitude.

[0123] (1.e) Preprocess and standardize environmental variables from multiple sources and at different resolutions.

[0124] (1.e.1) Unify the coordinate system of environmental variables from multiple sources, convert geographic coordinates to GCS_WGS_1984, and convert projection coordinates to WGS_1984_Albers.

[0125] (1.e.2) All environmental variable data were clipped to a consistent range using the vector boundary of the study area.

[0126] (1.e.3) Resample the raster data using the bilinear interpolation algorithm and set the resolution to 100 meters, 250 meters, 500 meters, 1000 meters, etc. as required.

[0127] (1.e.4) Standardize the units of environmental variables. For example, variables related to precipitation use the standard unit mm, and variables related to temperature use the standard unit ℃. The environmental layer is described as layer = (X1, X2, ..., X m ),X i Represents the environment layer, and m represents the number of environment layers.

[0128] (1.f) Use GIS software to extract the raster values ​​of the environmental variable layer corresponding to the species distribution points. The data set is described as data = (site, z1, z2, …, zm), which is an n × (m + 2) vector.

[0129] Step 2: Collect multiple available species distribution models, compare model performance, and select the optimal species distribution model.

[0130] (2.a) Collect a set of species distribution models available for the target plants {A1, A2, A3, ..., A n}, and obtain the average simulation accuracy {C1, C2, C3, ..., C n}, and select the model with the highest accuracy as the optimal species distribution model.

[0131] (2.b) The maximum entropy model (Maxent) is the most widely used model in species distribution research. It performs well under small sample and multivariate conditions and has a high accuracy in prediction results. The implementation process of the species distribution model is introduced using the maximum entropy model as an example. The basic process of conducting species distribution research is as follows: Figure 4 The detailed process is as follows

[0132] (2.b.1) The Maxent model is based on the maximum entropy principle and maximizes entropy under the constraints. Mathematically, entropy H(P) is defined by multiplying the probability distribution by its logarithm and then integrating (or summing) it. The specific formula is as follows:

[0133] H(P)=-∑ x∈X P(x)logP(x) (4)

[0134] P(x) represents the probability distribution of any grid value in the study area, and X represents the set of all grid points in the study area. In the absence of constraints, that is, when there are no known species distribution points, the most reasonable assumption to satisfy the maximum entropy is that the probability of the existence of each grid point is equal, that is, P(x) is a uniform distribution. For any grid point, When there are constraints in the model, the constraints must be met first and then the entropy is maximized.

[0135] (2.b.2) Establish constraints. Based on the collected environmental variables, define the characteristic function f i (x), which means the value of any grid x on the i-th environmental variable. Then, for any feature, there is an expected value constraint:

[0136] E p [f i (x)]=Σ x∈X f i (x)P(x)=c i (5)

[0137] Among them, c i is the expected value of all rasters for the ith environmental variable.

[0138] (2.b.3) Construct the Lagrangian function and solve the probability distribution with maximum entropy under given constraints.

[0139]

[0140] In the formula, L represents the Lagrangian function; -∑ x∈XP(x)logP(x) is the expression of entropy, and the goal is to maximize this term; construct λ0(∑ x∈X P(x)-1) is to ensure that the sum of the probabilities P(x) is equal to 1, where λ0 represents the Lagrange multiplier, which is used to adjust the weight of this term; Contains m environmental variable constraints, λ i It means that each constraint corresponds to a Lagrange multiplier.

[0141] In order to maximize the Lagrangian function L, the partial derivative of each P(x) is calculated and made equal to 0.

[0142]

[0143] Solving for this yields:

[0144]

[0145] Since the sum of all probabilities is equal to 1, for x∈X, the formula for calculating the sum of probabilities is as follows:

[0146]

[0147] definition but The above formula is sorted out to get:

[0148]

[0149] The above formula (10) is the basic formula for species distribution probability, where λ i It can be regarded as the weight parameter of the i-th environmental variable, and Z(λ) is the normalization constant.

[0150] (2.b.4) Perform parameter estimation and find the model parameter λ i , so that the predicted probability distribution l(λ) of the model is closest to the distribution of the actual observed data, which can be achieved by maximizing the log-likelihood function:

[0151]

[0152] Substituting the probability expression P(x) into the above formula yields:

[0153]

[0154] In the formula, n represents the number of species distribution points, x (j) represents the set of environmental conditions at the jth species distribution point.

[0155] (2.b.5) Establish the gradient function, that is, for each λ iFind the partial derivative and find the direction in which to adjust λ to make l(λ) change the fastest.

[0156]

[0157] In the formula, E p [f i ] indicates f i The expected value predicted by the model under the probability distribution P(x|λ) with given parameter λ. The first half of this formula indicates that for each observed x (j) f i The second half represents the sum of the values ​​of , and the second half represents the comprehensive expected value under the model prediction. The difference between the two reflects the deviation between the actual data and the model prediction.

[0158] (2.b.6) Use the gradient ascent method to iterate and update the parameter λ i ,make Close to 0 to maximize l(λ) and find an optimal set of λ i , so that the predicted probability distribution of the model best fits the observed data.

[0159]

[0160] In the formula, α is the learning rate, which determines the magnitude of each update step.

[0161] Step 3: Train the species distribution model and calculate the distribution probability of characteristic crops. Collect a variety of available classification thresholds, classify the distribution probability into two categories, and select the optimal threshold to obtain the potential distribution area of ​​characteristic crops.

[0162] (3.a) Divide the n rows of data in data (i.e., data obtained from n distribution points) into a training set and a validation set. Depending on the amount of data, you can choose simple random partitioning, K-fold cross validation, etc. The training set is used to build the model, and the validation set is used to test the prediction performance of the model.

[0163] (3.b) Run the species distribution model and output the species distribution probability. If you want to convert the species distribution probability into the potential distribution area of ​​characteristic crops, you need to select the optimal threshold (t) to convert the distribution probability into the presence or absence of the species to achieve binary classification.

[0164] (3.c) The threshold selection should ensure that the model has high sensitivity and specificity. Sensitivity is the ability of the model to correctly identify the points where species exist; specificity is the ability of the model to correctly identify the points where species do not exist.

[0165]

[0166] Among them, TP means that the model correctly predicts the existence of a species at a certain location; FN means that the model incorrectly predicts the absence of a species at a certain location, but in fact the species exists; TN means that the model correctly predicts the absence of a species at a certain location; FP means that the model incorrectly predicts the existence of a species at a certain location, but in fact the species does not exist.

[0167] (3.d) When performing data calculations, maxent software outputs multiple thresholds, such as the minimum training presence threshold (MTP), which represents the predicted probability value of all records in the training data; the maximum training sensitivity plus specificity (MSS), which represents the threshold that maximizes the sum of sensitivity and specificity; the balanced training omission, prediction area and threshold (BTOAT), which represents the threshold that optimizes the balance between the training omission rate, prediction area and threshold; and the fixed threshold of 0.05, which represents a low probability event. Using multiple threshold calculation methods, a set of classification thresholds {T1, T2, T3, ..., T n}, each classification threshold T i , you can calculate a potential distribution area R i , we get the potential distribution area set {R1, R2, R3, ..., R n}.

[0168] (3.e) Verify the actual boundaries of the potential distribution areas of specialty crops by combining literature search, Internet search, field survey and other methods, determine the optimal classification threshold, and obtain the most accurate potential distribution areas of specialty crops.

[0169] Step 4: cluster the growth environment of the specialty crop at multiple recording points using cluster analysis, and summarize the production environment characteristics and key quality characteristics of the specialty crop under each cluster environment.

[0170] (4.a) Assume that the quality indicators of specialty agricultural products of f record points (f≤n, is all or part of n) are obtained, and Pearson correlation analysis is performed on multiple quality indicators of specialty agricultural products obtained in the test to understand the relationship between the quality indicators.

[0171]

[0172] In the formula, qa i ,qb i are the values ​​corresponding to the two quality indicators qa and qb at any recording point; f represents the number of recording points where the quality of the characteristic crop is obtained, qa mean ,qb mean Represents the mean of the two quality indicators. If the correlation between the two indicators is large, such as r>0.8, one of the qualities can be used to replace the other to reduce the amount of subsequent calculations.

[0173] (4.b) Cluster the m environmental variables (growth environment factors) of f wolfberry recording points with quality indicators. Since the data dimension is relatively high, it is necessary to perform principal component analysis before clustering to reduce the collinearity of the data, increase the interpretability of the clustering results, improve the operation efficiency and clustering quality. The process is as follows Figure 5 .

[0174] (4.b.1) For the f×m matrix composed of f points and m environmental variables, first standardize the data, and convert the quality indicator data of each column to a mean of 0 and a standard deviation of 1.

[0175] (4.b.2) Calculate the covariance matrix, perform eigenvalue decomposition on the covariance matrix to obtain m principal components. In order to effectively reduce the dimension, select the first e principal components whose cumulative variance contribution rate reaches a certain threshold (such as 85%), and the set is denoted as {PC1, PC2, PC3, ……, PC e}}.

[0176] (4.b.3) Use the eigenvectors corresponding to the selected principal components to transform the original data into a low-dimensional space, that is, the principal component space, and thus the meaning represented by each principal component can be interpreted.

[0177] (4.c) Use K-means clustering to cluster the above e principal components to obtain the clustering results of the origin environment of this special crop, and then the quality characteristics of the special crops in different clusters can be statistically analyzed.

[0178] (4.c.1) Use the "elbow method" to determine the optimal number of clusters Ka. Its basic idea is that as the number of clustering clusters k (k < f) increases, the sum of the squares of the distances (WSS) from all data points to the center of their respective clusters will decrease. But at a certain number of clusters k (k = Ka), the decrease of WSS slows down, and this k is the optimal number of clusters. Ka can be selected by plotting and observing the WSS curve.

[0179]

[0180] In the formula, C i is all the data points in the i-th cluster; G i is the cluster center in the i-th cluster.

[0181] (4.c.2) For a specific number of clusters Ka, randomly select Ka points from all f points as the cluster centers, and assign each data point to the cluster where the nearest cluster center is located. Use the Euclidean distance to calculate the distance between each data point and the cluster center. The formula is as follows:

[0182]

[0183] In the formula, y represents any one of the f points, yj represents the value corresponding to this point on the jth principal component; G represents any cluster center among the Ka cluster centers, G j Represents the component of the cluster center on the jth principal component.

[0184] (4.c.3) Optimize to obtain the optimal cluster center. Re-use the mean of all data points in each cluster as the cluster center, and repeat the iteration until the cluster center no longer changes, and the K-means clustering result is obtained.

[0185] (4.c.4) Based on the clustering results, calculate the mean value μ and standard deviation σ of the key quality indicators of the specialty crops in each cluster to obtain the quality statistical characteristics of each cluster.

[0186] (4.d) Based on the clustering results of the production environment, the natural environment gradient of the growth environment of this characteristic crop can be constructed.

[0187] Step 5: Analyze the response pattern of key quality of specialty crops to each environmental factor based on the natural environment gradient, and establish the suitability classification of key environmental factors on a large scale. The process is as follows: Figure 6 .

[0188] (5.a) Preliminary screening of environmental variables, screening and removing some environmental factors with strong correlation.

[0189] (5.a.1) Based on the data from the entire study area, use the method in step (4.a) to calculate the correlation coefficient between each growth environment factor of the specialty crops and their quality indicators.

[0190] (5.a.2) Based on the data from the entire study area, use the method in step (4.a) to calculate the correlation coefficient between the growth environment factors of the specialty crop.

[0191] (5.a.3) Find pairs of environmental factors with correlation greater than 0.8, retain the one with higher correlation with the quality index, and delete the other one.

[0192] (5.b) Based on the natural environmental gradient established in step (4.d), through correlation analysis, the response of the key quality indicators of the specialty crops to different environmental factors under different environmental gradients is studied, and the suitability level of each environmental factor is divided.

[0193] (5.b.1) Based on the production environment of each cluster, i.e., the environmental gradient, use the correlation analysis method in step (4.a) to establish the correlation between the key quality indicators of specialty crops and various environmental factors.

[0194] (5.b.2) Combine the correlation coefficient and scatter plot to compare the responses of key quality indicators of specialty crops to various environmental factors under different environmental gradients, and explore the response patterns of specialty crops to each environmental factor in a large area.

[0195] (5.b.3) Based on the response pattern, find the optimal threshold and classify the suitability of characteristic crops for each environmental factor.

[0196] Step 6: Use the multi-layer weighted overlay method to construct a growth suitability evaluation model for the characteristic crops and obtain large-scale suitability evaluation results.

[0197] (6.a) The potential distribution area of ​​characteristic crops identified in step (3.e) is used as the study area.

[0198] (6.b) With the help of stepwise linear regression or nonlinear regression such as random forest, identify the environmental factors that have a significant impact on the growth of the characteristic crops as evaluation indicators, and establish the weight of the evaluation indicator based on the regression coefficient.

[0199] (6.b.1) Based on the data from the recording points in the entire study area, the regression relationship between the key quality indicators of specialty crops and environmental factors was established using stepwise linear regression. Using the Akaike Information Criterion (AIC) as the selection criterion, p environmental factors that have a significant impact on the growth of specialty crops were screened out and used as evaluation indicators.

[0200] (6.b.2) Based on step (5.b.3), establish the suitability classification of the evaluation indicators.

[0201] (6.b.3) Determine the weight based on the impact of the evaluation index on the key quality index of the specialty crop, that is, take the average value of the standardized regression coefficient of the evaluation index as the weight of the index, and normalize all weights to ensure that the weight of all evaluation indicators is W i The sum of is equal to 1.

[0202]

[0203] Among them, p represents the number of evaluation indicators used to construct the suitability evaluation model (i.e. the total number of environmental factors).

[0204] (6.c) Unify the dimensions of the suitability levels of each evaluation indicator, and perform spatial overlay operations on the evaluation indicators of multiple layers in ArcGIS software to obtain a comprehensive evaluation score.

[0205] (6.c.1) Reclassify each evaluation index to unify the dimensions. For example, if there are three levels, the most suitable level is set at 100 points, the relatively suitable level is 80 points, and the second most suitable level is 60 points. The total weighted score is 60 to 100 points.

[0206] (6.c.2) ArcGIS software is used to perform weighted overlay operations on the raster layers of the evaluation indicators to obtain the comprehensive score S (x, y) of each corresponding raster in the study area. The basic operation process is as follows:

[0207]

[0208] Among them, W i is the weight of the i-th layer, S i (x, y) represents the suitability score of the ith evaluation indicator at location (x, y).

[0209] (6.d) Use the natural breakpoint method (Formula (22)) to divide the comprehensive evaluation score into b categories, i.e. b levels, as the growth suitability level of the specialty crops, and finally complete the growth suitability evaluation of the specialty agricultural products in the study area.

[0210]

[0211] Where S is the sum of all intra-class variances, b is the number of levels, and n is j is the number of data points in the jth level, x ij is the value of the i-th data point in the j-th level, μ j is the average value of the jth level. The basic principle is to iterate continuously and find a classification method that satisfies the minimum s.

[0212] The following is an example of a study on the suitability zoning of potential wolfberry planting areas in my country using the method of the present invention.

[0213] Wolfberry is a traditional Chinese medicinal material, and Ningxia is the authentic production area of ​​wolfberry. As the market demand for wolfberry increases sharply, many places across the country have introduced and tested wolfberry. However, due to the lack of large-scale wolfberry planting suitability zoning guidance, many regions have blindly introduced wolfberry, resulting in poor planting returns and high risks.

[0214] Through systematic surveys and sampling tests, we obtain information on the distribution of wolfberry production areas and key quality data, and collect a variety of environmental variables covering the country, such as climate, topography, and soil, to build a database. The specific implementation steps are as follows:

[0215] The first step is to obtain information on wolfberry growth distribution points and key quality data in multiple production areas across the country through systematic surveys and sampling tests, and to collect a variety of environmental variables covering the country, such as climate, topography, and soil, to build a database.

[0216] (A1) The information of 255 wolfberry planting sites was collected and preprocessed, including 71 in Ningxia, 104 in Gansu, 51 in Qinghai, 14 in Inner Mongolia and 15 in Xinjiang. The longitude and latitude coordinates of the sites were converted into decimal. 52 national-scale environmental variables were collected, including 28 climate variables, 9 soil variables, 12 terrain variables, and 3 vegetation remote sensing indices. All of them were resampled to 1 km resolution, and the geographic coordinates were uniformly converted to GCS_WGS_1984.

[0217] (A2) The extract function in the raster package in R was used to batch extract the raster values ​​of the environmental variable layer corresponding to 255 wolfberry planting sites. The dataset was described as data = (long, lati, x1, x2, …, x52), which is a 255 × 54 vector.

[0218] In the second step, multiple species distribution models were screened, and the maximum entropy model was finally determined as the optimal model. The maximum entropy model was trained through data to identify the potential distribution areas of wolfberry in my country.

[0219] (B1) Load the Maxent software and perform model training. The first step is data preparation. Save the wolfberry distribution point data in .csv format and import it into the sample column; save the environmental variable layer in .asc format and import it into the environmental layer column; create an output folder at the output path.

[0220] (B2) Parameter settings. Check Generate response curve, Make prediction graph, and Execute Jackknife method to measure variable importance. Select 10-fold cross validation for data segmentation.

[0221] (B3) Environmental variable screening to reduce multicollinearity and prevent overfitting. After running the model three times, all variables with a contribution rate of 0 were removed in turn. This step removed a total of 14 variables. Pearson correlation analysis was performed on the remaining 38 variables. With the absolute value of correlation > 0.8 as the threshold, one of them was retained and the other highly correlated variable was removed. Finally, 22 ecological factors remained as input environmental variables.

[0222] (B4) Model output and model accuracy verification. The prediction accuracy of the model was verified using the ROC curve (receiver operating characteristic curve) and AUC value (area under the curve). Figure 7a , 7b As shown in the figure, using 10-fold cross validation, the average AUC of the test set is 0.98, indicating that the model is very accurate; the standard deviation of the results of 10 repeated runs ranges from 0 to 0.11, indicating that the model is very stable.

[0223] (B5) Select the classification threshold to convert the continuous distribution probability of wolfberry into two categories, that is, suitable for growth or unsuitable for growth. Two different thresholds were selected, threshold 1 is "maximum training sensitivity plus specificity (MSS)", and threshold 2 is determined based on "balanced training omission rate, prediction area and threshold". From the binary classification results ( Figure 8 ), the distribution patterns of the potential growth areas of wolfberry divided by the two thresholds are generally consistent, and the main difference is reflected in the area. The potential growth areas of wolfberry divided by threshold 1 and threshold 2 account for 4.09% and 9.68% of the national land area (excluding water areas), respectively.

[0224] (B6) Query verification, field investigation and threshold determination. The areas demarcated by the two thresholds were verified through network query and field investigation. The results showed that the areas demarcated by threshold 1 were more conservative, while the results selected by threshold 2 were more representative of the actual potential wolfberry planting areas.

[0225] In the third step, correlation analysis, cluster analysis and other methods were used to cluster the quality indicators and growth environments of wolfberries from multiple production areas, and the characteristics of quality classification and environmental classification were summarized.

[0226] (C1) A total of 211 fresh wolfberry samples were collected from the main wolfberry producing areas across the country. Five wolfberry quality indices, including longitudinal diameter, transverse diameter, 100-grain weight, soluble solids (SSC), and fruit shape index, were tested and obtained. The relationship between the five quality indices was studied using correlation analysis. The results showed that there was a strong positive correlation between the appearance indices, while there was a significant negative correlation between the appearance indices and the internal indices.

[0227] (C2) The same K-means clustering method was used to cluster the environmental factors of the sampling points. 12 principal components with eigenvalues ​​greater than 1 were extracted (the cumulative variance explanation rate was 84%), and the optimal cluster number k=3 was determined using the "elbow rule".

[0228] From the clustering results, the first category is in the Qaidam Basin in Qinghai, which is a high-altitude cold production area. The characteristics of the fruit are large grains, long fruit shape, and moderate sugar content. The second category is distributed in Zhangye City, Gansu, the southern part of Wuwei City, and Baiyin City in Ningxia and Gansu. It belongs to the traditional wolfberry production area at medium altitude. The characteristics of the fruit are medium grain size, long fruit shape, and moderately low sugar content. The third category is distributed in Jinghe County, Xinjiang and its surrounding areas, Jiuquan City, Gansu, northern Ningxia, and the Hetao area of ​​Inner Mongolia. The altitude of this area is relatively low and it belongs to the warm temperature production area. The characteristics of the fruit in this area are large grain roundness and high sugar content. Comparing the differences in wolfberry quality under different environmental clusters, it was found that the difference in environmental gradients in the production area has a significant impact on the quality of wolfberry.

[0229] The fourth step is to use the results of the origin environment classification as the natural environment gradient, analyze the response pattern of the core quality of specialty crops to each environmental factor based on the natural environment gradient, and establish a suitability classification of key environmental factors on a large scale.

[0230] (D1) According to step (B3) in the second step, the environmental variables are initially screened to reduce data redundancy.

[0231] (D2) The correlation between wolfberry quality indicators and environmental factors under different environmental gradients was established. The results showed that wolfberry quality indicators had a nonlinear response to environmental factors under different environmental gradients. By comparing the scatter plots and fitting curves, the response pattern and key threshold of wolfberry quality to environmental factors were determined.

[0232] The fifth step is to screen the evaluation indicators, establish the suitability classification and weights of the indicators, and use the multi-layer weighted overlay method to construct the evaluation system for the suitability of wolfberry growth in my country.

[0233] (E1) In the second step (B6), in addition to the potential distribution area for wolfberry planting, the study area for wolfberry growth suitability evaluation is also divided.

[0234] (E2) Evaluation index screening. The regression relationship between the quality indexes and environmental factors of wolfberry in the study area was established by stepwise regression, and the environmental factors that have an impact on the quality indexes of wolfberry were screened and obtained. Finally, 16 environmental factors were selected as evaluation indexes.

[0235] (E3) Based on the potential planting area of ​​Chinese wolfberry and the corresponding environmental factor range obtained in the second step, and according to the response pattern and key threshold of wolfberry quality and environmental factors in the fourth step (D2), the suitability classification of evaluation indicators is established, divided into three levels, namely, the most suitable, relatively suitable and sub-suitable. In order to unify the dimensions, the percentage system is used, and the three levels are represented by 100 points, 80 points and 60 points respectively.

[0236] (E4) Obtaining the weight of the evaluation factor. The standardized regression coefficient between the wolfberry quality index and the evaluation index reflects the importance of the index. The wolfberry fruit shape index is used as a representative of the appearance index, and the wolfberry soluble solids are used as a representative of the internal index. The average of the absolute values ​​of the standardized regression coefficients between the two wolfberry indexes and the evaluation index is used as the weight of the corresponding index.

[0237] (E5) Calculate the comprehensive score. The comprehensive suitability score is calculated by weighted superposition of multiple indicators, and the comprehensive score ranges from 60 to 100 points.

[0238] (E6) Establish suitability zones. The comprehensive scores were graded using the natural breakpoint method to obtain three suitability levels, namely the most suitable planting area, the relatively suitable planting area, and the second most suitable planting area.

[0239] Based on the above process, the prediction results of the example case are as follows Fig. 9 As shown, the present invention uses a "top-down" species distribution model, based on the wolfberry distribution point data and a variety of environmental variable data training model, to identify the potential distribution area of ​​wolfberry on a national scale; and through correlation analysis, cluster analysis and other methods, the natural environment gradient is divided, and the relationship between wolfberry quality and environmental factors studied on this basis is identified. The response mode and threshold of environmental factors to wolfberry growth are identified, and evaluation indicators are screened out at the same time, and an indicator classification and weight system is established, and finally a "bottom-up" multi-layer weighted superposition plant growth suitability evaluation model on a national scale is constructed, and a large-scale wolfberry growth suitability zoning is completed. This method can make full use of survey data and easily accessible multi-source environmental variables, and quickly and low-costly realize the growth suitability growth zoning of specialty crops on a large scale, which has high practical value.

Claims

1. A method for rapidly evaluating the growth suitability of specialty crops in a large region, characterized in that: The steps include: S1, through systematic survey and sampling test, obtain the growth distribution point information and key quality data of characteristic crops in high-yield areas, collect environmental variables covering the study area, and build a database; S2, construct species distribution model based on the constructed database and maximum entropy principle; S3, training species distribution model and calculating the distribution probability of characteristic crops; Collect multiple classification thresholds, perform binary classification on the distribution probability, and select the optimal threshold to obtain the potential distribution area of ​​characteristic crops; S4, cluster analysis is used to cluster the growth environments of characteristic crops at multiple recording points to obtain quality classification and origin environment classification results; S5, using the results of the origin environment classification as the natural environment gradient, analyzing the response pattern of the key qualities of characteristic crops to each environmental factor based on the natural environment gradient, and establishing the suitability classification of key environmental factors on a large scale; S6, using the multi-layer weighted overlay method, constructs a growth suitability evaluation model for characteristic crops and obtains large-scale suitability evaluation results.

2. The method for rapidly evaluating the growth suitability of characteristic crops in a large region according to claim 1, characterized in that: The implementation steps for building a database are as follows: S11, complete the systematic sampling by collecting the growth and distribution information of characteristic crops; S12, based on the distribution map, carry out actual sampling work and record the actual distribution points of characteristic crops {P1, P2, P3, ..., P n }, where P n Indicate the distribution point of the nth species, and conduct surveys and sampling tests to obtain product quality and soil properties at the planting site; S13, collects multiple environmental variables covering the study area, including climate, topography, soil, and organisms; S14, preprocessing the collected species distribution point data: The species distribution points were exported to the same coordinate system, the geographic coordinates were uniformly converted to GCS_WGS_1984, and the projection coordinates were uniformly converted to WGS_1984_Albers; Set a minimum geographic distance threshold D to filter out other points whose geographic location is less than the threshold D; set the spatial resolution of the environmental variable grid to r, and the following conditions must be met: Then the minimum geographic distance threshold D is expressed as: Where a is an arbitrary constant greater than 1; By screening the points, we ensure that the Euclidean distance d ≥ D between all the record points, and finally obtain n species distribution points site: site = (long, lati), which is an n × 2 vector, where long represents longitude and lati represents latitude; S15, preprocessing and standardizing the environmental variables; S16, using GIS software to extract the raster values ​​of the environmental variable layer corresponding to the species distribution points. The data set is described as data = (site, z1, z2, ..., zm), which is an n × (m + 2) vector, where m represents the number of environmental layers.

3. The method for rapidly evaluating the growth suitability of characteristic crops in a large region according to claim 2, characterized in that: In step S11, the growth distribution information of characteristic crops is collected through multiple channels to complete the system point sampling, and the implementation steps are as follows: S111, obtaining growth and distribution information of characteristic crops as reference information for site placement through literature search and database access; S112, using satellite image maps to observe the image features of the characteristic crops at the reference point, and based on this, continue to search for the planting distribution of the characteristic crops around the reference point, and preliminarily obtain the precise distribution map set of characteristic crops {S1, S2, S3, ..., S n }, where S n represents the nth image spot; S113, collect high-resolution remote sensing images of the study area, use distribution patch sets, combine remote sensing images, and identify similar image patches through supervised learning methods as characteristic crop planting points as actual point distribution references; S114, confirm the total number of points according to the distribution range of characteristic crops and the complexity of the terrain in the distribution area, and complete the sampling of points based on the planned point density of the terrain and soil types in different areas.

4. The method for rapidly evaluating the growth suitability of characteristic crops in a large region according to claim 1, characterized in that: The implementation steps for building a species distribution model are as follows: S21, based on the collected environmental variables, define the characteristic function f i (x), represents the value of any grid x on the i-th environmental variable; then for any feature, there is an expected value constraint E p [f i (x)]: E p [f i (x)]=∑ x∈X f i (x)P(x)=c i Among them, c i is the expected value of all grids of the ith environmental variable; P(x) represents the probability distribution of any grid value in the study area; For any grid point, X represents the set of all grid points in the study area; S22, construct the Lagrangian function and solve the probability distribution with the maximum entropy under given constraints: Where L represents the Lagrangian function; -∑ x∈X P(x)logP(x) is the expression of entropy; λ0(∑ x∈X P(x)-1) is to ensure that the sum of the probabilities P(x) is equal to 1, and λ0 represents the Lagrange multiplier, which is used to adjust the weight of this item; Contains m environmental variable constraints, λ i Indicates that each constraint corresponds to a Lagrange multiplier; Take the partial derivative of each P(x) and make it equal to 0: Solving for this yields: Since the sum of all probabilities is equal to 1, for x∈X, the formula for calculating the sum of probabilities is as follows: definition but After finishing, we get: As the species distribution probability, then λ i is regarded as the weight parameter of the i-th environmental variable, and Z(λ) is the normalization constant; S23, perform parameter estimation and find the model parameter λ i , by maximizing the log-likelihood function to achieve the predicted probability distribution of the crop distribution model The distribution is closest to the actual observed data: In the formula, n represents the number of species distribution points, x (j) represents the set of environmental conditions at the jth species distribution point; substituting the probability expression P(x) into the above formula, we get: S24, establish the gradient function, for each λ i Find partial derivatives: In the formula, E p [f i ] indicates f i The expected value predicted by the model under the probability distribution P(x|λ) with given parameter λ; For each observed x (j) f i The sum of the values ​​of nE p [f i ] represents the comprehensive expected value under the model prediction; S25, using the gradient ascent method to iterate and update the parameter λ i ,make Close to 0 to maximize Find an optimal set of λ i , so that the predicted probability distribution of the model fits the observed data best, the expression is as follows: In the formula, α is the learning rate, which determines the magnitude of each update step.

5. The method for rapidly evaluating the growth suitability of characteristic crops in a large region according to claim 1, characterized in that: In step S3, the steps for obtaining the potential distribution area of ​​characteristic crops are as follows: S31, dividing the database into a training set and a validation set, the training set is used to train the constructed crop distribution model, and the validation set is used to test the prediction performance of the crop distribution model; S32, running the crop distribution model and outputting species distribution probability; selecting the optimal threshold t, converting the species distribution probability into the potential distribution area of ​​characteristic crops, and converting the distribution probability into the presence or absence of the species to achieve binary classification; S33, select threshold t based on sensitivity and specificity; S34, using a threshold calculation method to obtain a set of classification thresholds {T1, T2, T3, ..., T n }, each classification threshold T i , calculate and obtain a potential distribution area R i , then we get the potential distribution area set {R1, R2, R3, ..., R n }; S35, combining literature query, network query and field investigation, verifies the actual boundaries of the potential distribution areas of characteristic crops, determines the optimal classification threshold, and obtains the most accurate potential distribution areas of characteristic crops.

6. The method for rapidly evaluating the growth suitability of characteristic crops in a large region according to claim 1, characterized in that: In step S4, cluster analysis is used to cluster the key qualities and growth environments of the characteristic crops at multiple recording points, and the implementation steps are as follows: S41, conduct Pearson correlation analysis on multiple quality indicators of specialty crops to understand the relationship between quality indicators: In the formula, qa i ,qb i are the values ​​corresponding to the two quality indicators qa and qb at any recording point; f represents the number of recording points where the quality of the characteristic crop is obtained, qa mean ,qb mean represents the mean of the two quality indicators; S42, principal component analysis is performed on the growth environment factors of characteristic crops. The specific steps are as follows: S421, for an f×m matrix consisting of f points and m environmental variables, perform data standardization to convert each column of quality indicator data into a mean equal to 0 and a standard deviation of 1; S422, calculate the covariance matrix, perform eigenvalue decomposition on the covariance matrix, obtain m principal components, select the first e principal components whose cumulative variance contribution rate reaches the set threshold, and record the set as {PC1, PC2, PC3, ..., PC e }; S423, using the eigenvector corresponding to the selected principal component, converting the original data into a principal component space; S43, clustering the e principal components using K-means clustering to obtain clustering results of the growth environment of characteristic crops, and counting the quality characteristics of characteristic crops in different clusters; S44, based on the clustering results of the production environment, construct the natural environment gradient of the growth environment of the characteristic crops.

7. The method for rapidly evaluating the growth suitability of characteristic crops in a large region according to claim 6, characterized in that: In step S5, the steps for establishing the suitability classification of key environmental factors on a large scale are as follows: S51, based on the data of the recording points in the entire study area, the Pearson correlation analysis method was used to calculate the correlation coefficient between each growth environment factor of the characteristic crops and their quality index, and the correlation coefficient between the growth environment factors of the characteristic crops; find the environmental factor pairs with a correlation greater than 0.8, retain the factor with a higher correlation with the quality index, and delete the other one; S52, for each production environment classification, use Pearson correlation analysis to establish the correlation between the key quality indicators of specialty crops and multiple environmental factors; combine the correlation coefficient and scatter plot to compare the response of the key quality indicators of specialty crops to various environmental factors under different environmental gradients, and explore the response pattern of specialty crops to each environmental factor in a large area; based on the response pattern, find the optimal threshold and divide the suitability of specialty crops to each environmental factor into grades.

8. The method for rapidly evaluating the growth suitability of specialty crops in a large region according to claim 6, characterized in that: In step S6, the implementation steps of constructing a characteristic crop growth suitability evaluation model are as follows: S61, the potential distribution area of ​​characteristic crops was taken as the study area; S62, using stepwise linear regression or random forest, identify environmental factors that have a significant impact on the growth of specialty crops as evaluation indicators, and establish the weights of the evaluation indicators based on standardized regression coefficients or variable importance; S63, unify the dimensions of the suitability levels of each evaluation index, and perform spatial superposition operations on the evaluation indexes of multiple layers in ArcGIS software to obtain a comprehensive evaluation score; S64, using the natural breakpoint method, divides the comprehensive evaluation score into b levels as the growth suitability level of the characteristic agricultural crop, and finally completes the growth suitability evaluation of the characteristic agricultural products in the study area.

Citation Information

Cited By

  • Low-maintenance garden configuration method for zingiberaceae plants

    CN120240266A

  • Intelligent division method for appropriate region of genuine medicinal material planting based on double-model coupling

    CN121960894A