Geographically identified agricultural product suitability evaluation method, device, equipment and medium
By using the connected maximum entropy model and random forest model in the geographic identification agricultural product suitability evaluation method and combining the GeoShapley method for correlation analysis, the problem of difficulty in accurately predicting the suitability distribution of geographical indication agricultural products in the existing technology is solved, and high-precision suitability distribution prediction and impact factor analysis are achieved.
Patent Information
- Application Number
- CN202411837654.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-13
AI Technical Summary
It is difficult for the prior art to accurately predict the suitability distribution of geographical indication agricultural products and clarify the relative contribution of influence factors to suitability, resulting in threats to production efficiency and quality.
A method for evaluating the suitability of geographical indication agricultural products is proposed. By obtaining the impact factor data of the target area and the distribution data of geographical indication agricultural products, the maximum entropy model and random forest model are used for prediction, and correlation analysis is performed in combination with the GeoShapley method.
The accuracy of suitability distribution prediction is significantly improved, and the specific role of spatial influence factors is revealed in detail, providing a scientific basis for the rational production layout of geographical indication agricultural products, and providing guidance for policy formulation and regional agricultural management.
Smart Images

Figure CN119940953A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a method, device, equipment and medium for evaluating the suitability of geographically indicated agricultural products. Background Art
[0002] Geographical indications of agricultural products (GIAP) are specialty agricultural products formed due to the unique natural conditions, historical culture or traditional skills of a specific geographical area. They have important brand value and market competitiveness and have become an important tool for promoting regional economy, protecting biodiversity and promoting sustainable agriculture. The suitability distribution of GIAP directly affects its production efficiency and quality. However, with the increasing environmental pressure brought about by climate change, soil degradation and dense population (Nuary et al., 2022), how to accurately predict the suitability distribution of GIAP and clarify the relative contribution of influencing factors to suitability has become an urgent problem to be solved. These challenges may threaten the quantity and quality of GIAP production, which in turn affects the realization of sustainable development goals and the implementation of relevant national policies. Therefore, establishing a methodological system that can accurately predict the suitability distribution of GIAP and analyze the contribution of influencing factors is of great significance to achieving food security and promoting regional sustainable development.
[0003] Although existing studies have made significant progress in suitability distribution modeling and provided important references for suitability research, suitability evaluation in different regions still faces challenges in standardization, especially when data availability and data quality are uneven, and there are still certain limitations. Summary of the invention
[0004] The main purpose of the embodiments of the present invention is to propose a method, device, equipment and medium for evaluating the suitability of geographically indicated agricultural products, in order to solve at least one problem of the prior art. The present invention can accurately realize the suitability evaluation of geographically indicated agricultural products.
[0005] To achieve the above purpose, one aspect of an embodiment of the present invention provides a method for evaluating the suitability of geographically indicated agricultural products, the method comprising:
[0006] Obtaining impact factor data of a target area; the impact factor data includes multiple impact factors;
[0007] Obtain the distribution data of geographical indication agricultural products in the target area; based on the influencing factor data and distribution data, obtain the target prediction value of the adaptive distribution of geographical indication agricultural products in the target area through a preset prediction model; the prediction model includes a maximum entropy model and a random forest model in series;
[0008] Based on the influencing factor data and target prediction values, the GeoShapley method was used to conduct an association analysis between the suitability of geographical indication agricultural products and influencing factors, and the associated evaluation data were obtained.
[0009] In some embodiments, the influencing factor data includes climate data, soil data, and human activity factor data, and the climate data, soil data, and human activity factor data all include several influencing factors; obtaining the influencing factor data of the target area includes the following steps:
[0010] Obtain global climate data from a preset climate database, perform mask extraction on the global climate data based on the boundary vector file of the target area, and obtain the climate data of the target area;
[0011] Acquire soil properties and trace element data of a target area from a preset soil database as soil data;
[0012] The human activity variables of the target area are obtained from the preset statistical data, the human activity variables are associated with the vector map of the target area, and then the data is converted into a raster format to obtain the human activity factor data.
[0013] In some embodiments, obtaining the distribution data of geographical indication agricultural products in a target area includes the following steps:
[0014] Obtain the latitude and longitude data of the distribution points of geographical indication agricultural products in the target area;
[0015] The distribution points are regarded as existence points, and background points are randomly generated in the area outside the target area based on the latitude and longitude data as non-existence points;
[0016] The distribution data of geographical indication agricultural products in the target area are obtained based on the existence points and non-existence points.
[0017] In some embodiments, based on the influencing factor data and the distribution data, obtaining the target prediction value of the adaptive distribution of the geographical indication agricultural products in the target area through a preset prediction model processing includes the following steps:
[0018] The distribution data and the impact factor data are spatially superimposed to construct an environmental data set;
[0019] Input the environmental data set into the maximum entropy model to obtain the initial prediction value;
[0020] The initial prediction value is added to the environmental data set as a new influencing factor, and then input into the random forest model to obtain the target prediction value of the adaptive distribution of geographical indication agricultural products in the target area.
[0021] In some embodiments, the method further comprises the following steps:
[0022] Based on the Pearson correlation coefficient, the factors are screened in the impact factor data by combining the correlation and factor contribution with the preset screening threshold, and the impact factor data is updated based on the results of factor screening.
[0023] In some embodiments, based on the influencing factor data and the target prediction value, the GeoShapley method is used to perform a correlation analysis between the suitability of the geographical indication agricultural products and the influencing factors to obtain the correlation evaluation data, including the following steps:
[0024] The SHAP value, Moran's index and spatial lag value are obtained by processing the impact factor data and the target prediction value;
[0025] The factor global contribution map and factor positive and negative contribution map are obtained according to the SHAP value. Based on the factor global contribution map and factor positive and negative contribution map, the spatial distribution characteristics and positive and negative contribution of the influencing factors to the suitability of geographical indication agricultural products are determined in turn.
[0026] The spatial correlation between influencing factors and the suitability of geographical indication agricultural products was determined based on the Moran index and spatial lag value.
[0027] In some embodiments, the SHAP value, the Moran's index and the spatial lag value are obtained according to the impact factor data and the target prediction value, including the following steps:
[0028] According to the influencing factor data and the target prediction value, the preset first spatial data analysis software is imported to synthesize a vector shp file, and the SHAP value is obtained based on the shp file by using the GeoShapley method;
[0029] The influencing factor data and target prediction values are imported into the preset second spatial data analysis software to obtain the Moran index and spatial lag value.
[0030] To achieve the above-mentioned purpose, another aspect of the embodiment of the present invention provides a device for evaluating the suitability of geographically indicated agricultural products, the device comprising:
[0031] The first module is used to obtain the impact factor data of the target area; the impact factor data includes multiple impact factors;
[0032] The second module is used to obtain the distribution data of geographical indication agricultural products in the target area; based on the influencing factor data and distribution data, the target prediction value of the adaptive distribution of geographical indication agricultural products in the target area is obtained through the preset prediction model; the prediction model includes a maximum entropy model and a random forest model in series;
[0033] The third module is used to conduct a correlation analysis between the suitability of geographical indication agricultural products and influencing factors based on the influencing factor data and target prediction values using the GeoShapley method to obtain correlation evaluation data.
[0034] To achieve the above objective, another aspect of an embodiment of the present invention provides an electronic device, the electronic device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the above method when executing the computer program.
[0035] To achieve the above objective, another aspect of an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above method is implemented.
[0036] The embodiments of the present invention include at least the following beneficial effects: the present invention provides a method, device, equipment and medium for evaluating the suitability of geographical indication agricultural products, the scheme obtains the influencing factor data of the target area; the influencing factor data includes multiple influencing factors; the distribution data of the geographical indication agricultural products in the target area is obtained; based on the influencing factor data and the distribution data, the target prediction value of the adaptive distribution of the geographical indication agricultural products in the target area is obtained by processing with a preset prediction model; the prediction model includes a maximum entropy model and a random forest model in series; based on the influencing factor data and the target prediction value, the GeoShapley method is used to perform the correlation analysis between the suitability of the geographical indication agricultural products and the influencing factors to obtain the correlation evaluation data. The present invention constructs a new series model as a model for predicting the suitability distribution of GIAP, integrating the advantages of multiple models to improve the accuracy of predicting the suitability distribution of GIAP. The present invention significantly improves the accuracy of the prediction of the suitability distribution, reveals in detail the specific role of each influencing factor in space, can provide a scientific basis for the reasonable production layout of geographical indication agricultural products (GIAP), and provides guidance for policy formulation and regional agricultural management. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 is a flow chart of a method for evaluating suitability of geographically indicated agricultural products provided by an embodiment of the present invention;
[0038] Figure 2 is a schematic diagram of an example of the correlation of the Pearson correlation coefficient factor provided by an embodiment of the present invention;
[0039] Figure 3 It is an overall flow chart of the geographical indication agricultural product suitability evaluation method provided by the embodiment of the present invention;
[0040] Figure 4 It is a data logic principle flow chart of the geographical indication agricultural product suitability evaluation method provided by an embodiment of the present invention;
[0041] Figure 5 is a schematic diagram of comparing performance indicators of models provided by an embodiment of the present invention;
[0042] Figure 6 is a schematic diagram of positive and negative contributions of SHAP values provided by an embodiment of the present invention;
[0043] Figure 7 is a schematic diagram of the global contribution of SHAP values provided by an embodiment of the present invention;
[0044] Figure 8 is a schematic diagram of the SHAP value space feature provided by an embodiment of the present invention;
[0045] Fig. 9 is a schematic diagram of the Moran's I index of suitability and environmental factors provided by an embodiment of the present invention;
[0046] Fig.10 It is a schematic diagram of the structure of a device for evaluating the suitability of geographically indicated agricultural products provided by an embodiment of the present invention;
[0047] Fig.11 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0048] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present invention. They are only examples of devices and methods consistent with some aspects of the embodiments of the present invention as detailed in the attached claims.
[0049] It is understood that the terms "first", "second", etc. used in the present invention can be used to describe various concepts in the present invention, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determination".
[0050] The terms "at least one", "multiple", "each", "any", etc. used in the present invention, at least one includes one, two or more, multiple includes two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.
[0051] Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the present invention are only for the purpose of describing the embodiments of the present invention and are not intended to limit the present invention.
[0052] To facilitate the understanding of the technical solution of the present invention, the technical means and features that may appear in the technical solution of the present invention are explained below:
[0053] SHAP value (SHapley Additive exPlanation value) is a method for explaining the prediction results of machine learning models, based on the Shapley value theory in cooperative game theory. SHAP value helps understand the behavior and decision-making process of the model by evaluating the contribution of each feature to the model's prediction results.
[0054] GeoShapley is a new method based on the concept of Shapley value, but designed for geospatial analysis, used to measure spatial effects in machine learning models. It extends the Shapley value framework, takes geographic neighborhood as the core, and takes geographic location information as an important analysis factor; it considers the interaction between spatial variables and the impact of spatial distribution characteristics on the target variable; it emphasizes the total contribution of variables to the target within the geographic region, and combines the Shapley value allocation rule with techniques such as Geographically Weighted Regression (GWR). It focuses on geospatial data, and the input data usually contains latitude and longitude or geographic region divisions. It is particularly suitable for analyzing variables related to spatial location (such as land use, climate conditions, pollution source distribution, etc.). It can capture the interaction and autocorrelation of spatial variables, provide an explanation of the spatial dimension, and is suitable for the fields of geographic information science and environmental science. Compared with the TreeSHAP method, the GeoShapley method expands the application field of Shapley value, focuses on the analysis of geospatial data, and emphasizes spatial correlation, while TreeSHAP is a machine learning model interpretation tool that focuses on the impact of features on prediction results.
[0055] The geographical indication agricultural product suitability evaluation method provided in the embodiment of the present invention relates to the field of data processing technology. The geographical indication agricultural product suitability evaluation method provided in the embodiment of the present invention can be applied to a terminal, can also be applied to a server, and can also be software running in a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and a car terminal, etc., but is not limited to this; the server side can be configured as an independent physical server, or a server cluster or distributed system composed of multiple physical servers, and can also be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements the geographical indication agricultural product suitability evaluation method, etc., but is not limited to the above forms.
[0056] The present invention can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0057] Figure 1 is an optional flow chart of the method for evaluating the suitability of geographically indicated agricultural products provided in an embodiment of the present invention. Figure 1 The method may include but is not limited to steps S100 to S300.
[0058] S100, obtaining impact factor data of a target area;
[0059] Among them, the impact factor data includes multiple impact factors;
[0060] It should be noted that the influencing factor data include climate data, soil data and human activity factor data, and the climate data, soil data and human activity factor data all include several influencing factors; in some embodiments, step S100 may include the following steps: obtaining global climate data from a preset climate database, performing mask extraction on the global climate data based on the boundary vector file of the target area, and obtaining the climate data of the target area; obtaining soil properties and trace element data of the target area as soil data from a preset soil database; obtaining human activity variables of the target area from preset statistical data, associating the human activity variables with the vector map of the target area, and then converting the data into a raster format to obtain human activity factor data.
[0061] For example, in some specific implementations, climate data can be obtained from the WorldClim 2.1 database (https: / / www.worldclim.org), the BCC-CSM2-MR global climate model (GCM) is selected, the spatial resolution is 2.5 arc minutes (about 5 kilometers), the time range is set to 2021-2040, and it is based on the shared socioeconomic pathway scenario SSP126. 19 bioclimatic variables (BIO1-BIO19) are downloaded in GeoTIFF format, and the relevant bioclimatic variables are extracted using the "Image Analysis" tool of ArcGIS 10.8.1 to obtain global climate data. In order to focus on the study area, the global climate data is masked and extracted using the boundary vector file of the target area (such as a national region or a provincial or municipal region, taking country A as an example), and the climate data of country A is obtained through the "Extract by Mask" function.
[0062] The soil data source can be obtained from the Global Harmonized Soil Database v2.0 (HWSD v2.0) (https: / / gaez.fao.org). Due to the large amount of data, R software version 4.2.0 was used for data processing and extraction. The soil properties of the D1 layer (0–20 cm) were selected as the main research object, including variables such as soil texture, organic matter content, pH value, and nutrient content. According to the FAO soil classification system, the soil of country A is divided into 22 soil types, including ferroalloy soil, eluvial soil, arid soil, and desert soil. The trace element data are from the Guangdong Provincial Soil Environment Monitoring Center.
[0063] The human activity factor data (taking Country A as an example) can be derived from the Statistical Yearbook of Country A and the Statistical Yearbook of County Areas of Country A in 2023. The selected human activity variables include population density, gross domestic product (GDP), etc. The sorted data are imported into ArcGIS 10.8.1 in CSV format and associated with the administrative division vector map of Country A. To meet the model requirements, spatial interpolation methods are used when necessary, and the data are finally converted to raster format through the "Feature to Raster" tool.
[0064] By supplementing multiple types of factors, the present invention enriches the research content and makes the predicted distribution of suitability of geographical indication agricultural products more convincing.
[0065] It should also be noted that in some embodiments, the method may also include the following steps: based on the Pearson correlation coefficient, factor screening is performed on the influencing factor data through correlation and factor contribution combined with a preset screening threshold, and the influencing factor data is updated based on the results of factor screening.
[0066] For example, in some specific implementations, the Pearson correlation coefficient is first tested before inputting the influencing factors. Based on the correlation and the contribution of the factors, the present invention performs a preliminary screening of the factors to eliminate the influence of multicollinearity to the greatest extent.
[0067] In some specific application scenarios, assuming that there are 40 influencing factors from C1 to C40, factor screening can be achieved as follows:
[0068] First, the correlation between factors was calculated using the Pearson correlation coefficient, and then factors with |r|>0.8 were screened. Figure 2 And the following Table 1 (Pearson correlation coefficient factor abbreviation table):
[0069] Table 1
[0070] Abbreviation Original Name Abbreviation Original Name Abbreviation Original Name Abbreviation Original Name C1 awc C11 bio18 C21 cec_clay C31 sand C2 bio1 C12 bio19 C22 cec_eff C32 silt C3 bio10 C13 bio2 C23 cec_soil C33 total_n C4 bio11 C14 bio3 C24 clay C34 road_density C5 bio12 C15 bio4 C25 cn_ratio C35 CACRP C6 bio13 C16 bio5 C26 coarse C36 NOE C7 bio14 C17 bio6 C27 Road C37 PI_GDP C8 bio15 C18 bio7 C28 org_carbon C38 SI_GDP C9 bio16 C19 bio8 C29 ph_water C39 TI_GDP C10 bio17 C20 bio9 C30 pop2020 C40 GDP
[0071] After screening, 23 impact factors were finally obtained, namely Bio3, GDP, Road, Pop2020, Bio15, CACRP, NOE, Sand, TI-GDP, SI-GDP, PI-GDP, Clay, Awc, Silt, pH_Water, Road_Density, Coarse, Cec_Clay, Cec_Eff, DORF, Cn_ratio, Root_Depth, and Add_Prop. Then, a series model was constructed in R language to calculate the contribution of related impact factors, and the top ten contribution rankings were selected for display.
[0072] According to the results of factor contribution, the top four influencing factors are: Bio3, which has the highest contribution of 10.06, is a climate factor (i.e., the influencing factor in climate data); followed by GDP, with a contribution of 9.90, which belongs to the human activity factor (i.e., the influencing factor in human activity factor data); Road and Pop2020, with contributions of 7.73 and 7.15, respectively, both of which belong to human activity factors. Among the soil factors (i.e., the influencing factors in soil data), Sand’s contribution ranks only 8th, at 4.81.
[0073] The specific contribution of influencing factors is shown in Table 2 (Contribution of key influencing factors):
[0074] Table 2
[0075]
[0076]
[0077] S200, obtaining distribution data of geographical indication agricultural products in the target area; based on the influencing factor data and the distribution data, obtaining a target prediction value of the adaptive distribution of the geographical indication agricultural products in the target area through a preset prediction model;
[0078] Among them, the prediction model includes the maximum entropy model and the random forest model in series;
[0079] It should be noted that, in some embodiments, obtaining the distribution data of geographical indication agricultural products in the target area may include the following steps: obtaining the latitude and longitude data of the distribution points of the geographical indication agricultural products in the target area; taking the distribution points as existence points, and randomly generating background points in the area outside the target area based on the latitude and longitude data as non-existence points; and obtaining the distribution data of the geographical indication agricultural products in the target area based on the existence points and non-existence points.
[0080] For example, in some specific implementations, the latitude and longitude data of geographical indication rice agricultural products (GIRAP) can be derived from the Anluyun platform (https: / / www.anluyun.com), and a total of 84 (adjustable according to actual needs) valid distribution points are collected. After the point data is exported in CSV format, it is imported into ArcGIS Desktop 10.8.1 software for spatial visualization and analysis of the Chinese map. Based on the latitude and longitude data of the existence points of geographical indication rice agricultural products, 2000 background points were randomly generated in RStudio and exported as CSV format files, relying on the existence points of geographical indication rice agricultural products and the background points generated in R language as the presence / absence points of species.
[0081] It should also be noted that in some embodiments, based on the influencing factor data and distribution data, the target prediction value of the adaptive distribution of geographical indication agricultural products in the target area is obtained through processing with a preset prediction model, which can include the following steps: spatially superimposing the distribution data and the influencing factor data to construct an environmental data set; inputting the environmental data set into the maximum entropy model to obtain an initial prediction value; adding the initial prediction value to the environmental data set as a new influencing factor, and then inputting it into the random forest model to obtain the target prediction value of the adaptive distribution of geographical indication agricultural products in the target area.
[0082] Exemplarily, in some specific implementations, the MaxEnt model (i.e., maximum entropy model) used is a species distribution model, which relies on the presence points of geographical indication rice agricultural products and the background points generated in the R language as the presence / absence points of species, and combines the three types of influencing factors in step S100 to generate a spatial asc format file, which is then input into the MaxEnt model to predict the distribution of geographical indication rice agricultural products suitability, and then the predicted value is derived as a new influencing factor, and the three types of environmental factors and the presence / absence points of species are input into the random forest model again, and then predicted, and the final predicted value of the distribution of geographical indication rice agricultural products suitability is obtained. This process constructs the MaxEnt-RF series model by integrated series connection, which greatly improves the accuracy of the prediction of geographical indication rice agricultural products suitability distribution.
[0083] In some specific application scenarios, in the present invention, firstly, 2000 background points can be randomly generated in RStudio based on the longitude and latitude data of the geographical indication rice agricultural products, and exported as CSV format files. Subsequently, the longitude and latitude data files are imported into ArcMap 10.8.1 software, and spatially superimposed with multidimensional environmental factors (climate, soil factors), thereby constructing a comprehensive environmental data set with a resolution of 90 meters.
[0084] In the suitability prediction stage, the MaxEnt model can be constructed in RStudio, and the environmental data set is input into the model to generate preliminary prediction results of suitability probability nationwide. Then, the prediction results of the MaxEnt model are introduced into the data set as new environmental factors to enhance the explanatory power of the model. To further improve the prediction accuracy, the present invention adopts the MaxEnt-RF series model, and the prediction results of the MaxEnt model are continuously input into the RF model constructed by RStudio, so as to obtain the final suitability probability prediction results of the whole country, and the contribution of each influencing factor to the suitability distribution is quantitatively analyzed by variable importance assessment. The model is verified by ten-fold cross validation, and the model performance is evaluated by AUC value, Kappa value and TSS value to ensure the accuracy and robustness of the prediction results.
[0085] In some specific application scenarios, in order to conduct suitability analysis and prediction of the maximum entropy model (MaxEnt) and random forest model (RF), the raster data of all environmental factors (climate, soil and human activities) were processed with a unified spatial resolution (90 m) and range, and converted into ASC format in ArcGIS10.8.1 to meet the data format requirements of the MaxEnt model (data format examples are shown in Table 3).
[0086] Table 3
[0087]
[0088]
[0089]
[0090] In some specific application scenarios, the maximum entropy model (MaxEnt) can be implemented as follows:
[0091] In the present invention, MaxEnt is used to identify the potential suitability distribution of GIAP. By combining the presence of GIAP with multiple environmental factors (including climate, soil and human activity factors), MaxEnt can estimate the relative contribution of each environmental variable to the suitability distribution and generate a nationwide suitability probability distribution. The advantage of this model is that it can infer the suitability probability distribution of unobserved areas in an optimal way under data constraints based on limited existing data, ensuring the robustness and scientificity of the prediction results.
[0092] Suppose there is a feature vector x=(x1, x2, ..., x n ), represents the environmental conditions at a specific location. The maximum entropy model aims to estimate the probability of the target distribution p(x) in this feature space. By maximizing the entropy value, MaxEnt ensures that the estimated distribution is unbiased while satisfying the constraints of the observed data. Entropy H(p) is defined as follows:
[0093]
[0094] In the present invention, MaxEnt provides a basis for the prediction of preliminary suitability distribution, enabling the present invention to more accurately identify potential suitable production areas of GIAP and laying an important foundation for subsequent model integration and factor contribution analysis.
[0095] In some specific application scenarios, the Random Forest model (RF) can be implemented as follows:
[0096] In the present invention, RF is used to further improve the prediction accuracy of GIAP suitability distribution. Based on the preliminary suitability probability generated by MaxEnt, the RF model constructs and integrates multiple decision trees to evaluate the contribution of each environmental factor to the GIAP suitability distribution. Through the variable importance evaluation of the RF model, the present invention can quantify the relative impact of factors such as climate, soil and human activities on the suitability distribution, thereby identifying the key factors that have a significant impact on the GIAP distribution. The ensemble learning characteristics of RF ensure the prediction accuracy and robustness of the model, so that the suitability prediction can maintain a high generalization ability in the environment of multidimensional heterogeneous factors.
[0097] The decision formula is as follows:
[0098]
[0099] in, Indicates h i In Category C j Output.
[0100] In the present invention, the RF model not only provides a more accurate prediction of the suitability distribution, but also provides an effective quantitative tool for factor contribution analysis.
[0101] S300. Based on the influencing factor data and target prediction values, the GeoShapley method is used to conduct a correlation analysis between the suitability of geographical indication agricultural products and influencing factors to obtain correlation evaluation data.
[0102] It should be noted that, in some embodiments, step S300 may include the following steps: obtaining SHAP value, Moran's index and spatial lag value according to the influencing factor data and target prediction value; drawing the factor global contribution map and the factor positive and negative contribution map according to the SHAP value, and determining the spatial distribution characteristics and positive and negative contribution of the influencing factor to the suitability of geographical indication agricultural products based on the factor global contribution map and the factor positive and negative contribution map; determining the spatial correlation between the influencing factor and the suitability of geographical indication agricultural products according to the Moran's index and the spatial lag value.
[0103] For example, in some specific implementations, the screened factors and the suitability results calculated by the model are merged into a complete vector shp file in ArcMap, and then the file is imported into Python, and a SHAP value is calculated using the GeoShapley method, and a factor global contribution map and a factor positive and negative contribution map are drawn according to the SHAP value. Based on these maps, we can have a deeper understanding of the spatial distribution characteristics of the factors for the suitability value and their positive and negative contributions, which can better guide the planting of geographical indication rice agricultural products and the formulation of relevant policies.
[0104] In addition, GeoDa software can be used to calculate Moran's I index (i.e. Moran's index) and LISA value (i.e. spatial lag value) to analyze the spatial correlation between suitability and influencing factors and the spatial clustering. At the same time, the spatial contribution of factors is generated based on the SHAP value, which can effectively analyze the positive and negative correlation between suitability and influencing factors and the accumulation in space, providing help for subsequent factor spatial analysis. According to the Moran's I index and LISA value, the spatial correlation between suitability and factors can be clearly understood, providing guidance for the suitability distribution of geographical indication rice agricultural products, and providing suggestions for policy formulation and rice planting.
[0105] In some preferred implementations, obtaining the SHAP value, the Moran index, and the spatial lag value based on the impact factor data and the target prediction value may include the following steps:
[0106] According to the influencing factor data and the target prediction value, the preset first spatial data analysis software is imported to synthesize a vector shp file, and the SHAP value is obtained based on the shp file by using the GeoShapley method;
[0107] For example, in some specific implementations, the application of machine learning and artificial intelligence technology in crop suitability evaluation is becoming more and more extensive. However, when revealing the complex environmental factor relationships and dynamic processes behind the suitability distribution, researchers often face the problem of "model black box". To meet this challenge, explainable artificial intelligence (XAI) has gradually become a research hotspot, and its core goal is to reveal the decision logic of machine learning models, thereby improving the transparency and interpretability of models. In the field of crop suitability evaluation, XAI has been proven to be able to effectively connect and integrate traditional statistical methods and machine learning techniques, especially in scenarios where environmental factors interact with each other in a complex manner and nonlinear relationships are significant. Traditional suitability evaluation methods are limited by linear assumptions, model structure specifications, and subjective weighting, and it is difficult to fully analyze the mechanism of action of multidimensional environmental factors. The machine learning model combined with XAI can not only overcome the limitations of traditional methods, but also finely analyze the positive and negative contributions of key factors, reveal the role of factors in spatial dimensions, and provide a more comprehensive perspective and scientific support for crop suitability distribution research, laying a data foundation for agricultural production layout optimization and policy formulation.
[0108] It should be noted that the GeoShapley method is the technical focus of the present invention. The GeoShapley method applied in the present invention is described in detail from multiple aspects as follows:
[0109] The GeoShapley algorithm is used as an XAI to explain the contribution of key factors in the GIAP suitability distribution prediction model, which can more accurately quantify the contribution of different geographical factors in the model. By calculating the contribution of each input feature (key factor) to the suitability results, GeoShapley helps reveal the positive and negative effects of factors such as climate, soil and human activities on the GIAP suitability distribution. The Shapley value of each feature represents its "fair contribution" to the prediction results, so that the specific role of different factors on the suitability distribution can be quantified, thereby revealing the prediction logic of the "black box" model.
[0110] Classic Shapley value formula:
[0111]
[0112] p is the total number of features, S is the feature subset, and f(S) is the predicted output of subset S.
[0113] GeoShapley value formula:
[0114]
[0115] GEO: a collection of all geographic features; g: the number of geographic features.
[0116] Expected calculation of Shaply Values:
[0117] f x (S) = E[f(X)|do(X s =x s )] (5)
[0118] S is the feature set we want to adjust, X s is a random variable representing the M input features of the model, x s is the model input vector for the current prediction.
[0119] The local accuracy is:
[0120]
[0121] are the feature attributes and f(x) is the original model output.
[0122] The advantage of GeoShapley in this invention is that it can not only analyze in detail the relative contribution of each environmental factor in the suitability model, but also clarify which factors have a positive or negative impact on the suitability of a specific area, thereby providing policy makers with a clear explanation of the factors and a basis for spatial layout. This interpretability enhances the credibility of the model and provides scientific support for the rational distribution and resource optimization of geographical indication agricultural products.
[0123] The influencing factor data and target prediction values are imported into the preset second spatial data analysis software to obtain the Moran index and spatial lag value.
[0124] For example, in some specific embodiments, in the present invention, the local Moran's I index is used to analyze the spatial similarity between the suitability of each region in the GIAP suitability distribution and the influencing factors of its surrounding regions.
[0125] The local Moran's I index is calculated by the following formula:
[0126]
[0127] I i: Local Moran's I value of the ith unit;
[0128] x i 、x j : The attribute values of the i-th and j-th regions;
[0129] Average property value across all regions;
[0130] w ij : The spatial weight between the i-th and j-th regions, usually used to represent the relationship between adjacent regions (adjacent is 1, non-adjacent is 0).
[0131] The local Moran's I value helps identify the relationship between the attribute value of each area and the adjacent areas. Through its value, you can determine:
[0132] Positive correlation (high-high or low-low clustering): When a region and its neighboring regions both exhibit high or low attribute values, the local Moran's I value is positive, indicating a clustering effect.
[0133] Negative correlation (high-low or low-high dispersion): When the attribute value in one area is high and the value in the adjacent area is low (or vice versa), the local Moran's I value is negative, indicating a dispersion effect.
[0134] In the present invention, LISA (Local Indicators of Spatial Association) is used to analyze the local spatial autocorrelation of GIAP suitability distribution. By calculating the LISA values of suitability and influencing factors for each region, it is possible to identify and visualize spatial aggregation and dispersion patterns in suitability distribution, such as "high-high" aggregation areas in high suitability areas and "low-low" aggregation areas in low suitability areas. The analysis results of LISA are usually presented in the form of heat maps, clearly showing the spatial hot spots and cold spots in GIAP suitability distribution.
[0135] The local Moran's I index and LISA value are powerful spatial autocorrelation analysis tools that can be used to understand the relationship between suitability and influencing factors in the data and neighboring areas in space. The local Moran's I index provides specific values to determine the type of autocorrelation, while the LISA value further expands the visualization of local autocorrelation and provides strong support for the analysis of spatial aggregation patterns.
[0136] In order to explain the principle of the technical solution of the present invention in detail, the overall process of the present invention is described below in combination with some specific embodiments. It is easy to understand that the following is an explanation of the technical principle of the present invention and cannot be regarded as a limitation of the present invention.
[0137] Taking China's rice-growing areas as the target area of the study, it covers China's major agricultural areas. China's rice-growing areas are widely distributed in the southeast coast (including Zhejiang, Fujian, Guangdong and Guangxi), the middle and lower reaches of the Yangtze River (including Jiangsu, Anhui, Hubei, Hunan, etc.), the Northeast Plain (including Heilongjiang, Jilin, Liaoning) and the Southwest Mountain Area (including Yunnan, Guizhou and Sichuan, etc.). These regions not only have diverse natural environments, but also show different human activity characteristics, such as economic development level, degree of agricultural mechanization and population density.
[0138] The geographical indication rice varieties in these regions not only reflect the diversity of natural conditions such as climate and soil in various places, but also show the profound impact of human activities on the distribution of agricultural suitability. The degree of mechanization, the advancement of agricultural technology and the change in population density have jointly shaped the pattern of rice production in China.
[0139] Specifically, the data source of the technical solution of the present invention can be realized through the following platforms or databases:
[0140] The latitude and longitude data of Geographical Indication Rice Agricultural Products (GIRAP) were obtained from the Anluyun platform (https: / / www.anluyun.com), and a total of 84 valid distribution points were collected. After the point data were exported in CSV format, they were imported into ArcGIS Desktop 10.8.1 software for spatial visualization and analysis of the China map.
[0141] Climate data can be obtained from the WorldClim 2.1 database (https: / / www.worldclim.org). The BCC-CSM2-MR global climate model (GCM) was selected with a spatial resolution of 2.5 arc minutes (about 5 kilometers), a time range of 2021–2040, and based on the shared socioeconomic pathway scenario SSP126. 19 bioclimatic variables (BIO1–BIO19) were downloaded in GeoTIFF format, and the “Image Analysis” tool of ArcGIS 10.8.1 was used to extract relevant bioclimatic variables and obtain global climate data. In order to focus on the study area, the global climate data was masked using the boundary vector file of the target area (such as a national region or a provincial or municipal region, taking country A as an example), and the climate data of country A was obtained through the “Extract by Mask” function.
[0142] The soil data source can be obtained from the Global Harmonized Soil Database v2.0 (HWSD v2.0) (https: / / gaez.fao.org). Due to the large amount of data, R software version 4.2.0 was used for data processing and extraction. The soil properties of the D1 layer (0–20 cm) were selected as the main research object, including variables such as soil texture, organic matter content, pH value, and nutrient content. According to the FAO soil classification system, the soil of country A is divided into 22 soil types, including ferroalloy soil, eluvial soil, arid soil, and desert soil. The trace element data are from the Guangdong Provincial Soil Environment Monitoring Center.
[0143] The human activity factor data (taking Country A as an example) can be derived from the Statistical Yearbook of Country A and the Statistical Yearbook of County Areas of Country A in 2023. The selected human activity variables include population density, gross domestic product (GDP), etc. The sorted data are imported into ArcGIS 10.8.1 in CSV format and associated with the administrative division vector map of Country A. To meet the model requirements, spatial interpolation methods are used when necessary, and the data are finally converted to raster format through the "Feature to Raster" tool.
[0144] In order to conduct suitability analysis and prediction of the maximum entropy model (MaxEnt) and random forest model (RF), the raster data of all environmental factors (climate, soil and human activities) were processed with a unified spatial resolution (90 m) and range, and converted into ASC format in ArcGIS10.8.1 to meet the data format requirements of the MaxEnt model (see Table 1 above for an example of the data format).
[0145] like Figure 3 and Figure 4 As shown, Figure 3 For details on examples of factor positive and negative contribution graphs, factor global contribution graphs, local Moran's I index, LISA value, and SHAP space graphs, see Figure 6 , Figure 7 , Figure 8 and Fig. 9 , Figure 3 Only vague examples are shown in Figure 3 and Figure 4 The parameters of the evaluation factors in the maximum entropy model and the data graph corresponding to the maximum entropy may change with different actual application scenarios, and the details of the relevant parameters do not affect the logical principle of the technical solution of the present invention, so no specific examples are given for the relevant parameters, and only vague examples are shown; the implementation process of the technical solution of the present invention can be specifically as follows:
[0146] In the present invention, 2000 background points were randomly generated in RStudio based on the longitude and latitude data of the geographical indication rice agricultural products, and exported as CSV format files. Subsequently, the longitude and latitude data files were imported into ArcMap10.8.1 software, and spatially superimposed with multidimensional environmental factors (climate, soil factors), thereby constructing a comprehensive environmental data set with a resolution of 90 meters.
[0147] In the suitability prediction stage, the MaxEnt model was constructed in RStudio, and the environmental data set was input into the model to generate preliminary prediction results of suitability probability across the country. Then, the prediction results of the MaxEnt model were introduced into the data set as new environmental factors to enhance the explanatory power of the model. To further improve the prediction accuracy, the present invention adopts the MaxEnt-RF series model, and the prediction results of the MaxEnt model are continuously input into the RF model constructed by RStudio. Before the factors are input into the RF model, the Pearson coefficient (i.e., Pearson correlation coefficient) is used to screen the factors, retain the key factors of |r|>0.8, and eliminate the variables with low contribution, so as to obtain the final suitability probability prediction results of the whole country, and quantitatively analyze the contribution of each influencing factor to the suitability distribution through variable importance assessment. The model is verified by ten-fold cross validation, and the model performance is evaluated by AUC value, Kappa value and TSS value to ensure the accuracy and robustness of the prediction results.
[0148] In order to further explore the positive and negative contribution of influencing factors to suitability distribution in spatial dimension, the present invention uses Python language to construct GeoShapley method in PyCharm, inputs the main contribution factors calculated by MaxEnt-RF model and suitability distribution results into the model, obtains and visualizes SHAP values, generates global contribution graphs and positive and negative contribution graphs of factors, thereby realizing the refined analysis of the positive and negative contributions of influencing factors in space. In addition, the study introduces local Moran's I index and LISA value, analyzes the spatial autocorrelation between factors and suitability distribution, intuitively displays the spatial correlation, spatial clustering characteristics and spatial contribution of factors, and reveals the aggregation or dispersion patterns in different regions. Through the visualization analysis of the contribution of factors in space, the present invention provides a more in-depth ecological explanation for the spatial characteristics of GIAP suitability distribution. Among them, the maximum entropy model (MaxEnt), random forest model (Random Forest, RF), GeoShapley method, Moran index and LISA value and the logical principle of factor screening used in the technical solution of the present invention refer to the embodiment of the aforementioned specific application scenario, which will not be repeated here.
[0149] It should be noted that in some specific application scenarios, the MaxEnt-RF model effects implemented by the technical solution of the present invention are as follows:
[0150] In the present invention, in order to predict the suitability distribution of geographical indication agricultural products, three different models are used, including MaxEnt, RF, and a series model formed by introducing the suitability probability value output by MaxEnt as a new environmental factor into the RF model. In order to comprehensively evaluate the prediction performance of each model, three indicators, AUC (area under the receiver operating characteristic curve), Kappa coefficient and TSS (true skill statistic), are used for model evaluation.
[0151] like Figure 5 As shown in the figure, by comparing the performance of the three models, in terms of the AUC index, the score of the MaxEnt-RF tandem model is 0.97, which is higher than 0.95 of the MaxEnt model and 0.93 of the RF model; in terms of the Kappa index, the score of the MaxEnt-RF tandem model is 0.32, which is higher than 0.31 of the RF model and 0.09 of the MaxEnt model; in terms of the TSS index, the score of the MaxEnt-RF tandem model is 0.78, which is higher than 0.59 of the RF model and 0.47 of the MaxEnt model.
[0152] It should also be noted that in some specific application scenarios, the spatial analysis of influencing factors can be implemented as follows:
[0153] The factors screened by the Pearson correlation coefficient were used to calculate their positive and negative contributions and SHAP values to suitability in space using the GeoShapley method. We selected the top four contributing factors, Bio3, GDP, Road, and Pop2020, to display their SHAP values.
[0154] From GeoShapley's positive and negative contributions Figure 6 The results show that:
[0155] The SHAP value of railway density shows a wide distribution, with a large number of distributions at both the positive and negative ends, and the high values are concentrated in the positive direction, indicating that areas with higher railway density have a positive impact on the suitability distribution. Bio3 is mainly distributed in the positive direction of the SHAP value, indicating that Bio3 has a positive effect on the suitability distribution when it is high. The SHAP values of GDP are mostly concentrated in the negative direction, indicating that higher GDP has a negative impact on the suitability distribution. The distribution range of the SHAP value of Pop is small, and there are relatively scattered positive and negative distributions at both ends, indicating that the impact of population density on the suitability distribution is small and relatively neutral.
[0156] From the global contribution Figure 7The results show that Road is the factor with the greatest impact on the suitability distribution, followed by Bio3 and GDP, and Pop has the least impact on the suitability distribution. This shows that in our model, the change in railway density contributes most significantly to the suitability distribution of GI rice.
[0157] according to Figure 8 The spatial characteristic maps of the four SHAP values can be used to analyze the spatial impact of population density, railway density, gross domestic product and climate factor BIO3 on the suitability distribution of geographical indication rice agricultural products.
[0158] The impact of GDP on suitability distribution is mainly concentrated in a few areas, mainly showing negative contributions. The negative areas are relatively extensive, indicating that areas with higher GDP have an inhibitory effect on suitability distribution. The impact of railway density on suitability distribution has significant positive and negative contributions in specific areas, showing a certain spatial heterogeneity. Areas with higher positive contributions are concentrated in the southeast and parts of the central region, while areas with low railway density are mostly negative contributions. The impact of Bio3 is mainly concentrated in the high-latitude areas in the north, and shows a significant positive contribution, indicating that the climatic conditions in these areas have a more positive effect on suitability distribution. The SHAP value distribution of population density is relatively scattered, with a small overall impact and low positive and negative contribution values, indicating that the impact of population density on suitability distribution is weak.
[0159] The spatial correlation between the suitability distribution of geographical indication rice agricultural products and the main influencing factors was analyzed based on the local bivariate Moran's I index and LISA value:
[0160] High-high concentration areas of GDP mainly appear in the southeastern coastal areas, such as Zhejiang, Jiangsu and Shanghai, showing the spatial aggregation of high GDP areas. Low-low concentration areas are distributed in the inland and western regions. High-high concentration areas of railway density are concentrated in the Yangtze River Delta, especially in Jiangsu and Zhejiang, and low-low concentration areas are located in Yunnan and Guizhou in the southwest, indicating that railway density has a significant spatial aggregation pattern in these areas. High-high concentration areas of Bio3 are distributed in the high latitudes of the north, and low-low concentration areas mainly appear in the southern coastal areas, reflecting the spatial differences in climatic conditions. High-high concentration areas of population density are concentrated in the middle and lower reaches of the Yangtze River, and low-low concentration areas are located in the sparsely populated western region.
[0161] from Fig. 9 It can be seen that:
[0162] The Moran's I of GDP is 0.104, indicating that the spatial autocorrelation of GDP is weak, but it still shows a certain positive correlation trend. The Moran's I of railway density is 0.431, showing a strong positive spatial autocorrelation, indicating that railway density has a strong aggregation effect in space. The Moran's I of Bio3 is -0.485, indicating that Bio3 shows obvious negative autocorrelation in space, indicating that high and low value areas are dispersed in space. The Moran's I of population density is 0.034, showing a weak positive spatial autocorrelation, indicating that the spatial aggregation of population density is not obvious.
[0163] In summary, the present invention proposes and verifies an improved MaxEnt-RF tandem model for predicting the suitability distribution of geographical indication agricultural products, and systematically compares the effects with the traditional single MaxEnt and random forest (RF) models. The results show that the MaxEnt-RF tandem model has significant advantages in key evaluation indicators such as AUC, Kappa and TSS, showing a significant improvement in the prediction accuracy and model stability of the model, thereby effectively solving the problem of insufficient accuracy in predicting suitability under complex influencing factors.
[0164] Specifically, in terms of the AUC index, the MaxEnt-RF tandem model scored the highest, indicating that it has a strong ability to discriminate in predicting the distribution of geographical indication agricultural products suitability, and can more effectively distinguish between suitable and unsuitable areas; in terms of the Kappa coefficient, the MaxEnt-RF tandem model and the RF model scored similarly, both significantly higher than the MaxEnt model, showing higher consistency, indicating that the model can still maintain a high degree of stability under variable environmental conditions. In addition, in terms of the TSS index, the MaxEnt-RF tandem model performed most outstandingly, significantly better than other models, and this result further demonstrated its ability to reduce false positive and false negative predictions.
[0165] By introducing the suitability probability of the MaxEnt model as a new environmental factor into the RF model, the MaxEnt-RF tandem model can more comprehensively integrate the complex relationships between different environmental factors, greatly enhancing the explanatory and predictive power of the model. This model not only captures the impact of a single factor on the suitability distribution, but also effectively combines the nonlinear feature extraction capability of the MaxEnt model with the powerful classification capability of the RF model in an integrated manner, making the prediction results more accurate and reliable.
[0166] In conclusion, the development and application of the MaxEnt-RF tandem model provides an effective method for the accurate prediction of the distribution of geographical indication agricultural products suitability at a high resolution of 90 m. The success of this model shows that integrating the advantages of different models can not only effectively improve the prediction accuracy, but also improve the interpretability of the model to a certain extent. Future research can continue to optimize the model to adapt to more complex environmental factors or larger geographical areas, providing a more scientific basis for policy makers and regional managers in the rational planning of the production layout of geographical indication agricultural products.
[0167] This paper introduces a large number of human activity factors into the suitability distribution model, and uses them together with the climate factors and soil factors commonly used in traditional research as key influencing factors. This innovative approach makes up for the lack of previous studies on the impact of human activities on suitability distribution, and helps to more comprehensively reveal the mechanism of human activities in suitability distribution. At the same time, the GeoShapley, Moran's I index and LISA value are innovatively combined to systematically analyze the positive and negative contributions of each influencing factor to suitability distribution in the spatial dimension.
[0168] GeoShapley is a method that extends TreeSHAP and combines the theory of Shapley Value to provide explanations for geospatial data. Compared with existing methods, GeoShapley can better reflect the differences in features in different geographical regions by combining spatial weighting; GeoShapley is based on Shapley value theory and can effectively decompose the contribution of feature interactions to prediction; traditional partial correlation analysis relies on the linear correlation assumption, while GeoShapley does not require this assumption and is suitable for nonlinear complex models; partial correlation analysis may lead to the underestimation or omission of some driving factors due to multicollinearity and significance threshold (P value), while GeoShapley can more robustly capture the overall and local contribution of features to prediction without relying on significance tests; GeoShapley's results have stronger spatial consistency, avoiding the subjectivity of the window moving method.
[0169] The factor analysis results obtained by the above method show that railway density has the highest contribution to the distribution of geographical indication rice agricultural product suitability, indicating that transportation infrastructure plays a key role in optimizing rice planting areas. Specifically, areas with high railway density provide convenient conditions for rice production and transportation, thereby effectively promoting the circulation and supply chain efficiency of agricultural products, especially in areas with high demand, where the positive impact of railway density on the distribution of rice suitability is particularly significant. The Moran's I index results show that the railway density factor has a strong positive spatial autocorrelation. This indicates that there are high-high clustering areas in the spatial distribution of railway density, that is, areas with denser railways tend to cluster with each other. This clustering shows that areas with convenient transportation are more suitable for the cultivation of geographical indication rice, and the advantages of transportation conditions can effectively improve the production and circulation efficiency of agricultural products. The LISA map further shows that in the Yangtze River Delta and some eastern coastal areas, railway-dense areas and areas with high rice suitability show a significant spatial clustering relationship.
[0170] Secondly, Bio3 has a high contribution, showing its importance in the suitability distribution. Bio3 represents the seasonal stability of temperature, and its high positive contribution indicates that the stability of climate has a positive effect on the suitability of rice cultivation. Especially in areas with small seasonal temperature differences and relatively stable climate, high values of Bio3 mean that the growing conditions of rice are relatively stable, which is conducive to improving the suitability distribution. This is consistent with the important role of climate factors in rice growth, showing the significant impact of temperature changes on the distribution of geographical indication rice. The Moran's I index and GeoShapley's spatial characteristic map show that Bio3 presents negative spatial autocorrelation, especially in high-latitude areas. This shows that the suitability of climate factors is unstable and uncertain in spatial distribution, and the seasonal fluctuations of climate in high-latitude areas are large, resulting in spatial discontinuity in suitability distribution. On the LISA map, high-latitude areas show obvious low-low clustering, which may be due to the unstable climate conditions in these areas, which have a negative impact on the suitability of rice. This climate instability leads to large differences in the distribution of rice suitability in these areas.
[0171] In contrast, the contribution of GDP and population density is relatively low, and high GDP regions show a negative impact on the suitability distribution. This may be because the land use in economically developed regions is mainly concentrated in urban expansion, industrial development, etc., which leads to the squeeze of agricultural land and is not conducive to the development of traditional rice cultivation. Therefore, although the regions with high GDP have active economies and rich resources, their land use patterns may not be consistent with the needs of agricultural production, which has an inhibitory effect on the suitability distribution of GI rice. Similarly, population density has a small contribution to the suitability distribution, which may be because rice cultivation mainly depends on natural conditions and transportation convenience, and areas with high population density do not necessarily have the best agricultural production environment. The Moran's I index value of the GDP factor is low, indicating that its spatial correlation in the suitability distribution of GI rice is weak. In the LISA map, high GDP regions do not show a significant spatial aggregation relationship with regions with high rice suitability. This may be because high GDP regions are mainly concentrated in cities and industrial areas, and their land use patterns are not conducive to the development of suitability of traditional agriculture, so their impact on rice suitability shows great heterogeneity in space. The Moran's I index value of population density is also low, indicating that its spatial correlation in the distribution of rice suitability is weak. In the LISA map, densely populated areas did not form a significant spatial aggregation pattern, indicating that the impact of population density on the distribution of rice suitability is relatively scattered and there is no obvious spatial consistency. This may be because the distribution of rice suitability mainly depends on natural conditions and transportation factors, rather than simply population distribution.
[0172] In summary, the present invention establishes the MaxEnt-RF model. Compared with its single model, the tandem model of the present invention has higher prediction accuracy and can apply complex environmental factors to complete the prediction at high spatial resolution. At the same time, the role and spatial characteristics of different environmental factors in the distribution of rice suitability are revealed. Railway density and climate stability are the main factors affecting the distribution of rice suitability, while the influence of economic and demographic factors shows large spatial heterogeneity. The research results provide a scientific basis for the future rational planning of the production layout of geographical indication rice agricultural products, and provide a reference for policy makers in transportation and land use management. In addition, the present invention uses the GeoShapley method to quantitatively evaluate the contribution of each factor in spatial distribution. The contribution of each factor to crop suitability is calculated, while considering the interaction and spatial heterogeneity between factors. And based on the calculation results, the global contribution distribution map of the factor and the positive and negative contribution analysis map of the factor are drawn. According to these figures, the spatial distribution characteristics of the factors for the suitability value and their positive and negative contributions can be more deeply understood, which can better guide the planting of geographical indication rice agricultural products and the formulation of relevant policies.
[0173] Specifically, the present invention constructs the MaxEnt-RF model to predict the suitability distribution of geographical indication rice agricultural products across the country. The model inputs climate factors, soil factors and human factors, and combines GeoShapley, Moran's I index and LISA values to deeply reveal the impact of human activities on suitability distribution, as well as the positive and negative contributions of each key factor to suitability distribution in space. The spatial positive and negative contribution analysis of factors is of great significance for optimizing the planting layout of geographical indication rice. At the same time, our research provides a complete research process system for establishing the suitability distribution and planting pattern of geographical indication agricultural products, clarifies the specific analysis steps, enriches our understanding of the influencing factors of geographical indication agricultural products, and proposes new research ideas and scientific hypotheses for in-depth exploration of the factor influence mechanism in the future. The research results show that the MaxEnt-RF model effectively improves the accuracy of the prediction, and the spatial analysis methods such as GeoShapley clearly show the positive and negative contributions of each influencing factor, which provides strong support for the realization of precise and intelligent planting of geographical indication agricultural products. The present invention provides a promising reference method for the planting planning and government decision-making of geographical indication agricultural products in the study area and other regions.
[0174] like Fig.10 As shown, the embodiment of the invention further provides a geographical indication agricultural product suitability evaluation device 900, which may include:
[0175] The first module 901 is used to obtain the impact factor data of the target area; the impact factor data includes multiple impact factors;
[0176] The second module 902 is used to obtain the distribution data of the geographical indication agricultural products in the target area; based on the influencing factor data and the distribution data, the target prediction value of the adaptive distribution of the geographical indication agricultural products in the target area is obtained through a preset prediction model; the prediction model includes a maximum entropy model and a random forest model connected in series;
[0177] The third module 903 is used to perform a correlation analysis between the suitability of geographical indication agricultural products and the influencing factors based on the influencing factor data and the target prediction value using the GeoShapley method to obtain correlation evaluation data.
[0178] The contents of the method embodiments of the present invention are all applicable to the device embodiments. The functions specifically implemented by the device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0179] The embodiment of the present invention further provides an electronic device, the electronic device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the above-mentioned geographical indication agricultural product suitability evaluation method when executing the computer program. The electronic device can be any intelligent terminal including a tablet computer, a car computer, etc.
[0180] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0181] See also Fig.11 , Fig.11 The hardware structure of an electronic device 1000 of another embodiment is illustrated. The electronic device 1000 includes:
[0182] The processor 1001 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present invention.
[0183] The memory 1002 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1002 can store an operating system and other application programs. When the technical solution provided in the embodiments of this specification is implemented by software or firmware, the relevant program code is stored in the memory 1002, and the processor 1001 calls and executes the geographical indication agricultural product suitability evaluation method of the embodiment of the present invention;
[0184] Input / output interface 1003, used to implement information input and output;
[0185] The communication interface 1004 is used to realize the communication interaction between the device and other devices. The communication can be realized through a wired manner (such as USB, network cable, etc.) or a wireless manner (such as mobile network, WIFI, Bluetooth, etc.);
[0186] A bus 1005 , which transmits information between various components of the device (e.g., the processor 1001 , the memory 1002 , the input / output interface 1003 , and the communication interface 1004 );
[0187] The processor 1001 , the memory 1002 , the input / output interface 1003 and the communication interface 1004 are connected to each other in communication within the device via the bus 1005 .
[0188] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned geographical indication agricultural product suitability evaluation method.
[0189] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiments, the functions specifically implemented by the present storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0190] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0191] The embodiment of the present invention provides a geographical indication agricultural product suitability evaluation method, a geographical indication agricultural product suitability evaluation device, an electronic device and a storage medium, which obtains the influencing factor data of the target area; the influencing factor data includes multiple influencing factors; the distribution data of the geographical indication agricultural product in the target area is obtained; based on the influencing factor data and the distribution data, the target prediction value of the adaptive distribution of the geographical indication agricultural product in the target area is obtained by processing with a preset prediction model; the prediction model includes a maximum entropy model and a random forest model in series; based on the influencing factor data and the target prediction value, the GeoShapley method is used to perform the correlation analysis between the suitability of the geographical indication agricultural product and the influencing factor to obtain the correlation evaluation data. The present invention constructs a new series model as a model for predicting the suitability distribution of GIAP, and integrates the advantages of multiple models to improve the accuracy of predicting the suitability distribution of GIAP. The present invention significantly improves the accuracy of the prediction of the suitability distribution, and reveals in detail the specific role of each influencing factor in space, which can provide a scientific basis for the reasonable production layout of geographical indication agricultural products (GIAP), and provides guidance for policy formulation and regional agricultural management.
[0192] The embodiments described in the embodiments of the present invention are intended to more clearly illustrate the technical solutions of the embodiments of the present invention, and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. Those skilled in the art can appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present invention are also applicable to similar technical problems.
[0193] Those skilled in the art will appreciate that the technical solutions shown in the figures do not limit the embodiments of the present invention and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0194] The system embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment.
[0195] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.
[0196] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0197] It should be understood that in the present invention, "at least one (item)" refers to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can represent: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0198] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of systems or units, which can be electrical, mechanical or other forms.
[0199] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0200] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0201] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including multiple instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods of various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store programs.
[0202] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the embodiments of the present invention is not limited thereby. Any modification, equivalent substitution and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present invention shall be within the scope of the rights of the embodiments of the present invention.
Claims
1. A method for evaluating the suitability of geographically indicated agricultural products, characterized in that: The method comprises the following steps: Acquire impact factor data of a target area; the impact factor data includes a plurality of impact factors; Obtaining the distribution data of the geographical indication agricultural products in the target area; based on the influencing factor data and the distribution data, obtaining the target prediction value of the adaptive distribution of the geographical indication agricultural products in the target area through a preset prediction model; the prediction model includes a maximum entropy model and a random forest model connected in series; Based on the influencing factor data and the target prediction value, the GeoShapley method is used to perform a correlation analysis between the suitability of the geographical indication agricultural products and the influencing factors to obtain correlation evaluation data.
2. The method for evaluating suitability of geographically indicated agricultural products according to claim 1, characterized in that: The influencing factor data includes climate data, soil data and human activity factor data, and the climate data, soil data and human activity factor data all include a number of the influencing factors; the step of obtaining the influencing factor data of the target area includes the following steps: Acquire global climate data from a preset climate database, perform mask extraction on the global climate data based on a boundary vector file of the target area, and obtain the climate data of the target area; Acquiring soil properties and trace element data of the target area from a preset soil database as the soil data; The human activity variables of the target area are obtained from preset statistical data, the human activity variables are associated with the vector map of the target area, and the data are converted into a raster format to obtain the human activity factor data.
3. The method for evaluating suitability of geographically indicated agricultural products according to claim 1, characterized in that: The step of obtaining the distribution data of the geographical indication agricultural products in the target area comprises the following steps: Obtaining the longitude and latitude data of the distribution points of the geographical indication agricultural products in the target area; Taking the distribution points as existence points, randomly generating background points in an area outside the target area based on the latitude and longitude data as non-existence points; The distribution data of the geographical indication agricultural products in the target area are obtained based on the existence points and the non-existence points.
4. The method for evaluating suitability of geographically indicated agricultural products according to claim 1, characterized in that: The method of obtaining the target predicted value of the adaptive distribution of the geographical indication agricultural product in the target area based on the influencing factor data and the distribution data by processing with a preset prediction model comprises the following steps: Spatially superimposing the distribution data and the influencing factor data to construct an environmental data set; Inputting the environmental data set into the maximum entropy model for processing to obtain an initial prediction value; The initial prediction value is added to the environmental data set as a new influencing factor, and then input into the random forest model for processing to obtain the target prediction value of the adaptive distribution of the geographical indication agricultural products in the target area.
5. The method for evaluating suitability of geographically indicated agricultural products according to claim 1, characterized in that: The method further comprises the following steps: Based on the Pearson correlation coefficient, the factors in the influencing factor data are screened by combining the correlation and the factor contribution with a preset screening threshold, and the influencing factor data is updated based on the result of the factor screening.
6. The method for evaluating suitability of geographically indicated agricultural products according to any one of claims 1 to 5, characterized in that: Based on the influencing factor data and the target prediction value, the GeoShapley method is used to perform correlation analysis between the suitability of the geographical indication agricultural product and the influencing factor to obtain correlation evaluation data, including the following steps: Obtaining a SHAP value, a Moran's index and a spatial lag value according to the impact factor data and the target prediction value; A factor global contribution map and a factor positive and negative contribution map are obtained according to the SHAP value, and the spatial distribution characteristics and positive and negative contribution of the influencing factors to the suitability of the geographical indication agricultural products are determined in turn based on the factor global contribution map and the factor positive and negative contribution map; The spatial correlation between the influencing factor and the suitability of the geographical indication agricultural product is determined based on the Moran's index and the spatial lag value.
7. The method for evaluating suitability of geographically indicated agricultural products according to claim 6, characterized in that: The step of obtaining the SHAP value, the Moran's index and the spatial lag value according to the impact factor data and the target prediction value comprises the following steps: According to the influencing factor data and the target prediction value, a preset first spatial data analysis software is imported to synthesize a vector shp file, and the SHAP value is obtained by processing the shp file using the GeoShapley method; The influencing factor data and the target prediction value are imported into a preset second spatial data analysis software for processing to obtain the Moran index and the spatial lag value.
8. A device for evaluating the suitability of geographically indicated agricultural products, characterized in that: The device comprises: The first module is used to obtain the impact factor data of the target area; the impact factor data includes multiple impact factors; The second module is used to obtain the distribution data of the geographical indication agricultural products in the target area; based on the influencing factor data and the distribution data, the target prediction value of the adaptive distribution of the geographical indication agricultural products in the target area is obtained through a preset prediction model; the prediction model includes a maximum entropy model and a random forest model connected in series; The third module is used to perform a correlation analysis between the suitability of the geographical indication agricultural products and the influencing factors based on the influencing factor data and the target prediction value using the GeoShapley method to obtain correlation evaluation data.
9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Urban development land suitability evaluation method based on principal component analysis
CN110472882A
Forest fire risk assessment method based on Maxent and GIS
CN112712275A
Method and system for evaluating nonlinear influence of ozone on productivity of farmland ecosystem
CN117726474A
GIS-based agricultural land suitability evaluation and analysis method
CN118037475A
Ecological source place identification method based on habitat suitability, risk and quality framework
CN118313985A
Cited By
Soybean planting suitable area evaluation method fused with yield prediction
CN120355272A
A soybean planting suitable area assessment method integrated with yield prediction
CN120355272B
Diversified rice planting method in composite environment based on big data analysis
CN120807195A
A diversified rice planting method in a composite environment based on big data analysis
CN120807195B
Digital agricultural comprehensive monitoring management system and method based on GIS (Geographic Information System)
CN121526521A