A method, device, equipment and medium for evaluating the suitability of geographically identified agricultural products
By combining the tandem maximum entropy model and the random forest model with the GeoShapley method, the problem of uneven data distribution in the prediction of the suitability distribution of geographical indication agricultural products was solved, achieving high-precision suitability analysis, providing scientific production and policy guidance, and improving the scientific nature and sustainability of agricultural management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGDONG INST OF ECO ENVIRONMENT & SOIL SCI
- Filing Date
- 2024-12-13
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies suffer from uneven data availability and quality in predicting the suitability distribution of geographical indication agricultural products, making it difficult to achieve evaluation standardization, affecting production efficiency and quality, and thus threatening the sustainable development of GIAP.
By combining the cascaded maximum entropy model and random forest model with the GeoShapley method, and by acquiring data on multiple influencing factors, we can predict the adaptive distribution and conduct correlation analysis of geographical indication agricultural products. We can also use SHAP values, Moran's index, and spatial lag values to reveal the contribution and spatial distribution characteristics of the influencing factors.
It significantly improves the accuracy of predicting the suitability distribution of geographical indication agricultural products, reveals in detail the spatial effects of various influencing factors, provides a scientific basis for rational production layout and policy formulation, and enhances the scientific nature and sustainability of agricultural management.
Smart Images

Figure CN119940953B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method, apparatus, equipment and medium for evaluating the suitability of geographical indication agricultural products. Background Technology
[0002] Geographical Indications of Agricultural Products (GIAPs) are distinctive agricultural products formed due to the unique natural conditions, historical culture, or traditional techniques of a specific geographical region. They possess significant brand value and market competitiveness, and have become important tools for promoting regional economies, protecting biodiversity, and fostering sustainable agriculture. The suitability distribution of GIAPs directly affects their production efficiency and quality. However, with the increasing environmental pressures brought about by climate change, soil degradation, and population density (Nuary et al., 2022), accurately predicting the suitability distribution of GIAPs and clarifying the relative contributions of influencing factors to suitability has become an urgent problem to be solved. These challenges may threaten the quantity and quality of GIAP production, thereby affecting the achievement of sustainable development goals and the implementation of relevant national policies. Therefore, establishing a methodology capable of accurately predicting the suitability distribution of GIAPs and analyzing the contributions of influencing factors is of great significance for achieving food security and promoting regional sustainable development.
[0003] Despite significant progress in suitability distribution modeling and the important references provided for suitability research, standardization of suitability assessments in different regions remains a challenge, especially given the uneven availability and quality of data. Summary of the Invention
[0004] The main objective of this invention is to provide a method, apparatus, equipment, and medium for evaluating the suitability of geographical indication agricultural products, in order to solve at least one problem in the prior art. This invention can accurately evaluate the suitability of geographical indication agricultural products.
[0005] To achieve the above objectives, one aspect of this invention proposes a method for evaluating the suitability of geographically indicated agricultural products, the method comprising:
[0006] Obtain impact factor data for the target region; the impact factor data includes multiple impact factors;
[0007] Obtain distribution data of geographical indication agricultural products in the target area; based on the influencing factor data and distribution data, obtain the target predicted value of the adaptive distribution of geographical indication agricultural products in the target area through a preset prediction model; the prediction model includes a cascaded maximum entropy model and a random forest model;
[0008] Based on the impact factor data and target predicted values, the GeoShapley method was used to conduct a correlation analysis between the suitability of geographical indication agricultural products and the impact factors, and the correlation evaluation data was obtained.
[0009] In some embodiments, the impact factor data includes climate data, soil data, and human activity factor data, each of which includes several impact factors; obtaining impact factor data for a target area includes the following steps:
[0010] Global climate data is obtained from a pre-set climate database. Based on the boundary vector file of the target region, the global climate data is masked and extracted to obtain the climate data of the target region.
[0011] Soil properties and trace element data of the target area are obtained from a pre-set soil database as soil data;
[0012] Human activity variables for the target area are obtained from pre-defined statistical data. These variables are then linked to a vector map of the target area, and the data is converted into a raster format to obtain human activity factor data.
[0013] In some embodiments, obtaining distribution data of geographical indication agricultural products in a target area includes the following steps:
[0014] Obtain the latitude and longitude data of the distribution points of geographical indication agricultural products in the target area;
[0015] The distribution points are taken as existing points, and background points are randomly generated in areas outside the target area based on latitude and longitude data as non-existent points;
[0016] The distribution data of geographical indication agricultural products in the target area are obtained by organizing the data based on the existence and non-existence of the points.
[0017] In some embodiments, based on influencing factor data and distribution data, a pre-defined prediction model is used to obtain the target predicted value of the adaptive distribution of geographical indication agricultural products in the target area, including the following steps:
[0018] By spatially overlaying distribution data and influencing factor data, an environmental dataset is constructed.
[0019] The environmental dataset is input into the maximum entropy model to obtain initial predicted values;
[0020] The initial predicted values are added to the environmental dataset as new influencing factors, and then input into the random forest model to obtain the target predicted values of the adaptive distribution of geographical indication agricultural products in the target region.
[0021] In some embodiments, the method further includes the following steps:
[0022] Based on the Pearson correlation coefficient, factor screening is performed on the impact factor data by combining correlation and factor contribution with a preset screening threshold, and the impact factor data is updated based on the results of factor screening.
[0023] In some embodiments, based on impact factor data and target predicted values, the GeoShapley method is used to conduct a correlation analysis between the suitability of geographical indication agricultural products and impact factors to obtain correlation evaluation data, including the following steps:
[0024] The SHAP value, Moran index, and spatial lag value are obtained by processing the impact factor data and target predicted values.
[0025] The global contribution map and positive and negative contribution map of the factors were drawn based on the SHAP values. Based on the global contribution map and the positive and negative contribution map of the factors, the spatial distribution characteristics and positive and negative contribution of the influencing factors on the suitability of geographical indication agricultural products were determined in turn.
[0026] The spatial correlation between influencing factors and the suitability of geographical indication agricultural products was determined based on the Moran index and spatial lag value.
[0027] In some embodiments, the SHAP value, Moran index, and spatial lag value are obtained by processing the impact factor data and target predicted value, including the following steps:
[0028] The impact factor data and target predicted values are imported into a preset first spatial data analysis software to synthesize a vector shape file. Based on the shape file, the GeoShapley method is used to process the SHAP value.
[0029] The impact factor data and target predicted values are imported into a pre-set second spatial data analysis software to obtain the Moran index and spatial lag value.
[0030] To achieve the above objectives, another aspect of the present invention provides a geographical indication agricultural product suitability evaluation device, the device comprising:
[0031] The first module is used to obtain impact factor data for the target region; the impact factor data includes multiple impact factors.
[0032] The second module is used to obtain the distribution data of geographical indication agricultural products in the target area; based on the influencing factor data and distribution data, the target predicted value of the adaptive distribution of geographical indication agricultural products in the target area is obtained through the pre-set prediction model; the prediction model includes the cascaded maximum entropy model and the random forest model;
[0033] The third module is used to conduct a correlation analysis between the suitability of geographical indication agricultural products and influencing factors based on influencing factor data and target predicted values, using the GeoShapley method to obtain correlation evaluation data.
[0034] To achieve the above objectives, another aspect of the present invention provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method.
[0035] To achieve the above objectives, another aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0036] The embodiments of this invention include at least the following beneficial effects: This invention provides a method, apparatus, equipment, and medium for evaluating the suitability of geographical indication agricultural products (GIAPs). This scheme involves acquiring influencing factor data for a target region; the influencing factor data includes multiple influencing factors; acquiring distribution data of GIAPs in the target region; based on the influencing factor data and distribution data, processing the data using a pre-set prediction model to obtain target predicted values for the adaptive distribution of GIAPs in the target region; the prediction model includes a cascaded maximum entropy model and a random forest model; based on the influencing factor data and target predicted values, using the GeoShapley method to perform correlation analysis between the suitability of GIAPs and influencing factors, obtaining correlation evaluation data. This invention constructs a new cascaded model as a model for predicting the suitability distribution of GIAPs, integrating the advantages of multiple models to improve the accuracy of predicting the suitability distribution of GIAPs. This invention significantly improves the accuracy of suitability distribution prediction, reveals in detail the specific roles of each influencing factor in space, provides a scientific basis for the rational production layout of GIAPs, and provides guidance for policy formulation and regional agricultural management. Attached Figure Description
[0037] Figure 1 This is a flowchart of the method for evaluating the suitability of geographically identified agricultural products provided in this embodiment of the invention;
[0038] Figure 2 This is a schematic diagram illustrating an example of the correlation of the Pearson correlation coefficient factor provided in an embodiment of the present invention;
[0039] Figure 3 This is a flowchart of the overall process for evaluating the suitability of geographically identified agricultural products provided in this embodiment of the invention.
[0040] Figure 4 This is a flowchart illustrating the data logic principle of the geographical indication agricultural product suitability evaluation method provided in this embodiment of the invention.
[0041] Figure 5 This is a schematic diagram comparing the model performance indicators provided in the embodiments of the present invention;
[0042] Figure 6 This is a schematic diagram illustrating the positive and negative contribution of the SHAP value provided in an embodiment of the present invention;
[0043] Figure 7 This is a schematic diagram illustrating the global contribution of the SHAP value provided in an embodiment of the present invention;
[0044] Figure 8 This is a schematic diagram of the SHAP value space features provided in an embodiment of the present invention;
[0045] Figure 9 This is a schematic diagram of Moran's I index for suitability and environmental factors provided in an embodiment of the present invention;
[0046] Figure 10 This is a schematic diagram of the structure of the geographical indication agricultural product suitability evaluation device provided in an embodiment of the present invention;
[0047] Figure 11 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation
[0048] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of this invention; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this invention as detailed in the appended claims.
[0049] It is understood that the terms "first," "second," etc., used in this invention may be used to describe various concepts, but unless specifically stated otherwise, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of embodiments of this invention, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words "if" or "when" as used herein may be interpreted as "when," "in response to determination," or "in the event of a determination."
[0050] The terms “at least one,” “multiple,” “each,” “any,” etc., used in this invention, “at least one” includes one, two, or more than two; “multiple” includes two or more than two; “each” refers to each of the corresponding multiple; and “any” refers to any one of the multiple.
[0051] Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used in this invention is for descriptive purposes only and is not intended to limit the invention.
[0052] To facilitate understanding of the technical solution of this invention, the technical means and features that may appear in the technical solution of this invention are explained below:
[0053] The SHAP (Shapley Additive exPlanation) value is a method for interpreting the predictions of machine learning models, based on the Shapley value theory in cooperative game theory. By evaluating the contribution of each feature to the model's predictions, the SHAP value helps to understand the model's behavior and decision-making process.
[0054] GeoShapley is a novel method based on the Shapley value concept but specifically designed for geospatial analysis, used to measure spatial effects in machine learning models. It extends the Shapley value framework, focusing on geographic neighborhoods and incorporating geographic location information as a crucial analytical factor. It considers the interactions between spatial variables and the impact of spatial distribution characteristics on the target variable; it emphasizes the total contribution of variables to the target within a geographic region, combining Shapley value assignment rules with techniques such as Geographically Weighted Regression (GWR). It focuses on geospatial data, with input data typically including latitude and longitude or geographic regional divisions. It is particularly suitable for analyzing variables related to spatial location (such as land use, climate conditions, and pollution source distribution). It captures the interactions and autocorrelation of spatial variables, providing an interpretation of spatial dimensions, making it suitable for geographic information science and environmental science. Compared to the TreeSHAP method, GeoShapley expands the application scope of Shapley values, focusing on geospatial data analysis and emphasizing spatial correlation, while TreeSHAP is a machine learning model interpretation tool that focuses on the impact of features on prediction results.
[0055] The method for evaluating the suitability of geographical indication agricultural products provided in this invention relates to the field of data processing technology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or vehicle-mounted terminal, but is not limited to these. The server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network. The software can be an application implementing the method for evaluating the suitability of geographical indication agricultural products, but is not limited to the above forms.
[0056] This invention can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0057] Figure 1 This is an optional flowchart of the geographical indication agricultural product suitability evaluation method provided in the embodiments of the present invention. Figure 1 The method may include, but is not limited to, steps S100 to S300.
[0058] S100. Obtain the impact factor data for the target region;
[0059] The impact factor data includes multiple impact factors;
[0060] It should be noted that the influencing factor data includes climate data, soil data, and human activity factor data, and each of the climate data, soil data, and human activity factor data includes several influencing factors. In some embodiments, step S100 may include the following steps: obtaining global climate data from a preset climate database, performing mask extraction on the global climate data based on the boundary vector file of the target area to obtain the climate data of the target area; obtaining soil properties and trace element data of the target area from a preset soil database as soil data; obtaining human activity variables of the target area from preset statistical data, associating the human activity variables with the vector map of the target area, and then converting the data into a raster format to obtain human activity factor data.
[0061] For example, in some specific implementations, climate data can be acquired from the WorldClim 2.1 database (https: / / www.worldclim.org), using the BCC-CSM2-MR global climate model (GCM) with a spatial resolution of 2.5 arcminutes (approximately 5 kilometers), a time range of 2021–2040, and based on the shared socioeconomic path scenario SSP126. Nineteen bioclimatic variables (BIO1–BIO19) are downloaded in GeoTIFF format, and relevant bioclimatic variables are extracted using the "Image Analysis" tool in ArcGIS 10.8.1 to obtain global climate data. To focus on the study area, the global climate data is masked using the boundary vector file of the target area (such as a country or province / city, taking country A as an example), and the climate data within country A is obtained using the "Extract by Mask" function.
[0062] Soil data was obtained from the Harmonized Global Soil Database version 2.0 (HWSD v2.0) (https: / / gaez.fao.org). Due to the large amount of data, R software version 4.2.0 was used for data processing and extraction. Soil properties in layer D1 (0–20 cm) were selected as the main research object, including variables such as soil texture, organic matter content, pH value, and nutrient content. According to the FAO soil classification system, the soils of country A were classified into 22 soil order types, including ferruginous soils, leached soils, arid soils, and desert soils. Trace element data were obtained from the Guangdong Provincial Soil Environmental Monitoring Center.
[0063] Human activity factor data (taking country A as an example) can be sourced from the "Country A Statistical Yearbook" and the "2023 Country A County Statistical Yearbook". Selected human activity variables include population density and Gross Domestic Product (GDP). The processed data is imported into ArcGIS 10.8.1 in CSV format and linked with the administrative division vector map of country A. To meet model requirements, spatial interpolation methods are used where necessary, and the data is finally converted to raster format using the "Feature to Raster" tool.
[0064] By supplementing the data with multiple factors, this invention enriches the research content and makes the predicted distribution of geographical indication agricultural products more convincing.
[0065] It should also be noted that in some embodiments, the method may further include the following steps: based on the Pearson correlation coefficient, factor screening is performed on the impact factor data by combining correlation and factor contribution with a preset screening threshold, and the impact factor data is updated based on the results of factor screening.
[0066] For example, in some specific implementations, the Pearson correlation coefficient was first tested before inputting the influencing factors. By considering the correlation and the level of factor contribution, the present invention performs preliminary screening of factors, thereby eliminating the impact of multicollinearity to the greatest extent.
[0067] In some specific application scenarios, assuming there are 40 influencing factors from C1 to C40, factor selection can be achieved as follows:
[0068] First, the correlation between factors was calculated using the Pearson correlation coefficient, and then factors with |r|>0.8 were selected. Specific correlation examples are shown below. Figure 2 And Table 1 below (abbreviated table of Pearson correlation coefficient factors):
[0069] Table 1
[0070] Abbreviation Original name Abbreviation Original name Abbreviation Original name Abbreviation Original name C1 awc C11 bio18 C21 cec_clay C31 sand C2 bio1 C12 bio19 C22 cec_eff C32 silt C3 bio10 C13 bio2 C23 cec_soil C33 total_n C4 bio11 C14 bio3 C24 clay C34 road_density C5 bio12 C15 bio4 C25 cn_ratio C35 CACRP C6 bio13 C16 bio5 C26 coarse C36 NOE C7 bio14 C17 bio6 C27 Road C37 PI_GDP C8 bio15 C18 bio7 C28 org_carbon C38 SI_GDP C9 bio16 C19 bio8 C29 ph_water C39 TI_GDP C10 bio17 C20 bio9 C30 pop2020 C40 GDP
[0071] After screening, 23 impact factors were finally obtained: Bio3, GDP, Road, Pop2020, Bio15, CACRP, NOE, Sand, TI-GDP, SI-GDP, PI-GDP, Clay, Awc, Silt, pH_Water, Road_Density, Coarse, Cec_Clay, Cec_Eff, DORF, Cn_ratio, Root_Depth, and Add_Prop. A cascade model was then constructed in R to calculate the contribution of each impact factor, and the top ten factors by contribution were displayed.
[0072] Based on the factor contribution results, the top four influencing factors are: Bio3, with the highest contribution of 10.06, is a climate factor (i.e., an influencing factor in climate data); followed by GDP with a contribution of 9.90, belonging to human activity factors (i.e., an influencing factor in human activity data); and Road and Pop2020, with contributions of 7.73 and 7.15 respectively, both belonging to human activity factors. Among soil factors (i.e., influencing factors in soil data), Sand ranks only 8th with a contribution of 4.81.
[0073] The specific contributions of the impact factors are shown in Table 2 below (contributions of key impact factors):
[0074] Table 2
[0075]
[0076]
[0077] S200. Obtain the distribution data of geographical indication agricultural products in the target area; based on the influencing factor data and distribution data, obtain the target predicted value of the adaptive distribution of geographical indication agricultural products in the target area through the preset prediction model.
[0078] The prediction models include the cascaded maximum entropy model and the random forest model;
[0079] It should be noted that, in some embodiments, obtaining the distribution data of geographical indication agricultural products in the target area may include the following steps: obtaining the latitude and longitude data of the distribution points of geographical indication agricultural products in the target area; using the distribution points as existing points, and randomly generating background points in areas outside the target area as non-existent points based on the latitude and longitude data; and organizing the distribution data of geographical indication agricultural products in the target area according to the existing points and non-existent points.
[0080] For example, in some specific implementations, the latitude and longitude data of Geographical Indication Rice Agricultural Products (GIRAP) can be obtained from the Anluyun platform (https: / / www.anluyun.com), collecting a total of 84 valid distribution points (which can be adjusted according to actual needs). After exporting the point data in CSV format, it is imported into ArcGIS Desktop 10.8.1 software for spatial visualization and analysis of the map of China. Specifically, based on the latitude and longitude data of the GIRAP locations, 2000 background points were randomly generated in RStudio and exported as CSV files. The presence points of GIRAP and the background points generated in R are used as the points indicating the presence / absence of the species.
[0081] It should also be noted that in some embodiments, the target predicted value of the adaptive distribution of geographical indication agricultural products in the target area is obtained by processing the influencing factor data and distribution data through a preset prediction model. This may include the following steps: spatially overlaying the distribution data and influencing factor data to construct an environmental dataset; inputting the environmental dataset into a maximum entropy model to obtain an initial predicted value; adding the initial predicted value as a new influencing factor to the environmental dataset, and then inputting it into a random forest model to obtain the target predicted value of the adaptive distribution of geographical indication agricultural products in the target area.
[0082] For example, in some specific implementations, the MaxEnt model (i.e., the maximum entropy model) used is a species distribution model. It relies on the presence points of geographical indication rice products and background points generated in R language as the presence / absence points of the species. Combined with the three types of influencing factors generated in step S100, an .asc format file is generated and then input into the MaxEnt model to predict the suitability distribution of geographical indication rice products. The predicted value is then exported as a new influencing factor, which, along with the aforementioned three types of environmental factors and the presence / absence points of the species, is input into a random forest model for further prediction, yielding the final predicted value for the suitability distribution of geographical indication rice products. This process, through an integrated cascade approach, constructs a MaxEnt-RF cascade model, significantly improving the accuracy of predicting the suitability distribution of geographical indication rice products.
[0083] In some specific application scenarios, in this invention, 2000 background points are first randomly generated in RStudio based on the latitude and longitude data of the geographical indication rice agricultural products, and then exported as a CSV file. Subsequently, the latitude and longitude data file is imported into ArcMap 10.8.1 software and spatially overlaid with multidimensional environmental factors (climate, soil factors) to construct a comprehensive environmental dataset with a resolution of 90 meters.
[0084] In the suitability prediction phase, a MaxEnt model was constructed in RStudio. The environmental dataset was input into the model to generate preliminary suitability probability predictions nationwide. Then, the predictions from the MaxEnt model were introduced into the dataset as new environmental factors to enhance the model's explanatory power. To further improve prediction accuracy, this invention employed a MaxEnt-RF cascade model. The predictions from the MaxEnt model were then input into an RF model constructed in RStudio to obtain the final suitability probability predictions for the entire country. The contribution of each influencing factor to the suitability distribution was quantitatively analyzed through variable importance assessment. Ten-fold cross-validation was used to validate the model, and model performance was evaluated using AUC, Kappa, and TSS values to ensure the accuracy and robustness of the prediction results.
[0085] In some specific application scenarios, in order to conduct suitability analysis and prediction of the maximum entropy model (MaxEnt) and the random forest model (RF), the raster data of all environmental factors (climate, soil and human activities) were processed with a unified spatial resolution (90m) and extent, and converted to ASCII format in ArcGIS 10.8.1 to meet the data format requirements of the MaxEnt model (data format examples are shown in Table 3).
[0086] Table 3
[0087]
[0088]
[0089]
[0090] In some specific application scenarios, the maximum entropy model (MaxEnt) can be implemented as follows:
[0091] In this invention, MaxEnt is used to identify the potential suitability distribution of GIAPs. By combining the locations of GIAPs with various environmental factors (including climate, soil, and human activity factors), MaxEnt can estimate the relative contribution of each environmental variable to the suitability distribution, generating a nationwide suitability probability distribution. The advantage of this model is that it can infer the suitability probability distribution of unobserved areas in an optimal manner under data constraints, based on limited existing data, ensuring the robustness and scientific validity of the prediction results.
[0092] Suppose we have an eigenvector x = (x1, x2, ..., xn) containing environmental factors (climate, soil, human activities). n Let H(p) represent the environmental conditions at a specific location. The maximum entropy model aims to estimate the probability of the target distribution p(x) in this feature space. By maximizing the entropy value, MaxEnt ensures that the estimated distribution is unbiased while satisfying the constraints of the observed data. The entropy H(p) is defined as follows:
[0093]
[0094] In this invention, MaxEnt provides a basis for the prediction of the initial suitability distribution, enabling the invention to more accurately identify potential suitable production areas for GIAP and laying an important foundation for subsequent model integration and factor contribution analysis.
[0095] In some specific application scenarios, the Random Forest (RF) model can be implemented as follows:
[0096] In this invention, RF (Enhanced Randomization) is used to further improve the prediction accuracy of GIAP (Growth-Induced Aptitude Test) suitability distribution. Based on the initial suitability probabilities generated by MaxEnt, the RF model constructs and integrates multiple decision trees to evaluate the contribution of each environmental factor to the GIAP suitability distribution. Through the variable importance assessment of the RF model, this invention can quantify the relative impact of factors such as climate, soil, and human activities on the suitability distribution, thereby identifying key factors that have a significant impact on the GIAP distribution. The ensemble learning characteristics of RF ensure the model's prediction accuracy and robustness, enabling suitability predictions to maintain high generalization ability in environments with multidimensional heterogeneous factors.
[0097] The decision formula is as follows:
[0098]
[0099] in, h i In category C j The output.
[0100] In this invention, the RF model not only provides more accurate predictions of fitness distributions, but also provides an effective quantitative tool for factor contribution analysis.
[0101] S300. Based on the impact factor data and target predicted values, the GeoShapley method is used to conduct a correlation analysis between the suitability of geographical indication agricultural products and the impact factors, and the correlation evaluation data is obtained.
[0102] It should be noted that, in some embodiments, step S300 may include the following steps: processing the impact factor data and target predicted values to obtain SHAP values, Moran's index, and spatial lag values; drawing a global contribution map and a positive and negative contribution map of the factors based on the SHAP values; determining the spatial distribution characteristics and positive and negative contributions of the impact factors to the suitability of geographical indication agricultural products based on the global contribution map and the positive and negative contribution map of the factors; and determining the spatial correlation between the impact factors and the suitability of geographical indication agricultural products based on the Moran's index and spatial lag values.
[0103] For example, in some specific implementations, the screened factors and the suitability results calculated by the model are merged into a complete vector shapefile in ArcMap. This file is then imported into Python, and a SHAP value is calculated using the GeoShapley method. Simultaneously, a global contribution map and a positive / negative contribution map of the factors are plotted based on the SHAP value. These maps provide a deeper understanding of the spatial distribution characteristics of factors on the suitability value and their positive / negative contributions, which can better guide the planting of geographical indication rice products and the formulation of related policies.
[0104] Furthermore, GeoDa software can be used to calculate Moran's I index and LISA value (spatial lag value) to analyze the spatial correlation and spatial clustering between suitability and influencing factors. Simultaneously, the spatial contribution of factors is generated based on the SHAP value, effectively analyzing the positive and negative correlations and spatial clustering between suitability and influencing factors, thus aiding subsequent factor spatial analysis. Based on the Moran's I index and LISA value, the spatial correlation between suitability and factors can be clearly understood, providing guidance for the suitability distribution of geographical indication rice products and offering suggestions for policy formulation and rice cultivation.
[0105] In some preferred embodiments, obtaining the SHAP value, Moran index, and spatial lag value based on impact factor data and target predicted values may include the following steps:
[0106] The impact factor data and target predicted values are imported into a preset first spatial data analysis software to synthesize a vector shape file. Based on the shape file, the GeoShapley method is used to process the SHAP value.
[0107] For example, in some specific implementations, machine learning and artificial intelligence technologies are increasingly widely used in crop suitability assessment. However, researchers often face the problem of "model black boxes" when revealing the complex relationships and dynamic processes of environmental factors behind suitability distribution. To address this challenge, interpretable artificial intelligence (XAI) has gradually become a research hotspot, with its core objective being to reveal the decision-making logic of machine learning models, thereby improving the transparency and interpretability of the models. In the field of crop suitability assessment, XAI has proven to effectively connect and integrate traditional statistical methods with machine learning techniques, especially in scenarios where the interactions of environmental factors are complex and nonlinear relationships are significant. Traditional suitability assessment methods are limited by linear assumptions, standardized model structures, and subjective weighting, making it difficult to fully analyze the mechanisms of action of multidimensional environmental factors. Machine learning models combined with XAI can not only overcome the limitations of traditional methods but also refine the analysis of the positive and negative contributions of key factors, revealing the role of factors in the spatial dimension. This provides a more comprehensive perspective and scientific support for crop suitability distribution research, laying a data foundation for agricultural production layout optimization and policy formulation.
[0108] It should be noted that the GeoShapley method is the key technical focus of this invention, and the following description elaborates on the GeoShapley method applied in this invention from several aspects:
[0109] The GeoShapley algorithm, used by XAI to explain the contributions of key factors in the GIAP suitability distribution prediction model, can more accurately quantify the contributions of different geographical factors in the model. By calculating the contribution of each input feature (key factor) to the suitability outcome, GeoShapley helps reveal the positive and negative impacts of factors such as climate, soil, and human activities on the GIAP suitability distribution. The Shapley value of each feature represents its "fair contribution" to the prediction result, enabling the quantification of the specific role of different factors in the suitability distribution, thereby unveiling the prediction logic of the "black box" model.
[0110] Classic Shapley value formula:
[0111]
[0112] p is the total number of features, S is the subset of features, and f(S) is the predicted output of subset S.
[0113] GeoShapley value formula:
[0114]
[0115] GEO: The set containing all geographic features; g: The number of geographic features.
[0116] Expected calculation of Shaply Values:
[0117] f x (S)=E[f(X)|do(X s =x s (5)
[0118] S is the feature set we want to adjust, X s Let x be a random variable representing the M input features of the model. s It is the current model input vector for prediction.
[0119] Local accuracy is:
[0120]
[0121] is the feature attributes and f(x) is the output of the original model.
[0122] The advantage of GeoShapley in this invention lies in its ability to not only analyze in detail the relative contributions of each environmental factor in the suitability model, but also to identify which factors have a positive or negative impact on the suitability of a specific region. This provides policymakers with a clear explanation of the factors and a basis for spatial planning. This interpretability enhances the credibility of the model and provides scientific support for the rational distribution and resource optimization of geographical indication agricultural products.
[0123] The impact factor data and target predicted values are imported into a pre-set second spatial data analysis software to obtain the Moran index and spatial lag value.
[0124] For example, in some specific embodiments, the Local Moran's I index is used to analyze the spatial similarity of the suitability of each region in the GIAP suitability distribution with the influence factors of its surrounding regions.
[0125] The local Moran's I exponent is calculated using the following formula:
[0126]
[0127] I i: The local Moran's I value of the i-th cell;
[0128] x i x j : Attribute values of the i-th and j-th regions;
[0129] Average attribute value across all regions;
[0130] w ij Spatial weight between the i-th and j-th regions, typically used to represent the relationship between adjacent regions (1 for adjacent regions, 0 for non-adjacent regions).
[0131] Local Moran's I values help identify the relationship between attribute values in each region and neighboring regions. Based on its value, one can determine:
[0132] Positive correlation (high-high or low-low clustering): When a region and its neighboring regions all exhibit high or low attribute values, the local Moran's I value is positive, indicating a clustering effect.
[0133] Negative correlation (high-low or low-high dispersion): When the attribute value of a region is high, while the value of neighboring regions is low (or vice versa), the local Moran's I value is negative, indicating a dispersion effect.
[0134] In this invention, LISA (Local Indicators of Spatial Association) is used to analyze the local spatial autocorrelation of GIAP suitability distribution. By calculating the LISA values of suitability and influencing factors for each region, spatial clustering and dispersion patterns in the suitability distribution can be identified and visualized, such as "high-high" clustering areas in high suitability regions and "low-low" clustering areas in low suitability regions. The analysis results of LISA are usually presented in the form of heatmaps, clearly showing the spatial hotspots and cold spots in the GIAP suitability distribution.
[0135] Local Moran's I index and LISA values are powerful tools for spatial autocorrelation analysis, revealing the spatial relationship between suitability and influencing factors in data and their neighboring regions. The local Moran's I index provides specific numerical values to determine the type of autocorrelation, while the LISA value further expands the visualization of local autocorrelation, offering strong support for the analysis of spatial clustering patterns.
[0136] To explain in detail the principle of the technical solution of the present invention, the overall process of the present invention will be described below with reference to some specific embodiments. It is easy to understand that the following is an explanation of the technical principle of the present invention and should not be regarded as a limitation of the present invention.
[0137] Taking China's rice-growing areas as the target region of this study as an example, it covers the major agricultural areas of China. China's rice-growing areas are widely distributed in the southeastern coastal areas (including Zhejiang, Fujian, Guangdong, and Guangxi), the middle and lower reaches of the Yangtze River (including Jiangsu, Anhui, Hubei, and Hunan), the northeastern plains (including Heilongjiang, Jilin, and Liaoning), and the southwestern mountainous areas (including Yunnan, Guizhou, and Sichuan). These regions not only have diverse natural environments but also exhibit different characteristics of human activities, such as economic development levels, agricultural mechanization levels, and population densities.
[0138] The geographical indication rice varieties in these regions not only reflect the diversity of natural conditions such as climate and soil, but also demonstrate the profound impact of human activities on the distribution of agricultural suitability. The level of mechanization, advancements in agricultural technology, and changes in population density have collectively shaped the pattern of rice production in China.
[0139] Specifically, the data for the technical solution of this invention can be obtained through the following platforms or databases:
[0140] The latitude and longitude data for Geographical Indication Rice Agricultural Products (GIRAP) were obtained from the Anluyun platform (https: / / www.anluyun.com), collecting a total of 84 valid distribution points. The point data was exported in CSV format and imported into ArcGIS Desktop 10.8.1 software for spatial visualization and analysis of the map of China.
[0141] Climate data was obtained from the WorldClim 2.1 database (https: / / www.worldclim.org). The BCC-CSM2-MR global climate model (GCM) was selected, with a spatial resolution of 2.5 arcminutes (approximately 5 kilometers) and a time range of 2021–2040, based on the shared socioeconomic path scenario SSP126. Nineteen bioclimatic variables (BIO1–BIO19) were downloaded in GeoTIFF format, and relevant bioclimatic variables were extracted using the "Image Analysis" tool in ArcGIS 10.8.1 to obtain global climate data. To focus on the study area, the boundary vector file of the target area (such as a country or province / city, taking country A as an example) was used to perform mask extraction on the global climate data, and the climate data within country A was obtained using the "Extract by Mask" function.
[0142] Soil data was obtained from the Harmonized Global Soil Database version 2.0 (HWSD v2.0) (https: / / gaez.fao.org). Due to the large amount of data, R software version 4.2.0 was used for data processing and extraction. Soil properties in layer D1 (0–20 cm) were selected as the main research object, including variables such as soil texture, organic matter content, pH value, and nutrient content. According to the FAO soil classification system, the soils of country A were classified into 22 soil order types, including ferruginous soils, leached soils, arid soils, and desert soils. Trace element data were obtained from the Guangdong Provincial Soil Environmental Monitoring Center.
[0143] Human activity factor data (taking country A as an example) can be sourced from the "Country A Statistical Yearbook" and the "2023 Country A County Statistical Yearbook". Selected human activity variables include population density and Gross Domestic Product (GDP). The processed data is imported into ArcGIS 10.8.1 in CSV format and linked with the administrative division vector map of country A. To meet model requirements, spatial interpolation methods are used where necessary, and the data is finally converted to raster format using the "Feature to Raster" tool.
[0144] To conduct suitability analysis and prediction for the Maximum Entropy (MaxEnt) model and the Random Forest (RF) model, the raster data of all environmental factors (climate, soil, and human activities) were processed with a uniform spatial resolution (90m) and extent, and converted to ASCII format in ArcGIS 10.8.1 to meet the data format requirements of the MaxEnt model (see Table 1 above for an example of the data format).
[0145] like Figure 3 and Figure 4 As shown, where Figure 3 For detailed examples of factor positive and negative contribution plots, factor global contribution plots, local Moran's I index, LISA values, and SHAP space plots, please refer to [link / reference]. Figure 6 , Figure 7 , Figure 8 and Figure 9 , Figure 3 This is just a vague example; in addition, Figure 3 and Figure 4 The evaluation factors and parameters for selecting the data graph corresponding to the maximum entropy in the maximum entropy model may change depending on the actual application scenario. Furthermore, the details of these parameters do not affect the logical principle of the technical solution of this invention. Therefore, no specific examples of these parameters will be provided; only a vague example will be shown. The specific implementation process of the technical solution of this invention can be as follows:
[0146] In this invention, 2000 background points were first randomly generated in RStudio based on the latitude and longitude data of geographical indication rice agricultural products, and then exported as CSV files. Subsequently, the latitude and longitude data files were imported into ArcMap 10.8.1 software and spatially overlaid with multidimensional environmental factors (climate, soil factors) to construct a comprehensive environmental dataset with a resolution of 90 meters.
[0147] In the suitability prediction phase, a MaxEnt model was constructed in RStudio. The environmental dataset was input into the model to generate preliminary suitability probability predictions nationwide. Then, the prediction results from the MaxEnt model were introduced into the dataset as new environmental factors to enhance the model's explanatory power. To further improve prediction accuracy, this invention employs a MaxEnt-RF cascade model, inputting the MaxEnt model's prediction results into the RF model constructed in RStudio. Before inputting factors into the RF model, Pearson correlation coefficients were used to screen factors, retaining key factors with |r|>0.8 and removing variables with low contribution, thus obtaining the final suitability probability predictions nationwide. The contribution of each influencing factor to the suitability distribution was quantitatively analyzed through variable importance assessment. Ten-fold cross-validation was used to validate the model, and AUC, Kappa, and TSS values were used to evaluate model performance to ensure the accuracy and robustness of the prediction results.
[0148] To further explore the spatial impact of influencing factors on the positive and negative contributions of fitness distribution, this invention utilizes Python within PyCharm to construct the GeoShapley method. The main contributing factors calculated by the MaxEnt-RF model and the fitness distribution results are input into the model to obtain and visualize the SHAP values, generating a global contribution map and a positive / negative contribution map of the factors. This enables a refined analysis of the spatial positive and negative contributions of influencing factors. Furthermore, the study introduces local Moran's I index and LISA values to analyze the spatial autocorrelation between factors and fitness distribution, visually demonstrating the spatial correlation, spatial clustering characteristics, and spatial contribution of factors, revealing aggregation or dispersion patterns in different regions. Through the visualization analysis of factor contributions in space, this invention provides a more in-depth ecological explanation of the spatial characteristics of GIAP fitness distribution. The maximum entropy model (MaxEnt), random forest model (RF), GeoShapley method, Moran's I index and LISA values, and the logical principles of factor selection used in this invention are described in the aforementioned specific application scenario examples and will not be repeated here.
[0149] It should be noted that in some specific application scenarios, the MaxEnt-RF model implemented by the technical solution of this invention exhibits the following effects:
[0150] In this invention, three different models were used to predict the suitability distribution of geographical indication agricultural products: MaxEnt, RF, and a cascade model that incorporates the suitability probability value output by MaxEnt as a new environmental factor into the RF model. To comprehensively evaluate the predictive performance of each model, three indicators were used for model evaluation: AUC (Area Under the Receiver Operating Characteristic), Kappa coefficient, and TSS (True Skill Statistic).
[0151] like Figure 5 As shown, by comparing the performance of the three models, the MaxEnt-RF cascade model scored 0.97 in AUC, higher than the MaxEnt model's 0.95 and the RF model's 0.93; in terms of Kappa, the MaxEnt-RF cascade model scored 0.32, higher than the RF model's 0.31 and the MaxEnt model's 0.09; and in terms of TSS, the MaxEnt-RF cascade model scored 0.78, higher than the RF model's 0.59 and the MaxEnt model's 0.47.
[0152] It should also be noted that in some specific application scenarios, spatial analysis of influence factors can be implemented as follows:
[0153] Factors selected using Pearson correlation coefficients were analyzed using the GeoShapley method to calculate their spatial positive and negative contributions to fitness and their SHAP values. We selected the top four contributing factors—Bio3, GDP, Road, and Pop2020—to display their SHAP values.
[0154] From GeoShapley's positive and negative contributions Figure 6 The results show that:
[0155] The SHAP value of railway density exhibits a broad distribution, with significant values at both positive and negative ends, and high values concentrated in the positive direction, indicating that areas with higher railway density have a positive impact on fitness distribution. Bio3 is mainly distributed in the positive direction of SHAP values, suggesting that higher Bio3 has a positive effect on fitness distribution. The SHAP value of GDP is mostly concentrated in the negative direction, indicating that higher GDP has a negative impact on fitness distribution. The SHAP value of Pop has a smaller distribution range, with relatively dispersed positive and negative distributions at both ends, indicating that population density has a relatively small and neutral impact on fitness distribution.
[0156] From the perspective of global contribution Figure 7The results show that Road density has the greatest impact on the suitability distribution, followed by Bio3 and GDP, while Pop density has the least impact. This indicates that in our model, changes in railway density contribute most significantly to the suitability distribution of the geographical indication rice.
[0157] according to Figure 8 The spatial characteristic maps of the four SHAP values can be used to analyze the spatial impact of population density, railway density, GDP and climate factor BIO3 on the suitability distribution of geographical indication rice agricultural products.
[0158] The impact of GDP on suitability distribution is mainly concentrated in a few areas, primarily showing a negative contribution. The negative value areas are relatively widespread, indicating that areas with higher GDP have a suppressive effect on suitability distribution. The impact of railway density on suitability distribution has significant positive and negative contributions in specific areas, exhibiting some spatial heterogeneity. Areas with higher positive contributions are concentrated in the southeast and parts of the central region, while areas with low railway density mostly show negative contributions. The impact of Bio3 is mainly concentrated in the high-latitude regions of the north, showing a significant positive contribution, indicating that climatic conditions in these areas have a relatively positive effect on suitability distribution. The SHAP value of population density is relatively dispersed, with a small overall impact and low positive and negative contribution values, indicating that population density has a weak impact on suitability distribution.
[0159] Based on the local bivariate Moran's I index and LISA values, the spatial correlation between the distribution of geographical indication rice agricultural product suitability and major influencing factors was analyzed:
[0160] High-high GDP clusters are mainly found in the southeastern coastal areas, such as Zhejiang, Jiangsu, and Shanghai, demonstrating the spatial clustering of high GDP regions. Low-low clusters are distributed in inland and western regions. High-high railway density clusters are concentrated in the Yangtze River Delta region, especially Jiangsu and Zhejiang, while low-low clusters are located in Yunnan and Guizhou in the southwest, indicating a significant spatial clustering pattern of railway density in these regions. High-high Bio3 clusters are distributed in the high-latitude northern regions, while low-low clusters are mainly found in the southern coastal areas, reflecting spatial differences in climatic conditions. High-high population density clusters are concentrated in the middle and lower reaches of the Yangtze River, while low-low clusters are located in the sparsely populated western regions.
[0161] from Figure 9 It can be seen that:
[0162] Moran's I for GDP is 0.104, indicating weak spatial autocorrelation, but still showing a certain positive correlation trend. Moran's I for railway density is 0.431, showing strong positive spatial autocorrelation, indicating a strong spatial clustering effect of railway density. Moran's I for Bio3 is -0.485, indicating a significant negative spatial autocorrelation, suggesting a spatially dispersed distribution of high and low value areas. Moran's I for population density is 0.034, showing weak positive spatial autocorrelation, indicating that the spatial clustering of population density is not significant.
[0163] In summary, this invention proposes and validates an improved MaxEnt-RF cascade model for predicting the suitability distribution of geographical indication agricultural products, and systematically compares its performance with traditional standalone MaxEnt and Random Forest (RF) models. The results show that the MaxEnt-RF cascade model exhibits significant advantages in key evaluation indicators such as AUC, Kappa, and TSS, demonstrating a significant improvement in prediction accuracy and model stability, thus effectively addressing the problem of insufficient accuracy in predicting suitability under complex influencing factors.
[0164] Specifically, regarding the AUC index, the MaxEnt-RF concatenated model scored the highest, indicating its strong discriminative ability in predicting the suitability distribution of geographical indication agricultural products, and its ability to more effectively distinguish between suitable and unsuitable areas. In terms of the Kappa coefficient, the MaxEnt-RF concatenated model and the RF model scored similarly, both significantly higher than the MaxEnt model, showing higher consistency and indicating that the model maintains high stability under changing environmental conditions. Furthermore, in terms of the TSS index, the MaxEnt-RF concatenated model performed the best, significantly outperforming other models. This result further demonstrates its ability to reduce false positives and false negatives.
[0165] By introducing the fitness probability from the MaxEnt model as a new environmental factor into the RF model, the MaxEnt-RF cascade model can more comprehensively integrate the complex relationships between different environmental factors, greatly enhancing the model's explanatory and predictive power. This model not only captures the impact of individual factors on the fitness distribution but also effectively combines the nonlinear feature extraction capability of the MaxEnt model with the powerful classification ability of the RF model through ensemble processing, making the prediction results more accurate and reliable.
[0166] In summary, the development and application of the MaxEnt-RF cascade model provides an effective method for accurately predicting the suitability distribution of geographical indication agricultural products at a high resolution of 90m. The success of this model demonstrates that integrating the advantages of different models can not only effectively improve prediction accuracy but also enhance model interpretability to some extent. Future research can further optimize the model to adapt to more complex environmental factors or larger geographical areas, providing policymakers and regional managers with a more scientific basis for rationally planning the production layout of geographical indication agricultural products.
[0167] This invention incorporates numerous human activity factors into a suitability distribution model, treating them alongside climate and soil factors commonly used in traditional studies as key influencing factors. This innovative approach addresses the shortcomings of previous research in understanding the impact of human activities on suitability distribution, contributing to a more comprehensive understanding of the mechanisms by which human activities affect suitability distribution. Furthermore, it innovatively combines GeoShapley, Moran's I index, and LISA values to systematically analyze the positive and negative contributions of each influencing factor to suitability distribution across spatial dimensions.
[0168] GeoShapley is an extension of TreeSHAP, incorporating Shapley Value theory to interpret geospatial data. Its advantages over existing methods include: GeoShapley, by incorporating spatial weighting, better reflects the differences in features across different geographic regions; based on Shapley Value theory, it effectively decomposes the contribution of feature interactions to predictions; traditional partial correlation analysis relies on the assumption of linear correlation, while GeoShapley does not, making it suitable for complex nonlinear models; partial correlation analysis may underestimate or omit some driving factors due to multicollinearity and significance thresholds (P-values), while GeoShapley more robustly captures the overall and local contributions of features to predictions without relying on significance tests; and GeoShapley's results exhibit stronger spatial consistency, avoiding the subjectivity of window-shifting methods.
[0169] Factor analysis results obtained using the above methods show that railway density contributes the most to the distribution of geographical indication rice suitability, indicating that transportation infrastructure plays a crucial role in optimizing rice-growing areas. Specifically, areas with high railway density provide convenient conditions for rice production and transportation, effectively promoting the circulation and supply chain efficiency of agricultural products. This positive impact of railway density on rice suitability distribution is particularly significant in areas with high demand. Moran's I index results show that the railway density factor has strong positive spatial autocorrelation. This indicates that there are high-high clustering areas in the spatial distribution of railway density, meaning that areas with denser railways tend to cluster together. This clustering shows that areas with convenient transportation are more suitable for the cultivation of geographical indication rice, and the advantages of transportation conditions can effectively improve the efficiency of agricultural product production and circulation. The LISA map further shows that in the Yangtze River Delta and some eastern coastal areas, areas with dense railways and areas with high rice suitability exhibit a significant spatial clustering relationship.
[0170] Secondly, the high contribution of Bio3 indicates its importance in suitability distribution. Bio3 represents seasonal temperature stability, and its high positive contribution suggests that climate stability plays a positive role in rice cultivation suitability. Especially in areas with small seasonal temperature differences and relatively stable climate, high Bio3 values indicate relatively stable rice growing conditions, which is beneficial for improving suitability distribution. This aligns with the important role of climatic factors in rice growth, demonstrating the significant impact of temperature changes on the distribution of geographical indication rice. The spatial characteristics maps of Moran's I index and GeoShapley show that Bio3 exhibits negative spatial autocorrelation, particularly in high-latitude regions. This indicates that the suitability of climatic factors is spatially unstable and uncertain; the greater seasonal fluctuations in climate in high-latitude regions lead to spatial discontinuities in suitability distribution. On the LISA map, high-latitude regions show a clear low-low clustering, which may be due to unstable climatic conditions in these areas, negatively impacting rice suitability. This climatic instability results in significant differences in rice suitability distribution in these regions.
[0171] In contrast, GDP and population density contribute relatively little, and high GDP areas have a negative impact on suitability distribution. This may be because land use in economically developed regions is mainly concentrated on urban expansion and industrial development, leading to the compression of agricultural land and hindering the development of traditional rice cultivation. Therefore, although regions with high GDP are economically active and resource-rich, their land use patterns may not be compatible with the needs of agricultural production, thus inhibiting the suitability distribution of geographical indication rice. Similarly, population density contributes little to suitability distribution, possibly because rice cultivation mainly depends on natural conditions and convenient transportation, and areas with high population density do not necessarily have the optimal agricultural production environment. The Moran's I index value of the GDP factor is low, indicating its weak spatial correlation in the suitability distribution of geographical indication rice. In the LISA map, high GDP areas and areas with high rice suitability do not show a significant spatial clustering relationship. This may be because high GDP areas are mainly concentrated in urban and industrial areas, and their land use patterns are not conducive to the suitable development of traditional agriculture, thus their impact on rice suitability shows significant spatial heterogeneity. The Moran's I index value for population density was also low, indicating a weak spatial correlation with rice suitability distribution. In the LISA map, densely populated areas did not form significant spatial clustering patterns, suggesting that the influence of population density on rice suitability distribution is relatively dispersed and lacks clear spatial consistency. This may be because rice suitability distribution mainly depends on natural conditions and transportation factors, rather than simply population distribution.
[0172] In summary, this invention establishes the MaxEnt-RF model. Compared to individual models, the cascaded model of this invention exhibits higher predictive accuracy and can apply complex environmental factors at high spatial resolution to complete predictions. It also reveals the roles and spatial characteristics of different environmental factors in the distribution of rice suitability. Railway density and climate stability are the main factors influencing the distribution of rice suitability, while economic and population factors show significant spatial heterogeneity. The research results provide a scientific basis for the future rational planning of the production layout of geographical indication rice products and offer a reference for policymakers in transportation and land use management. Furthermore, this invention uses the GeoShapley method to quantitatively assess the contribution of each factor in spatial distribution. The contribution of each factor to crop suitability is calculated, considering the interactions and spatial heterogeneity between factors. Based on the calculation results, a global contribution distribution map of the factors and a positive / negative contribution analysis map of the factors are drawn. These maps provide a deeper understanding of the spatial distribution characteristics of factors on suitability values and their positive / negative contributions, which can better guide the planting of geographical indication rice products and the formulation of related policies.
[0173] Specifically, this invention constructs a MaxEnt-RF model to predict the suitability distribution of geographical indication rice products nationwide. The model takes climate, soil, and human factors as inputs and combines GeoShapley, Moran's I index, and LISA values to reveal the impact of human activities on suitability distribution, as well as the spatial positive and negative contributions of each key factor. The spatial positive and negative contribution analysis of these factors is of great significance for optimizing the planting layout of geographical indication rice. Simultaneously, our research provides a complete research process system for establishing the suitability distribution and planting patterns of geographical indication agricultural products, clarifies specific analytical steps, enriches our understanding of the influencing factors of geographical indication agricultural products, and proposes new research ideas and scientific hypotheses for future in-depth exploration of the factor influence mechanisms. The results show that the MaxEnt-RF model effectively improves the accuracy of predictions, and spatial analysis methods such as GeoShapley clearly demonstrate the positive and negative contributions of each influencing factor, providing strong support for achieving precise and intelligent planting of geographical indication agricultural products. This invention provides a promising reference method for the planning of geographical indication agricultural product planting and government decision-making in the study area and other regions.
[0174] like Figure 10 As shown, the invention also provides a geographical indication agricultural product suitability evaluation device 900, which may include:
[0175] The first module, 901, is used to acquire impact factor data for the target region; the impact factor data includes multiple impact factors.
[0176] The second module 902 is used to acquire the distribution data of geographical indication agricultural products in the target area; based on the influencing factor data and distribution data, the target predicted value of the adaptive distribution of geographical indication agricultural products in the target area is obtained through the processing of the preset prediction model; the prediction model includes the cascaded maximum entropy model and the random forest model;
[0177] The third module, 903, is used to conduct a correlation analysis between the suitability of geographical indication agricultural products and influencing factors based on influencing factor data and target predicted values, using the GeoShapley method to obtain correlation evaluation data.
[0178] The content of the method embodiments of the present invention is applicable to the device embodiments. The specific functions implemented by the device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above methods.
[0179] This invention also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned method for evaluating the suitability of geographical indication agricultural products. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0180] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0181] Please see Figure 11 , Figure 11 The hardware structure of an electronic device 1000 according to another embodiment is illustrated. The electronic device 1000 includes:
[0182] The processor 1001 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present invention.
[0183] The memory 1002 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 1002 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1002 and is called and executed by the processor 1001 to execute the geographical indication agricultural product suitability evaluation method of the embodiments of this invention.
[0184] Input / output interface 1003 is used to implement information input and output;
[0185] The communication interface 1004 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0186] Bus 1005 transmits information between various components of the device (e.g., processor 1001, memory 1002, input / output interface 1003, and communication interface 1004);
[0187] The processor 1001, memory 1002, input / output interface 1003 and communication interface 1004 are connected to each other within the device via bus 1005.
[0188] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for evaluating the suitability of geographical indication agricultural products.
[0189] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0190] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0191] The present invention provides a method, device, electronic device, and storage medium for evaluating the suitability of geographical indication agricultural products (GIAPs). This method acquires influencing factor data for a target region, including multiple influencing factors. It also acquires distribution data of GIAPs within the target region. Based on the influencing factor data and distribution data, a pre-defined prediction model is used to obtain a target predicted value for the adaptive distribution of GIAPs in the target region. The prediction model includes a cascaded maximum entropy model and a random forest model. Based on the influencing factor data and the target predicted value, the GeoShapley method is used to perform a correlation analysis between the suitability of GIAPs and the influencing factors, yielding correlation evaluation data. This invention constructs a novel cascaded model as a model for predicting the suitability distribution of GIAPs, integrating the advantages of multiple models to improve the accuracy of predicting GIAP suitability distribution. This invention significantly improves the accuracy of suitability distribution prediction, reveals in detail the specific roles of each influencing factor in space, provides a scientific basis for the rational production layout of GIAPs, and provides guidance for policy formulation and regional agricultural management.
[0192] The embodiments described in this invention are for the purpose of more clearly illustrating the technical solutions of the embodiments of this invention, and do not constitute a limitation on the technical solutions provided by the embodiments of this invention. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this invention are also applicable to similar technical problems.
[0193] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present invention, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0194] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0195] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0196] The terms "first," "second," "third," "fourth," etc. (if present) in the specification and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0197] It should be understood that in this invention, "at least one (item)" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0198] In the embodiments provided by this invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between systems or units may be electrical, mechanical, or other forms.
[0199] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0200] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0201] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0202] The preferred embodiments of the present invention have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and spirit of the present invention should be within the scope of the claims of the present invention.
Claims
1. A method for evaluating the suitability of geographical indication agricultural products, characterized in that, The method includes the following steps: Obtain impact factor data for the target region; the impact factor data includes multiple impact factors; Obtain the distribution data of geographical indication agricultural products in the target area; based on the influencing factor data and the distribution data, process the data using a preset prediction model to obtain the target predicted value of the adaptive distribution of the geographical indication agricultural products in the target area; the prediction model includes a cascaded maximum entropy model and a random forest model; The step of obtaining the target predicted value of the adaptive distribution of the geographical indication agricultural product in the target area based on the influencing factor data and the distribution data through a preset prediction model includes the following steps: The distribution data and the influencing factor data are spatially overlaid to construct an environmental dataset; the influencing factor data includes various environmental factors such as climate, soil, and human activities. The environmental dataset is input into the maximum entropy model to obtain initial predicted values; the initial predicted values represent the preliminary suitability probability distribution. The initial predicted values are added to the environmental dataset as new influencing factors, and then input into the random forest model to obtain the target predicted values of the adaptive distribution of the geographical indication agricultural products in the target area; the target predicted values represent the contribution of each environmental factor. Based on the impact factor data and the target predicted value, the GeoShapley method is used to conduct a correlation analysis between the suitability of the geographical indication agricultural product and the impact factors, and the correlation evaluation data is obtained. The step of conducting a correlation analysis between the suitability of the geographical indication agricultural product and the influencing factors based on the influencing factor data and the target predicted value, using the GeoShapley method to obtain correlation evaluation data, includes the following steps: The SHAP value, Moran index, and spatial lag value are obtained by processing the impact factor data and the target predicted value. Based on the SHAP value, a global contribution map and a positive and negative contribution map of the factors are drawn. Based on the global contribution map and the positive and negative contribution map of the factors, the spatial distribution characteristics and positive and negative contribution degrees of the influencing factors for the suitability of the geographical indication agricultural products are determined in sequence. The spatial correlation between the influencing factors and the suitability of the geographical indication agricultural products is determined based on the Moran index and the spatial lag value.
2. The method for evaluating the suitability of geographical indication agricultural products according to claim 1, characterized in that, The influencing factor data includes climate data, soil data, and human activity factor data, each of which includes several influencing factors; obtaining the influencing factor data for the target area includes the following steps: Global climate data is obtained from a preset climate database, and the global climate data is masked and extracted based on the boundary vector file of the target region to obtain the climate data of the target region; Soil properties and trace element data of the target area are obtained from a preset soil database as the soil data; Human activity variables for the target area are obtained from preset statistical data, the human activity variables are associated with a vector map of the target area, and the data is then converted into a raster format to obtain the human activity factor data.
3. The method for evaluating the suitability of geographical indication agricultural products according to claim 1, characterized in that, The process of obtaining the distribution data of geographical indication agricultural products in the target area includes the following steps: Obtain the latitude and longitude data of the distribution points of the geographical indication agricultural products in the target area; The distribution points are taken as existing points, and background points are randomly generated in areas outside the target area based on the latitude and longitude data as non-existent points; The distribution data of the geographical indication agricultural products in the target area are obtained by organizing the points that exist and the points that do not exist.
4. The method for evaluating the suitability of geographical indication agricultural products according to claim 1, characterized in that, The method further includes the following steps: Based on the Pearson correlation coefficient, factor screening is performed on the impact factor data by combining correlation and factor contribution with a preset screening threshold, and the impact factor data is updated based on the results of the factor screening.
5. The method for evaluating the suitability of geographical indication agricultural products according to claim 1, characterized in that, The process of obtaining the SHAP value, Moran index, and spatial lag value based on the impact factor data and the target predicted value includes the following steps: The impact factor data and the target predicted value are imported into a preset first spatial data analysis software to synthesize a vector shape file. The shape file is then processed using the GeoShapley method to obtain the SHAP value. The impact factor data and the target predicted value are imported into a preset second spatial data analysis software for processing to obtain the Moran index and the spatial lag value.
6. A device for evaluating the suitability of geographical indication agricultural products, characterized in that, The device includes: The first module is used to acquire impact factor data for the target region; the impact factor data includes multiple impact factors. The second module is used to acquire the distribution data of geographical indication agricultural products in the target area; based on the influencing factor data and the distribution data, the target predicted value of the adaptive distribution of the geographical indication agricultural products in the target area is obtained through a preset prediction model; the prediction model includes a cascaded maximum entropy model and a random forest model; The step of obtaining the target predicted value of the adaptive distribution of the geographical indication agricultural product in the target area based on the influencing factor data and the distribution data through a preset prediction model includes the following steps: The distribution data and the influencing factor data are spatially overlaid to construct an environmental dataset; the influencing factor data includes various environmental factors such as climate, soil, and human activities. The environmental dataset is input into the maximum entropy model to obtain initial predicted values; the initial predicted values represent the preliminary suitability probability distribution. The initial predicted values are added to the environmental dataset as new influencing factors, and then input into the random forest model to obtain the target predicted values of the adaptive distribution of the geographical indication agricultural products in the target area; the target predicted values represent the contribution of each environmental factor. The third module is used to perform a correlation analysis between the suitability of the geographical indication agricultural product and the influencing factors based on the influencing factor data and the target predicted value, using the GeoShapley method, to obtain correlation evaluation data. The step of conducting a correlation analysis between the suitability of the geographical indication agricultural product and the influencing factors based on the influencing factor data and the target predicted value, using the GeoShapley method to obtain correlation evaluation data, includes the following steps: The SHAP value, Moran index, and spatial lag value are obtained by processing the impact factor data and the target predicted value. Based on the SHAP value, a global contribution map and a positive and negative contribution map of the factors are drawn. Based on the global contribution map and the positive and negative contribution map of the factors, the spatial distribution characteristics and positive and negative contribution degrees of the influencing factors for the suitability of the geographical indication agricultural products are determined in sequence. The spatial correlation between the influencing factors and the suitability of the geographical indication agricultural products is determined based on the Moran index and the spatial lag value.
7. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 5.