A system and method for processing and analyzing data on uncultivated arable land

By constructing a data processing and analysis system for uncultivated arable land, and combining multi-model collaborative analysis and significance index screening, the problem of a single perspective in the analysis of uncultivated arable land was solved. This enabled the accurate identification of influencing factors of uncultivated arable land and the prediction of future trends, thereby improving the effectiveness and accuracy of the analysis results.

CN120725290BActive Publication Date: 2025-10-31SICHUAN PROVINCIAL INST OF LAND SCI & TECH (SICHUAN PROVINCIAL SATELLITE APPL TECH CENT)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511163439.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-19
Publication Date
2025-10-31
Estimated Expiration
2045-08-19

AI Technical Summary

Technical Problem

Existing technologies for analyzing uncultivated arable land suffer from limitations such as a single perspective in factor analysis, a lack of multi-dimensional spatiotemporal data fusion and cross-model collaborative analysis, and an inability to fully consider the coupling mechanism of socio-economic systems, farmer behavior characteristics, and natural endowment conditions. This results in insufficient explanatory power for the complex process of uncultivated arable land and a lack of in-depth analysis of the dynamic evolution of uncultivated rates.

Method used

A data processing and analysis system for uncultivated farmland was constructed, which includes multivariate analysis modules such as geographic detector model, panel model, spatial model and spatiotemporal geographic weighted regression model. Through the collaboration of multiple models, combined with the screening and comprehensive analysis of significant indicators, the system can scientifically identify and accurately extract key influencing factors, and build predictive models to predict future trends.

Benefits of technology

It enables a comprehensive and in-depth analysis and refined quantitative analysis of the factors affecting uncultivated arable land, improves the accuracy of factor identification and spatial prediction precision, reduces time series prediction errors, and provides a basis for zoned and classified governance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120725290B_ABST
    Figure CN120725290B_ABST
Patent Text Reader

Abstract

This invention provides a system and method for processing and analyzing data on uncultivated farmland, including a module for analyzing factors influencing uncultivated farmland rates and a module for predicting uncultivated farmland rates. The module for analyzing factors influencing uncultivated farmland rates integrates modules such as a geographic detector model, panel model, spatial model, and spatiotemporal geographic weighted regression model to obtain results on significant influencing factors of uncultivated farmland in a target area. The module for predicting uncultivated farmland rates includes a panel model prediction module and a time t prediction module to predict the uncultivated farmland rates for each administrative region in the following year. Based on the technical solution of this invention, multi-source heterogeneous data such as land use and climate environment are integrated; the indicator classification is more detailed and comprehensive, making the analysis results more effective; temporal and spatial factors are integrated to conduct more accurate analysis of variables in different regions; the accuracy of factor identification is significantly improved, the spatial prediction accuracy is significantly improved, and the time series prediction error is significantly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of big data analysis and mining technology, and in particular to a system and method for processing and analyzing data on uncultivated farmland. Background Technology

[0002] While the current technical system for analyzing factors contributing to uncultivated arable land at the national, provincial, and municipal levels has broad coverage, it still suffers from significant shortcomings. Existing technologies are mostly based on single-factor analysis, focusing on exploring the impact of social and economic factors on uncultivated land, primarily using linear regression models. For individualized factors affecting farmers, binary choice models are mainly used for analysis; and for natural factors, qualitative or descriptive statistical analysis is employed. Overall, although the methods are diverse, the perspectives are limited, failing to consider the overall impact of uncultivated land and neglecting the intrinsic connections between different factors. In other words, current analysis of uncultivated arable land has significant shortcomings in its technical application, relying heavily on single-factor analysis models and lacking a multi-dimensional spatiotemporal data fusion and cross-model collaborative analysis system. Existing technologies often employ traditional regression models or single geographic detectors and lack a systematic analysis of the coupling mechanisms between socioeconomic systems, farmer behavior characteristics, and natural endowments. Current technologies fail to fully reveal the spatiotemporal interaction effects between urban and rural land use and topographic constraints, resulting in insufficient explanatory power for the complex process of uncultivated arable land.

[0003] Furthermore, existing technologies generally fail to simultaneously consider adjacency relationships (such as the land use contagion effect between villages), geographical distance (such as the impact of transportation accessibility on farmland transfer), and economic linkages (such as the traction effect of regional industrial agglomeration on farmland use). This leads to an inability to analyze the structural contradictions between specialty industrial land and grain production land, and to reveal the evolutionary patterns and key driving factors of uncultivated farmland through spatiotemporal data mining techniques. At the methodological level, the widespread application of traditional econometric models (such as ordinary least squares) struggles to capture the spatial heterogeneity of data. Although existing technologies attempt to introduce spatial autocorrelation analysis, the application of advanced spatial econometric methods such as geographically weighted regression and spatiotemporal autoregressive models remains in the exploratory stage. Moreover, the lack of multi-source heterogeneous data fusion technologies hinders the collaborative analysis of socioeconomic statistics, remote sensing monitoring data, and farmer survey data, limiting the refined identification of driving factors for uncultivated farmland. For example, while remote sensing image overlay analysis can identify suspicious areas, it lacks dynamic prediction capabilities.

[0004] On the other hand, there are currently few existing technologies in China for predicting the future trend of arable land remaining uncultivated. Existing technical solutions mainly predict the total area of ​​arable land in a certain region in the future, without predicting the uncultivated rate. Existing prediction technologies mainly predict arable land area by establishing multiple regression models, including using multiple regression and neural networks to build arable land area prediction models and comparing and evaluating these models; and using Markov models to predict the structural proportions and areas of arable land, forest land, grassland, water areas, construction land, and unused land.

[0005] In statistics, p-values ​​and q-values ​​are important indicators for measuring the significance of hypothesis testing results. The p-value represents the probability of the observed outcome or a more extreme outcome if the null hypothesis is true. The q-value is the p-value after correction. Current technologies largely focus on predicting changes in total arable land, lacking in-depth analysis of the dynamic evolution of the core indicator of uncultivated land rate. Because changes in total arable land cannot reflect deep adjustments in land use structure, such as the conversion of fragmented arable land into large-scale construction land, protection measures proposed based on total land analysis often fail to accurately address the regionally differentiated risks of arable land loss. For example, the uncultivated land patterns differ significantly across different terrains, requiring a zoned and categorized governance strategy, but existing methods struggle to achieve such refined analysis. Summary of the Invention

[0006] To address the problems in the existing technologies, this application proposes a data processing and analysis system for uncultivated farmland, comprising an influencing factor analysis model and a prediction model. The influencing factor analysis model includes multivariate analysis modules such as a geographic detector model module, a panel model module, a spatial model module, and a spatiotemporal geographic weighted regression model module. By performing modeling, solving, and statistical analysis on each module separately, common feature screening is conducted based on the significance indicators output by each module. Specifically, the p-value and q-value of the indicators are used, with the q-value representing the degree of influence as a quantifiable parameter. A multidimensional comprehensive judgment is also made based on the positive or negative direction of the influence, thereby achieving the scientific identification and precise extraction of key influencing factors. The prediction model is built upon the organic combination of panel data modeling methods and quadratic regression analysis techniques. By integrating the advantages of panel models in dynamically analyzing time series and cross-sectional data, and combining the fitting ability of quadratic regression for nonlinear relationships, a prediction analysis framework that combines temporal dynamic characteristics and spatial heterogeneity is formed. This invention constructs a comprehensive and detailed indicator system. Through multi-model collaboration, it conducts a comprehensive and in-depth analysis of the influencing factors of uncultivated arable land from multiple dimensions such as geospatial and time series, and realizes a comprehensive, refined quantitative analysis of the influencing factors of uncultivated arable land in a specified area and prediction of future trends.

[0007] The present invention provides a data processing and analysis system for uncultivated farmland, comprising a module for analyzing factors influencing the uncultivated rate and a module for predicting the uncultivated rate.

[0008] The module for analyzing factors affecting uncultivated rates integrates a geographic detector model module, a panel model module, a spatial model module, a spatiotemporal geographic weighted regression model module, and a comprehensive analysis module.

[0009] The uncultivated rate prediction module includes a panel model prediction module and a time t prediction module;

[0010] The geographic detector model module collects uncultivated index data and preprocesses missing values; classifies different data types; calculates the q and p values ​​of each variable based on the geographic detector; determines influencing factors through two key factor screening mechanisms; and obtains the influencing factors year by year.

[0011] The panel model module performs initial screening of uncultivated indicators and establishes different types of panel models through F-tests and / or t-tests and / or Hausman tests to obtain different significant influencing factors. The initial screening methods include screening based on the correlation between independent variables, screening based on the correlation between independent variables and uncultivated rate, and, based on the correlation between independent variables and uncultivated rate, successively removing variables that are weakly correlated with uncultivated rate; and, based on the correlation between independent variables and uncultivated rate, successively removing insignificant variables.

[0012] The spatial model module first uses five spatial weight matrices—spatial adjacency weight matrix, spatial inverse distance weight matrix, squared spatial inverse distance weight matrix, distance-based spatial adjacency matrix, and economic distance matrix—to sequentially perform spatial correlation tests on the uncultivated rate. If the tests pass, a spatial model is established; otherwise, no spatial model needs to be established.

[0013] The spatiotemporal weighted regression model module constructs an influencing factor system containing multiple primary indicators and multiple secondary indicators. It uses panel data of the uncultivated rate in the target region and the target year, selects highly correlated indicators through correlation coefficients, selects indicator variables by combining variance inflation factor, establishes a spatiotemporal weighted regression model and solves it to obtain the coefficients of each indicator in the target region.

[0014] The comprehensive analysis module organizes the results of the spatiotemporal geographic weighted regression model of the target area, filters significant features, and combines the solution results of each model to determine the degree of influence and positive or negative of the analysis results of each significant indicator, and finally obtains the results of significant influencing factors of uncultivated areas in the target area.

[0015] The panel model prediction module constructs different types of regression models based on the changing characteristics of the independent variables to make a preliminary prediction of the uncultivated rate for the next year. After completing the prediction of the independent variables, based on the solution results of the panel model and combined with the obtained predicted values ​​of the independent variables, a time series prediction model is constructed based on the parameter estimation results of the panel model to further predict the uncultivated rate of each administrative region for the next year.

[0016] The time t prediction module classifies regions into types using Euclidean clustering and hierarchical clustering. For all cities in each category, it establishes a dynamic prediction model for uncultivated rates based on the time dimension and performs regression prediction using a regression prediction mechanism. The regression models used include linear regression, quadratic regression, exponential regression, logarithmic regression, exponential decay, and trend extrapolation. A prediction model that fits each city category is selected to obtain the corresponding accurate predicted value for uncultivated rates.

[0017] In one implementation, the formulas for differentiation and factor detection in the geographic detector model module are as follows:

[0018] (1-1)

[0019] in, Represents the i-th factor The uncultivated rate of cultivated land in year t The explanatory power of is defined in the range [0,1]. The larger the value, the stronger the explanatory power of the i-th factor for the uncultivated rate of cultivated land in year t; This represents the classification of the i-th factor in year t, and is divided into several categories. If the factor is a continuous variable, it needs to be discretized; N represents the total number of samples in the entire region, and the total number of samples is the same every year. This represents the number of samples corresponding to the h-th class of the i-th factor in year t; This represents the variance of the uncultivated rate of cultivated land in the entire region in year t. Let represent the variance of the uncultivated farmland rate corresponding to the h-th category of the i-th factor in year t;

[0020] Will After a simple transformation, it will satisfy the following non-central F distribution:

[0021]

[0022] ;

[0023] in, For non-central parameters, This represents the average rate of uncultivated farmland corresponding to category h of factor i in year t.

[0024] Verifiable Is it significant? At a significance level of 0.1, if If the result is significant, then the factor is considered to be a factor influencing the fact that arable land is not cultivated.

[0025] In one implementation, the panel model module establishes different types of panel models through F-tests and / or t-tests and / or Hausman tests, specifically including mixed models, fixed effects models, and random effects models.

[0026] In one implementation, the spatiotemporal geographic weighted regression model module constitutes a spatiotemporal three-dimensional coordinate system. The model's expression is as follows:

[0027]

[0028] Where n is the number of sample points, and r is the number of independent variables. , , Let be the longitude, latitude, and time of the i-th sample point, respectively. It is the dependent variable of the i-th sample point. It is the k-th independent variable of the i-th sample point. For the intercept term, It is the regression coefficient of the k-th independent variable for the i-th sample point. It is the random error of the i-th sample point that is independently and identically normally distributed.

[0029] This application discloses a method for processing and analyzing data on uncultivated arable land, comprising the following steps:

[0030] Step 1: Analyze the data using the geographic detector module to obtain relevant indicators of uncultivated land rate. Collect indicator data and preprocess the data. Use the equal interval method, quantile method, and hierarchical clustering method to classify the data of one set of data corresponding to one indicator of uncultivated land in the target area. Use hierarchical clustering method to classify the data of multiple sets of data corresponding to one indicator of uncultivated land in the target area. After classifying the independent variables using different methods, calculate the q-value and p-value of each independent variable according to the geographic detector method. Determine the factors affecting uncultivated farmland each year and their explanatory power based on the p-value and / or q-value.

[0031] Step 2: Through panel model analysis, establish an indicator system for influencing factors and preprocess the indicator data; after processing some indicator data proportionally, conduct initial screening of indicators using four methods, and then establish different types of panel models using F-test and / or t-test and / or Hausman test; after processing some indicator data by area, conduct initial screening of indicators using four methods, and then establish different types of panel models using F-test and / or t-test and / or Hausman test; the four methods are: screening based on the correlation between independent variables; screening based on the correlation between independent variables and uncultivated rate; screening based on the correlation between independent variables and uncultivated rate, and then successively removing variables weakly correlated with uncultivated rate; screening based on the correlation between independent variables and uncultivated rate, and then successively removing insignificant variables.

[0032] Step 3: Through spatial panel analysis, obtain panel data on uncultivated land rate and influencing factors for the target region and target year interval. Preprocess the sample data to construct a spatial adjacency weight matrix, a spatial inverse distance weight matrix, the square of the spatial inverse distance weight matrix, a distance-based spatial adjacency matrix, and an economic distance matrix. Perform spatial autocorrelation tests on the uncultivated land rate for the target region and target year interval in sequence. Analyze the test results.

[0033] Step 4: Analyze the spatiotemporal weighted regression module to obtain panel data on uncultivated rates and influencing factors for the target region in the target year; construct an influencing factor indicator system including multiple primary and secondary indicators; filter the indicators based on the variance inflation factor; prepare a spatiotemporal weighted regression model; and obtain the coefficients of each indicator in the target region based on the established spatiotemporal weighted regression model.

[0034] Step 5: Through comprehensive analysis, the results of the spatiotemporal weighted regression model are analyzed and organized, and p-values ​​and q-values ​​are extracted from the relevant calculation results. Simultaneously, the outputs of the panel model, spatial model, and spatiotemporal weighted regression model are integrated, and the influence direction and degree of each indicator are recorded in detail. The integrated data is systematically organized, and significant indicator results are selected from the summary table based on the rounding up strategy. The selected significant influence indicators are then sorted in descending order according to their q-values. Finally, the results of significant influencing factors for uncultivated areas in the target region are obtained.

[0035] Step 6, Uncultivated Land Prediction Analysis; Based on panel model uncultivated land rate prediction, after completing the independent variable prediction, the uncultivated land rate corresponding to each administrative region in the next year is predicted using the solution results of the panel model and combined with the obtained independent variable prediction values; Uncultivated land rate prediction based on time t, regional types are divided through Euclidean clustering and hierarchical clustering; For all cities in each category, a dynamic prediction model of uncultivated land rate based on the time dimension is established and regression prediction is performed using a regression prediction mechanism; The regression models used include linear regression model, quadratic regression model, exponential regression model, logarithmic regression model, exponential decay model, and trend extrapolation; A prediction model that fits each city category is selected to obtain accurate uncultivated land rate prediction values; Corresponding evaluation criteria are constructed, and the accuracy of the prediction results is measured by calculating the absolute error and relative error of each uncultivated land rate prediction value.

[0036] In one implementation, the comprehensive analysis module analyzes, specifically including:

[0037] Step 5.1 Organize the spatiotemporal weighted regression model results. Based on the output data of the spatiotemporal weighted regression model, according to the pre-set three-level target region division rules, the significant indicators in each year are systematically merged, and finally, an independent and complete solution result data set for each target region is formed. First, the number of three-level target regions is accurately counted, and the significant indicators obtained are filtered based on the screening threshold. The filtered indicators and their related data are arranged and recorded in an orderly manner according to the preset data format and arrangement rules, and finally, a spatiotemporal weighted regression model solution result table is constructed.

[0038] Step 5.2 Summarize the model solution results, specifically including:

[0039] Step 5.2.1: Extract p-values ​​and q-values ​​from the relevant calculation results. Simultaneously, comprehensively integrate the outputs of the panel model, spatial model, and spatiotemporal geographic weighted regression model; record in detail the direction and degree of influence of each indicator; and construct a summary table of the influencing factor model solution results using the integrated data.

[0040] Step 5.2.2: This table summarizes the results of each model and related indicator information; based on the rounding up strategy, the significant indicator results are selected from the summary table.

[0041] Step 5.2.3: Sort the selected significant impact indicators in descending order according to their q values. At the same time, classify them into the positive impact factor set and the negative impact factor set according to the direction of each indicator's influence.

[0042] Step 5.2.4: Using the rounding-up strategy, calculate the new screening threshold and screen for indicators with significant impact;

[0043] Step 5.2.5: Use only the q-value results as the screening criterion to select indicators that have a key impact on the target variable;

[0044] Step 5.2.6: Systematically organize the results obtained from the screening and analysis in the above steps, and finally obtain the results of the significant influencing factors of the uncultivated areas in the target region.

[0045] In one implementation, step 6, uncultivated prediction analysis, specifically includes:

[0046] Step 6.1, based on panel model prediction of uncultivated rate, firstly, the independent variables are predicted; based on the changing characteristics of each independent variable, different types of regression models are constructed to achieve a preliminary prediction of the uncultivated rate for the next year; based on the completion of the independent variable prediction, the uncultivated rate of each administrative region for the next year is further predicted by using the solution results of the panel model and combining the obtained predicted values ​​of the independent variables.

[0047] Step 6.2, prediction of uncultivated rate based on time t, includes:

[0048] Step 6.2.1: Transform the three years of uncultivated data for each city into... ,in, This represents the uncultivated rate of the i-th city in year t. This represents the uncultivated rate of the i-th city in year t+1. This represents the uncultivated rate of the i-th city in year t+2;

[0049] Step 6.2.2, city cluster analysis operation, using Euclidean clustering for classification; at the same time, hierarchical clustering is used to select clusters to be merged by minimizing the objective function;

[0050] Step 6.2.3, Regression Prediction: For all cities in each category, regression prediction is performed separately. The regression models used include six types: linear regression, quadratic regression, exponential regression, logarithmic regression, exponential decay, and trend extrapolation. Through the above operations, each city can obtain 6 predicted values ​​for the uncultivated rate.

[0051] Step 6.2.4, Model selection: Select a suitable prediction model for each city category to obtain an accurate predicted value of the uncultivated rate;

[0052] Step 6.3: Establish evaluation criteria. Construct corresponding evaluation criteria and measure the accuracy of the prediction results by calculating the absolute error and relative error of each uncultivated rate prediction value.

[0053] The present invention provides a data processing and analysis system and method for uncultivated arable land, which, compared with the prior art, has at least the following beneficial effects:

[0054] By establishing a unified spatiotemporal data warehouse and integrating heterogeneous data indicators from multiple sources, such as land use, climate environment, production conditions, and natural factors, and through the construction of a multi-dimensional analysis framework and the organic integration of a five-dimensional model, a comprehensive analysis method integrating influencing factor diagnosis, spatial differentiation analysis, and time series trend prediction is formed. The analysis results will be more effective, comprehensive, and complete. Compared with traditional analysis methods, the accuracy of factor identification is significantly improved, the accuracy of spatial prediction is significantly improved, and the error of time series prediction is significantly reduced.

[0055] The geographic detector model provides a more detailed classification of indicators, with a more comprehensive and detailed classification of primary and secondary indicators, resulting in more detailed and effective analysis results. The panel model automatically generates differentiated variable combinations to address the specific circumstances of different regions, which significantly reduces the model's prediction error rate and greatly improves the relevance of regional analysis.

[0056] For the first time, the spatial model utilizes five spatial weight matrices, simultaneously considering three dimensions: adjacency and geography, thus avoiding single-variability test errors and significantly improving effectiveness and reliability. The accuracy of spatial autocorrelation detection is significantly improved, effectively avoiding test bias. By combining geospatial coordinates with time and integrating temporal and spatial factors, more precise analysis of variables in different regions is achieved. The prediction model first predicts the independent variables, then uses the dependent variable to predict the independent variables, and uses variations in time scale and independent variables to enrich the data and avoid data gaps. A two-way prediction model for time series and spatial patterns is established. By dynamically adjusting the time window length and spatial grid precision, the prediction bias caused by data gaps is addressed, significantly improving data utilization. Attached Figure Description

[0057] The invention will now be described in more detail with reference to embodiments and the accompanying drawings.

[0058] Figure 1 This diagram illustrates the complete technical solution flow of the present invention.

[0059] Figure 2 This diagram illustrates the panel model selection process of the present invention.

[0060] Figure 3 The flowchart of the geographic detector module of the present invention is shown;

[0061] Figure 4 The flowchart of the panel model module of the present invention is shown;

[0062] Figure 5The flowchart of the spatial panel module and the spatiotemporal geographic weighted regression module of the present invention is shown. Detailed Implementation

[0063] The invention will now be further described with reference to the accompanying drawings.

[0064] This invention provides a general system and method for analyzing and predicting factors affecting uncultivated land rates.

[0065] In one embodiment, such as Figure 1 As shown, the analysis model for factors affecting the uncultivated rate first includes the following modules:

[0066] 1. Geographic Detector Model Module: First, collect data on uncultivated indicators and preprocess missing values. Classification is implemented for different data types: for single-indicator, single-group data, nine classification combinations are formed using equal intervals, quantiles, and hierarchical clustering; for single-indicator, multi-group data, only hierarchical clustering is used for classification. Based on the principle of the geographic detector, the q-values ​​and p-values ​​of each variable are calculated. Influencing factors are determined through two key factor screening mechanisms: one is to sort by q-value when p < 0.1; the other is to directly sort by q-value to obtain the annual influencing factors and their explanatory power.

[0067] 2. Panel Model Module: Based on preprocessed data from the geographic detector, and addressing issues of strong correlation and multicollinearity among indicators, two data transformation modes (proportional data / area data) are designed, along with four variable selection strategies: 1. Correlation coefficient method between independent variables; 2. Correlation coefficient method between independent variables and uncultivated rate; 3. Removal of weakly correlated variables based on strategy 2; 4. Removal of variables with insignificant regression based on strategy 2. The time-fixed effects model is determined using F-test, t-test, and Hausman test to identify different significant influencing factors.

[0068] 3. Spatial Model Module: Innovatively constructs five types of spatial weight matrices (spatial adjacency, inverse distance, squared inverse distance, distance-based spatial adjacency, and economic distance matrix), simultaneously considering three dimensions: adjacency relationship, geographical distance, and economic correlation, to achieve multi-dimensional testing of the spatial effect of uncultivated rate;

[0069] 4. Spatiotemporal Geographic Weighted Regression Model Module: Constructs an influencing factor system containing 5 primary indicators and 42 secondary indicators. Highly correlated indicators are selected through correlation coefficients, and multicollinearity is eliminated by combining variance inflation factor. The spatial distribution characteristics and influencing mechanisms of uncultivated rate are analyzed from the spatiotemporal dimension.

[0070] 5. Comprehensive Analysis: The results of the spatiotemporal geographic weighted regression model of the tertiary target areas under the primary target area are compiled. Significant features are selected according to the screening criteria. Then, based on the solution results of different models, the influence degree and positive or negative of each significant indicator are judged, and finally a comprehensive conclusion is drawn.

[0071] Secondly, the uncultivated rate prediction module includes the following modules:

[0072] Panel model prediction module: integrates independent variable prediction screening mechanism and prediction mechanism, and constructs time series prediction model based on panel model parameter estimation results.

[0073] Time t prediction module: It classifies regional types through cluster analysis and establishes a dynamic prediction model of uncultivated rate based on time dimension by combining regression prediction mechanism.

[0074] In one embodiment, the detailed description of each module is as follows:

[0075] Geographic detectors are a statistical method used to detect spatial heterogeneity and reveal its underlying driving forces. The basic idea is that if an independent variable has a significant impact on a dependent variable, then the spatial distributions of the independent and dependent variables should tend to be consistent.

[0076] Different administrative regions have different uncultivated rates, and the economic level, population factors, and production conditions corresponding to different administrative regions also have certain differences. The independent and dependent variables have their own spatial distribution characteristics. Therefore, a geographic detector is used to find the factors that affect the uncultivated rate of arable land.

[0077] In one embodiment, such as Figure 2 As shown, it includes a panel model module, which mainly includes mixed models, fixed effects models and random effects models.

[0078] In one embodiment, a spatial model module is included. For the spatial model, five spatial weight matrices are first used to perform spatial correlation tests on the uncultivated rate in sequence: spatial adjacency weight matrix, spatial inverse distance weight matrix, square of spatial inverse distance weight matrix, distance-based spatial adjacency matrix, and economic distance matrix.

[0079] Spatial models refer to the introduction of spatial factors into traditional econometrics. Before building a spatial model, a spatial weight matrix needs to be established, followed by a spatial autocorrelation test. If the test passes, a spatial model is built; otherwise, it is not necessary. The global Moran index is usually used to reflect the spatial clustering of the entire study area.

[0080] For spatial models with fixed effects, the within-group deviation transformation method is first used to remove individual or time effects, and then maximum likelihood estimation is performed. For spatial models with random effects, the generalized deviation transformation is first performed, and then maximum likelihood estimation is performed.

[0081] In one embodiment, a comprehensive analysis module is included. Since the spatiotemporal weighted model performs analysis year by year and yields the influencing factors of uncultivated farmland in third-level target areas, it cannot be compared with the results of other models. Therefore, data interface standardization is required. If a variable significantly affects the uncultivated rate of an administrative region in any year, this variable is considered an influencing factor of uncultivated farmland in that administrative region. After processing, the solution results for each third-level target area of ​​the spatiotemporal weighted model can be obtained. Next, significant indicators affecting more than half of the targets are selected, yielding the solution selection results of the spatiotemporal weighted model, which are then incorporated into the comprehensive analysis framework. Finally, based on the p-value and q-value analysis of the geographic detector model, indicators considered significant in more than half of the modules in the influencing factor model are selected category by category, sorted according to q-value, and classified according to their direction of influence, ultimately yielding the comprehensive analysis results of the uncultivated farmland influencing factors.

[0082] In one embodiment, the system includes an uncultivated land prediction module and a panel model uncultivated land prediction module. To effectively address the challenge of predicting the uncultivated land rate, this application employs a panel model for relevant prediction operations. In this process, the prediction of independent variables is initiated first. Based on the changing characteristics of each independent variable, different types of regression models are constructed to achieve a preliminary prediction of the uncultivated land rate for the following year. Building upon the completed independent variable predictions, the results of the panel model, combined with the obtained predicted values ​​of the independent variables, allow for further prediction of the uncultivated land rate for each administrative region in the following year.

[0083] In one embodiment, a time-t-based uncultivated rate prediction module is included. When performing time-t-based uncultivated rate prediction, Euclidean clustering is first used to accurately measure the differences between cities. For all cities in each category, regression prediction is performed separately. The regression models used cover six types: linear regression, quadratic regression, exponential regression, logarithmic regression, exponential decay, and trend extrapolation. Through these operations, six uncultivated rate prediction values ​​can be obtained for each city. Based on this, a suitable prediction model is carefully selected for each city category to obtain the corresponding accurate uncultivated rate prediction value.

[0084] In one embodiment, an uncultivated factor analysis module is included, such as... Figure 3 As shown, the first step, the analysis of the geographic detector module, includes step 1.1: the construction of the influencing factor index system and data preprocessing.

[0085] Step 1.1.1 Based on the current research status of domestic and international studies on the influencing factors of uncultivated arable land, this application constructs an indicator system for the influencing factors of uncultivated arable land from five aspects: economic level, population factors, production conditions, policy regulation, and natural factors. Among these, population mobility refers to the difference between the year-end registered population and the year-end resident population; the proportion of employees in the primary industry refers to the ratio of employees in the primary industry to the total employed population; total grain output per unit area refers to the ratio of total grain output to total sown area; and distance from the city center refers to the driving distance from the city center of each prefecture-level city to the city center of the provincial capital.

[0086] Step 1.1.2 Collect indicator data. Through data collection and organization, data for 50 indicators were obtained. Since all 50 indicators are continuous variables, the independent variables first need to be discretized. Based on the data type, the 50 indicators can be divided into two categories: one category has only one set of data for each indicator, such as per capita GDP, rural employment, and arable land water resource coverage rate; this category contains 40 indicators. The other category has multiple sets of data for each indicator, such as arable land grade (excellent / high / medium / low), soil texture (loam / clay / sandy), and soil biodiversity (abundant / average / not abundant); this category contains 10 indicators.

[0087] Step 1.1.3 Data Preprocessing. For missing data, imputation methods are used to fill in all missing values.

[0088] Step 1.2 Data discretization for a single indicator corresponding to a set of data:

[0089] Step 1.2.1 Classify continuous index data using the equal interval method (divided into 3 / 4 / 5 categories);

[0090] The data of an indicator type corresponding to a set of data for an uncultivated indicator in the target area is classified into three, four and five categories with equal intervals. The classification results of the uncultivated indicators in the target area using the equal interval method are shown. The classification results of the equal interval method are greatly affected by extreme values, which may reduce the accuracy and reliability of the results of the geographic detector model.

[0091] Step 1.2.2 Quantile method (divided into 3 / 4 / 5 categories) classification index data;

[0092] The data of an indicator type corresponding to a set of data related to uncultivated land in the target area is classified by quantile method into 3, 4 and 5 categories. The classification results of uncultivated land indicators in the target area are shown by quantile method. The classification results of quantile method are not affected by extreme values ​​and the number of samples in each category is relatively uniform, and there will be no case where the variance of a certain category is 0. However, this classification method cannot well reflect the distribution characteristics of the data.

[0093] Step 1.2.3 Systematic clustering (classified into 3 / 4 / 5 categories) classification index data;

[0094] Systematic clustering with 3, 4, and 5 classes was used to classify the data of an indicator type corresponding to a set of data for an uncultivated indicator in the target area. The classification results of uncultivated indicators in various regions in 2022 using the quantile method are shown. Systematic clustering results in small intra-class differences and large inter-class differences. However, like the equal interval method, some categories correspond to only one sample, which may reduce the accuracy and reliability of the geographic detector model results.

[0095] Step 1.3 Discretization of data corresponding to multiple sets of data for one indicator;

[0096] The system clustering method, which classifies data into three, four, and five categories, is used to classify the data of multiple sets of data corresponding to one indicator of uncultivated land in the target area. The classification results of the uncultivated land indicators in the target area using the quantile method are presented.

[0097] Step 1.4 Geographic Detector Model Construction:

[0098] After classifying the independent variables using different methods, the q-value and p-value of each independent variable are calculated sequentially according to the principle of the geospatial detector. After classifying continuous index data using the equal interval method, the classified data is substituted into the geospatial detector model, and the q-value and p-value of each independent variable are calculated sequentially, yielding the solution results for some variables in the target region. Similarly, after classifying continuous index data using the quantile method, the classified data is substituted into the geospatial detector model, and the q-value and p-value of each independent variable are calculated sequentially, yielding the solution results for some variables in the target region. Finally, after classifying continuous index data using the hierarchical clustering method, the classified data is substituted into the geospatial detector model, and the q-value and p-value of each independent variable are calculated sequentially, yielding the solution results for some variables in the target region.

[0099] Step 1.5 Identify the key factors affecting the uncultivated rate;

[0100] Step 1.5.1 Determine the results of the geospatial detector based on the p-value and q-value:

[0101] Based on the principle of geographic detectors, the data for each target year is solved sequentially, and then the results obtained from different classification methods are compared and organized. If there is more than one classification method that results in a p-value less than 0.1 for a certain variable in a given year, the classification method with the largest q-value is selected. The factors affecting uncultivated arable land and their explanatory power for each year are determined by ranking the q-values ​​corresponding to p-values ​​less than 0.1.

[0102] Step 1.5.2 Determine the result based solely on the q-value:

[0103] If we do not perform a significance test on the q-value, it still has physical meaning. Therefore, we only focus on the q-value, and when q > 0.5, we choose the classification method with the largest q-value. We then rank the factors affecting uncultivated farmland each year based on their q-values, determining their explanatory power.

[0104] In one embodiment, such as Figure 4 As shown, it includes panel model module analysis, step 2.1 construction of the influencing factor index system and data preprocessing.

[0105] Step 2.2 Some indicator data are processed in two ways: proportion and area.

[0106] Step 2.2.1 Data for some indicators are processed proportionally. The seven indicators—the proportion of arable land within the arable land protection target area, arable land water resource coverage rate, arable land urban and rural road coverage rate, arable land rural road coverage rate, arable land rural settlement coverage rate, arable land urban agglomeration area coverage rate, and soil pH value—are represented by their corresponding arable land area proportions. Some similar indicators are subjectively filtered out, and then an influencing factor indicator system is constructed.

[0107] Step 2.2.2 Some indicator data are processed by area. The seven indicators—the proportion of arable land within the arable land protection target area, arable land water resource coverage rate, arable land urban and rural road coverage rate, arable land rural road coverage rate, arable land rural settlement coverage rate, arable land urban agglomeration area coverage rate, and soil pH value—are expressed using their corresponding arable land areas. Some similar indicators are subjectively filtered out, and then an influencing factor indicator system is constructed.

[0108] Step 2.3 Establish the model based on the indicators corresponding to the factors in proportion.

[0109] After processing a portion of the indicator data proportionally, the indicators are initially screened within the indicator system. Through various tests, different types of panel models are established.

[0110] Method 1 for initial screening of indicators: Screening based on the correlation between independent variables;

[0111] Step 2.3.1 Filter based on the correlation between independent variables. If the correlation coefficient is greater than 0.7, it indicates a strong correlation between the two indicators. Only retain the indicator that has a greater impact on the uncultivated farmland rate. See the correlation coefficient graph for the remaining indicators after filtering.

[0112] Step 2.3.2 Perform F-tests and t-tests. The F-test determines whether to use a mixed model or a fixed-effects model. The F-test results show that only the p-value for the time effect is less than 0.05, indicating the existence of a time effect and no individual effect. The t-test is used to test the significance of each time effect; the corresponding p-values ​​are generally less than 0.05, indicating the existence of a time effect. The Hausman test is used to determine whether it is a fixed effect or a random effect.

[0113] Step 2.3.3 Determine the type of panel model to be built. Based on the F-test and Hausman test, a time-fixed effects model should be used.

[0114] Step 2.3.4 Solve the model and obtain the results.

[0115] Method 2 for initial screening of indicators: Screening based on the correlation between independent variables and uncultivated rates;

[0116] Step 2.3.1 For screening based on the correlation between the independent variable and the uncultivated rate, when the correlation between two factors is greater than 0.7, their correlation coefficients with the uncultivated rate can be calculated, and factors with higher correlation coefficients with the uncultivated rate are retained. Correlation coefficient graph of the remaining indicators after screening;

[0117] Step 2.3.2 Perform F-tests and t-tests. The F-test determines whether a mixed-effects model or a fixed-effects model is used. The F-test results show that only the time effect has a p-value less than 0.05, indicating the presence of a time effect and no individual effect. The t-test is used to test the significance of each time effect; the corresponding p-values ​​are generally less than 0.05, indicating the presence of a time effect.

[0118] Step 2.3.3 Use the Hausman test to determine whether it is a fixed effect or a random effect;

[0119] Step 2.3.4 Determine the type of panel model to be built. Based on the F-test and Hausman test, a time-fixed effects model should be used.

[0120] Step 2.3.5 Solve the model and obtain the results.

[0121] Since the two methods mentioned above yield relatively few manifest variables affecting uncultivated land, we consider using a combination of strong correlation coefficients between independent variables and weak correlation coefficients with the uncultivated land rate to screen variables.

[0122] Method 3 for initial screening of indicators: Based on method 2, remove insignificant variables one by one.

[0123] Step 2.3.1 After retaining one of the indicators with a correlation coefficient greater than 0.7 with the uncultivated rate, which has a greater impact on the uncultivated rate of cultivated land, remove the variables with insignificant regression coefficients in the regression results in turn;

[0124] Steps 2.3.2-2.3.5 perform F-tests and Hausman tests to determine the panel model to use and solve the model. After each solution, remove the variable with the largest p-value until all remaining variables are significant, and then obtain the regression results.

[0125] Step 2.4 Establish the model based on the area index corresponding to the factors.

[0126] First, using some of the obtained indicator data, a corresponding data visualization chart is constructed. Second, based on the corresponding image area in the chart, and combined with the indicator system and actual needs, the indicators are initially screened in four ways within the established indicator system. Finally, through various verification methods, different types of corresponding panel models are established.

[0127] Method 1 for initial screening of indicators: Screening based on the correlation between independent variables;

[0128] Step 2.4.1 Filtering is performed based on the correlation between independent variables. If the correlation coefficient is greater than 0.7, it indicates a strong correlation between the two indicators. Only the indicator that has a greater impact on the uncultivated farmland rate is retained. The correlation coefficient graph of the remaining indicators after filtering is shown.

[0129] Step 2.4.2 Perform F-tests and t-tests. The F-test determines whether a mixed-effects model or a fixed-effects model is used. The F-test results show that only the p-value for the time effect is less than 0.05, indicating the presence of a time effect and no individual effect. The t-test is used to test the significance of each time effect; the corresponding p-values ​​are generally less than 0.05, indicating the presence of a time effect.

[0130] Step 2.4.2 Perform the Hausman test. The Hausman test is then used to determine whether it is a fixed effect or a random effect. The result shows that p < 0.05 at a significance level of 0.05, therefore a fixed effects model should be used.

[0131] Step 2.4.3 Determine the type of panel model to be built. Based on the F-test and Hausman test, a time-fixed effects model should be used.

[0132] Step 2.4.4 Solve the model and obtain the results.

[0133] Method 2 for initial screening of indicators: Screening based on the correlation between independent variables and uncultivated rates;

[0134] Step 2.4.1 For screening based on the correlation between the independent variable and the uncultivated rate, when the correlation between two factors is greater than 0.7, their correlation coefficients with the uncultivated rate can also be calculated, and factors with higher correlation coefficients with the uncultivated rate can be retained.

[0135] Step 2.4.2 Perform F-tests and t-tests. The F-test determines whether a mixed-effects model or a fixed-effects model is used. The F-test results show that only the p-value for the time effect is less than 0.05, indicating the presence of a time effect and no individual effect. The t-test is used to test the significance of each time effect; the corresponding p-values ​​are generally less than 0.05, indicating the presence of a time effect.

[0136] Step 2.4.3 Perform the Hausman test. The Hausman test is then used to determine whether it is a fixed effect or a random effect. The result shows that p < 0.05 at a significance level of 0.05, therefore a fixed effects model should be used.

[0137] Step 2.4.4 Determine the type of panel model to be built. Based on the F-test and Hausman test, a time-fixed effects model should be used.

[0138] Step 2.4.5 Solve the model and obtain the results.

[0139] Since the two methods mentioned above yield relatively few manifest variables affecting uncultivated land, we consider using a combination of strong correlation coefficients between independent variables and weak correlation coefficients with the uncultivated land rate to screen variables.

[0140] Method 3 for initial screening of indicators: Based on method 2, remove insignificant variables one by one;

[0141] Step 2.4.1 After retaining one of the indicators with a correlation coefficient greater than 0.7 with the uncultivated rate, which has a greater impact on the uncultivated rate of cultivated land, remove the variables with insignificant regression coefficients in the regression results in turn;

[0142] Steps 2.4.2-2.4.5 perform F-tests and Hausman tests to determine the panel model to use and solve the model. After each solution, remove the variable with the largest p-value until all remaining variables are significant, and then obtain the regression results.

[0143] Step 2.5 Identify the key factors affecting the uncultivated rate.

[0144] A panel model was established to select significant variables related to uncultivated arable land. It was found that in order to retain more variables while achieving better results, proportional data should be used. After preliminary screening based on the correlation between independent variables and uncultivated land rate, variables that are not significant and have the largest p-values ​​were removed in turn.

[0145] In one embodiment, such as Figure 5 The flowcharts for the Spatial Panel Module and the Spatiotemporal Geographic Weighted Regression Module are shown. The analysis in the Spatial Panel Module includes:

[0146] Step 3.1 Sample selection: Obtain non-quantitative rate panel data and influencing factor panel data for the target region and target year interval from relevant departments or official data websites;

[0147] Step 3.2 Data Preprocessing; Preprocess the sample data. For example, in the indicator system used by the geographic detector model, cultivated land elevation and cultivated land slope each correspond to multiple sets of data for one indicator. Before solving the panel model, these multiple sets of data for these indicators need to be merged into one set. Multiply the median of each interval by the proportion of cultivated land area corresponding to that interval, and then add the results from different intervals.

[0148] Step 3.3 Construct five types of spatial weight matrices: spatial adjacency weight matrix, spatial inverse distance weight matrix, square of spatial inverse distance weight matrix, distance-based spatial adjacency matrix, and economic distance matrix.

[0149] Step 3.4 Spatial autocorrelation test; Using the five spatial weight matrices, perform spatial autocorrelation tests on the uncultivated farmland rate in the target region for each target year interval;

[0150] Step 3.5 Analysis of test results: When using a spatial adjacency matrix, if the p-value of the global Moran index for the uncultivated rate of the target year is less than 0.05, but the p-value of the global Moran index for the other weight matrices and years is greater than 0.05, it means that there is no spatial autocorrelation.

[0151] In one embodiment, analysis is performed using a spatiotemporal geographic weighted regression module;

[0152] Step 4.1 Sample selection: Obtain panel data of non-quantitative rates for the target region and panel data of influencing factors for the target year from relevant departments or official data websites;

[0153] Step 4.2 Data Preprocessing; Preprocess the sample data. For example, in the indicator system used by the geographic detector model, cultivated land elevation and cultivated land slope each correspond to multiple sets of data for one indicator. Before solving the panel model, these multiple sets of data for these indicators need to be merged into one set. Multiply the median of each interval by the proportion of cultivated land area corresponding to that interval, and then add the results from different intervals.

[0154] Step 4.3 Construct an indicator system and screen influencing factors;

[0155] Step 4.3.1 Preliminary construction of the influencing factor indicator system;

[0156] The preliminary influencing factor index system constructed by this model includes 5 primary indicators and 24 secondary indicators;

[0157] Step 4.3.2 Screen influencing factors based on the correlation coefficient between the indicator variable and the dependent variable;

[0158] To reduce the impact of multicollinearity, the correlation between indicator variables needs to be reduced, requiring the removal of variables with strong correlations. When the correlation between two indicator variables is greater than 0.7, the correlation coefficient between each indicator variable and the uncultivated rate is calculated, and the indicator variable with the larger correlation coefficient to the uncultivated rate is retained, thereby reducing multicollinearity. A graph showing the uncultivated rate with each indicator variable and the correlation coefficients between each indicator variable is plotted; when the correlation coefficient between two indicators is greater than 0.7, the indicator variable with the larger correlation coefficient to the uncultivated rate is retained, and the others are deleted.

[0159] Step 4.3.3: Screening is performed based on the variance inflation factor between the indicator variables;

[0160] Step 4.4 Preparation of the spatiotemporal geographic weighted regression model;

[0161] Step 4.5 Establish a spatiotemporal geographic weighted regression model;

[0162] The spatiotemporal geographic weighted regression model incorporates time and space dimensions into the regression model, fully considering temporal and spatial factors and enabling more precise analysis of variables from different regions. This model combines existing geographic spatial coordinates with time to form a three-dimensional spatiotemporal coordinate system. .

Claims

1. A data processing and analysis system for uncultivated arable land, characterized in that, Includes a module for analyzing factors affecting uncultivated land rate and a module for predicting uncultivated land rate; The analysis module for factors affecting the uncultivated rate includes a geographic detector model module, a panel model module, a spatial model module, a spatiotemporal geographic weighted regression model module, and a comprehensive analysis module. The uncultivated rate prediction module includes a panel model prediction module and a time t prediction module; The geographic detector model module collects uncultivated data and preprocesses missing values; classifies different data types; calculates the q-values ​​and p-values ​​of each variable based on the geographic detector, determines influencing factors, and obtains annual influencing factors. The panel model module performs initial screening of uncultivated data and establishes different types of panel models by using F-test and / or t-test and / or Hausman test to obtain different significant influencing factors. The spatial model module uses the spatial adjacency weight matrix, the spatial inverse distance weight matrix, the square of the spatial inverse distance weight matrix, the distance-based spatial adjacency matrix, and the economic distance matrix to sequentially perform spatial correlation tests on the uncultivated rate; if the tests pass, a spatial model is established; if the tests fail, a spatial model is not established. The spatiotemporal weighted regression model module constructs an influencing factor system containing multiple primary indicators and multiple secondary indicators. It uses panel data of the uncultivated rate of the target region and the target year, filters related indicators through correlation coefficients, and filters indicator variables by combining variance inflation factor. It then establishes and solves the spatiotemporal weighted regression model to obtain the coefficients of each indicator in the target region. The comprehensive analysis module organizes the results of the spatiotemporal geographic weighted regression model of the target area, filters significant features, and combines the solution results of each model to determine the degree of influence of each significant indicator, and finally obtains the results of significant influencing factors of uncultivated areas in the target area. The panel model prediction module constructs different types of regression models based on the changing characteristics of each variable to make a preliminary prediction of the uncultivated rate for the following year. After completing the prediction of independent variables, based on the solution results of the panel model and combined with the obtained predicted values ​​of independent variables, a time series prediction model is constructed based on the parameter estimation results of the panel model to further predict the uncultivated rate of each region in the following year. The time t prediction module classifies regions into types using Euclidean clustering and hierarchical clustering. For all cities in each category, it establishes a dynamic prediction model for uncultivated rates based on the time dimension by combining a regression prediction mechanism, and performs regression prediction. The regression models used include linear regression, quadratic regression, exponential regression, logarithmic regression, exponential decay, and trend extrapolation. A prediction model that fits each city category is selected to obtain the corresponding predicted value of uncultivated rates.

2. The data processing and analysis system for uncultivated arable land according to claim 1, characterized in that, Different types of panel models include mixed models, fixed effects models, and random effects models.

3. The data processing and analysis system for uncultivated arable land according to claim 1, characterized in that, The spatiotemporal three-dimensional coordinates in the spatiotemporal geographic weighted regression model module are: The expression for the spatiotemporal geographic weighted regression model is as follows: ; Where n is the number of sample points, and r is the number of independent variables. , , Let be the longitude, latitude, and time of the i-th sample point, respectively. It is the dependent variable of the i-th sample point. It is the k-th independent variable of the i-th sample point. For the intercept term, It is the regression coefficient of the k-th independent variable for the i-th sample point. It is the random error of the i-th sample point that is independently and identically normally distributed.

4. A method for processing and analyzing data on uncultivated arable land, characterized in that, The data processing and analysis system for uncultivated arable land based on any one of claims 1-3 includes the following steps: Step 1: Analyze and collect relevant indicator data on uncultivated land rate through the geographic detector module, preprocess the indicator data, and classify the indicator types of uncultivated land data in the target area using the equal interval method, quantile method, and hierarchical clustering method; calculate the q value and p value of each independent variable in turn according to the geographic detector method; determine the factors affecting uncultivated farmland in each year and their explanatory power based on the p value and / or q value. Step 2: Analyze the panel model module to establish an index system of influencing factors, and establish different types of panel models through F-test and / or t-test and / or Hausman test. Step 3: Analyze the spatial model module to obtain panel data on uncultivated land rates and influencing factors for the target region and target year. Use the spatial adjacency weight matrix to perform spatial autocorrelation tests on the uncultivated land rates for the target region and target year. Step 4: Analyze using the spatiotemporal weighted regression module to obtain panel data on uncultivated rates and influencing factors for the target region and target year; construct an influencing factor indicator system including multiple primary and secondary indicators; filter indicators based on variance inflation factors; and construct a spatiotemporal weighted regression model to obtain the coefficients of each indicator in the target region. Step 5: Analyze and organize the results data output by the spatiotemporal weighted regression model through the comprehensive analysis module, extracting p-values ​​and q-values. Simultaneously, integrate the output results of the panel model, spatial model, and spatiotemporal weighted regression model, recording the influence direction and degree of each indicator in detail. Systematically organize the integrated data, using an upward rounding strategy to filter out significant indicators. Sort the selected significant indicators in descending order based on their q-values. Finally, obtain the results of significant influencing factors for uncultivated areas in the target region. Step 6, Uncultivated Farming Prediction Analysis; Uncultivated Farming Rate Prediction Based on Panel Model: Based on the prediction of independent variables, the uncultivated farming rate of each administrative region in the next year is predicted by using the solution results of the panel model and combining the obtained predicted values ​​of independent variables; Uncultivated Farming Rate Prediction Based on Time t: Regional Types are Divided by Euclidean Clustering and Hierarchical Clustering. For all cities in each category, a dynamic prediction model for uncultivated rate based on the time dimension is established by combining a regression prediction mechanism and regression prediction is performed; corresponding evaluation criteria are constructed, and the accuracy of the prediction results is measured by calculating the absolute error and relative error of each uncultivated rate prediction value.

5. The method for processing and analyzing uncultivated farmland data according to claim 4, characterized in that, Based on the results data output by the spatiotemporal geographic weighted regression model, and according to the pre-defined three-level target region division rules, the significant indicators in each year are systematically merged and processed to finally form an independent and complete set of solution results data for each target region. First, the number of three-level target regions is accurately counted, and the significant indicators obtained are filtered based on the screening threshold. The selected indicators and their related data are arranged and recorded in an orderly manner according to the preset data format and arrangement rules, and finally a spatiotemporal geographic weighted regression model solution result table is constructed.

Citation Information

Patent Citations

  • Decision optimization method of multi-land seed selection based on combination optimization

    CN107832892A

  • Forecasting national crop yield during the growing season using weather indices

    US20170213141A1