Agricultural productivity evaluation method based on Newey-West standard error

By using the Newey-West standard error corrected OLS model, the problems of autocorrelation and heteroscedasticity in agricultural time series data were solved, enabling accurate assessment of the relationship between agricultural inputs and crop yields, and improving the scientific rigor and reliability of agricultural production optimization.

CN121503902APending Publication Date: 2026-02-10NORTHEAST INST OF GEOGRAPHY & AGRIECOLOGY C A S
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511688097.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-18
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively handle autocorrelation and heteroscedasticity issues in agricultural time series data, leading to decreased reliability of OLS regression results and an inability to accurately identify the causal relationship between agricultural inputs and crop yields.

Method used

An OLS regression model based on the Newey-West standard error was adopted. By correcting the standard error of traditional OLS, the autocorrelation and heteroscedasticity of time series data were handled, a robust covariance matrix was constructed, and the t-statistic and p-value were calculated to assess agricultural productivity.

Benefits of technology

Accurate assessment of the impact of agricultural mechanization and chemical inputs on the yield of different crops provides a scientific basis for optimizing agricultural production and improves the accuracy and reliability of quantitative analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121503902A_ABST
    Figure CN121503902A_ABST
Patent Text Reader

Abstract

The invention relates to an agricultural productivity evaluation method based on Newey-West standard error. Relates to the technical field of agricultural economic data processing and metering analysis, in particular to an agricultural productivity evaluation method based on Newey-West standard errors. By correcting a traditional standard error, synchronous processing of time sequence data autocorrelation and heterovariance is realized. The method comprises the following steps: obtaining crop yield data over the years of each province and agricultural input data over the years of each province; constructing a least square regression model by taking sum as a dependent variable and an independent variable; calculating a vector estimation value of a regression coefficient; constructing a robust covariance matrix, and standardizing error SE; the constructed estimation covariance matrix is corrected and obtained; and calculating a t statistic and a P value according to the SE, and evaluating the agricultural productivity according to the t statistic and the P value.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of agricultural economic data processing and measurement analysis, and particularly relates to an agricultural productivity evaluation method based on Newey-West standard error. BACKGROUND

[0002] In agricultural economic research, accurately identifying the causal relationship between agricultural inputs (such as agricultural mechanization and chemical inputs) and crop yield is the core prerequisite for guiding the optimal allocation of agricultural production resources and improving crop yield per unit. In the current mainstream quantitative analysis methods, OLS regression is widely used in empirical research on the relationship between "input-output" due to its intuitive principle and simple calculation.

[0003] However, agricultural production data (such as crop yield, mechanization level, and chemical fertilizer and pesticide application amount in a certain region over the years) are mostly time series data, which generally have two key problems of autocorrelation and heteroscedasticity, resulting in a significant decrease in the reliability of the traditional OLS regression results. Agricultural production has obvious seasonality and continuity, and the production inputs (such as the continuous use of mechanized equipment) or natural conditions (such as the cumulative effect of soil fertility) in the previous period will have an impact on the crop yield in the next period, so that there is correlation between adjacent observations in time series data. The traditional OLS assumes that "observations are independent of each other", and autocorrelation will lead to underestimation of the standard error of the regression coefficient, and thus exaggerate the statistical significance of the estimation result. In addition, there are differences in agricultural production scale, technology level, and input intensity in different periods (such as the rapid increase in mechanization level and the improvement in chemical fertilizer application accuracy in recent years), resulting in changes in the error term variance of time series data over time, i.e. heteroscedasticity. The traditional OLS assumes that "the error term variance is constant", and heteroscedasticity will cause estimation bias of the standard error of the regression coefficient, and cannot accurately determine the real significance of the impact of input factors on yield.

[0004] In existing research, some methods attempt to solve the above problems through data transformation (such as difference method) or model replacement (such as mixed OLS), but the former may lose data information, and the latter is difficult to adapt to the production characteristics of different crops (such as the significant difference in fertilizer requirement between corn and rice), and cannot provide accurate and robust measurement evidence for the "input-output" relationship of different crops. Therefore, there is an urgent need for a quantitative evaluation method that can handle time series autocorrelation and heteroscedasticity and adapt to the analysis of multiple crops. SUMMARY

[0005] To address the bottlenecks in the quantitative analysis of agricultural inputs and crop yields in existing technologies, this invention aims to propose an agricultural productivity assessment method based on an OLS regression model with Newey-West standard errors. This invention corrects the standard errors of traditional OLS, simultaneously processing autocorrelation and heteroscedasticity in time-series data. This allows for the accurate output of the impact coefficients and statistical significance of agricultural mechanization and chemical inputs on the yields of different crops such as corn, soybeans, and rice, providing a scientific tool for optimizing agricultural production.

[0006] The method includes the following steps: S1. Obtain historical crop yield data for each province within the target region. and agricultural input data of various provinces over the years ; S2. Perform preprocessing, stationarity test and multicollinearity diagnosis on the data obtained in step S1 in sequence; S3, respectively with and Construct a least squares OLS regression model with the dependent and independent variables as the dependent and independent variables: ,in, , The index value of the province in the table. , The index value representing the year. , This represents the index value of agricultural input data. , Index value representing crop yield data, Represents the intercept value. to These represent the regression coefficients of the corresponding variables. Indicates the first The province in the The random error term for the year; S4. Calculate the regression coefficients to vector estimate ; S5. Construct the Newey-West robust covariance matrix And solve The square root of the main diagonal element gives the standard error SE; S6, Pass , build Estimated covariance matrix , correct The revised Recorded as ; S7, according to Calculate the t-statistic and p-value with SE, and assess agricultural productivity based on the t-statistic and p-value.

[0007] Furthermore, in step S4, The formula for calculation is: ,in, Represent the independent variable The vector matrix, express transpose, express A vector matrix.

[0008] Furthermore, in step S5, the Newey-West robust covariance matrix... Specifically: ,in, This represents the basis matrix that takes heteroscedasticity into account. Indicates the maximum lag order. The index value represents the lag order. Represents the weighting function. This represents the estimation of the original autocovariance matrix. express The transpose of .

[0009] Furthermore, in step S6, Estimated covariance matrix Specifically: .

[0010] Furthermore, in step S7, the formula for calculating the t-statistic is: The t-statistic follows a t-distribution, and the degrees of freedom of the t-distribution are based on the total number of years. Total agricultural input data calculate.

[0011] Furthermore, in step S7, the P value is equal to the tail area of ​​the t distribution.

[0012] An electronic device includes a memory and a processor, the memory storing a computer program, characterized in that the processor executes the computer program to implement the steps of the above-described method.

[0013] A computer-readable storage medium for storing computer instructions, characterized in that the computer instructions, when executed by a processor, implement the steps of the above-described method.

[0014] The beneficial effects of the method described in this invention are as follows: (1) The method described in this invention corrects the traditional OLS model by using the Newey-West standard error, which effectively solves the problems of autocorrelation and heteroscedasticity in agricultural time series data and improves the accuracy and reliability of the quantitative results of crop yield influencing factors.

[0015] (2) The method described in this invention can accurately distinguish the different impacts of agricultural mechanization and chemical input on the yield of different crops such as corn, soybeans and rice, and provide a scientific basis for formulating personalized agricultural input optimization schemes for different crops.

[0016] (3) The method described in this invention is simple to operate, can be combined with mainstream statistical analysis software, and is applicable to data analysis in various agricultural production areas. It has a wide range of application scenarios and practical value. Attached Figure Description

[0017] Figure 1 This is a flowchart of the method described in this invention. Detailed Implementation

[0018] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] This embodiment provides an agricultural productivity assessment method based on the Newey-West standard error, the flowchart of which is as follows: Figure 1 As shown: The study area of ​​this embodiment is the three northeastern provinces of China, including Heilongjiang, Jilin, and Liaoning. In 2024, Heilongjiang's grain output was 80.017 million tons, Jilin's was 42.66 million tons, and Liaoning's was 25.003 million tons, accounting for approximately 21% of the national grain output, making them the "stabilizer" and "ballast" of China's food security. In the agricultural planting structure of the three northeastern provinces, corn production accounts for about 60% of the region's total, rice production accounts for about 30%, and soybeans account for about 10%.

[0020] All data are sourced from authoritative publicly available statistical data and academic research to ensure reliability and comparability. Agricultural mechanization data, including total power of agricultural machinery, agricultural diesel consumption, and agricultural electricity consumption, are derived from the *China Rural Statistical Yearbook* (2002-2024) and the statistical yearbooks of each province (Heilongjiang, Jilin, and Liaoning). These indicators collectively characterize the scale of mechanization and energy consumption intensity in agricultural production. Agricultural chemical input data, including fertilizer application (pure amount), nitrogen fertilizer, phosphate fertilizer, potash fertilizer, compound fertilizer use, and pesticide use, are compiled from the *China Rural Statistical Yearbook* (2002-2024) and the *China Statistical Yearbook*. Agricultural output and economic data, including total grain output, yield per unit area of ​​rice, corn, and soybeans, and total agricultural output value, are sourced from the *China Statistical Yearbook* (2002-2024) and the statistical yearbooks of each province. To eliminate the impact of price fluctuations, the total agricultural output value data is deflated using the agricultural production materials price index of each province and calculated uniformly at constant 2000 prices. Agricultural carbon emissions data, including methane ( ), nitrous oxide ( Agricultural carbon emission data, including greenhouse gas emissions (in carbon dioxide equivalents), are cited from a standardized grid dataset published in the top international journal Nature Food. This dataset is a long-term time series (1993-2020) collected from panel surveys of rural households (200,000 households per year) across China.

[0021] This embodiment collects historical crop yield data for each province within the target area, including yield per unit area of ​​grain, rice, corn, and soybeans, agricultural output value per unit area (actual value), and kilograms of carbon dioxide equivalent per hectare (agricultural carbon emission intensity). This includes agricultural mechanization input data for each province during the corresponding period (such as total agricultural machinery power and agricultural machinery operation area), chemical input data (such as fertilizer application and pesticide use), and control variable data such as natural conditions and policy environment. Specifically, it includes annual agricultural input data for each province, such as diesel fuel consumption for mechanization, electricity consumption for mechanization, total agricultural machinery power, agricultural water consumption, application of agricultural fertilizers, nitrogen fertilizer use, phosphate fertilizer use, potash fertilizer use, compound fertilizer use, and pesticide use. .

[0022] This embodiment cleans the collected data, removes outliers, fills in missing values ​​using interpolation or regression methods, and standardizes all variables to ensure data consistency and usability. Secondly, a stationarity test (unit root test) is performed. The purpose is twofold: first, to avoid "spurious regression," where two completely independent but trending non-stationary time series may show a high degree of divergence during regression. and significant The first is the statistical measure, but this relationship is spurious; the second is to ensure the validity of statistical inference, that is, in standard OLS regression. test, Tests rely on asymptotic distribution theory, one of the fundamental assumptions of which is the stationarity of the data. Non-stationary data alters the distribution of these test statistics, rendering significance assessments ineffective. The typical test method is the Augmented Dickey-Fuller (ADF) test. Finally, multicollinearity is diagnosed to determine the stability of model coefficient estimates. Small data fluctuations can lead to significant changes in coefficient estimates, even altering their signs. The method typically uses the Variance Inflation Factor (VIF). If VIF = 1, the variable is completely unrelated to other independent variables; if 1 < VIF < 5, a moderate degree of correlation exists, which is generally acceptable; if 5 < VIF < 10, high multicollinearity is present and requires attention; if VIF ≥ 10, severe multicollinearity exists, and the coefficient estimates for this variable are unreliable and must be addressed.

[0023] This embodiment effectively solves the problems of autocorrelation and heteroscedasticity in agricultural time series data through the following three steps: Step 1: In this embodiment, respectively using and Construct a least squares OLS regression model with the dependent and independent variables as the dependent and independent variables: ,in, , The index value of the province in the table. , The index value representing the year. This represents the total number of years. , This represents the index value of agricultural input data. This represents the total amount of agricultural input data. , Index value representing crop yield data, This represents the total number of crop yield data; The intercept value represents the baseline level of the dependent variable when all independent variables are zero. to These represent the regression coefficients of the corresponding variables, which represent the marginal impact of each independent variable on the dependent variable (i.e., the direction and magnitude of the impact of the independent variable on the dependent variable). Indicates the first The province in the The annual random error term includes all unobserved factors affecting the dependent variable; Step 2: Calculate the regression coefficients to vector estimate ; The formula for calculation is: ,in, Represent the independent variable A vector matrix (including constant terms). express transpose, express The vector matrix, The purpose is to calculate the sum of cross products between the independent variables to form the coefficient matrix of the normal system of equations. The goal is to calculate the relationship between each independent variable and the dependent variable. The covariance is used to form the constant term on the right side of the normal system of equations.

[0024] Step 3: Considering the common autocorrelation and heteroscedasticity problems in agricultural economic time series data, this embodiment introduces the Newey-West standard error adjustment mechanism (hereinafter referred to as NW) into the traditional OLS model. By optimizing the covariance matrix estimation and setting a reasonable lag order, the autocorrelation and heteroscedasticity problems of the time series data are corrected, resulting in robust standard error estimation results. Specifically, this involves: constructing the Newey-West robust covariance matrix. , This refers to the estimated Newey-West robust covariance matrix. The standard error SE required in the model is the square root of the elements on the main diagonal of this matrix.

[0025] Newey-West robust covariance matrix Specifically: ,in, This represents a basis matrix that takes into account heteroscedasticity (similar to the White heteroscedasticity-uniform covariance matrix). Indicates the maximum lag order. The index value represents the lag order. Represents the weighting function. It is the weighting function (usually using the Bartlett kernel function): Its function is to ensure the positive definiteness of the covariance matrix and to give higher weights to more recent lag terms; This represents the estimate of the original autocovariance matrix, used to capture autocorrelation; express The transpose of , because mathematically, for any covariance matrix They all But alone It is usually not symmetrical, so The purpose of addition is to create a symmetric matrix, ensuring the symmetry of the covariance matrix (i.e., simultaneously considering both the "influence of the past on the present" and the "influence of the present on the past"). This example calculates the Newey-West heteroscedasticity and autocorrelation consistent (HAC) standard errors. Maximum lag order. Based on sample size (In this embodiment) =23, meaning that in this embodiment, crop yield data and agricultural input data for 23 years from each province are selected. Common rules are automatically determined to ensure the robustness of standard error estimation, among which, This represents the floor function. All econometric analyses are implemented using Python's statsmodels library.

[0026] This embodiment is illustrated by... , build Estimated covariance matrix , correct The revised Recorded as .

[0027] Then according to Calculate the t-statistic and p-value with SE, and assess agricultural productivity based on the t-statistic and p-value: The formula for calculating the t-statistic is: The t-statistic follows a t-distribution. The shape of the t-distribution is determined by one parameter, namely the degrees of freedom. Degrees of freedom are based on the total number of years. Total agricultural input data calculate: In this embodiment, , According to calculations in this embodiment, the t-statistic is based on degrees of freedom. It follows a t-distribution of 12.

[0028] The p-value is defined as the probability of observing the current t-statistic or a more extreme value, assuming the null hypothesis is true. Mathematically, the p-value is the area under the tails of the t-distribution, and the formula for calculating the p-value is: ,in, Let x represent the x-coordinate of any point on the t-distribution. express When, the area (probability) value of the right tail.

[0029] In this embodiment, Newey-West robust standard error Statistic (Right now ), degrees of freedom Based on the above values, first find the t-distribution with 12 degrees of freedom, and find... The corresponding cumulative distribution function value; Secondly, due to Calculate the right tail probability Finally, because it is a symmetric two-sided matrix test, therefore , .

[0030] The t-statistic is a core statistical concept, referring to the ratio of signal to noise. For example, the coefficients estimated by an OLS model... The independent variable represents the strength of its influence on the dependent variable, i.e., the "signal," while the uncertainty of this coefficient estimate, i.e., the Newey-West robust standard error, represents the fluctuation range of the "signal" and the measurement error, i.e., the "noise." In this embodiment... =0.601, and its Newey-West robust standard error SE=0.134, then t=0.601 / 0.134≈4.49. This result indicates that the coefficient estimate (0.601) is 4.49 times its standard error (0.134). This is a very strong "signal-to-noise ratio," suggesting that the independent variable has a significant positive impact on the dependent variable. In this embodiment, the summary of the OLS regression model with Newey-West standard error is shown in Table 1: The significance matrix of the OLS regression model with Newey-West standard error is shown in Table 2: Table 1

[0031] Table 2

[0032] In Tables 1 and 2, Y1, Y2, Y3, Y4, Y5, and Y6 represent crop yield data when the year and province remain constant. The value of ( That is, the yield per unit area of ​​grain, the yield per unit area of ​​rice, the yield per unit area of ​​corn, the yield per unit area of ​​soybean, the agricultural output value per unit area, and the kilograms of carbon dioxide equivalent per hectare; X1, X2, X3, X4, X5, X6, X7, X8, X9, and X10 represent the input data when the year and province remain unchanged. The value of ( That is, the amount of diesel fuel used for mechanization, the amount of electricity consumed for mechanization, the total power of agricultural machinery, the amount of water used for agriculture, the amount of fertilizer applied for agriculture, the amount of nitrogen fertilizer used, the amount of phosphate fertilizer used, the amount of potash fertilizer used, the amount of compound fertilizer used, and the amount of pesticides used; In Table 2: if p < 0.01, it means that the effect of the independent variable on the dependent variable is extremely significant (indicated by ***), if p < 0.05, it means that the effect of the independent variable on the dependent variable is significant (indicated by **), and if p < 0.10, it means that the effect of the independent variable on the dependent variable is weakly significant (indicated by *).

[0033] The goodness-of-fit evaluation results (Table 1) show that there are significant differences in the overall fit of the various regression models. The agricultural output model has the highest goodness of fit ( (Rounded to 3 decimal places), indicating that the model can explain 98.9% of the variation in agricultural output, reflecting the high explanatory power of agricultural mechanization and input factors on agricultural output. Rice yield per unit area model ( ), grain yield model ( ) and carbon emission models ( The model also showed a good fit, indicating that the selected independent variables have a strong explanatory power for these three key indicators. In contrast, the corn yield per unit area model ( The fit of the model was moderate, while the soybean yield per unit area model ( The fitting effect of ) is relatively poor, and adjustments are needed. Even a negative value (-0.332) indicates that the response mechanism of soybean production to mechanization and chemical inputs is more complex.

[0034] The results in Tables 1 and 2 show that the p-values ​​of the F-statistics for all models are less than 0.05, indicating that the regression models are statistically significant overall. In particular, the agricultural output model (Fp = 3.91e-16) and the grain yield model (Fp = 4.81e-11) exhibit extremely high significance levels, further validating the rationality of the model specification. The residual autocorrelation test reveals that the Durbin-Watson statistic shows that the DW values ​​for all models fluctuate around 2 (1.911-2.367), indicating that residual autocorrelation is effectively controlled. Furthermore, the p-values ​​for the residual autocorrelation test are all greater than 0.05, further confirming that after correction using the Newey-West standard error, there is no significant autocorrelation in the model residuals, ensuring the effectiveness of the parameter estimation.

[0035] This embodiment also performs residual diagnosis and robustness analysis on the method described herein. The residual diagnosis includes: autocorrelation test (to test whether there is serial correlation in the residuals) and heteroscedasticity test (to test whether the residual variance is constant). The verification results are as follows: Autocorrelation test: The Ljung-Box Q test (automatically generated by Python) yielded a p-value of 0.32, which is greater than 0.05. This demonstrates that the Newey-West method has effectively corrected for autocorrelation, and the model specification is appropriate. (If the p-value is greater than 0.05, the null hypothesis cannot be rejected, indicating that there is no significant autocorrelation in the residuals; if the p-value is less than or equal to 0.05, the null hypothesis is rejected, indicating that there is significant autocorrelation in the residuals.) Heteroscedasticity test: The Breusch-Pagan test (automatically generated by Python) yielded a p-value of 0.41, which is greater than 0.05. This indicates that the Newey-West method has effectively corrected for heteroscedasticity. (If the p-value is greater than 0.05, the null hypothesis cannot be rejected, and the residuals are considered homoscedastic; if the p-value is less than or equal to 0.05, the null hypothesis is rejected, and the residuals are considered heteroscedastic.)

[0036] The robustness analysis specifically involves verifying that the conclusions of the method described in this invention are not driven by specific variable measures or sample intervals, and the verification results show that the method described in this invention has reliability and stability.

[0037] In the three steps described above, the first step aims to establish a basic model and define the analytical framework; the second step aims to provide the best estimates of the independent variables on the dependent variable through OLS calculations. The third step aims to introduce the Newey-West standard error adjustment mechanism, ensuring that the judgment of whether these effects are "significant" is truly reliable when facing common statistical problems in time series data. Therefore, the method in this embodiment is scientific and robust precisely because of this interconnected three-step design.

Claims

1. A method for assessing agricultural productivity based on the Newey-West standard error, characterized in that, The method includes the following steps: S1. Obtain historical crop yield data for each province within the target region. and agricultural input data of various provinces over the years ; S2. Perform preprocessing, stationarity test and multicollinearity diagnosis on the data obtained in step S1 in sequence; S3, respectively with and Construct a least squares OLS regression model with the dependent and independent variables as the dependent and independent variables: ,in, , The index value of the province in the table. , The index value representing the year. , This represents the index value of agricultural input data. , Index value representing crop yield data, Represents the intercept value. to These represent the regression coefficients of the corresponding variables. Indicates the first The province in the The random error term for the year; S4. Calculate the regression coefficients to vector estimate ; S5. Construct the Newey-West robust covariance matrix and solve The square root of the main diagonal element gives the standard error SE; S6, Pass , build Estimated covariance matrix , correct The revised Recorded as ; S7, according to Calculate the t-statistic and p-value with SE, and assess agricultural productivity based on the t-statistic and p-value.

2. The agricultural productivity assessment method based on the Newey-West standard error according to claim 1, characterized in that, In step S4, The formula for calculation is: ,in, Represent the independent variable The vector matrix, express transpose, express A vector matrix.

3. The agricultural productivity assessment method based on the Newey-West standard error according to claim 2, characterized in that, In step S5, the Newey-West robust covariance matrix Specifically: ,in, This represents the basis matrix that takes heteroscedasticity into account. Indicates the maximum lag order. The index value represents the lag order. Represents the weighting function. This represents the estimation of the original autocovariance matrix. express The transpose of .

4. The agricultural productivity assessment method based on the Newey-West standard error according to claim 3, characterized in that, In step S6, Estimated covariance matrix Specifically: .

5. The agricultural productivity assessment method based on the Newey-West standard error according to claim 4, characterized in that, In step S7, the formula for calculating the t-statistic is: The t-statistic follows a t-distribution, and the degrees of freedom of the t-distribution are based on the total number of years. Total agricultural input data calculate.

6. The agricultural productivity assessment method based on the Newey-West standard error according to claim 5, characterized in that, In step S7, the P value is equal to the tail area of ​​the t distribution.

7. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-6.

8. A computer-readable storage medium for storing computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the steps of the method according to any one of claims 1-6.