Intelligent diagnostic method for coordinated optimization of comprehensive benefits of black soil protection and utilization
By constructing a multi-source spatiotemporal data diagnostic chain and using one-hot coding and multi-objective optimization models, the problem of coordinated optimization of ecological, economic and social benefits in black soil protection and utilization was solved, the comprehensive benefit evaluation and optimization strategy generation in the black soil area were realized, and the scientific nature and adaptability of agricultural management were improved.
Patent Information
- Application Number
- CN202510508837.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-04-22
Smart Images

Figure CN120373559B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of geography and agricultural information technology, and specifically relates to an intelligent diagnosis method for coordinated optimization of the comprehensive benefits of black soil protection and utilization. Background Art
[0002] As a globally important agricultural resource, black soil, with its excellent soil properties and high yield potential, has become crucial for agricultural production. Northeast China's black soil region is a core production area for national food security, playing a vital role in the agricultural economy, social development, and ecological protection, and has a significant impact on national food supply and sustainable agricultural development. However, due to factors such as global climate change, long-term high-intensity tillage, excessive fertilization, and improper straw handling, black soil regions face multiple challenges, including declining soil organic matter, increased erosion, increased carbon emissions, and diminishing marginal agricultural economic benefits. The sustainability of black soil conservation and utilization is becoming increasingly prominent.
[0003] Extensive research has been conducted both domestically and internationally to address these issues. Internationally, research on black soil conservation has primarily focused on improving ecological benefits, focusing on optimizing tillage practices to enhance black soil quality and reduce soil erosion and carbon emissions. For example, the U.S. Department of Agriculture (USDA) promotes crop rotation and fallow, reduced tillage and no-tillage technologies, and establishes a comprehensive monitoring system to prevent soil erosion and reduce nutrient loss. Russia's black soil regions are enhancing the ecological functions of black soil through rational land use planning, fallow systems, and soil improvement techniques. Canada is using cover cropping to improve soil quality and promote sustainable black soil utilization. However, these measures largely focus on restoring black soil's ecological functions and enhancing its climate adaptability, often overlooking the synergistic relationship between black soil conservation and agricultural economic and social benefits. This has resulted in some conservation measures being difficult to effectively promote in practice, and has even led to reduced economic benefits and low farmer adoption. Regarding black soil conservation and utilization, China has developed models such as the "Lishu Model" and the "Longjiang Model," which have introduced black soil conservation measures such as straw incorporation, conservation tillage, and deep tillage. However, existing research often relies on plot-scale experiments and lacks a comprehensive benefit evaluation system based on multi-source data fusion. Furthermore, most studies employ static assessment methods, ignoring the dynamic changes in black soil systems under different scenarios and the spatiotemporal heterogeneity of their benefits. Therefore, the urgent challenge facing sustainable agricultural development is to develop a scientific, comprehensive, and intelligent system for benefit evaluation and problem diagnosis in black soil conservation and utilization, analyze the benefit trade-offs in black soil conservation from a data-driven perspective, and provide targeted optimization paths. Summary of the Invention
[0004] In response to the above problems, the present invention proposes an intelligent diagnostic method for the coordinated optimization of the comprehensive benefits of black soil protection and utilization, and constructs a technical chain of "data collection and processing-spatiotemporal heterogeneity diagnosis-deep correlation mining-multi-objective strategy generation" to achieve a comprehensive assessment of the ecological, economic and social benefits in the process of black soil protection, tracing of problems and generation of optimization strategies.
[0005] The present invention discloses an intelligent diagnosis method for the coordinated optimization of the comprehensive benefits of black soil protection and utilization, which comprises the following steps:
[0006] Step 1: construct a multi-source spatiotemporal simulation data input and output dataset and convert the category information into a numerical vector form;
[0007] Step 2: Draw a spatial distribution heat map and time series curve based on the simulated data input and output data sets;
[0008] Step 3, constructing a correlation matrix between input variables and output variables based on the simulated data input and output data sets;
[0009] Step 4: construct a scoring system based on the historical optimal value of each output variable in the simulated data input and output dataset;
[0010] Step 5: Based on the scores of different dimensions in the scoring system, optimize the weight coefficient of each target dimension and the contribution weight within the benefit to build a multi-objective optimization model;
[0011] Step 6: Based on the spatial distribution heat map, time series curve, correlation matrix, dynamic scoring system and multi-objective optimization model, a closed-loop automated operation chain of data-driven problem diagnosis and strategy generation is formed.
[0012] Furthermore, in step 1, the following steps are also included:
[0013] Step 11: Construct a multi-source spatiotemporal simulated data input and output dataset. The simulated data includes an input dataset X, an output dataset Y, and a sample set A, where:
[0014] Input dataset X = {x (1) ,x (2) ,…,x (12)}include:
[0015] Agricultural operation parameters: tillage mode x (1) , Fertilization type x (2) , irrigation method x (3) , straw processing method x (4) ;
[0016] Climate parameters: temperature x (5), precipitation x (6) , sunshine duration x (7) , wind speed x (8) ;
[0017] Soil parameters: soil pH x (9) , black soil layer thickness x 910) , soil moisture x (11) , soil texture x (12) ;
[0018] Output data set Y = {y (econ1) 、y (econ2) 、y (soc1) 、y (soc2) 、y (soc3) 、y (soc4) 、y (soc5) 、y (eco1) 、y (eco2) 、y (eco3)},include:
[0019] Economic benefit indicators: farmers' net income (econ1) , policy subsidy investment (econ2) ;
[0020] Social benefit indicators: total grain output y (soc1) , soybean yield (soc2) , corn yield (soc3) , rice yield (soc4) , food equivalent y (soc5) ;
[0021] Ecological benefit indicators: soil quality (eco1) , net carbon emissions (eco2) , soil erosion amount y (eco3) ;
[0022] Sample set A={a1,a2,…,a m}, where each sample represents the observation record of a spatial unit at a specific time point, the data volume is m, and each sample contains the input dataset X, the output dataset Y, and the year and longitude and latitude information corresponding to the dataset.
[0023] Furthermore, in step 1, the following steps are also included:
[0024] Step 12, converting the category information in the agricultural operation parameters into a numerical vector form through one-hot encoding;
[0025] Step 121: Encode each categorical variable separately:
[0026] Categorical variables include the following four items: x (1) Farming mode, there are n1 categories; x(2) Fertilization type, there are n2 categories; x (3) Irrigation method, there are n3 categories; x (4) There are n4 categories of straw treatment methods;
[0027] For each variable x (j) (j=1,2,3,4), all possible values of which constitute the category set:
[0028]
[0029] C (j) Represents the jth categorical variable x in the agricultural operation parameters (j) (j=1,2,3,4) The set of categories consisting of all possible values; Represents the categorical variable x (j) The category set C of (j=1,2,3,4) (j) The kth category in the equation, the value range of k is 1≤k≤n j ;n j Represents the variable x (j) Total number of categories;
[0030] For each sample A i In the variable x (j) The corresponding code on (j=1,2,3,4) is recorded as:
[0031]
[0032] Represents the i-th sample A i In the jth categorical variable x (j) The one-hot encoding vector on the variable has a length equal to the total number of categories n of the variable j , Represents the i-th sample A i In the jth categorical variable x (j) The encoding value of the kth category on the index value k is in the range of 1≤k≤n j , Represents a vector is in n j Vectors in dimensional real space;
[0033]
[0034] Step 122, concatenate the sample category encoding vectors:
[0035]
[0036] In step 123, the one-hot encoding matrix of all samples is:
[0037] Stack the category vectors of all m samples row by row to get the one-hot encoding matrix:
[0038]
[0039] Where: m is the total number of samples, each row is z i Represents the encoding result of the categorical variable of the i-th sample; each column corresponds to the position of a certain category value, a value of 1 indicates that it belongs to the category, and 0 indicates that it does not belong to the category.
[0040] Furthermore, in step 2, the Moran's I index is used to measure the autocorrelation of spatial data to determine whether the data has a trend of aggregation or dispersion in space:
[0041]
[0042] Where: m is the total number of samples, v p and v q are the values of the pth and qth samples on a certain variable, is the mean of all samples; T pq is the spatial weight matrix, which represents the spatial relationship between positions p and q;
[0043] For each variable at each time point, a spatial distribution heat map is drawn, and the color gradient is used to represent the size of the variable value; for each spatial unit, a time series curve of a variable in different years is drawn, and the historical average and standard deviation are calculated by sliding.
[0044] Furthermore, in step 3, the following steps are also included:
[0045] Step 31, for the correlation between continuous variables, the Pearson correlation coefficient was used to measure the linear relationship;
[0046] In the Pearson correlation coefficient, Expressed as:
[0047]
[0048] Input dataset X = {x (1) ,x (2) ,…,x (12)}, where e is the input variable index, e=1,2,…,12,x (e) Represents input variables; Y={y (econ1) ,y (econ2) ,…,y (eci3)}, o is the output variable index, y (o) Represents the output variable; the sample set is A={a1,a2,…,a m}, where the kth sample ak The input and output variables are and The variables x are (e) 、y (o) The mean of is the Pearson correlation coefficient,
[0049] Step 32: Combine all input-output variables in pairs to generate a correlation matrix
[0050] Correlation Matrix The dimension is 12×10, the rows correspond to 12 input variables, the columns correspond to 10 output indicators, and the elements are The absolute value of the correlation coefficient of the input and output variables is used. The continuous variable takes the absolute value of the Pearson coefficient, and the one-hot encoded variable takes the absolute value of the point-binary coefficient. The value range is [0,1].
[0051] Step 33, obtain all input and output combinations to construct a causal test score matrix Γ;
[0052] F-statistic formula:
[0053]
[0054] Among them, F e,o Represents the input variable x (e) With the output variable y (o) Statistics, RSS res and RSS unres The residual sum of squares of the restricted model and the unrestricted model, L is the number of lagged terms, and pcount is the number of explanatory variables in the unrestricted model;
[0055] Based on F e,o The corresponding P value P e,o And construct the causal score:
[0056]
[0057] represents the causal score;
[0058] Get all input and output combinations to construct a causal test score matrix with a dimension of 12×10
[0059] Step 34: Build an XGBoost model for each output variable to train it with all input variables x (1) ,x (2) ,…,x (12)The nonlinear relationship between the two variables is analyzed, and the SHAP method is used to quantify the relative contribution of each input variable:
[0060]
[0061] represents the combination of the e-th input variable and the o-th output variable, is the kth sample variable x (e) For output y (o) SHAP value, e′ is the sum index variable, represents the combination of the e′th input variable and the oth output variable, Represents x (e) y (o) The normalized contribution of
[0062] Finally, the final contribution matrix S with a dimension of 12×10 is obtained (shap) , 12 corresponds to the 12 input variables in the input dataset X, and 10 corresponds to the 10 output variables in the input dataset Y;
[0063] The equal weight method is used to perform weighted summation on the relationship matrix of correlation, causality and SHAP analysis results to obtain the deep correlation matrix:
[0064]
[0065] represents the correlation matrix, Γ represents the causal test score matrix, and Ω represents the deep correlation matrix;
[0066] Step 35: Save the correlation matrix, SHAP matrix, and deep correlation matrix as CSV files, and save the negative correlation prompts to a text file.
[0067] Furthermore, in step 4, the score of any sample on the output data of the type item is:
[0068]
[0069] k represents the sample number, Indicates the score of the k-th sample on the output data of the type item, is the output variable of the kth sample, is the historical optimal value among all samples;
[0070] The score of the kth sample in this class is:
[0071]
[0072] Θ (t) is the output variable type, t∈{econ,soc,eco}; Indicates the score of the k-th sample in the econ, soc, eco class;
[0073] The total score of sample k is:
[0074]
[0075] TotalScore k represents the total score of the kth sample, represents the score of the kth sample in the econ class, represents the score of the kth sample in the soc class, Represents the score of the k-th sample in the eco class.
[0076] Furthermore, in step 5, the multi-objective function is defined as a weighted combination of the three types of benefit scores:
[0077]
[0078] in, is the weight coefficient of the three types of targets;
[0079] So that the score of the kth sample in class t Optimal under different strategic objectives:
[0080]
[0081] Indicates the optimality under different strategic objectives; is the score of the k-th sample on the output data of type item, Internal variable representing benefit t Contribution weight;
[0082] The following constraints are introduced during the optimization process to ensure the balance and rationality of different benefit dimensions and avoid imbalance caused by extreme optimization of one dimension. The constraints can be formally expressed as follows:
[0083] g(x (e) ,y (o) )≤0, y (o)
[0084] Among them, g(x (e) ,y (o) ) represents the constraint relationship between decision variables;
[0085] Finally, a multi-objective optimization algorithm is used to obtain an optimal trade-off solution for comprehensive optimization and scenario simulation; the goal of the optimization process is to maximize the comprehensive score and obtain the optimal benefits under different strategies while satisfying the constraints.
[0086] Furthermore, in step 6, the closed-loop automated operation chain of data-driven problem diagnosis and strategy generation includes multi-dimensional short board identification;
[0087] Multi-dimensional short board identification by comparing regional benefit scores and the global optimal value The difference between the two values is used to locate the areas with poor overall performance; the generated spatial distribution heat map and time series curve are combined for analysis to mark high-priority governance areas that have been inefficient for a long time.
[0088] Furthermore, in step 6, the closed-loop automated operation chain of data-driven problem diagnosis and strategy generation includes root cause tracing analysis;
[0089] Root cause tracing analysis is based on the deep correlation matrix Ω, which quantifies the impact of each variable on the target indicator, locks in the key driving variables, and identifies the dominant constraints.
[0090] Furthermore, in step 6, the closed-loop automated operation chain of data-driven problem diagnosis and strategy generation includes dynamic strategy generation;
[0091] Dynamic strategy generation generates a customized partitioning solution based on the results of the multi-objective optimization algorithm and the deep correlation matrix Ω:
[0092] For ecologically weak areas, we formulate farming adjustment rules based on the driving relationships in the deep correlation matrix Ω. For economically inefficient areas, we design a subsidy-price linkage optimization algorithm based on the constraints and synergies in the deep correlation matrix Ω.
[0093] The final output is a visual decision report, including strategy priority ranking, constraints and expected benefit increase.
[0094] The beneficial effects achieved by the present invention are:
[0095] This invention achieves iterative strategy optimization through the synergistic effect of four core modules. A spatiotemporal diagnostic library identifies abnormal regions and identifies optimization targets. An input-output relationship matrix analyzes nonlinear relationships between variables, generating a deep correlation matrix and dynamic constraints. A benefit scoring system quantifies multidimensional target deviations to guide optimization. A multi-objective optimization model outputs an optimal solution set to support precise decision-making.
[0096] This invention breaks through the limitations of traditional empirical decision-making and realizes the full-process data-driven technology chain of "data collection and processing-spatiotemporal heterogeneity diagnosis-deep correlation mining-multi-objective strategy generation", which significantly improves the scientificity and adaptability of agricultural management strategies.
[0097] The results of this invention can be applied to agricultural management and farmer income optimization strategies in Northeast China's black soil regions, promoting the precise management of black soil conservation and providing a replicable reference paradigm for sustainable development in black soil regions worldwide. In the future, this technology can also be expanded to other soil degradation issues, such as grassland degradation and wetland protection, providing intelligent solutions for ecological and environmental governance. BRIEF DESCRIPTION OF THE DRAWINGS
[0098] Figure 1 .It is a flow chart of the diagnostic process method for black soil protection and utilization;
[0099] Figure 2 .Spatiotemporal problem diagnosis heat map;
[0100] Figure 3 SHAP influence diagram of each input variable for output data;
[0101] Figure 4 .Input-output relationship depth correlation matrix;
[0102] Figure 5 .3D visualization of optimization scenarios in different regions. DETAILED DESCRIPTION
[0103] The present invention will be further described below with reference to specific embodiments, and the advantages and features of the present invention will become clearer as the description proceeds. However, these embodiments are merely exemplary and do not constitute any limitation to the scope of the present invention. It should be understood by those skilled in the art that the details and forms of the technical solutions of the present invention may be modified or replaced without departing from the spirit and scope of the present invention, and such modifications and replacements fall within the scope of protection of the present invention.
[0104] As attached Figure 1 As shown, the present invention provides an intelligent diagnosis method for the coordinated optimization of the comprehensive benefits of black soil protection and utilization, which mainly includes the following steps:
[0105] Step 1: construct a multi-source spatiotemporal simulation data input and output dataset and convert the category information into a numerical vector form;
[0106] This study uses scenario-based approaches to quantify indicators such as different tillage patterns, fertilization methods, and irrigation methods, and uses these as multidimensional input datasets to simulate the performance of different regions and agricultural operation modes in the black soil region. Furthermore, relevant indicators such as climate change, soil properties, and farmer income are used as output datasets. This simulated data provides a foundation for subsequent benefit analysis, relationship exploration, and problem diagnosis.
[0107] Specifically, step 1 also includes the following steps:
[0108] Step 11: construct a multi-source spatiotemporal simulation data input and output dataset to provide a data basis for subsequent analysis;
[0109] The simulated data includes the input dataset X, the output dataset Y, and the sample set A, where:
[0110] Input dataset X = {x (1) ,x (2) ,…,x (12)}include:
[0111] Agricultural operation parameters (categorical variables): tillage mode x (1) (e.g. corn continuous cropping / rotation, rice continuous cropping / rotation), fertilization type x (2) (Organic fertilizer / chemical fertilizer), irrigation method x (3) (drip irrigation / flood irrigation), straw treatment method x (4) ;
[0112] Climate parameters (numeric variables): temperature x (5) , precipitation x (6) , sunshine duration x (7) , wind speed x (8) ;
[0113] Soil parameters (numeric variables): soil pH x (9) , black soil layer thickness x (10) , soil moisture x (11) , soil texture x (12) .
[0114] Output data set Y = {y (econ1) 、y (econ2) 、y (soc1) 、y (soc2) 、y (soc3) 、y (soc4) 、y (soc5) 、y (eco1) 、y (eco2) 、y (eco3)},include:
[0115] Economic benefit indicators: Net income of farmers (y (econ1) ), policy subsidy investment (y (econ2) );
[0116] Social benefit indicators: total grain output (y (soc1) ), soybean yield (y (soc2) ), corn yield (y (soc3) ), rice yield (y (soc4) ), food equivalent (y (soc5) );
[0117] Ecological benefit indicators: soil quality (y(eco1) ), net carbon emissions (y (eco2) ), soil erosion amount (y (eco3) ).
[0118] Sample set A={a1,a2,…,a m}, where each sample represents an observation record of a spatial unit at a specific point in time, with a data volume of m. Each sample in the set contains all X and Y (input and output sets) and additional year and latitude and longitude information.
[0119]
[0120] Table 1 Data diagram
[0121] Step 12: Convert the category information into a numerical vector form through one-hot encoding;
[0122] Considering the existence of non-digital data in multi-source data (x (1) 、x (2) 、x (3) 、x (4) ) cannot be directly used for modeling analysis, so this patent innovatively proposes the application of one-hot encoding technology to convert category information into a numerical vector form that is easy to process by machine learning models. The specific method is as follows:
[0123] Step 121: Encode each categorical variable separately:
[0124] Suppose the categorical variables include the following 4 items: x (1) Farming mode, there are n1 categories; x (2) Fertilization type, there are n2 categories; x (3) Irrigation method, there are n3 categories; x (4) There are n4 categories of straw processing methods.
[0125] For each variable x (j) (j=1,2,3,4), all possible values of which constitute the category set:
[0126]
[0127] C (j) Represents the jth categorical variable x in the agricultural operation parameters (j) (j=1,2,3,4) The set of categories consisting of all possible values; Represents the categorical variable x (j) The category set C of (j=1,2,3,4) (j) The kth category in the equation, the value range of k is 1≤k≤n j ;n j Represents the variable x(j) The total number of categories.
[0128] For each sample A i In the variable x (j) The corresponding code on (j=1,2,3,4) is recorded as:
[0129]
[0130] in, Represents a vector is in n j Vectors in dimensional real space. ;
[0131]
[0132] Step 122, concatenate the sample category encoding vectors:
[0133]
[0134] In step 123, the one-hot encoding matrix of all samples is:
[0135] Stack the category vectors of all m samples row by row to get the one-hot encoding matrix:
[0136]
[0137] Where: m is the total number of samples (i.e., sample set A = {a1, a2, ..., a m} size); each row z i Represents the encoding result of the categorical variable of the i-th sample; each column corresponds to the position of a certain category value, a value of 1 indicates that it belongs to the category, and 0 indicates that it does not belong to the category.
[0138] Step 2: Draw a spatial distribution heat map and time series curve based on the simulated data input and output data sets;
[0139] Based on the dataset processed in step 1, this paper applies global spatial autocorrelation analysis to detect spatiotemporal heterogeneity in the data. By plotting spatial distribution heat maps and time series curves, the paper focuses on anomalies in regions or time periods during the black soil conservation process. The diagnostic results provide preliminary clues for subsequent relationship analysis, helping to understand the spatiotemporal distribution characteristics of the data and analyze the factors that lead to fluctuations in benefits. The formula is as follows:
[0140]
[0141] Where: m is the total number of samples, v p and v q are the values of the pth and qth samples on a certain variable, is the mean value over all samples. p,qis a spatial weight matrix that represents the spatial relationship between locations p and q. I represents the Moran's I index, which is used to measure the autocorrelation of spatial data and determine whether the data has a tendency to cluster or disperse spatially.
[0142] Moran's I index was used to determine the spatial clustering of input and output variables (significant when p < 0.05). For each variable at each time point, a spatial distribution heat map was drawn, and the color gradient was used to represent the size of the variable value. If a sample variable value v p If the value exceeds the range of (mean ± 2 times the standard deviation), it is marked as a "spatial outlier." For each spatial unit, a time series curve of a variable in different years is plotted, and the historical mean and standard deviation are calculated by sliding. If the variable value in consecutive years exceeds the sliding standard deviation range of ± 20%, it is considered a "temporal outlier segment."
[0143] Step 3, constructing a correlation matrix between input variables and output variables based on the simulated data input and output data sets;
[0144] Spatial autocorrelation only analyzes the clustering of variables in the spatial distribution within the study area. In order to deeply explore the correlation between data, this patent uses three methods: correlation coefficient analysis, causal inference, and machine learning model to explore the relationship between input and output data respectively, and comprehensively integrates the correlation to generate a more authoritative deep correlation matrix between input and output data.
[0145] The deep correlation matrix of this patent mainly integrates Pearson correlation analysis, Granger causality test and XGBoost model.
[0146] The specific steps include:
[0147] In step 31, the Pearson correlation coefficient is used to measure linear relationships between continuous variables. The point-biserial correlation coefficient is used to measure the correlation between one-hot encoded variables and continuous variables. This is used to measure the linear or nonlinear relationship between each input factor and the output benefit. Using the correlation coefficients between each pair of variables, a correlation matrix between the input and output variables is constructed.
[0148] Assume that the input data set X = {x (1) ,x (2) ,…,x (12)}, where e is the input variable index, e=1,2,…,12,x (e) Indicates the input variable. Y={y (ecin1) ,y (econ2) ,…,y (eco3)}, o is the output variable index, y (o) Represents the output variable. The sample set is A={a1,a2,…,a m}, where the kth sample ak The input and output variables are and The variables x are (e) 、y (o) The mean of .
[0149] In the Pearson correlation coefficient, The calculation formula is:
[0150]
[0151] In the point biserial correlation coefficient, if the input variable x (type) is a categorical variable. After one-hot encoding, the binary variable generated by the hth class is x (type,h) , whose value is 0 or 1, indicating whether it belongs to the category. Let an output variable be y (o) , is a continuous variable. In the entire sample A, μ1 is y (o) In x (type,h) =1 (i.e., belongs to this class); μ0 is the sample average value of y (o) In x (type,h) = 0 (i.e. does not belong to this class); σ y is the output variable y (o) The standard deviation of the entire sample; π is the number of x in the sample (type,h) =1 (i.e. the proportion of this category). Then the point-biserial correlation coefficient is defined as:
[0152]
[0153] This coefficient is used to measure the binary variable x obtained by one-hot encoding. (type,h) And continuous output variable y (o) The correlation between them is in the range of [-1,1]. The larger the absolute value, the stronger the correlation.
[0154] Step 32: Combine all input-output variables in pairs to generate a correlation matrix
[0155] The matrix dimension is 12×10, with rows corresponding to 12 input variables (index e), columns corresponding to 10 output indicators (index o), and elements are It is the absolute value of the correlation coefficient between the input and output variables (the continuous variable takes the absolute value of the Pearson coefficient, and the one-hot encoded variable takes the absolute value of the point-binary coefficient), and the value range is [0,1].
[0156] Finally, the correlation matrix is defined as Participate in the weighted fusion of the final deep correlation matrix. Negative correlations are marked and recorded separately and are not included in the matrix numerical calculation.
[0157] Step 33, obtain all input and output combinations to construct a causal test score matrix Γ;
[0158] Granger causality test is performed on each pair of input-output variables, and the maximum lag order maxlag is set to 2. Restricted and unconstrained models are constructed respectively, and the residual sum of squares RSS is calculated. res , RSS unres The F statistic formula is:
[0159]
[0160] Among them, F e,o Represents the input variable x (e) With the output variable y (o) Statistics, RSS res and RSS unres The residual sum of squares of the restricted model and the unrestricted model, L is the number of lagged terms, and pcount is the number of explanatory variables in the unrestricted model (here 2L+1, including the constant term and the lagged terms of the two variables). e,o The corresponding P value P e,o And construct the causal score:
[0161]
[0162] represents the causal score. This formula linearly maps significant results within 0.1 to a score interval of [0.5, 1.0], while forcing non-significant results to zero to prevent weak causal noise from interfering with the overall correlation matrix. A higher score indicates a more significant causal influence of the input variable on the output variable.
[0163] All input and output combinations construct a causal test score matrix with a dimension of 12×10 Will participate in the weighted fusion of the final deep correlation matrix.
[0164] Step 34, for each output variable y Io) Build an XGBoost model to train it with all input variables x (1) ,x (2) ,…,x (12) The nonlinear relationship between the two variables is analyzed, and the SHAP method is used to quantify the relative contribution of each input variable.
[0165] The SHAP value is based on the Shapley value in cooperative game theory and is used to quantify the value of each input variable x. (e) The average marginal gain in model prediction over all variable combinations. For the kth sample, variable x (e) For output y (o) The SHAP value is recorded as For all samples x (e) The absolute SHAP values are averaged to obtain the y effect of the variable on the model (o) Overall contribution
[0166] Normalize the contribution of the input variables corresponding to each output variable o:
[0167]
[0168] Where e' is the sum index variable, which traverses all input variables (e'=1~12), and the denominator represents the sum of all input variables to y (o) Contribution to ensure is a relative proportion (the sum of each column is 1). Represents x (e) y (o) The normalized contribution of . Finally, the final contribution matrix S with a dimension of 12×10 is obtained (shap) , 12 corresponds to the 12 input variables in the input dataset X, and 10 corresponds to the 10 output variables in the input dataset Y.
[0169] Finally, the equal weight method is used to perform weighted summation on the relationship matrix of correlation, causality and SHAP analysis results to obtain the deep correlation matrix.
[0170]
[0171] represents the correlation matrix, Γ represents the causal test score matrix, and Ω represents the deep association matrix.
[0172] Step 35: Save the correlation matrix, SHAP matrix, and deep correlation matrix as CSV files, and save the negative correlation prompts to a text file. Draw a heat map to display the comprehensive score matrix and intuitively present the comprehensive relationship between the input variables and the output variables.
[0173] Step 4: construct a scoring system based on the historical optimal value of each output variable in the simulated data input and output dataset;
[0174] To comprehensively evaluate the performance of different samples across multiple indicators, this step constructs a dynamic scoring system. Each sample is scored based on its historical best value for each output indicator. The scores are then averaged across groups (economic, social, ecological) for a final total score.
[0175] For each sample patent requirement, each output is scored first. Assume that the output variable of the kth sample is In general, the historical optimal value of this variable in all samples is represents the score of the kth sample on the output data of the type item, then the formula is:
[0176]
[0177] Then, scores are calculated for different benefit aspects. Assume that the output variables are divided into Θ (t) , where t∈{econ,soc,eco}, the score of the k-th sample in this class is:
[0178]
[0179] Indicates the score of the k-th sample in the econ, soc, eco class;
[0180] The total score of sample k is obtained by adding the scores of three aspects:
[0181]
[0182] TotalScore k represents the total score of the kth sample, represents the score of the kth sample in the econ class, represents the score of the kth sample in the soc class, Represents the score of the k-th sample in the eco class.
[0183] Step 5: Based on the scores of different dimensions in the scoring system, optimize the weight coefficient of each target dimension and the contribution weight within the benefit to build a multi-objective optimization model;
[0184] Given that the scores on different dimensions are known in the previous step, in order to find the optimal balance between ecological, economic, and social benefits, as well as the optimal balance of different outputs within each benefit, so that the management strategy can achieve the optimal solution in multiple dimensions, this step uses a multi-objective optimization method to automatically adjust the weight of each benefit and optimize the contribution of different factors within each benefit to obtain the optimal strategy.
[0185] Assume that each sample a k Scores for the three dimensions were obtained in step 4 The multi-objective function is defined as a weighted combination of three types of benefit scores:
[0186]
[0187] in, are the weight coefficients of the three types of objectives, and the weights are automatically adjusted through the Pareto multi-objective optimization method to ensure the balance of the three aspects.
[0188] Then, based on the score of the k-th sample calculated in step 4 on the output data of the type item Internal variable representing benefit t The contribution weight of The model will automatically find the optimal Make The optimal strategy under different objectives. The specific rules are as follows:
[0189]
[0190] Indicates the optimality under different strategic objectives;
[0191] For example, if the contribution of a farming model to ecological benefits is calculated to be 0.25, its weight would be 25% after considering the total contribution of all variables influencing ecological benefits. This weighting method intuitively reflects the relative importance of each variable in influencing different benefits, ensuring that the role of each variable is appropriately reflected in the overall benefit calculation.
[0192] During the optimization process, the following constraints need to be introduced to ensure the balance and rationality of different benefit dimensions and avoid imbalance caused by extreme optimization of one dimension. The constraints can be formally expressed as:
[0193] g(x (e) ,y (o) )≤0, y (o)
[0194] Among them, g(x (e) ,y (o) ) represents the constraint relationship between decision variables.
[0195] Finally, a multi-objective optimization algorithm is used to find an optimal trade-off solution for comprehensive optimization and scenario simulation. The goal of the optimization process is to maximize the overall score while achieving the optimal benefits under different strategies while satisfying constraints (such as policy restrictions and environmental sustainability).
[0196] This invention generates different optimization scenarios for each region, including an "optimal scenario," a "worst scenario," and an "average scenario." Through three-dimensional visualization, it demonstrates the trade-offs between benefits under each scenario, helping policymakers intuitively understand the effects of various management strategies and make informed decisions.
[0197] Step 6: Based on the spatial distribution heat map, time series curve, correlation matrix, dynamic scoring system and multi-objective optimization model, a closed-loop automated operation chain of data-driven problem diagnosis and strategy generation is formed.
[0198] The specific implementation includes the following steps:
[0199] a) Multi-dimensional short board identification:
[0200] By comparing regional benefit scores (Economic / Societal / Ecological) and Global Optimum The difference between the two groups is used to locate the areas with poor overall performance; combined with the results of step 2, a spatiotemporal hotspot analysis (spatial clustering I + "time anomaly period") is performed to mark high-priority governance areas that have been inefficient for a long time.
[0201] b) Root cause analysis:
[0202] Based on the input-output deep correlation matrix Ω obtained in step 3, the degree of influence of each variable on the target indicator is quantified, the key driving variables are locked, and the dominant constraints are identified.
[0203] c) Dynamic strategy generation:
[0204] Based on the multi-objective optimization results obtained in step 5, combined with the input-output deep correlation matrix Ω obtained in step 3, a customized zoning plan is generated. For ecologically weak areas, farming adjustment rules are formulated based on the driving relationships in Ω (e.g., mandatory implementation of straw return and crop rotation in high-erosion areas). For economically inefficient areas, a subsidy-price linkage optimization algorithm is designed based on the constraints and synergy factors in Ω (e.g., dynamically binding the subsidy coefficient to market price volatility).
[0205] Output a visual decision report (PDF format) including strategy priority ranking, constraints, and expected benefit increase.
[0206] The above are only specific steps of the present invention and do not constitute any limitation to the scope of protection of the present invention; any technical solutions formed by equivalent transformation or equivalent replacement fall within the scope of protection of the present invention; the parts not elaborated in detail in the present invention belong to the common knowledge of those skilled in the art.
Claims
1. An intelligent diagnostic method for the coordinated optimization of the comprehensive benefits of black soil protection and utilization, characterized by: The intelligent diagnosis method for coordinated optimization of comprehensive benefits of black soil protection and utilization includes the following steps: Step 1: construct a multi-source spatiotemporal simulation data input and output dataset and convert the category information into a numerical vector form; Step 2: Draw a spatial distribution heat map and time series curve based on the simulated data input and output data sets; Step 3, constructing a correlation matrix between input variables and output variables based on the simulated data input and output data sets; Step 4: construct a scoring system based on the historical optimal value of each output variable in the simulated data input and output dataset; Step 5: Based on the scores of different dimensions in the scoring system, optimize the weight coefficient of each target dimension and the contribution weight within the benefit to build a multi-objective optimization model; Step 6: Based on spatial distribution heat maps, time series curves, correlation matrices, dynamic scoring systems, and multi-objective optimization models, a closed-loop automated operation chain for data-driven problem diagnosis and strategy generation is formed. In step 1, the following steps are also included: Step 11: Construct a multi-source spatiotemporal simulated data input and output dataset. The simulated data includes an input dataset X, an output dataset Y, and a sample set A, where: Input dataset X = {x (1) ,x (2) ,…,x (12) }include: Agricultural operation parameters: tillage mode x (1) , Fertilization type x (2) , irrigation method x (3) , straw processing method x (4) ; Climate parameters: temperature x (5) , precipitation x (6) , sunshine duration x (7) , wind speed x (8) ; Soil parameters: soil pH x (9) , black soil layer thickness x (10) , soil moisture x (11) , soil texture x (12) ; Output data set Y = {y (econ1) , y (econ2) , y (soc1) , y (soc2) , y (soc3) , y (soc4) , y (soc5) , y (eco1) , y (eco2) , y (eco3)}, including: Economic benefit indicators: farmers' net income (econ1) , policy subsidy investment (econ2) ; Social benefit indicators: total grain output y (soc1) , soybean yield (soc2) , corn yield (soc3) , rice yield (soc4) , food equivalent y (soc5) ; Ecological benefit indicators: soil quality (eco1) , net carbon emissions (eco2) , soil erosion amount y (eco3) ; Sample set A={a1,a2,…,a m }, where each sample represents an observation record of a spatial unit at a specific time point, the data volume is m, and each sample contains the input dataset X, the output dataset Y, and the year and longitude and latitude information corresponding to the dataset; In step 3, the following steps are also included: Step 31, for the correlation between continuous variables, the Pearson correlation coefficient was used to measure the linear relationship; In the Pearson correlation coefficient, Expressed as: Input dataset X = {x (1) ,x (2) ,…,x (12) }, where e is the input variable index, e=1,2,…,12,x (e) Represents input variables; Y={y (econ1) ,y (econ2) ,…,y (eco3) }, o is the output variable index, y (o) Represents the output variable; the sample set is A={a1,a2,…,a m }, where the kth sample a k The input and output variables are and The variables x are (e) 、y (o) The mean of is the Pearson correlation coefficient, Step 32: Combine all input-output variables in pairs to generate a correlation matrix Correlation Matrix The dimension is 12×10, the rows correspond to 12 input variables, the columns correspond to 10 output indicators, and the elements are is the absolute value of the correlation coefficient of the input and output variables. Continuous variables take the absolute value of the Pearson coefficient, and one-hot encoded variables take the absolute value of the point-binary coefficient. The range is [0,1]. Step 33, obtain all input and output combinations to construct a causal test score matrix Γ; F-statistic formula: Among them, F e,o Represents the input variable x (e) With the output variable y (o) Statistics, RSS res and RSS unres The residual sum of squares of the restricted model and the unrestricted model, L is the number of lagged terms, and pcount is the number of explanatory variables in the unrestricted model; Based on F e,o The corresponding P value P e,o And construct the causal score: represents the causal score; Get all input and output combinations to construct a causal test score matrix with a dimension of 12×10 Step 34: Build an XGBoost model for each output variable to train it with all input variables x (1) ,x (2) ,…,x (12) The nonlinear relationship between the two variables is analyzed, and the SHAP method is used to quantify the relative contribution of each input variable: represents the combination of the e-th input variable and the o-th output variable, is the kth sample variable x (e) For output y (o) SHAP value, e ′ To sum index variables, Indicates the e ′ The combination of the input variable and the oth output variable, Represents x (e) y (o) The normalized contribution of Finally, the final contribution matrix S with a dimension of 12×10 is obtained (shap) , 12 corresponds to the 12 input variables in the input dataset X, and 10 corresponds to the 10 output variables in the input dataset Y; The equal weight method is used to perform weighted summation on the relationship matrix of correlation, causality and SHAP analysis results to obtain the deep correlation matrix: represents the correlation matrix, Γ represents the causal test score matrix, and Ω represents the deep correlation matrix; Step 35: Transform the correlation matrix Causality test score matrix Γ, SHAP matrix S (shap) The deep correlation matrix Ω is saved as a CSV file, and the negative correlation prompt is saved to a text file.
2. The intelligent diagnostic method for coordinated optimization of comprehensive benefits of black soil protection and utilization according to claim 1 is characterized in that: In step 1, the following steps are also included: Step 12, converting the category information in the agricultural operation parameters into a numerical vector form through one-hot encoding; Step 121: Encode each categorical variable separately: Categorical variables include the following four items: x (1) Farming mode, there are n1 categories; x (2) Fertilization type, there are n2 categories; x (3) Irrigation method, there are n3 categories; x (4) There are n4 categories of straw treatment methods; For each variable x (j) , j = 1, 2, 3, 4, all possible values of which constitute the category set: C (j) Represents the jth categorical variable x in the agricultural operation parameters (j) , the category set consisting of all possible values of j=1,2,3,4; Represents a categorical variable x (j) , the category set C with j = 1, 2, 3, 4 (j) The kth category in the equation, the value range of k is 1≤k≤n j ;n j Represents the variable x (j) Total number of categories; For each sample A i In the variable x (j) , the corresponding codes for j=1,2,3,4 are recorded as: Represents the i-th sample A i In the jth categorical variable x (j) The one-hot encoding vector on the variable has a length equal to the total number of categories n of the variable j , Represents the i-th sample A i In the jth categorical variable x (j) The encoding value of the kth category, the value range of index k is 1≤k≤n j , Represents a vector is in n j Vectors in dimensional real space; Step 122, concatenate the sample category encoding vectors: In step 123, the one-hot encoding matrix of all samples is: Stack the category vectors of all m samples row by row to get the one-hot encoding matrix: Where: m is the total number of samples, each row is z i Represents the encoding result of the categorical variable of the i-th sample; each column corresponds to the position of a certain category value, a value of 1 indicates that it belongs to the category, and 0 indicates that it does not belong to the category.
3. The intelligent diagnostic method for coordinated optimization of comprehensive benefits of black soil protection and utilization according to claim 1 is characterized in that: In step 2, the Moran's I index is used to measure the autocorrelation of spatial data to determine whether the data has a trend of aggregation or dispersion in space: Where: m is the total number of samples, v p and v q are the values of the pth and qth samples on a certain variable, is the mean of all samples; T p,q is the spatial weight matrix, which represents the spatial relationship between positions p and q; For each variable at each time point, a spatial distribution heat map is drawn, and the color gradient is used to represent the size of the variable value; for each spatial unit, a time series curve of a variable in different years is drawn, and the historical average and standard deviation are calculated by sliding.
4. The intelligent diagnostic method for coordinated optimization of comprehensive benefits of black soil protection and utilization according to claim 1 is characterized in that: In step 4, the score of any sample on the output data of type item is: k represents the sample number, Indicates the score of the k-th sample on the output data of the type item, is the output variable of the kth sample, is the historical optimal value among all samples; The score of the kth sample in this class is: Θ (t) is the output variable type, t∈{econ,soc,eco}; Indicates the score of the k-th sample in the econ, soc, eco class; The total score of sample k is: TotalScore k represents the total score of the kth sample, represents the score of the kth sample in the econ class, represents the score of the kth sample in the soc class, Represents the score of the k-th sample in the eco class.
5. The intelligent diagnosis method for coordinated optimization of comprehensive benefits of black soil protection and utilization according to claim 1 is characterized in that: In step 5, the multi-objective function is defined as a weighted combination of three types of benefit scores: in, is the weight coefficient of the three types of targets; represents the score of the kth sample in the econ class, represents the score of the kth sample in the soc class, represents the score of the kth sample in the eco class; So that the score of the kth sample in class t Optimal under different strategic objectives: Indicates the optimality under different strategic objectives; is the score of the k-th sample on the output data of type item, Internal variable representing benefit t Contribution weight; The following constraints are introduced during the optimization process to ensure the balance and rationality of different benefit dimensions and avoid imbalance caused by extreme optimization of one dimension. The constraints can be formally expressed as follows: Among them, g(x (e) ,y (o) ) represents the constraint relationship between decision variables; Finally, a multi-objective optimization algorithm is used to obtain an optimal trade-off solution for comprehensive optimization and scenario simulation; the goal of the optimization process is to maximize the comprehensive score and obtain the optimal benefits under different strategies while satisfying the constraints.
6. The intelligent diagnostic method for coordinated optimization of comprehensive benefits of black soil protection and utilization according to claim 1 is characterized in that: In step 6, the closed-loop automated operation chain of data-driven problem diagnosis and strategy generation includes multi-dimensional short board identification; Multi-dimensional short board identification by comparing regional benefit scores and the global optimal value The difference between the two values is used to locate the areas with lagging comprehensive performance; the generated spatial distribution heat map and time series curve are combined for analysis to mark high-priority governance areas that have been inefficient for a long time.
7. The intelligent diagnostic method for coordinated optimization of comprehensive benefits of black soil protection and utilization according to claim 1 is characterized in that: In step 6, the closed-loop automated operation chain of data-driven problem diagnosis and strategy generation includes root cause tracing analysis; Root cause tracing analysis is based on the deep correlation matrix Ω, which quantifies the degree of influence of each variable on the target indicator, locks in the key driving variables, and identifies the dominant restrictive factors.
8. The intelligent diagnostic method for coordinated optimization of comprehensive benefits of black soil protection and utilization according to claim 1 is characterized in that: In step 6, the closed-loop automated operation chain of data-driven problem diagnosis and strategy generation includes dynamic strategy generation; Dynamic strategy generation generates a customized partitioning solution based on the results of the multi-objective optimization algorithm and the deep correlation matrix Ω: For ecologically weak areas, we formulate farming adjustment rules based on the driving relationships in the deep correlation matrix Ω. For economically inefficient areas, we design a subsidy-price linkage optimization algorithm based on the constraints and synergies in the deep correlation matrix Ω. The final output is a visual decision report, including strategy priority ranking, constraints and expected benefit increase.
Citation Information
Patent Citations
Land utilization type spatial distribution prediction method and system
CN118229162A