Intelligent diagnosis method for black soil protection and utilization comprehensive benefit collaborative optimization

By constructing intelligent diagnostic methods for multi-source spatiotemporal data, the problem of failure to effectively evaluate comprehensive benefits in the protection and utilization of black soil is solved, and a comprehensive evaluation and optimization strategy generation of ecological, economic and social benefits is achieved, which improves the scientificity and adaptability of agricultural management strategies in black soil area.

CN120373559AActive Publication Date: 2025-07-25INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510508837.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-25
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

In the protection and utilization of black soil, existing research has failed to effectively build a comprehensive benefit evaluation system for multi-source data fusion, and ignores the dynamic changes of black soil system in different scenarios and the spatiotemporal heterogeneity of benefits, which makes it difficult to promote some protection measures in practice, and even causes the problems of decline in economic benefits and low willingness of farmers to adopt.

Method used

A technical chain of "data collection and processing-spatial-temporal heterogeneity diagnosis-deep correlation mining-multi-objective strategy generation" was constructed. The input and output data sets were constructed through multi-source spatial-temporal simulation data, a spatial distribution heat map and time series curve were drawn, and the correlation matrix between input variables and output variables was constructed, and the weight coefficients and internal contribution weights of each target dimension were optimized, forming a closed-loop automated operation chain for data-driven problem diagnosis and strategy generation.

Benefits of technology

The comprehensive assessment and optimization strategy of ecological, economic and social benefits in the process of black soil protection have been achieved, the scientificity and adaptability of agricultural management strategies have been improved, and the sustainable development of black soil areas has been promoted.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373559A_ABST
    Figure CN120373559A_ABST
Patent Text Reader

Abstract

The invention discloses a black soil protection and utilization comprehensive benefit collaborative optimization-oriented intelligent diagnosis method. The method comprises the steps of 1, constructing a multi-source time-space analog data input and output data set; 2, drawing a spatial distribution thermodynamic diagram and a time sequence curve based on the data set; step 3, constructing a correlation matrix between the input variable and the output variable based on the data set; 4, constructing a scoring system based on the historical optimal value of each output variable in the data set; 5, optimizing the weight coefficient of each target dimension and the contribution weight in the benefit based on the scoring system, and obtaining a multi-target optimization model; and step 6, based on the spatial distribution thermodynamic diagram, the time sequence curve, the correlation matrix, the dynamic scoring system and the multi-objective optimization model, forming a closed-loop automatic operation chain of data-driven problem diagnosis and strategy generation. The method breaks through the limitation of traditional experience decision, realizes a full-process data-driven technical chain, and significantly improves the scientificity and adaptability of an agricultural management strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of geography and agricultural informatization, and specifically relates to an intelligent diagnosis method for collaborative optimization of the comprehensive benefits of black soil protection and utilization. Background Technique

[0002] As an important global agricultural resource, the excellent soil properties and high-yield potential of black soil have become the key to agricultural production. The Northeast Black Soil Region is the core production area of the country's food security, playing a crucial role in agricultural economy, social development and ecological protection, and having a decisive impact on the country's food supply and agricultural sustainable development. However, due to factors such as global climate change, long-term intensive cultivation, excessive fertilization, and improper straw treatment, the black soil area faces multiple challenges such as declining soil organic matter, increasing erosion, increasing carbon emissions, and diminishing marginal agricultural economic benefits. The sustainability issue of black soil protection and utilization has become increasingly prominent.

[0003] A large number of studies have been carried out at home and abroad to address the above problems. Internationally, the research on black soil protection mainly focuses on the improvement of ecological benefits. They are more concerned with how to improve the quality of black soil, reduce soil erosion and carbon emissions by optimizing tillage methods. For example, the United States Department of Agriculture (USDA) promotes crop rotation and fallow, reduced tillage and no-till technologies, and constructs a three-dimensional monitoring system to prevent soil erosion and reduce soil nutrient loss. The Russian black soil area realizes the improvement of the ecological function of black soil by reasonably planning land use, implementing a fallow system and improving soil technology. Canada improves soil quality through cover cropping to promote the sustainable use of black soil. However, most of these measures focus on the restoration of the ecological function of black soil and the improvement of climate adaptability, often ignoring the synergistic relationship between black soil protection and agricultural economic and social benefits, resulting in some protection measures being difficult to effectively promote in practice, and even causing problems such as declining economic benefits and low willingness of farmers to adopt them to a certain extent. In terms of black soil protection and utilization, China has formed models such as the "Lishu Model" and the "Longjiang Model", and put forward black soil protection measures such as straw returning to the field, conservation tillage, and deep loosening of soil. However, the existing research mostly relies on plot-scale experiments, has not yet formed a comprehensive benefit evaluation system based on multi-source data fusion, and mostly uses static evaluation methods, ignoring the dynamic changes of the black soil system under different scenarios and the spatio-temporal heterogeneity of its benefits. Therefore, how to construct a scientific, comprehensive and intelligent black soil protection and utilization benefit evaluation and problem diagnosis system, analyze the benefit trade-off mechanism in black soil protection from a data-driven perspective, and provide targeted optimization paths is an urgent problem to be solved in the current field of agricultural sustainable development. Summary of the Invention

[0004] In view of the above problems, the present invention proposes an intelligent diagnosis method for the collaborative optimization of the comprehensive benefits of black soil protection and utilization, constructs a technical chain of "data collection and processing - spatio-temporal heterogeneity diagnosis - deep correlation mining - multi-objective strategy generation", and realizes the comprehensive evaluation, problem tracing and optimization strategy generation of the ecological, economic and social benefits in the process of black soil protection.

[0005] The present invention discloses an intelligent diagnosis method for the collaborative optimization of the comprehensive benefits of black soil protection and utilization, and the intelligent diagnosis method for the collaborative optimization of the comprehensive benefits of black soil protection and utilization includes the following steps:

[0006] Step 1, construct a multi-source spatio-temporal simulated data input-output data set, and convert the category information therein into a numerical vector form;

[0007] Step 2, draw a spatial distribution heat map and a time series curve based on the simulated data input-output data set;

[0008] Step 3, construct a correlation matrix between the input variables and the output variables based on the simulated data input-output data set;

[0009] Step 4, construct a scoring system based on the historical optimal values of each output variable in the simulated data input-output data set;

[0010] Step 5, optimize the weight coefficients of each target dimension and the contribution weights within the benefits based on the scores of different dimensions in the scoring system to construct a multi-objective optimization model;

[0011] Step 6, form a closed-loop automated operation chain for data-driven problem diagnosis and strategy generation based on the spatial distribution heat map, time series curve, correlation matrix, dynamic scoring system and multi-objective optimization model.

[0012] Furthermore, in Step 1, the following steps are further included:

[0013] Step 11, construct a multi-source spatio-temporal simulated data input-output data set, where the simulated data includes an input data set X, an output data set Y, and a sample set A, where:

[0014] The input data set X = {x (1) , x (2) , …, x (12)} includes:

[0015] Agricultural operation parameters: tillage pattern x (1) , fertilization type x (2) , irrigation method x (3) , straw treatment method x (4) ;

[0016] Climate parameters: temperature x (5), precipitation x (6) , sunshine duration x (7) , wind speed x (8) ;

[0017] Soil parameters: soil pH x (9) , black soil layer thickness x 910) , soil moisture x (11) , soil texture x (12) ;

[0018] Output data set Y = {y (econ1) , y (econ2) , y (soc1) , y (soc2) , y (soc3) , y (soc4) , y (soc5) , y (eco1) , y (eco2) , y (eco3)}, including:

[0019] Economic benefit indicators: net income of farmers y (econ1) , policy subsidy input y (econ2) ;

[0020] Social benefit indicators: total grain output y (soc1) , soybean output y (soc2) , corn output y (soc3) , rice output y (soc4) , food equivalent y (soc5) ;

[0021] Ecological benefit indicators: soil quality y (eco1) , net carbon emissions y (eco2) , soil erosion amount y (eco3) ;

[0022] Sample set A = {a1, a2,..., a m} where each sample represents the observation record of a certain spatial unit at a certain specific time point, the data volume is m, and each sample contains the input data set X, the output data set Y, and the year and longitude and latitude information corresponding to the data set.

[0023] Furthermore, in step 1, the following steps are also included:

[0024] Step 12, convert the categorical information in the agricultural operation parameters into a numerical vector form through one-hot encoding;

[0025] Step 121, encode each category variable separately:

[0026] The categorical variables include the following 4 items: x (1) Cultivation pattern, with a total of n1 categories; x(2) Fertilization type, with a total of n2 categories; x (3) Irrigation method, with a total of n3 categories; x (4) Straw treatment method, with a total of n4 categories;

[0027] For each variable x (j) (j = 1, 2, 3, 4), all its possible values form a category set:

[0028]

[0029] C (j) Represents the category set formed by all possible values of the j-th categorical variable x (j) (j = 1, 2, 3, 4) in the agricultural operation parameters; Represents the category set of the categorical variable x (j) (j = 1, 2, 3, 4); C (j) The k-th category in the set C, where the value range of k is 1 ≤ k ≤ n j ; n j Represents the total number of categories of the variable x (j) ;

[0030] For each sample A i The encoding corresponding to it on the variable x (j) (j = 1, 2, 3, 4) is denoted as:

[0031]

[0032] Represents the one-hot encoding vector of the i-th sample A i On the j-th categorical variable x (j) The length of which is the total number of categories n of this variable j , Represents the i-th sample A i The encoding value of the k-th category on the j-th categorical variable x (j) The value range of the index value k is 1 ≤ k ≤ n j , Represents the vector Is a vector in the n j -dimensional real number space;

[0033]

[0034] Step 122, concatenate the sample category encoding vectors:

[0035]

[0036] Step 123, the one-hot encoding matrix of all samples is:

[0037] Stack the class vectors of all m samples row by row to obtain a one-hot encoding matrix:

[0038]

[0039] where: m is the total number of samples, and each row z i represents the encoding result of the categorical variable of the i-th sample; each column corresponds to the position of a certain category value, with a value of 1 indicating belonging to that category and 0 indicating not belonging.

[0040] Furthermore, in step 2, the Moran's I index is used to measure the autocorrelation of spatial data to determine whether there is a tendency of clustering or dispersion in space:

[0041]

[0042] where: m is the total number of samples, v p and v q are the values of the p-th and q-th samples on a certain variable respectively, is the mean value over all samples; T pq is the spatial weight matrix, representing the spatial relationship between positions p and q;

[0043] For each variable at each time point, draw a spatial distribution heat map, using a color gradient to represent the variable value size; for each spatial unit, draw a time series curve of a certain variable in different years, and slide to calculate the historical mean and standard deviation.

[0044] Furthermore, in step 3, the following steps are also included:

[0045] Step 31, for the correlation between continuous variables, use the Pearson correlation coefficient to measure the linear relationship;

[0046] In the Pearson correlation coefficient, is expressed as:

[0047]

[0048] Input dataset X = {x (1) , x (2) , …, x (12)}, where e is the input variable index, e = 1, 2, …, 12, and x (e) represents the input variable; Y = {y (econ1) , y (econ2) , …, y (eci3)}, o is the output variable index, and y (o) represents the output variable; the sample set is A = {a1, a2, …, a m}, where the k-th sample ak The input and output variables in and are the mean values of variables x (e) , y (o) respectively; is the Pearson correlation coefficient,

[0049] Step 32: Combine all input-output variables in pairs to generate a correlation matrix

[0050] Correlation matrix The dimension is 12×10. The rows correspond to 12 input variables, the columns correspond to 10 output indicators, and the elements are the absolute values of the correlation coefficients of the input-output variables. For continuous variables, the absolute value of the Pearson coefficient is taken; for one-hot encoded variables, the absolute value of the point-biserial coefficient is taken. The value range is [0,1].

[0051] Step 33: Obtain all input-output combinations to construct a causal test score matrix Γ;

[0052] F-statistic formula:

[0053]

[0054] where F e,o represents the statistic of the input variable x (e) and the output variable y (o) , RSS res and RSS unres are the residual sum of squares of the restricted model and the unrestricted model respectively. L is the number of lag terms, and pcount is the number of explanatory variables in the unrestricted model;

[0055] Based on the F e,o value, calculate the corresponding P value P e,o and construct a causal score:

[0056]

[0057] represents the causal score;

[0058] Obtain all input-output combinations to construct a causal test score matrix with a dimension of 12×10

[0059] Step 34: For each output variable, construct an XGBoost model to train its relationship with all input variables x (1) , x (2) ,…, x (12)The non-linear relationship between them is considered, and the SHAP method is used to quantify the relative contribution of each input variable:

[0060]

[0061] denotes the combination of the \(e\)-th input variable and the \(o\)-th output variable, is the \(k\)-th sample variable \(x\) (e) to the output \(y\) (o) of the SHAP value, where \(e'\) is the summation index variable, denotes the combination of the \(e'\)-th input variable and the \(o\)-th output variable, denotes \(x\) (e) to \(y\) (o) of the normalized contribution;

[0062] Finally, the final contribution matrix \(S\) with dimensions \(12\times10\) is obtained (shap) , where 12 corresponds to the 12 input variables in the input dataset \(X\), and 10 corresponds to the 10 output variables in the input dataset \(Y\);

[0063] The equal-weight method is used to perform weighted summation on the relationship matrices of the correlation, causality, and SHAP analysis results to obtain the deep association matrix:

[0064]

[0065] represents the correlation matrix, \(\Gamma\) represents the causality test score matrix, and \(\Omega\) represents the deep association matrix;

[0066] Step 35, save the correlation matrix, SHAP matrix, and deep association matrix as CSV files, and save the negative correlation prompts to text files.

[0067] Furthermore, in step 4, the score of any sample on the output data of the type item is:

[0068]

[0069] where \(k\) represents the sample number, represents the score of the \(k\)-th sample on the output data of the type item, is the output variable of the \(k\)-th sample, is the historical optimal value among all samples;

[0070] The score of the \(k\)-th sample on this class is:

[0071]

[0072] \(\Theta\) (t) is the output variable type, \(t\in\{econ, soc, eco\}\); Denote the score of the k-th sample on the econ, soc, and eco classes;

[0073] The total score of sample k is:

[0074]

[0075] TotalScore k Denote the total score of the k-th sample, Denote the score of the k-th sample on the econ class, Denote the score of the k-th sample on the soc class, Denote the score of the k-th sample on the eco class.

[0076] Furthermore, in step 5, define the multi-objective function as a weighted combination of the three types of benefit scores:

[0077]

[0078] where, are the weight coefficients of the three types of objectives;

[0079] Make the score of the k-th sample on the t class optimal under different strategic objectives:

[0080]

[0081] Denote being optimal under different strategic objectives; is the score of the k-th sample on the output data of the type item, Denote the internal variable of benefit t contribution weight;

[0082] Introduce the following constraint conditions during the optimization process to ensure the trade-off and rationality of different benefit dimensions and avoid imbalance caused by extreme optimization of a certain dimension. The constraint conditions can be formally expressed as:

[0083] g(x (e) ,y (o) ) ≤ 0, y (o)

[0084] where, g(x (e) ,y (o) ) represents the constraint relationship between decision variables;

[0085] Finally, use the multi-objective optimization algorithm to obtain a trade-off optimal solution for comprehensive optimization and scenario simulation; the goal during the optimization process is to maximize the comprehensive score, and at the same time, obtain the optimal benefits under different strategies while satisfying the constraint conditions.

[0086] Furthermore, in step 6, the closed-loop automated operation chain of data-driven problem diagnosis and strategy generation includes multi-dimensional shortcoming identification;

[0087] Multi-dimensional shortcoming identification locates the regions with poor comprehensive performance by comparing the difference between the regional benefit score and the global optimal value ; analyzes by combining the generated heat map of spatial distribution and time series curve, and marks the high-priority governance regions that have been in an inefficient state for a long time.

[0088] Furthermore, in step 6, the closed-loop automated operation chain of data-driven problem diagnosis and strategy generation includes root cause tracing and analysis;

[0089] Root cause tracing and analysis is based on the depth correlation matrix Ω, quantifies the influence degree of each variable on the target index, locks the key driving variables, and identifies the dominant restrictive factors.

[0090] Furthermore, in step 6, the closed-loop automated operation chain of data-driven problem diagnosis and strategy generation includes dynamic strategy generation;

[0091] Dynamic strategy generation generates a customized solution for each region according to the results of the multi-objective optimization algorithm, in combination with the depth correlation matrix Ω:

[0092] For ecological shortcoming regions, formulate tillage method adjustment rules in combination with the driving relationships in the depth correlation matrix Ω; for economically inefficient regions, design a subsidy-price linkage optimization algorithm in combination with the restrictive and collaborative factors in the depth correlation matrix Ω;

[0093] Finally, output a visual decision-making report, including the priority ranking of strategies, constraint conditions, and expected benefit increase.

[0094] The beneficial effects achieved by the present invention are:

[0095] The present invention realizes strategy iterative optimization through the synergistic effect of four core modules. The spatio-temporal diagnosis library marks abnormal regions and determines optimization targets; the input-output relationship matrix analyzes the non-linear correlation between variables, generates a depth correlation matrix and dynamic constraint conditions; the benefit scoring system quantifies the deviation degree of multi-dimensional objectives and guides the optimization direction; the multi-objective optimization model outputs the optimal solution set to support accurate decision-making.

[0096] The present invention breaks through the limitations of traditional experience-based decision-making, realizes a full-process data-driven technology chain of "data collection and processing - spatio-temporal heterogeneity diagnosis - deep correlation mining - multi-objective strategy generation", and significantly improves the scientificity and adaptability of agricultural management strategies.

[0097] The achievements of this invention can be applied to aspects such as agricultural management in the black soil region of Northeast China and the optimization of farmers' income increase strategies, promoting the precise management of black soil protection, and providing a replicable reference paradigm for the sustainable development of global black soil regions. In the future, this technology can also be extended to other soil degradation problems, such as grassland degradation and wetland protection, providing intelligent solutions for ecological environment governance. Brief Description of the Drawings

[0098] Figure 1 . is a flowchart of a diagnostic process method for black soil protection and utilization;

[0099] Figure 2 . Spatial and temporal problem diagnosis heat map;

[0100] Figure 3 . SHAP influence diagram for each input variable of the output data;

[0101] Figure 4 . Input-output relationship depth correlation matrix;

[0102] Figure 5 . Three-dimensional visualization of optimization scenarios in different regions. Detailed Implementation Modes

[0103] The following will further describe the present invention in combination with specific embodiments, and the advantages and features of the present invention will become clearer with the description. However, these embodiments are merely exemplary and do not constitute any limitation to the scope of the present invention. Those skilled in the art should understand that the details and forms of the technical solutions of the present invention can be modified or replaced without departing from the spirit and scope of the present invention, but these modifications and replacements all fall within the protection scope of the present invention.

[0104] As shown in the Figure 1 accompanying drawings, the present invention provides an intelligent diagnosis method for synergistically optimizing the comprehensive benefits of black soil protection and utilization, and the intelligent diagnosis method mainly includes the following steps:

[0105] Step 1, construct a multi-source spatio-temporal simulation data input-output dataset, and convert the category information therein into a numerical vector form;

[0106] The present invention quantifies indicators such as different tillage patterns, fertilization methods, and irrigation methods through scenario setting and uses them as a multi-dimensional input dataset to simulate the performance of different regions and different agricultural operation modes in the black soil region. At the same time, relevant indicators such as climate change, soil characteristics, and farmers' income are used as the output dataset. These simulation data provide a basis for subsequent benefit analysis, relationship exploration, and problem diagnosis.

[0107] Specifically, Step 1 further includes the following steps:

[0108] Step 11: Construct a simulation data input and output dataset for multi-source spatio-temporal data to provide a data basis for subsequent analysis;

[0109] The simulation data includes an input dataset X, an output dataset Y, and a sample set A, where:

[0110] The input dataset X = {x (1) , x (2) , …, x (12)} includes:

[0111] Agricultural operation parameters (categorical variables): tillage pattern x (1) (such as continuous cropping / rotation of corn, continuous cropping / rotation of rice), fertilization type x (2) (organic fertilizer / chemical fertilizer), irrigation method x (3) (drip irrigation / flood irrigation), straw treatment method x (4) ;

[0112] Climate parameters (numerical variables): temperature x (5) , precipitation x (6) , sunshine duration x (7) , wind speed x (8) ;

[0113] Soil parameters (numerical variables): soil pH x (9) , black soil layer thickness x (10) , soil moisture x (11) , soil texture x (12) .

[0114] The output dataset Y = {y (econ1) , y (econ2) , y (soc1) , y (soc2) , y (soc3) , y (soc4) , y (soc5) , y (eco1) , y (eco2) , y (eco3)}, includes:

[0115] Economic benefit indicators: net income of farmers (y (econ1) ), policy subsidy input (y (econ2) );

[0116] Social benefit indicators: total grain output (y (soc1) ), soybean output (y (soc2) ), corn output (y (soc3) ), rice output (y (soc4) ), food equivalent (y (soc5) );

[0117] Ecological benefit indicators: soil quality (y(eco1) ) Net carbon emissions (y (eco2) ) Soil erosion (y (eco3) ).

[0118] Sample set A = {a1, a2, …, a m}, where each sample represents the observation record of a certain spatial unit at a certain specific time point, and the data volume is m. Each sample in the set contains all X and Y (input-output set) and additional year and longitude / latitude information.

[0119]

[0120] Schematic diagram of data in Table 1

[0121] Step 12: Convert the category information into a numerical vector form through one-hot encoding;

[0122] Considering that there are non-digital data (x (1) , x (2) , x (3) , x (4) ) in the multi-source data that cannot be directly used for modeling analysis, this patent innovatively proposes to apply one-hot encoding technology to convert the category information into a numerical vector form that is convenient for machine learning models to process. The specific method is as follows:

[0123] Step 121: Encode each category variable separately:

[0124] Assume that the categorical variables include the following 4 items: x (1) Cultivation mode, with a total of n1 categories; x (2) Fertilizer type, with a total of n2 categories; x (3) Irrigation method, with a total of n3 categories; x (4) Straw treatment method, with a total of n4 categories.

[0125] For each variable x (j) (j = 1, 2, 3, 4), all its possible values form a category set:

[0126]

[0127] C (j) represents the category set formed by all possible values of the j-th categorical variable x (j) (j = 1, 2, 3, 4) in the agricultural operation parameters; represents the k-th category in the category set C (j) of the categorical variable x (j) (j = 1, 2, 3, 4), and the value range of k is 1 ≤ k ≤ n j ; n j represents the variable x(j) The total number of categories.

[0128] For each sample A i On variable x (j) (j = 1, 2, 3, 4), the corresponding encoding is denoted as:

[0129]

[0130] Where Denotes the vector Is a vector in the n j Dimensional real number space.;

[0131]

[0132] Step 122, concatenate the sample category encoding vectors:

[0133]

[0134] Step 123, the one - hot encoding matrix of all samples is:

[0135] Stack the category vectors of all m samples row - by - row to obtain the one - hot encoding matrix:

[0136]

[0137] Where: m is the total number of samples (i.e., the size of the sample set A = {a1, a2, …, a m}); Each row z i Represents the encoding result of the categorical variable of the i - th sample; Each column corresponds to the position of a certain category value, and the value 1 indicates belonging to the category, and 0 indicates not belonging.

[0138] Step 2, draw a spatial distribution heat map and a time - series curve based on the simulated data input - output data set;

[0139] Based on the data set processed in Step 1, the present invention applies global spatial autocorrelation analysis to detect the spatio - temporal heterogeneity of the data, and by drawing a spatial distribution heat map and a time - series curve, focuses on the anomalies shown in the black soil protection process in a region or time period. The diagnostic results provide preliminary clues for subsequent relationship analysis, help understand the spatio - temporal distribution characteristics in the data, and analyze the factors causing the fluctuations in benefits. The formula is as follows:

[0140]

[0141] Where: m is the total number of samples, v p And v q Are the values of the p - th and q - th samples on a certain variable respectively, Is the mean value over all samples. T p,qis the spatial weight matrix, representing the spatial relationship between positions p and q. I represents Moran's I index, which is used to measure the autocorrelation of spatial data and determine whether there is a tendency of clustering or dispersion in the space of the data.

[0142] Using Moran's I index, determine the spatial clustering of input and output variables (significant when p < 0.05). For each variable at each time point, draw a spatial distribution heat map, using color gradient to represent the variable value size. If the variable value v of a certain sample p exceeds the range of (mean ± 2 times the standard deviation), it is marked as a "spatial outlier". For each spatial unit, draw the time series curve of a certain variable in different years, and slide to calculate the historical average and standard deviation. If the variable value exceeds the ±20% sliding standard deviation interval for consecutive years, it is regarded as a "time anomaly segment".

[0143] Step 3, construct a correlation matrix between input variables and output variables based on the simulated data input-output data set;

[0144] Spatial autocorrelation only analyzes the clustering of variables in the spatial distribution within the study area. To deeply explore the correlation between data, this patent uses three methods: correlation coefficient analysis, causal inference, and machine learning models to respectively explore the relationship between input and output data and generate a more authoritative depth correlation matrix between input and output data by synthesizing the correlation.

[0145] The depth correlation matrix of this patent mainly synthesizes Pearson correlation analysis, Granger causality test, and XGBoost model.

[0146] Specifically, it includes the following steps:

[0147] Step 31, for the correlation between continuous variables, use the Pearson correlation coefficient to measure the linear relationship; for the correlation between one-hot encoded variables and continuous variables, use the point-biserial correlation coefficient. In this way, measure the linear or non-linear relationship between each input factor and the output benefit. Through the correlation coefficient between pairwise variables, construct a correlation matrix between input and output variables.

[0148] Let the input data set X = {x (1) , x (2) , …, x (12)}, where e is the input variable index, e = 1, 2, …, 12, and x (e) represents the input variable. Y = {y (ecin1) , y (econ2) , …, y (eco3)}, o is the output variable index, and y (o) represents the output variable. The sample set is A = {a1, a2, …, a m}, where the kth sample ak The input and output variables in and are the mean values of variables x (e) , y (o) respectively.

[0149] In the Pearson correlation coefficient, the calculation formula is:

[0150]

[0151] In the point-biserial correlation coefficient, if the input variable x (type) is a categorical variable, the binary variable generated for the h-th category after one-hot encoding is x (type,h) , and its value is 0 or 1, indicating whether it belongs to this category. Let a certain output variable be y (o) , which is a continuous variable. In the full sample A, μ1 is the sample mean of y (o) when x (type,h) =1 (i.e., belonging to this category); μ0 is the sample mean of y (o) when x (type,h) =0 (i.e., not belonging to this category); σ y is the standard deviation of the output variable y (o) in the full sample; π is the proportion of x (type,h) =1 in the sample (i.e., the proportion of this category). Then the point-biserial correlation coefficient is defined as:

[0152]

[0153] This coefficient is used to measure the correlation between the binary variable x (type,h) obtained by one-hot encoding and the continuous output variable y (o) . Its value range is [-1, 1], and the larger the absolute value, the stronger the correlation.

[0154] Step 32: Combine all input-output variables in pairs to generate a correlation matrix

[0155] The dimension of this matrix is 12×10. The rows correspond to 12 input variables (index e), the columns correspond to 10 output indicators (index o), and the elements are the absolute values of the correlation coefficients of the input-output variables (the absolute value of the Pearson coefficient for continuous variables and the absolute value of the point-biserial coefficient for one-hot encoded variables), and the value ranges are all [0, 1].

[0156] Finally, define the correlation matrix as to participate in the weighted fusion of the final depth correlation matrix. The negative correlation relationships are separately marked and recorded and not included in the matrix numerical calculation.

[0157] Step 33: Obtain all input-output combinations to construct the causal test score matrix Γ;

[0158] Perform Granger causality tests on each pair of input-output variables, setting the maximum lag order maxlag to 2. Respectively construct a restricted model and an unrestricted model, and calculate the residual sum of squares RSS res 、RSS unres . The F-statistic formula:

[0159]

[0160] where F e,o represents the statistic of the input variable x (e) and the output variable y (o) . RSS res and RSS unres are the residual sums of squares of the restricted model and the unrestricted model respectively. L is the number of lag terms, and pcount is the number of explanatory variables in the unrestricted model (here it is 2L + 1, including the constant term and the lag terms of the two variables). Calculate the corresponding P-value P e,o according to the above F e,o value and construct the causal score:

[0161]

[0162] represents the causal score. This formula linearly maps the significance result within 0.1 to the score interval [0.5, 1.0], and at the same time forces the non-significant result to zero to avoid the interference of weak causal noise on the overall correlation matrix result. The higher the score, the more significant the causal impact of the input variable on the output variable.

[0163] Construct a causal test score matrix with dimensions of 12×10 for all input-output combinations which will be weighted and fused with the final depth correlation matrix.

[0164] Step 34: For each output variable y Io) construct an XGBoost model to train the non-linear relationship between it and all input variables x (1) ,x (2) ,…,x (12) and use the SHAP method to quantify the relative contribution of each input variable.

[0165] The SHAP value is based on the Shapley value in cooperative game theory and is used to quantify the average marginal gain of each input variable x (e) in all variable combinations on the model prediction. For the k-th sample, the SHAP value of the variable x (e) to the output y (o) is denoted as For x in all samples (e) Average the absolute SHAP values to obtain the overall contribution of this variable to the model y (o) Overall contribution

[0166] Perform column normalization on the contribution of the input variable corresponding to each output variable o:

[0167]

[0168] where e′ is the summation index variable, traversing all input variables (e′ = 1 to 12), and the denominator represents the contribution of all input variables to y (o) To ensure is a relative ratio (the sum of each column is 1). Represents the normalized contribution of x (e) to y (o) Finally, obtain the final contribution matrix S with dimensions 12×10 (shap) , where 12 corresponds to the 12 input variables in the input dataset X, and 10 corresponds to the 10 output variables in the input dataset Y

[0169] Finally, use the equal - weight method to perform weighted summation on the relationship matrices of correlation, causality, and SHAP analysis results to obtain the deep association matrix

[0170]

[0171] represents the correlation matrix, Γ represents the causal test score matrix, and Ω represents the deep association matrix

[0172] Step 35, save the correlation matrix, SHAP matrix, and deep association matrix as CSV files, and save the negative - correlation prompts to text files. Draw a heatmap to display the comprehensive score matrix, visually presenting the comprehensive relationship between input variables and output variables

[0173] Step 4, construct a scoring system based on the historical optimal values of each output variable in the simulated data input - output dataset

[0174] To achieve a comprehensive evaluation of different samples on multi - indicator output results, this step constructs a dynamic scoring system. For each sample on each output indicator, score based on the historical optimal value of this indicator, then group and average according to the benefit type (economic, social, ecological), and finally calculate the total score

[0175] For each sample patent, first score each output. Assume that for the output variable of the k - th sample, the historical optimal value of the value of this variable in all samples is Let the score of the k-th sample on the output data of the type item be represented by the formula:

[0176]

[0177] Subsequently, the scores are calculated for different benefit aspects. Let the output variables be divided into Θ by type (t) , where t ∈ {econ, soc, eco}, and the score of the k-th sample on this class is:

[0178]

[0179] represents the score of the k-th sample on the econ, soc, eco classes;

[0180] The total score of sample k is obtained by adding the scores from three aspects:

[0181]

[0182] TotalScore k represents the total score of the k-th sample, represents the score of the k-th sample on the econ class, represents the score of the k-th sample on the soc class, represents the score of the k-th sample on the eco class.

[0183] Step 5: Based on the scores of different dimensions in the scoring system, optimize the weight coefficients of each target dimension and the contribution weights within the benefits to construct a multi-objective optimization model;

[0184] Given the scores on different dimensions obtained in the previous step, in order to find the optimal balance between ecological, economic, and social benefits and the optimal balance among different outputs within each benefit so that the management strategy can achieve the optimal solution in multiple dimensions, therefore, in this step, the multi-objective optimization method is used to automatically adjust the weights of each benefit and optimize the contribution degrees of different factors within each benefit to obtain the optimal strategy.

[0185] Let each sample a k has obtained the scores of three dimensions in Step 4 Define the multi-objective function as a weighted combination of the scores of three types of benefits:

[0186]

[0187] where, are the weight coefficients of the three types of objectives, and the weights are automatically adjusted through the Pareto multi-objective optimization method to ensure the balance of the three aspects.

[0188] Subsequently, based on the score of the k-th sample calculated in Step 4 on the output data of the type item The internal variable representing the benefit t The contribution weight of, satisfying The model will automatically find the optimal such that is optimal under different strategic objectives. The specific rules are as follows:

[0189]

[0190] Indicates being optimal under different strategic objectives;

[0191] For example, if the contribution degree of the farming mode to the ecological benefit is calculated as 0.25, after comprehensively considering the total contribution degree of all variables affecting the ecological benefit, its weight ratio is 25%. This way of weight distribution can intuitively reflect the relative importance of each variable in affecting different benefits, enabling the role of each variable to be appropriately reflected in the comprehensive benefit calculation.

[0192] During the optimization process, the following constraint conditions also need to be introduced to ensure the trade-off and rationality of different benefit dimensions and avoid imbalance caused by extreme optimization of a certain dimension. The constraint conditions can be formally expressed as:

[0193] g(x (e) ,y (o) ) ≤ 0, y (o)

[0194] where g(x (e) ,y (o) ) represents the constraint relationship between decision variables.

[0195] Finally, use the multi-objective optimization algorithm to find a trade-off optimal solution for comprehensive optimization and scenario simulation. The goal during the optimization process is to maximize the comprehensive score, and at the same time, under the satisfaction of constraint conditions (such as policy restrictions, environmental sustainability, etc.), obtain the optimal benefits under different strategies.

[0196] The present invention generates different optimization scenarios for each region, including "optimal scenario", "worst scenario" and "average scenario". Through three-dimensional visualization, it shows the trade-off of benefits under each scenario, helps policy makers intuitively understand the effects of various management strategies, and make scientific decisions.

[0197] Step 6, based on the spatial distribution heat map, time series curve, correlation matrix, dynamic scoring system and multi-objective optimization model, form a closed-loop automated operation chain for data-driven problem diagnosis and strategy generation.

[0198] The specific implementation includes the following links:

[0199] a) Multi-dimensional short board identification:

[0200] By comparing the regional benefit scores (economy / society / environment) with the global optimal value of the difference, locate the regions with backward comprehensive performance; combine the results of step 2 for spatio-temporal hotspot analysis (spatial clustering I + "time anomaly segment"), and mark the high-priority governance regions that have been in an inefficient state for a long time.

[0201] b) Root cause tracing analysis:

[0202] Based on the input-output depth correlation matrix Ω obtained in step 3, quantify the degree of influence of each variable on the target index, lock the key driving variables, and identify the dominant restrictive factors.

[0203] c) Dynamic strategy generation:

[0204] According to the multi-objective optimization results obtained in step 5, combined with the input-output depth correlation matrix Ω obtained in step 3, generate a customized zoning plan: for the ecological short board regions, formulate tillage method adjustment rules in combination with the driving relationships in Ω (such as mandating straw return to the field + rotation mode in high erosion areas); for the economically inefficient regions, design a subsidy-price linkage optimization algorithm in combination with the restrictive and collaborative factors in Ω (such as dynamically binding the subsidy coefficient to the market price volatility);

[0205] Output a visual decision-making report (in PDF format), including the priority ranking of strategies, constraint conditions, and expected benefit increase.

[0206] The above are only the specific steps of the present invention, which do not constitute any limitation to the protection scope of the present invention; all technical solutions formed by equivalent transformation or equivalent substitution fall within the scope of the rights protection of the present invention; the parts not detailed in the present invention belong to the well-known technologies in the art.

Claims

1. An intelligent diagnosis method for the collaborative optimization of the comprehensive benefits of black soil protection and utilization, characterized in that The intelligent diagnosis method for the collaborative optimization of the comprehensive benefits of black soil protection and utilization includes the following steps: Step 1: Construct an input-output dataset of multi-source spatio-temporal simulation data, and convert the category information therein into the form of numerical vectors; Step 2: Draw a spatial distribution heat map and a time series curve based on the input-output dataset of simulation data; Step 3: Construct a correlation matrix between input variables and output variables based on the input-output dataset of simulation data; Step 4: Construct a scoring system based on the historical optimal values of each output variable in the input-output dataset of simulation data; Step 5: Optimize the weight coefficients of each target dimension and the contribution weights within the benefits based on the scores of different dimensions in the scoring system to construct a multi-objective optimization model; Step 6: Based on the spatial distribution heat map, time series curve, correlation matrix, dynamic scoring system and multi-objective optimization model, form a closed-loop automated operation chain for data-driven problem diagnosis and strategy generation.

2. The intelligent diagnosis method for collaborative optimization of comprehensive benefits of black soil protection and utilization according to claim 1, wherein, In Step 1, the following steps are further included: Step 11: Construct an input-output dataset of multi-source spatio-temporal simulation data. The simulation data includes an input dataset X, an output dataset Y, and a sample set A, where: Input data set X = {x (1) , x (2) , …, x (12)} includes: Agricultural operation parameters: tillage mode x (1) , fertilization type x (2) , irrigation method x (3) , straw treatment method x (4) ; Climate parameters: temperature x (5) , precipitation x (6) , sunshine duration x (7) , wind speed x (8) ; Soil parameters: soil pH value x (9) , thickness of black soil layer x (10) , soil moisture x (11) , soil texture x (12) ; Output data set Y = {y (econ1) , y (econ2) , y (soc1) , y (soc2) , y (soc3) , y (soc4) , y (soc5) , y (eco1) , y (eco2) , y (eco3)}, including: Economic benefit indicators: net income of farmers y (econ1) , policy subsidy input y (econ2) ; Social benefit indicators: total grain output y (soc1) , soybean output y (soc2) , corn output y (soc3) , rice output y (soc4) , food equivalent y (soc5) ; Ecological benefit indicators: Soil quality y (eco1) , Net carbon emissions y (eco2) , Soil erosion amount y (eco3) ; Sample set A = {a1, a2, …, a m}, where each sample represents the observation record of a certain spatial unit at a certain specific time point, the data volume is m, and each sample contains the input data set X, the output data set Y, and the year, longitude and latitude information corresponding to the data set.

3. The intelligent diagnosis method for collaborative optimization of the comprehensive benefits of black soil protection and utilization according to claim 2, wherein In Step 1, the following steps are further included: Step 12: Convert the category information in the agricultural operation parameters into the form of numerical vectors through one-hot encoding; Step 121: Encode each category variable separately: The categorical variables include the following four items: x (1) Cultivation patterns, with a total of n1 categories; x (2) Fertilizer types, with a total of n2 categories; x (3) Irrigation methods, with a total of n3 categories; x (4) Straw treatment methods, with a total of n4 categories; For each variable x (j) (j = 1, 2, 3, 4), all its possible values form a category set: C (j) represents the set of all possible values of the j-th categorical variable x among the agricultural operation parameters (j) (j = 1, 2, 3, 4); represents the categorical variable x (j) (j = 1, 2, 3, 4) of the categorical set C (j) the k-th category in, where the value range of k is 1 ≤ k ≤ n j ; n j represents the total number of categories of the variable x (j) ; For each sample A i In the variable x (j) (j = 1, 2, 3, 4), the corresponding encoding is denoted as: Denote the i-th sample A i One-hot encoding vector on the j-th categorical variable x (j) whose length is the total number of categories n of this variable j , Denote the i-th sample A i The encoding value of the k-th category on the j-th categorical variable x (j) where the index value k ranges from 1 ≤ k ≤ n j , Denote the vector is a vector in the n j -dimensional real space. Step 122: Concatenate the sample category encoding vectors: Step 123: The one-hot encoding matrix of all samples is: Stack the category vectors of all m samples by rows to obtain the one-hot encoding matrix: where: m is the total number of samples, and each row z i represents the encoding result of the categorical variable of the i-th sample; each column corresponds to the position of a certain category value, with the value 1 indicating belonging to that category and 0 indicating not belonging.

4. The intelligent diagnosis method for collaborative optimization of comprehensive benefits of black soil protection and utilization according to claim 1, characterized in that In Step 2, use Moran's I index to measure the autocorrelation of spatial data and judge whether there is a tendency of aggregation or dispersion of the data in space: Where: m is the total number of samples, v p and v q are the values of the p-th and q-th samples on a certain variable respectively, is the mean value over all samples; T pq is the spatial weight matrix, representing the spatial relationship between positions p and q; For each variable at each time point, draw a spatial distribution heat map, and use a color gradient to represent the magnitude of the variable value; for each spatial unit, draw a time series curve of a certain variable in different years, and slide to calculate the historical average value and standard deviation.

5. The intelligent diagnosis method for collaborative optimization of comprehensive benefits of black soil protection and utilization according to claim 1, wherein In Step 3, the following steps are further included: Step 31: For the correlation between continuous variables, use the Pearson correlation coefficient to measure the linear relationship; In the Pearson correlation coefficient, It is expressed as: The input data set X = {x (1) , x (2) , …, x (12)}, where e is the input variable index, e = 1, 2, …, 12, and x (e) represents the input variable; Y = {y (econ1) , y (econ2) , …, y (eco3)}, where o is the output variable index, and y (o) represents the output variable; the sample set is A = {a1, a2, …, a m}, where the input and output variables in the k-th sample a k are respectively and are respectively the means of the variables x (e) and y (o) ; is the Pearson correlation coefficient, Step 32: Combine all input-output variables in pairs to generate a correlation matrix Correlation Matrix The dimension is 12×10, where the rows correspond to 12 input variables, the columns correspond to 10 output metrics, and the elements are the absolute values of the correlation coefficients of the input and output variables. For continuous variables, the absolute value of the Pearson coefficient is taken, and for one-hot encoded variables, the absolute value of the point-biserial coefficient is taken. The value range is all [0,1]. Step 33: Obtain all input-output combinations to construct a causal test score matrix Γ; F statistic formula: Among them, F e,o represents the statistic of the input variable x (e) and the output variable y (o) Rss res and RSS unres are the residual sum of squares of the restricted model and the unrestricted model respectively, L is the number of lag terms, and pcount is the number of explanatory variables in the unrestricted model; Based on F e,o Calculate the corresponding P value P based on the value e,o And construct a causal score: ζ e,o represents the causal score; Obtain all input-output combinations to construct a causal test score matrix Γ of dimension 12×10 = Step 34, construct an XGBoost model for each output variable to train its non-linear relationship with all input variables x (1) , x (2) , …, x (12) , and use the SHAP method to quantify the relative contribution of each input variable: represents the combination of the e-th input variable and the o-th output variable, is the k-th sample variable x (e) for the output y (o) the SHAP value of, where e′ is the summation index variable, represents the combination of the e′-th input variable and the o-th output variable, represents x (e) for y (o) the normalized contribution of; Finally, the final contribution matrix S with dimensions 12×10 is obtained. (shap) , where 12 corresponds to the 12 input variables in the input data set X, and 10 corresponds to the 10 output variables in the input data set Y; Use the equal-weight method to perform weighted summation on the relationship matrix of the results of correlation, causality and SHAP analysis to obtain a deep association matrix: denotes the correlation matrix, Γ denotes the causal test score matrix, and Ω denotes the deep association matrix; Step 35, save the correlation matrix causal test score matrix Γ, SHAP matrix S (shap) and the deep association matrix Ω as a CSV file, and save the negative correlation prompt to a text file.

6. The intelligent diagnosis method for collaborative optimization of comprehensive benefits of black soil protection and utilization according to claim 1, characterized in that, In Step 4, the score of any sample on the output data of the type item is: k represents the sample serial number, indicating the score of the k-th sample on the output data of the type item, is the output variable of the k-th sample, is the historical optimal value among all samples; The score of the k-th sample on this category is: Θ (t) is the output variable type, where t ∈ {econ, soc, eco}; represents the score of the k-th sample on the econ, soc, eco classes; The total score of sample k is: TotalScore k Denotes the total score of the k-th sample, Denotes the score of the k-th sample in the econ class, Denotes the score of the k-th sample in the soc class, Denotes the score of the k-th sample in the eco class.

7. The intelligent diagnosis method for collaborative optimization of comprehensive benefits of black soil protection and utilization according to claim 1, wherein In Step 5, define the multi-objective function as a weighted combination of the scores of three types of benefits: Among them, is the weight coefficient of three types of targets; The score of the k-th sample on the t-th class Optimal under different strategic objectives: Indicates optimal under different strategic objectives; is the score of the k-th sample on the output data of the type item, represents the internal variable of benefit t contribution weight; In the optimization process, introduce the following constraint conditions to ensure the trade-off and rationality of different benefit dimensions and avoid imbalance caused by extreme optimization of a certain dimension. The constraint conditions can be formally expressed as: where, g(x (e) , y (o) ) represents the constraint relationship between decision variables; Finally, use the multi-objective optimization algorithm to obtain a trade-off optimal solution for comprehensive optimization and scenario simulation; the goal in the optimization process is to maximize the comprehensive score, and at the same time, under the condition of meeting the constraint conditions, obtain the optimal benefits under different strategies.

8. The intelligent diagnosis method for collaborative optimization of comprehensive benefits of black soil protection and utilization according to claim 1, characterized in that, In Step 6, the closed-loop automated operation chain for data-driven problem diagnosis and strategy generation includes multi-dimensional short-board identification; Multi-dimensional shortcoming identification locates regions with lagging comprehensive performance by comparing the regional benefit scores with the global optimal value and analyzes the generated heat map of spatial distribution and time series curve to mark high-priority governance regions that have been in an inefficient state for a long time.

9. The intelligent diagnosis method for collaborative optimization of comprehensive benefits of black soil protection and utilization according to claim 1, wherein In step 6, the closed-loop automated operation chain of data-driven problem diagnosis and strategy generation includes root cause tracing and analysis; Based on the depth correlation matrix Ω, the root cause tracing and analysis quantifies the influence degree of each variable on the target index, locks the key driving variables, and identifies the dominant restrictive factors.

10. The intelligent diagnosis method for collaborative optimization of comprehensive benefits of black soil protection and utilization according to claim 1, wherein, In step 6, the closed-loop automated operation chain of data-driven problem diagnosis and strategy generation includes dynamic strategy generation; The dynamic strategy generation generates a customized regional plan according to the results of the multi-objective optimization algorithm and in combination with the depth correlation matrix Ω: For the ecological short-board areas, tillage method adjustment rules are formulated in combination with the driving relationships in the depth correlation matrix Ω; for the economically inefficient areas, a subsidy-price linkage optimization algorithm is designed in combination with the restrictive and synergistic factors in the depth correlation matrix Ω; Finally, a visual decision-making report is output, including the strategy priority ranking, constraint conditions, and expected benefit increase.

Citation Information

Patent Citations

  • Land utilization type spatial distribution prediction method and system

    CN118229162A