A method and system for tracing the cause of cultivated land quality evaluation anomalies by fusing machine learning
By combining multi-scale feature extraction and fusion with Bayesian neural networks and evidence-based deep learning, the problems of insufficient multi-scale analysis and uncertainty quantification in farmland quality assessment are solved. This enables accurate location and source tracing of farmland quality anomalies, improving the credibility of source tracing and resource utilization efficiency.
Patent Information
- Application Number
- CN202511021099.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-07-24
AI Technical Summary
Existing technologies for evaluating arable land quality suffer from insufficient multi-scale analysis, lack of uncertainty quantification, and low reliability of traceability results, leading to inadequate accuracy in anomaly identification and traceability, and failing to provide effective risk assessment and decision support.
This study employs a combination of multi-scale feature extraction and fusion, Bayesian neural networks, evidence deep learning, and gradient boosting trees. Through wavelet transformation, attention mechanism, Bayesian neural network, Dempster-Shafer evidence theory, and causal graph network, it achieves multi-scale analysis, uncertainty quantification, and source tracing analysis of cultivated land quality data, and combines adaptive source tracing strategies for differentiated processing.
It has improved the accuracy of capturing multi-scale characteristics and identifying anomalies in farmland quality evaluation indicators, distinguished between cognitive and random uncertainties, provided a basis for risk assessment, accurately located and tracked the causes of anomalies, and improved the credibility of source tracing and resource utilization efficiency.
Smart Images

Figure CN120524352B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of cultivated land quality evaluation and agricultural information technology, more particularly, it relates to a cultivated land quality evaluation abnormality tracing method and system fusing machine learning. BACKGROUND
[0002] Cultivated land quality evaluation is an important basis for scientific formulation of agricultural production strategies and soil improvement measures in modern agricultural management. With the increasing attention of the state to the protection of cultivated land quality, accurately identifying cultivated land quality abnormalities and tracing their causes has become a core link in the development of precision agriculture. At present, cultivated land quality evaluation mainly adopts multi-index comprehensive evaluation method, which comprehensively analyzes multi-dimensional data of soil physical and chemical properties, environmental factors and management measures.
[0003] Traditional cultivated land quality abnormality detection and tracing methods are mainly based on statistical analysis and experience judgment, such as using Mahalanobis distance method, cluster analysis and principal component analysis statistical method to identify abnormal points, and then tracing through expert experience or simple correlation analysis. In recent years, with the development of machine learning technology, some researches have begun to use classification models, regression models and deep learning methods for abnormality detection, such as random forest, support vector machine and deep neural network. However, the existing technology has the following technical problems in cultivated land quality abnormality tracing: first, cultivated land quality evaluation indexes show different characteristics at different scales, and single scale analysis cannot fully capture multi-scale abnormal phenomena; second, existing methods are usually based on deterministic models, lack of quantification of uncertainty, and cannot provide risk assessment basis for decision-making; third, traditional methods are difficult to distinguish between cognitive uncertainty and accidental uncertainty, making the credibility evaluation of the tracing result lack scientific basis; finally, the abnormality tracing result lacks reliability measurement, which cannot provide risk level evaluation for decision-makers, restricting the precision management ability of cultivated land quality monitoring and management.
[0004] Therefore, there is an urgent need for a cultivated land quality evaluation abnormality tracing method that can integrate multi-scale analysis, uncertainty quantification and causal inference to improve the accuracy of abnormality identification and the reliability of tracing decision. SUMMARY
[0005] The present application provides a cultivated land quality evaluation abnormality tracing method and system fusing machine learning, which solves the technical problems of insufficient precision of cultivated land quality evaluation abnormality identification and tracing, inability to quantify uncertainty and lack of reliability measurement in related technologies.
[0006] The present application discloses a cultivated land quality evaluation abnormality tracing method fusing machine learning, comprising the following steps:
[0007] Multi-scale feature extraction and fusion: Utilize Wavelet transform to perform multi-scale decomposition on the cultivated land quality evaluation data, obtain feature representations at different scales, and realize adaptive weighted fusion of features through attention mechanism;
[0008] Bayesian neural network construction: Construct a Bayesian neural network model, introduce parameter prior distribution and variational inference method, model the fused features, and output prediction results containing uncertainty information;
[0009] Evidence deep learning implementation: Implement evidence deep learning method, generate evidence vector, explicitly quantify and decompose uncertainty, and distinguish cognitive uncertainty and random uncertainty;
[0010] Anomaly detection and tracing: Based on gradient boosting tree and causal graph network, realize anomaly detection and tracing analysis, locate and track the detected anomalies, and identify potential causes and transmission paths;
[0011] Adaptive tracing strategy execution: Based on uncertainty decomposition and tracing result reliability, develop and execute adaptive tracing strategy, adopt differentiated processing schemes according to different types of uncertainty, and improve resource utilization efficiency.
[0012] Further, the multi-scale feature extraction and fusion specifically includes:
[0013] Data preprocessing: Standardize the original cultivated land quality evaluation data matrix , handle missing values, and obtain the normalized data matrix ;
[0014] Wavelet transform multi-scale decomposition: Apply multi-layer Wavelet transform to the normalized data matrix , obtain the approximate coefficient matrix and detail coefficient matrix at different scales;
[0015] Attention mechanism feature fusion: Through attention mechanism, adaptively weight and fuse features at different scales to obtain fused features .
[0016] Further, the Bayesian neural network construction specifically includes:
[0017] Bayesian neural network architecture: Based on multi-layer perception, construct the basic architecture of Bayesian neural network, including hidden layers;
[0018] Parameter prior distribution introduction: Introduce Gaussian prior distribution to the weight parameters in the neural network;
[0019] Variational Inference Implementation: Using variational inference method to approximate the posterior distribution
[0020] Predictive Distribution Output: Estimate the predictive distribution of new samples through Monte Carlo sampling method, output the prediction results containing uncertainty information.
[0021] Further, the evidence deep learning implementation specifically includes:
[0022] Evidence Theory Model Construction: Based on Dempster-Shafer evidence theory, introduce Dirichlet distribution to model the confidence of classification results;
[0023] Evidence Vector Generation Network: Add evidence generation layer based on Bayesian neural network, map network output to evidence vector;
[0024] Uncertainty Quantification and Decomposition: Calculate the expected probability of categories through evidence vector, and decompose uncertainty into cognitive uncertainty and random uncertainty;
[0025] Loss Function of Evidence Deep Learning: Define a comprehensive loss function containing negative log-likelihood loss and evidence regularization term for model training.
[0026] Further, the anomaly detection and tracing specifically includes:
[0027] Gradient Boosting Tree Anomaly Detection: Based on fusion features, cognitive uncertainty and random uncertainty, use gradient boosting tree algorithm to build anomaly detection model;
[0028] Anomaly Feature Contribution Analysis: Use SHAP value to analyze the contribution of each feature, identify the key features causing anomalies;
[0029] Causal Diagram Construction and Analysis: Construct the causal relationship graph between the quality of cultivated land indicators, including structure learning and parameter learning;
[0030] Anomaly Tracing Path Generation: Based on anomaly feature contribution analysis and causal graph, generate anomaly tracing path, track the cause and impact of anomalies.
[0031] Further, the adaptive tracing strategy execution specifically includes:
[0032] Uncertainty Threshold Adaptive Adjustment: According to historical tracing results and expert feedback, establish uncertainty threshold adaptive adjustment mechanism;
[0033] A differentiated traceability strategy is designed based on the uncertainty decomposition result, which is differentiated for samples with high cognitive uncertainty, high random uncertainty and low uncertainty, wherein samples with cognitive uncertainty greater than a cognitive uncertainty threshold are identified as samples with high cognitive uncertainty, samples with random uncertainty greater than a random uncertainty threshold are identified as samples with high random uncertainty, and samples with total uncertainty lower than a total uncertainty threshold are identified as samples with low uncertainty.
[0034] A traceability result credibility score is defined in combination with the total uncertainty and the influence strength.
[0035] A resource optimization allocation algorithm is designed based on the traceability result credibility and the abnormal emergency level to maximize the traceability effect.
[0036] Further, the uncertainty decomposition formula is:
[0037] Cognitive uncertainty , and the calculation formula is:
[0038]
[0039] wherein is the sum of the evidence vectors, is the Dirichlet parameter of the i-th category.
[0040] Random uncertainty , and the calculation formula is:
[0041]
[0042] wherein is the prediction probability of the i-th category.
[0043] Further, the abnormal traceability path generation includes upward tracking and downward tracking, and the specific steps are as follows:
[0044] Features with a contribution degree higher than a preset contribution degree threshold identified in the causal relationship diagram are taken as a starting point, and upstream factors that may cause abnormal feature changes are tracked upward, i.e., parent nodes of the starting point in the causal relationship diagram are taken as the upstream factors; downstream indicators that may be affected by the abnormality are tracked downward, i.e., child nodes of the starting point in the causal relationship diagram are taken as the downstream indicators.
[0045] Further, the differentiated traceability strategy adopts different processing schemes for different types of uncertainty:
[0046] For samples with high cognitive uncertainty: prefer to collect more relevant data, introduce expert knowledge for auxiliary judgment, and increase the training of the model on this type of sample;
[0047] For samples with high random uncertainty: increase the sampling frequency or density, use integrated methods to improve the robustness of prediction, and consider introducing additional sensors or measurement methods;
[0048] For samples with low uncertainty: directly use the causal tracing result to execute the standardized problem solving process.
[0049] The application provides a cultivated land quality evaluation anomaly tracing system fused with machine learning, which is used to execute the steps in the cultivated land quality evaluation anomaly tracing method fused with machine learning.
[0050] The multi-scale feature extraction and fusion module is used for multi-scale decomposition of cultivated land quality evaluation data by using Wavelet transformation, obtaining feature representations at different scales, and realizing adaptive weighted fusion of features through an attention mechanism.
[0051] The Bayesian neural network module is used to construct a Bayesian neural network model, model the fused features by introducing a parameter prior distribution and a variational inference method, and output prediction results containing uncertainty information.
[0052] The evidence deep learning module is used to implement an evidence deep learning method, explicitly quantify and decompose uncertainty by generating an evidence vector, and distinguish cognitive uncertainty from random uncertainty.
[0053] The anomaly detection and tracing module is used to realize anomaly detection and tracing analysis based on gradient boosting trees and causal graph networks, locate and track detected anomalies, and identify potential causes and propagation paths.
[0054] The adaptive tracing strategy execution module is used to develop and execute adaptive tracing strategies based on uncertainty decomposition and tracing result reliability, adopt differentiated processing schemes according to different types of uncertainty, and improve resource utilization efficiency.
[0055] The application has the following advantages:
[0056] The multi-scale feature extraction and fusion method can comprehensively capture the characteristic changes and abnormal performance of cultivated land quality evaluation indicators at different scales, improve the accuracy and comprehensiveness of anomaly identification.
[0057] The method combining Bayesian neural networks and evidence deep learning realizes the quantification and decomposition of uncertainty, can distinguish cognitive uncertainty from random uncertainty, and provides risk assessment basis for anomaly tracing decision-making.
[0058] The abnormality detection and tracing method based on gradient boosting tree and causal graph network can not only accurately locate the abnormality, but also track the cause and propagation path of the abnormality, and realize the explainable tracing of the abnormality.
[0059] The adaptive tracing strategy can take different processing schemes according to different types of uncertainty, improve resource utilization efficiency, and realize precise agricultural management.
[0060] Through the credibility evaluation of the tracing result and the resource optimization allocation algorithm, the risk level evaluation and resource allocation suggestion are provided for the decision maker, and the precise management ability of the cultivated land quality monitoring and management is improved. BRIEF DESCRIPTION OF DRAWINGS
[0061] Figure 1 is a flowchart of a cultivated land quality evaluation abnormality tracing method fusing machine learning of the present application;
[0062] Figure 2 is a flowchart of a multi-scale feature extraction and fusion step of the present application;
[0063] Figure 3 is a flowchart of a Bayesian neural network construction step of the present application;
[0064] Figure 4 is a flowchart of an evidence deep learning implementation step of the present application;
[0065] Figure 5 is a flowchart of an abnormality detection and tracing step of the present application;
[0066] Figure 6 is a flowchart of an adaptive tracing strategy execution step of the present application. DETAILED DESCRIPTION
[0067] The subject matter described herein will now be discussed with reference to example implementations. It should be understood that the discussion of these implementations is merely meant to provide a better understanding of the subject matter described herein and can include changes, modifications, or additions of elements to the functions and arrangements of elements discussed without departing from the scope of the present disclosure. Various examples can omit, substitute, or add various procedures or components as appropriate, or in appropriate combination. Also, some features described for one example can be combined with features described for another example.
[0068] In at least one embodiment of the present application, a cultivated land quality evaluation abnormality tracing method fusing machine learning is disclosed, as shown in Figures 1 to 6 includes the following steps:
[0069] Step 1, multi-scale feature extraction and fusion: Wavelet transform is used to decompose the cultivated land quality evaluation data at multiple scales to obtain feature representations at different scales, and an attention mechanism is used to achieve adaptive weighted fusion of features.
[0070] Step 1.1, data preprocessing;
[0071] The original cultivated land quality evaluation data is standardized and missing value processed. Let the original cultivated land quality evaluation data matrix be wherein denotes the number of samples, denotes the number of indicators. The standardization process uses the following formula:
[0072] ;
[0073] wherein is the mean vector of each indicator, is the standard deviation vector of each indicator, is the standardized data.
[0074] Step 1.2, Wavelet transform multi-scale decomposition;
[0075] The standardized data is decomposed at multiple scales by Wavelet transform to extract feature representations in different frequency domains. Let be the decomposition level, , be the number of decomposition levels, then the multi-scale decomposition can be represented as:
[0076] ;
[0077] wherein denotes the th layer Wavelet transform operation, denotes the feature representation after the th layer decomposition. The specific decomposition process uses discrete Wavelet transform. For a one-dimensional signal , the decomposition formula is:
[0078] ;
[0079] ;
[0080] wherein is the scaling function, is the wavelet function, is the initial scale, is the translation parameter, is the signal length, , They represent the approximate coefficient matrix and detail coefficient matrix of different scales respectively. For multi-dimensional data, tensor product is used for expansion.
[0081] Step 1.3, feature fusion of attention mechanism;
[0082] The features of different scales are adaptively weighted and fused through the attention mechanism to obtain multi-granularity feature representation. The features obtained by layer decomposition are , then the fused feature representation It can be expressed as:
[0083] ;
[0084] in, is the attention weight, which is calculated as follows:
[0085] ;
[0086] ;
[0087] in, 、 and is a learnable parameter, is the hyperbolic tangent activation function, Represents a transpose operation. 、 Respectively represent 、 The importance score of layer features is calculated. In this way, the model can adaptively focus on important features at different scales and improve the ability to identify multi-scale anomalies.
[0088] The multi-scale feature extraction and fusion method combining Wavelet transform and attention mechanism can effectively capture the characteristic changes of cultivated land quality evaluation indicators at different scales, and provide a more comprehensive feature representation for subsequent anomaly detection and traceability.
[0089] Step 2: Bayesian neural network construction: Build a Bayesian neural network model, introduce parameter prior distribution and variational inference methods to model the fusion features, and output prediction results containing uncertainty information;
[0090] Step 2.1, Bayesian neural network architecture;
[0091] The basic architecture of Bayesian neural network is constructed based on multi-layer perceptron. Assume that the network contains hidden layer, The weight matrix of the layer is , the bias vector is , the activation function is Then the forward propagation process is as follows:
[0092] ;
[0093] ;
[0094] ;
[0095] where, is the input layer of the network, is the fusion feature obtained in step 1, is the output of the first hidden layer, is the output of the second hidden layer, is the output of the third hidden layer, is the output of the network, which is used to represent the class prediction of normal or abnormal, , and respectively represent the weight matrix and the bias vector of the output layer.
[0096] Step 2.2, parameter prior distribution introduction;
[0097] The weight parameters in the neural network are introduced with a prior distribution, and a Gaussian prior is adopted, and its calculation formula is:
[0098] ;
[0099] where, is the prior distribution of the weight parameter ; is a Gaussian distribution; is the mean vector, indicating that the expected value of the prior hypothesis weight is zero; is the variance; is the unit matrix;
[0100] The complete expression of the Bayesian neural network is:
[0101] ;
[0102] where, represents the prediction distribution of the output under the condition of given input and training data set , represents the training data set; is the posterior distribution of the parameter; is the output distribution under the condition of given parameter and input .
[0103] Step 2.3, variational inference implementation;
[0104] Since the posterior distribution It is usually difficult to compute directly, we use variational inference method to approximate:
[0105] ;
[0106] where is a parameterized approximate posterior distribution, which can be chosen as follows:
[0107] ;
[0108] where is a Gaussian distribution, is the mean vector of weights, is the variance parameter, is the identity matrix.
[0109] The optimization objective is to minimize the KL divergence, which is calculated as:
[0110] ;
[0111] where is the loss function of variational inference; denotes the KL divergence, which measures the difference between two probability distributions;
[0112] Through the variational lower bound (ELBO) transformation, the optimization objective is equivalent to:
[0113] ;
[0114] where denotes the evidence lower bound loss function, denotes the expectation with respect to the distribution is the log-likelihood function, which measures the fitting degree of the model to the training data, denotes the KL divergence between the approximate posterior distribution and the prior distribution, which serves as a regularization term to prevent overfitting. Step 2.4, predict the distribution output;
[0115] After the model training is completed, the prediction distribution of new samples
[0116] can be estimated by the Monte Carlo sampling method:
[0117] ;
[0118] where denotes the output given the new input sample and the training data set the predictive distribution; the number of Monte Carlo samples; denotes the prediction of the summation operation of the denotes the output probability under the given weight parameter and the new input ; denotes the parameter sample drawn from the approximate posterior distribution .
[0119] Bayesian neural networks can not only give prediction results, but also quantify the uncertainty of prediction, providing a foundation for subsequent evidence deep learning and uncertainty decomposition.
[0120] Step 3, evidence deep learning implementation: implement the evidence deep learning method, and explicitly quantify and decompose the uncertainty by generating evidence vectors, and distinguish cognitive uncertainty and random uncertainty;
[0121] Step 3.1, construction of evidence theory model;
[0122] Based on Dempster-Shafer evidence theory, an evidence deep learning model is constructed. For binary classification problems (normal samples and abnormal samples), Dirichlet distribution is introduced to model the confidence of classification results, and the calculation formula is:
[0123] ;
[0124] where, denotes the category probability vector, denotes the probability that the sample belongs to the normal class, denotes the probability that the sample belongs to the abnormal class, denotes the evidence vector, , correspond to the evidence parameters of the normal class and the abnormal class respectively, , is the parameter of Dirichlet distribution, which represents the confidence of the model that the sample belongs to the th class, is the evidence quantity of the th class.
[0125] The probability density function of Dirichlet distribution is:
[0126] ;
[0127] where, denotes the probability density function of Dirichlet distribution with parameter , is the gamma function, denotes the sum of all evidence parameters, denotes the product of the gamma function values of each evidence parameter, denotes the weighted product of the probabilities of each class, where, denotes the normal class, denotes the abnormal class.
[0128] Step 3.2, evidence vector generation network;
[0129] On the basis of Bayesian neural network, add evidence generation layer, map the output of the last layer of the network to evidence vector , the calculation formula is:
[0130] ;
[0131] where, denotes the output of the feature extraction network, is the evidence generation function, which can be an exponential function to ensure that the evidence is non-negative, and the calculation formula is:
[0132] ;
[0133] where, is the evidence vector, is the activation value of the network output.
[0134] Step 3.3, uncertainty quantification and decomposition;
[0135] Calculate the expected value of the class probability through the evidence vector, the calculation formula is:
[0136] ;
[0137] where, is the expected value of the prediction probability of the th class, denotes the evidence parameter of the th class, is the sum of the evidence vector.
[0138] The uncertainty of class prediction can be decomposed into two parts:
[0139] Cognitive uncertainty : Uncertainty caused by insufficient knowledge or data, which can be reduced by collecting more data. The variance of Dirichlet distribution is used to represent:
[0140] ;
[0141] Random uncertainty : Uncertainty caused by the inherent randomness of the data cannot be reduced by increasing the amount of data. It is expressed using the prediction entropy, which is calculated as follows:
[0142] ;
[0143] in, Indicates the The expected value of the predicted probability of the class;
[0144] Total uncertainty , and its calculation formula is:
[0145] ;
[0146] Step 3.4, loss function of evidence deep learning;
[0147] Defining the loss function for evidence-based deep learning consists of two parts:
[0148] Negative log-likelihood loss , and its calculation formula is:
[0149] ;
[0150] in, For samples The true label, is the Dirichlet parameter corresponding to the label, For samples The sum of all parameters.
[0151] Evidence regularization term : Encourage the model to generate less evidence for uncertain samples. The calculation formula is:
[0152] ;
[0153] in, is the indicator function, For samples No. Class evidence parameters.
[0154] The comprehensive loss function is , and its calculation formula is:
[0155] ;
[0156] in, is the regularization coefficient.
[0157] Through the evidence-based deep learning method, the uncertainty of the anomaly detection results of cultivated land quality evaluation can be quantified and decomposed in a refined manner, providing a more reliable decision-making basis for subsequent anomaly tracing.
[0158] Step 4, anomaly detection and tracing: based on gradient boosting tree and causal graph network, realize anomaly detection and tracing analysis, locate and track the detected anomalies, identify potential causes and propagation paths;
[0159] Analyze the contribution of abnormal features to locate the detected anomalies, and generate a tracing path based on the causal graph to track the potential causes and propagation paths.
[0160] Step 4.1, gradient boosting tree anomaly detection;
[0161] Based on the output of the aforementioned evidence deep learning, use the gradient boosting tree algorithm to build an anomaly detection model. For each sample , use the following feature vector as input:
[0162] ;
[0163] Where, is the feature vector of the sample, is the fusion feature of the sample, is the prediction probability, and are cognitive uncertainty and random uncertainty, respectively.
[0164] The training goal of the gradient boosting tree model is to minimize the following loss function, whose calculation formula is:
[0165] ;
[0166] Where, is the loss function (such as cross-entropy), is the true label of the sample, is the predicted value, is the complexity penalty term of the th tree, is the total number of samples, is the number of trees.
[0167] After the model training is completed, the anomaly detection is performed on the new sample, and the anomaly probability score is output, and the anomaly is determined based on the set threshold :
[0168] ;
[0169] Step 4.2, anomaly feature contribution analysis;
[0170] For the detected abnormal samples, use SHAP value to analyze the contribution of each feature to identify the key features causing the anomaly:
[0171] ;
[0172] where, denotes the feature contribution to the prediction of the sample , is the full set of features, denotes the subset of features not containing the feature , denotes the prediction of the sample using only the subset of features , denotes the number of parameters in the set, denotes the model output using the subset of features plus the feature , denotes the factorial operation.
[0173] Ranking the features by their contribution to generate a feature importance list , , , denote the top , , important features, is the total number of features, where, , denote the contribution of the feature , , to the prediction of the sample .
[0174] Step 4.3. Causal graph construction and analysis;
[0175] Constructing the causal graph between the quality indicators of cultivated land , where the node set denotes each quality indicator, and the edge set denotes the causal relationship between indicators. The construction of the causal graph includes two steps:
[0176] Structure learning: learn the causal relationship between variables using the PC algorithm or score-based method:
[0177] ;
[0178] where, is the optimal graph structure, is the score function of the graph structure under the given data set.
[0179] Parameter learning: determining conditional probability distributions wherein is the set of parent nodes of node .
[0180] Step 4.4, abnormality trace path generation;
[0181] Based on the constructed causal graph and the abnormal feature contribution analysis result, an abnormality trace path is generated:
[0182] For the detected abnormal sample , a feature subset with the highest contribution degree is identified wherein, is the contribution degree threshold, represents a feature importance ranking list of sample .
[0183] In the causal graph, the node in is taken as a starting point to track potential causal paths:
[0184] Upward tracking: finding upstream factors (causes) that may cause abnormal feature changes;
[0185] Downward tracking: evaluating downstream indicators (results) that may be affected by the abnormality;
[0186] The specific steps are: taking the features identified in the causal relationship graph and having a contribution degree higher than the preset contribution degree threshold as a starting point, tracking the parent nodes of the starting point in the causal relationship graph as upstream factors, and tracking the child nodes of the starting point in the causal relationship graph as downstream indicators.
[0187] The impact strength of each node on the path is calculated , based on the conditional probability and intervention evaluation, and the calculation formula is:
[0188] ;
[0189] wherein, represents the probability of feature taking value under the condition that feature takes value , represents that variable is intervened to take value , represents the marginal probability of feature taking value .
[0190] According to the impact strength ranking, a set of most possible abnormality trace paths is generated wherein, 、 、 respectively represent the first 、 、 path set, is the total number of paths, and each path is a sequence of nodes arranged in causal order.
[0191] By combining gradient boosting trees and causal graph networks, not only can the anomalies in cultivated land quality evaluation be accurately detected, but also the causes and possible impact paths of the anomalies can be analyzed in depth, providing decision basis for subsequent adaptive tracing strategies.
[0192] Step 5, adaptive tracing strategy execution: based on uncertainty decomposition and tracing result reliability, adaptive tracing strategy is formulated and executed, and different processing schemes are adopted according to different types of uncertainty to improve resource utilization efficiency;
[0193] Step 5.1, adaptive adjustment of uncertainty threshold;
[0194] According to the historical tracing results and expert feedback, an adaptive adjustment mechanism of uncertainty threshold is established, which includes cognitive uncertainty threshold , random uncertainty threshold and total uncertainty threshold :
[0195] ;
[0196] ;
[0197] ;
[0198] wherein, and are the thresholds of cognitive uncertainty and random uncertainty in the first round of iteration, and are the thresholds of cognitive uncertainty and random uncertainty in the first round of iteration, is the learning rate, is the performance change amount, and respectively represent the total uncertainty threshold in the first round of iteration and the first round of iteration.
[0199] If the cognitive uncertainty is greater than the cognitive uncertainty threshold is identified as a high cognitive uncertainty sample, and the random uncertainty is greater than a random uncertainty threshold is identified as a high random uncertainty sample, and the total uncertainty is lower than a total uncertainty threshold is identified as a low uncertainty sample.
[0200] Step 5.2, design of differentiated traceability strategy;
[0201] Based on the uncertainty decomposition results, design differentiated traceability strategies:
[0202] High cognitive uncertainty sample processing strategy:
[0203] Prioritize collecting more relevant data;
[0204] Introduce expert knowledge for auxiliary judgment;
[0205] Increase the training of the model on this type of sample;
[0206] For samples with cognitive uncertainty exceeding the threshold ( ), perform the following operations:
[0207] ;
[0208] wherein and respectively represent the cognitive uncertainty value and the cognitive uncertainty threshold of the sample , represents the operation set taken for the high cognitive uncertainty sample , represents data collection, represents expert consultation, represents targeted training;
[0209] High random uncertainty sample processing strategy:
[0210] Increase the sampling frequency or density;
[0211] Adopt integrated methods to improve prediction robustness;
[0212] Consider introducing additional sensors or measurement means;
[0213] For samples with random uncertainty exceeding the threshold ( ), perform:
[0214] ;
[0215] wherein and respectively represent the cognitive uncertainty value and the cognitive uncertainty threshold of the sample a random uncertainty value and a random uncertainty threshold, represents high random uncertainty samples a set of operations taken, represents increasing sampling, represents integrated prediction, represents sensor fusion;
[0216] Low uncertainty sample processing strategy:
[0217] Directly adopt the causal tracing result;
[0218] Perform standardized problem solving process;
[0219] For samples with low total uncertainty, directly follow the tracing path Perform standardized processing.
[0220] Step 5.3, Tracing result credibility evaluation;
[0221] Combined with the total uncertainty and the influence strength , define the tracing result credibility score :
[0222] ;
[0223] wherein, and are weight coefficients, represents the total uncertainty of the th sample, represents the set of all possible tracing paths of the th sample, represents a causal link from node to node in path , represents the product of the influence strength of all causal links on path .
[0224] Step 5.4, Resource optimization allocation algorithm;
[0225] Based on the tracing result credibility and the uncertainty type, design a resource optimization allocation algorithm to maximize the tracing effect:
[0226] ;
[0227] wherein, is the amount of resources allocated to the sample , and is the urgency of the anomaly, the total number of exceptions that need to be handled at present, the total amount of available resources, and respectively represent the inverse of the traceability result reliability score of the sample and .
[0228] Through the execution of the adaptive traceability strategy, resource allocation and processing schemes can be flexibly adjusted according to different types of uncertainty and traceability result reliability, improving the efficiency and accuracy of cultivated land quality evaluation exception traceability, and ultimately achieving the goal of precision agriculture management.
[0229] In one embodiment of the present application, an application example of the aforementioned cultivated land quality evaluation exception traceability method based on machine learning is provided:
[0230] Application scenario description: Taking the cultivated land quality evaluation data of a certain agricultural demonstration area as an example, the data set contains 1000 cultivated land samples, each sample contains 38 indicators, covering soil physical and chemical properties (such as organic matter content, pH value, total nitrogen, total phosphorus, total potassium, etc.), environmental factors (such as precipitation, light duration, temperature, etc.), management measures (such as fertilizer amount, irrigation frequency, etc.) and output indicators (such as unit area yield).
[0231] The data set is divided into two parts: 700 samples as the training set and 300 samples as the test set. Data preprocessing includes: missing value filling (using mean or median), outlier processing (based on 3σ principle) and feature standardization (Z-score standardization). The specific operation is as follows:
[0232] Four-layer Wavelet decomposition is used to perform multi-scale analysis on the original cultivated land quality evaluation data to obtain feature information in different frequency domains:
[0233] The first layer (highest frequency): captures the short-term rapid change characteristics of the cultivated land quality evaluation indicators;
[0234] The second layer: captures the medium and short-term change characteristics;
[0235] The third layer: captures the medium and long-term change characteristics;
[0236] The fourth layer (lowest frequency): captures the long-term stable trend characteristics.
[0237] For the organic matter content indicator of a certain cultivated land sample, the different scale decomposition results are shown in the following table:
[0238]
[0239] An example of the multi-scale feature fusion weight calculated by the attention mechanism is shown in the following table:
[0240]
[0241] It can be seen that for mutation-type abnormal samples, the model automatically assigns higher weights to high-frequency features; while for trend-type abnormal samples, the model pays more attention to medium- and low-frequency features, reflecting the advantages of adaptive fusion.
[0242] The Bayesian neural network adopts a three-hidden layer structure, with the number of neurons in each layer being 64, 32, and 16 respectively. The network parameters are configured as follows:
[0243] Activation function: ReLU;
[0244] Prior distribution: zero-mean Gaussian distribution, precision parameter λ=1.0;
[0245] Initial variational parameters in variational inference: μ_θ is initialized to the standard normal distribution sampling value, σ_θ is initialized to 0.1;
[0246] Optimizer: Adam, learning rate set to 0.001;
[0247] Batch size: 32;
[0248] Number of training rounds: 100 rounds.
[0249] After training the Bayesian neural network, its prediction performance is shown in the following table:
[0250]
[0251] By Monte Carlo sampling ( =30) shows a clear bimodal feature in the distribution of the estimated prediction uncertainty, reflecting the uncertainty differences of the model on different types of samples.
[0252] Based on the output of the Bayesian neural network, we build an evidence deep learning layer. The main configuration is as follows:
[0253] Evidence generation function: Softplus function is used to ensure that the evidence is non-negative;
[0254] Dirichlet distribution parameter initialization: α_1=α_2=1.01 (close to uninformative prior);
[0255] Loss function weight: λ = 0.1 (weight of evidence regularization term);
[0256] Number of training rounds: 50 rounds.
[0257] The evidence vectors and uncertainty decomposition results of some typical samples are shown in the following table:
[0258]
[0259] From the results, it can be seen that: typical normal samples and typical abnormal samples have higher evidence and lower total uncertainty; mild abnormal samples have moderate evidence and moderate uncertainty; fuzzy samples located at the classification boundary have lower evidence and the highest uncertainty.
[0260] Based on the output of evidence deep learning, gradient boosting trees are used for anomaly detection, with the following configurations:
[0261] Number of trees: 100;
[0262] Maximum depth: 5;
[0263] Learning rate: 0.1;
[0264] Subsampling ratio: 0.8;
[0265] Feature sampling ratio: 0.7.
[0266] For the abnormal sample with sample ID #156, the SHAP value analysis result shows that the top five features with the highest contribution degree are: organic matter content (SHAP value: 0.37); pH value (SHAP value: 0.29); total phosphorus content (SHAP value: 0.21); irrigation frequency (SHAP value: 0.18); groundwater level (SHAP value: 0.15)
[0267] Through causal diagram analysis, the main abnormal source paths generated are as follows:
[0268] Path 1 (impact strength: 0.61):
[0269] Irrigation frequency (too high) → groundwater level (rise) → pH value (rise) → organic matter content (decrease) → farmland quality abnormal;
[0270] Path 2 (impact strength: 0.43):
[0271] Fertilizer application amount (excessive) → total phosphorus content (too high) → microbial activity (decrease) → organic matter content (decrease) → farmland quality abnormal;
[0272] Path 3 (impact strength: 0.28):
[0273] Pesticide use (improper) → soil heavy metal content (rise) → organic matter content (decrease) → farmland quality abnormal;
[0274] For the above detected abnormal sample #156, the system performs adaptive source tracing strategy as follows:
[0275] Uncertainty analysis: cognitive uncertainty (0.14) is lower than the threshold (0.2), and random uncertainty (0.35) is higher than the threshold (0.3);
[0276] Based on the uncertainty type, a high random uncertainty processing strategy is selected:
[0277] Increase the number of sampling points in the irrigation area (from 3 to 7);
[0278] Deploy additional real-time pH monitoring sensors;
[0279] Use integrated prediction methods to enhance the monitoring accuracy of the area;
[0280] Traceability result credibility score calculation:
[0281] ;
[0282] Resource allocation algorithm allocation result:
[0283] Total available traceability resource units: 100;
[0284] Current number of abnormal samples to be processed: 5;
[0285] Emergency score of sample #156: 0.8 (based on the degree of organic matter content abnormality);
[0286] Resources allocated to sample #156: 32 resource units;
[0287] Specific traceability actions performed:
[0288] Adjust the irrigation strategy to reduce the irrigation frequency (consume 15 resource units);
[0289] Implement soil improvement measures to adjust the pH value (consume 10 resource units);
[0290] Deploy real-time soil organic matter monitoring system (consume 7 resource units);
[0291] After executing the above traceability strategies, the sample #156 area is tracked and monitored for 3 months, and the results show that the organic matter content gradually rises, the pH value tends to be within the normal range, and the cultivated land quality evaluation index returns to normal level, which proves the effectiveness of the proposed abnormal traceability method.
[0292] In an embodiment of the present application, the cultivated land quality evaluation abnormal traceability method based on machine learning can also be applied to the verification and review process of the cultivated land quality grade evaluation database, effectively improving the accuracy and reliability of cultivated land quality evaluation. This embodiment mainly includes two core links: evaluation database review and AI-based report generation:
[0293] 1. Evaluation database review;
[0294] The abnormality tracing method of the application first examines the county-level cultivated land quality grade evaluation spatial database, including two levels of formal examination and compliance examination:
[0295] Formal examination:
[0296] The county-level cultivated land quality grade evaluation spatial database is formally examined, mainly investigating the integrity, standardization and accuracy of the database content in three dimensions:
[0297] Integrity examination: using the evidence depth learning model of the application, it is automatically detected whether the spatial data is complete, including 1:10,000 cultivated land resource management unit map, 1:10,000 cultivated land quality grade survey and evaluation point map, 1:10,000 quality change area range map (including: high-standard farmland construction, degraded cultivated land management, and quality construction area, and occupation and compensation balance area) and attribute data including cultivated land quality grade survey and evaluation point data table, cultivated land resource management unit attribute data table, and cultivated land quality evaluation result data table. Through multi-scale feature extraction, the missing data can be accurately identified, and the potential influence of missing data on the evaluation result can be quantified.
[0298] Standardization examination: using the abnormality detection module of the application, the coordinate system, scale, storage format of the spatial data and the field name, field type, field length, field content of the attribute data are automatically inspected whether they meet the data dictionary specification requirements. Through the uncertainty quantification mechanism constructed by the Bayesian neural network, the data items that do not meet the standard can be risk rated, and the different degrees of standardization abnormalities can be distinguished.
[0299] Accuracy examination: combined with causal graph network analysis, it is examined whether the spatial position of the survey and evaluation point is consistent with the administrative division information in the attribute data field, and whether it is located on the cultivated land map patch. Through the adaptive tracing strategy, the root cause of the inconsistent position problem can be traced, such as coordinate system conversion error, data entry deviation or equipment precision problem, etc.
[0300] Compliance examination:
[0301] On the basis of formal examination, the method of the application is further applied to the compliance examination of the spatial data reported by the county, checking the logical relationship of the fields in the attribute table, and judging whether the attribute data is scientific and reasonable:
[0302] Based on the regional scale, the multi-scale feature extraction and fusion technology of the application is applied to review whether the investigation sample points in the county space are laid out according to the four types of regions of conventional utilization area, quality construction area, cultivated land occupation compensation area and damaged area, and whether the sample point density and quantity meet the requirements of the 'Annual Cultivated Land Quality Grade Change Investigation Technical Specification of County NY / T4322-2023'. Through attention mechanism feature fusion, the feature differences of different types of regions can be adaptively focused, and the accuracy of sample point layout anomaly detection can be improved.
[0303] Based on the regional scale, the Bayesian neural network model is used to review whether the county-level evaluation cultivated land area is consistent with the area announced by the Natural Resources Department of Guizhou Province. The Bayesian model can quantify the uncertainty source of the inconsistent area and distinguish whether it is caused by random error or systematic deviation.
[0304] Based on the regional scale, the evidence deep learning method is used to audit the annual comparison and analysis of the cultivated land quality grade evaluation results. Through the generation of evidence vectors, the evaluation results with large changes between years are decomposed for uncertainty, and it is distinguished which changes belong to the normal fluctuation range and which belong to potential abnormalities.
[0305] Based on the spatial position of the plot, the anomaly detection and tracing module of the application is used to realize the comparison with the evaluation index and result of the previous year, and the plot with a large change trend or not conforming to the logic is labeled. Using SHAP value analysis, the key indicators causing abnormal changes are identified, and the logical relationship between the indicators of the same plot and the same year is analyzed through the cause-effect diagram, such as the scientific correlation between organic matter and total nitrogen, texture and bulk density, etc.
[0306] In the review process, the adaptive tracing strategy execution module of the application plays an important role, and different processing schemes are adopted for different types of abnormalities and uncertainties:
[0307] For review results with high cognitive uncertainty (such as evaluation uncertainty caused by insufficient sample point layout), priority is given to collecting supplementary data and introducing expert knowledge to assist in judgment;
[0308] For review results with high random uncertainty (such as fluctuations caused by measurement errors), the sampling frequency is increased, and the integrated method is used to improve the robustness of the judgment;
[0309] For clear abnormalities with low uncertainty, the cause-effect tracing result is directly used to execute the standardized problem solving process.
[0310] 2. Generate report based on AI;
[0311] After completing the above review, the method of the application uses the accumulated analysis results to automatically generate two types of reports:
[0312] Quality inspection report: after completing the spatial quality inspection, based on the abnormality traceability result credibility evaluation mechanism of the application, a quality inspection report is generated, and the basis for systematic analysis and evaluation of unreasonable results is provided. The report includes:
[0313] Statistical distribution and spatial distribution characteristics of various abnormalities
[0314] Abnormality traceability path analysis, including upward tracking of cause chain and downward tracking of impact chain
[0315] Uncertainty quantification and decomposition results, distinguishing evaluation uncertainty from different sources
[0316] Abnormal severity classification and priority processing suggestions
[0317] Evaluation result report: after completing the quality inspection, combined with the prediction results of evidence deep learning and causal graph analysis, a county-level cultivated land quality grade evaluation report is generated. The report not only includes traditional statistical analysis of quality grade area and proportion, but also combines the evaluation results of the previous year, applies the causal inference ability of the application, analyzes the reasons for the improvement or decrease of cultivated land quality grade, and based on the confidence evaluation of evidence deep learning, puts forward targeted countermeasures and suggestions.
[0318] By applying the cultivated land quality evaluation abnormality traceability method of the application to the cultivated land quality grade evaluation verification process, the reliability of the evaluation data and the scientificity of the results can be significantly improved. Compared with traditional manual review, the method has the following significant advantages:
[0319] Multi-scale feature extraction and fusion capability, which can simultaneously focus on abnormal features of different spatial and temporal scales;
[0320] Bayesian neural network combined with evidence deep learning uncertainty quantification mechanism provides reliable confidence evaluation for evaluation results;
[0321] Abnormality traceability method based on gradient boosting tree and causal graph network provides interpretable evaluation bias source analysis;
[0322] Adaptive traceability strategy can take differentiated verification and error correction measures according to different types of evaluation abnormalities, improving resource utilization efficiency.
[0323] In the cultivated land quality grade evaluation verification practice of a certain province, compared with the traditional review process, the abnormality detection rate of the application method is increased by 35%, the review efficiency is improved by 60%, and the consistency of the verification results is improved by 28%, effectively reducing human judgment bias, providing scientific and efficient technical support for cultivated land quality evaluation work.
[0324] The above describes the embodiments of the present application, but the embodiments are not limited to the above specific embodiments, and the above specific embodiments are only illustrative but not restrictive, and the ordinary skilled in the art can make more equivalent embodiments under the inspiration of the embodiments, which are all within the protection scope of the embodiments.
Claims
1. A method for abnormality backtracking of cultivated land quality evaluation fused with machine learning, characterized in that, The method comprises the following steps: Multi-scale feature extraction and fusion: using Wavelet transform to perform multi-scale decomposition on the cultivated land quality evaluation data, obtaining feature representations at different scales, and realizing adaptive weighted fusion of features through attention mechanism; Bayesian neural network construction: constructing a Bayesian neural network model, introducing parameter prior distribution and variational inference method to model the fused features, and outputting prediction results containing uncertainty information; Evidence deep learning implementation: implementing evidence deep learning method, explicitly quantifying and decomposing uncertainty by generating evidence vector, and distinguishing cognitive uncertainty and random uncertainty; Abnormality detection and tracing: based on gradient boosting tree and causal graph network, realizing abnormality detection and tracing analysis, positioning and tracking the detected abnormality, and identifying potential causes and propagation path; Adaptive tracing strategy execution: based on uncertainty decomposition and tracing result reliability, formulating and executing adaptive tracing strategy, and taking differentiated processing scheme according to different types of uncertainty; The evidence deep learning implementation specifically comprises: Evidence theory model construction: based on Dempster-Shafer evidence theory, introducing Dirichlet distribution to model the confidence of classification results; Evidence vector generation network: adding an evidence generation layer based on the Bayesian neural network, mapping the network output to an evidence vector; Uncertainty quantification and decomposition: calculating the class probability expectation through the evidence vector, and decomposing the uncertainty into cognitive uncertainty and random uncertainty; Loss function of evidence deep learning: defining a comprehensive loss function containing negative log-likelihood loss and evidence regularization term for model training; The abnormality detection and tracing specifically comprises: Gradient boosting tree anomaly detection: based on fused features, cognitive uncertainty and random uncertainty, using gradient boosting tree algorithm to construct an anomaly detection model; Abnormal feature contribution analysis: using SHAP value to analyze the contribution of each feature, and identifying the key features causing the abnormality; Causal graph construction and analysis: constructing a causal relationship graph between the cultivated land quality indicators, including structure learning and parameter learning; Abnormality tracing path generation: based on abnormal feature contribution analysis and causal graph, generating abnormality tracing path to trace the cause and influence of the abnormality; The adaptive tracing strategy execution specifically comprises: Uncertainty threshold adaptive adjustment: establishing an uncertainty threshold adaptive adjustment mechanism according to historical tracing results and expert feedback; A differentiated traceability strategy is designed based on the uncertainty decomposition result, and the differentiated traceability strategies for samples with high cognitive uncertainty, high random uncertainty and low uncertainty are designed, wherein the sample with cognitive uncertainty greater than a cognitive uncertainty threshold is identified as a sample with high cognitive uncertainty , the sample with random uncertainty greater than a random uncertainty threshold is identified as a sample with high random uncertainty , and the sample with total uncertainty lower than a total uncertainty threshold is identified as a sample with low uncertainty . Tracing result reliability evaluation: combining total uncertainty and influence intensity to define the reliability score of the tracing result; Resource optimization allocation algorithm: based on the reliability of the tracing result and the emergency level of the abnormality, designing a resource optimization allocation algorithm to maximize the tracing effect.
2. The method according to claim 1, wherein, The multi-scale feature extraction and fusion specifically comprises: Data preprocessing: the original data matrix of cultivated land quality evaluation is standardized, missing values are processed, and a standardized data matrix is obtained ; Wavelet transform multi-scale decomposition: on the normalized data matrix Applying multi-level wavelet transform, get different scale approximation coefficient matrix and detail coefficient matrix ; Attention mechanism feature fusion: through attention mechanism, different scale features are adaptively weighted and fused to obtain fused features .
3. The method of claim 1, wherein the method comprises: The Bayesian neural network construction specifically comprises: Bayesian neural network architecture: the basic architecture of building a Bayesian neural network based on a multilayer perception mechanism, including layer hidden layer; Parameter prior distribution introduction: to the weight parameters in neural networks Introducing Gaussian prior distribution ; Variational inference implementation: using variational inference method to approximate the posterior distribution Introducing variational lower bound optimization objective; Predictive distribution output: estimating the predictive distribution of new samples through Monte Carlo sampling method, and outputting prediction results containing uncertainty information.
4. The method of claim 1, wherein the method comprises: The uncertainty decomposition formula is: Cognitive uncertainty The formula for which is: ; wherein is the sum of the evidence vectors, is the Dirichlet parameter of the class. Stochastic uncertainty The formula for which is: ; wherein is the predicted probability of the class .
5. The method of claim 1, wherein the method is characterized by, The abnormality tracing path generation includes upward tracing and downward tracing, and the specific steps are: The features with a contribution degree higher than a preset contribution degree threshold identified in the causal relationship diagram are taken as a starting point, and upstream factors that may cause abnormal feature changes are tracked upwards, i.e., parent nodes of the starting point in the causal relationship diagram are taken as upstream factors; downstream indicators that may be affected by the anomaly are tracked downwards, i.e., child nodes of the starting point in the causal relationship diagram are taken as downstream indicators.
6. The method of claim 1, wherein the method further comprises: The differentiated traceability strategy takes different processing schemes for different types of uncertainty: For high cognitive uncertainty samples: prefer to collect more relevant data, introduce expert knowledge for auxiliary judgment, and increase the training of the model on this type of sample; For high random uncertainty samples: increase the sampling frequency or density, use ensemble methods to improve the robustness of prediction, and consider introducing additional sensors or measurement methods; For low uncertainty samples: directly use the causal traceability result to execute the standardized problem solving process.
7. A system for anomaly backtracking in cultivated land quality assessment fusing machine learning, configured to perform the steps of a method for anomaly backtracking in cultivated land quality assessment fusing machine learning according to any one of claims 1 to 6, characterized in that, Comprise: A multi-scale feature extraction and fusion module is used to perform multi-scale decomposition on the cultivated land quality evaluation data using Wavelet transform, obtain feature representations at different scales, and realize adaptive weighted fusion of features through an attention mechanism; A Bayesian neural network module is used to construct a Bayesian neural network model, model the fused features by introducing a parameter prior distribution and a variational inference method, and output prediction results containing uncertainty information; An evidence deep learning module is used to implement an evidence deep learning method, explicitly quantify and decompose uncertainty by generating an evidence vector, and distinguish between cognitive uncertainty and random uncertainty; An anomaly detection and traceability module is used to realize anomaly detection and traceability analysis based on gradient boosting trees and causal graph networks, locate and track detected anomalies, and identify potential causes and propagation paths; An adaptive traceability strategy execution module is used to develop and execute an adaptive traceability strategy based on uncertainty decomposition and traceability result reliability, and take different processing schemes for different types of uncertainty.
Citation Information
Patent Citations
Data shortage estuary composite flood disaster research method considering climatic change
CN118760899A
Active learning classifier engine using beta approximation
US20230142131A1