Machine learning-fused cultivated land quality evaluation abnormity traceability method and system
Through multi-scale feature extraction, Bayesian neural network and deep learning of evidence, the problems of multi-scale analysis and uncertainty quantification in cultivated land quality evaluation are solved, efficient traceability of cultivated land quality abnormalities is achieved, scientific risk assessment and decision-making basis is provided, and the accuracy of cultivated land quality monitoring and management is improved.
Patent Information
- Application Number
- CN202511021099.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-07-24
AI Technical Summary
The existing technology has problems such as insufficient multi-scale analysis, lack of uncertainty quantification, and low credibility in the quality evaluation of cultivated land, resulting in insufficient abnormal identification and traceability accuracy, making it difficult to provide scientific risk assessment and decision-making basis.
Multi-scale feature extraction and fusion, Bayesian neural network, deep learning of evidence and gradient improvement tree are used, combined with causal graph networks, multi-scale analysis, uncertainty quantification and traceability of cultivated land quality data, and the accuracy and reliability of abnormal identification and traceability are improved through adaptive traceability strategies.
It has achieved a comprehensive capture of cultivated land quality evaluation indicators under different scales, distinguished between cognition and random uncertainty, provided a basis for risk assessment, improved the accuracy of abnormal traceability and resource utilization efficiency, and supported precise agricultural management.
Smart Images

Figure CN120524352A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of cultivated land quality evaluation and agricultural information technology, and more specifically, to a method and system for tracing abnormalities in cultivated land quality evaluation integrated with machine learning. Background Art
[0002] Cultivated land quality assessment is a crucial foundation for formulating scientific agricultural production strategies and soil improvement measures in modern agricultural management. With the increasing national emphasis on protecting cultivated land quality, accurately identifying abnormal cultivated land quality and tracing its causes have become key components of precision agriculture development. Currently, cultivated land quality assessment primarily utilizes a multi-index comprehensive evaluation approach, collecting and analyzing multidimensional data on soil physical and chemical properties, environmental factors, and management measures.
[0003] Traditional methods for detecting and tracing farmland quality anomalies primarily rely on statistical analysis and empirical judgment. For example, they employ the Mahalanobis distance method, cluster analysis, and principal component analysis to identify outliers, followed by attribution based on expert experience or simple correlation analysis. In recent years, with the advancement of machine learning, some studies have begun employing classification models, regression models, and deep learning methods, such as random forests, support vector machines, and deep neural networks, for anomaly detection. However, existing technologies for tracing farmland quality anomalies face the following technical challenges: First, farmland quality evaluation indicators exhibit distinct characteristics at different scales, making single-scale analysis incapable of fully capturing multi-scale anomalies. Second, existing methods are typically based on deterministic models and lack the quantification of uncertainty, making them unable to provide a risk assessment basis for decision-making. Third, traditional methods struggle to distinguish between epistemic and aleatory uncertainty, making the credibility assessment of attribution results lacking a scientific basis. Finally, attribution results lack a reliability metric, preventing decision-makers from providing risk assessments, thus hindering the precise governance capabilities of farmland quality monitoring and management.
[0004] Therefore, there is an urgent need for an anomaly tracing method for cultivated land quality evaluation that can integrate multi-scale analysis, uncertainty quantification, and causal inference to improve the accuracy of anomaly identification and the reliability of tracing decisions. Summary of the Invention
[0005] The present invention provides a method and system for tracing the source of abnormalities in cultivated land quality evaluation by integrating machine learning, which solves the technical problems in related technologies of insufficient accuracy in identifying and tracing abnormalities in cultivated land quality evaluation, inability to quantify uncertainty, and lack of reliability measurement.
[0006] The present invention discloses a method for tracing the source of abnormalities in cultivated land quality evaluation by integrating machine learning, comprising the following steps: Multi-scale feature extraction and fusion: Wavelet transform is used to perform multi-scale decomposition of cultivated land quality evaluation data to obtain feature representations at different scales, and adaptive weighted fusion of features is achieved through the attention mechanism; Bayesian neural network construction: Build a Bayesian neural network model, introduce parameter prior distribution and variational inference methods to model the fusion features, and output prediction results containing uncertainty information; Implementation of Evidence Deep Learning: Implementing evidence deep learning methods to explicitly quantify and decompose uncertainty by generating evidence vectors, distinguishing between epistemic uncertainty and random uncertainty; Anomaly detection and traceability: Based on gradient boosting trees and causal graph networks, anomaly detection and traceability analysis are implemented to locate and track detected anomalies and identify potential causes and propagation paths. Adaptive traceability strategy execution: Based on uncertainty decomposition and the credibility of traceability results, formulate and execute adaptive traceability strategies, adopt differentiated processing solutions according to different types of uncertainties, and improve resource utilization efficiency.
[0007] Furthermore, the multi-scale feature extraction and fusion specifically includes: Data preprocessing: original farmland quality evaluation data matrix Perform standardization, handle missing values, and obtain a normalized data matrix ; Wavelet transform multi-scale decomposition: normalized data matrix Apply multi-layer Wavelet transform to obtain approximate coefficient matrices of different scales and detail coefficient matrix ; Attention mechanism feature fusion: Adaptively weighted fusion of features of different scales is performed through the attention mechanism to obtain fusion features .
[0008] Furthermore, the Bayesian neural network construction specifically includes: Bayesian neural network architecture: The basic architecture of Bayesian neural network based on multi-layer perceptron, including layer hidden layer; Introduction of parameter prior distribution: weight parameters in neural networks Introducing Gaussian prior distribution ; Variational inference implementation: Using variational inference method to posterior distribution Make an approximation and introduce the variational lower bound optimization objective; Prediction distribution output: Estimate the prediction distribution of new samples through the Monte Carlo sampling method and output the prediction results containing uncertainty information.
[0009] Furthermore, the deep learning of evidence specifically includes: Evidence theory model construction: Based on the Dempster-Shafer evidence theory, the Dirichlet distribution is introduced to model the confidence of the classification results; Evidence vector generation network: Based on the Bayesian neural network, an evidence generation layer is added to map the network output into evidence vectors; Uncertainty quantification and decomposition: Calculate the expected probability of a category through the evidence vector and decompose uncertainty into epistemic uncertainty and random uncertainty; Loss function for evidence-based deep learning: A comprehensive loss function consisting of negative log-likelihood loss and evidence regularization term is defined for model training.
[0010] Furthermore, the anomaly detection and tracing specifically includes: Gradient boosted tree anomaly detection: Based on fusion features, epistemic uncertainty and random uncertainty, the gradient boosted tree algorithm is used to build an anomaly detection model; Abnormal feature contribution analysis: Use SHAP value to analyze the contribution of each feature and identify the key features that cause abnormalities; Causal graph construction and analysis: Constructing causal relationship graphs between farmland quality indicators, including structure learning and parameter learning; Anomaly traceability path generation: Based on anomaly feature contribution analysis and causal diagrams, anomaly traceability paths are generated to track the causes and impacts of anomalies.
[0011] Furthermore, the adaptive tracing strategy execution specifically includes: Adaptive adjustment of uncertainty thresholds: Establish an adaptive adjustment mechanism for uncertainty thresholds based on historical traceability results and expert feedback; Differentiated traceability strategy design: Based on the uncertainty decomposition results, differentiated traceability strategies are designed for samples with high epistemic uncertainty, high random uncertainty, and low uncertainty, where the epistemic uncertainty is greater than the epistemic uncertainty threshold. The samples with high epistemic uncertainty are identified as samples with high epistemic uncertainty, and the random uncertainty is greater than the random uncertainty threshold. The samples are identified as high random uncertainty samples, and the total uncertainty is lower than the total uncertainty threshold The samples are identified as low uncertainty samples; Credibility assessment of traceability results: Combine the total uncertainty and impact intensity to define the credibility score of the traceability results; Resource optimization allocation algorithm: Based on the credibility of traceability results and the urgency of exceptions, a resource optimization allocation algorithm is designed to maximize the traceability effect.
[0012] Furthermore, the uncertainty decomposition formula is: Epistemic uncertainty , and its calculation formula is: ; in is the sum of the evidence vectors, For the Dirichlet parameter of the class; Random uncertainty , and its calculation formula is: ; in For category The predicted probability of .
[0013] Furthermore, the generation of the anomaly tracing path includes two directions: upward tracing and downward tracing. The specific steps are as follows: The feature identified in the causal relationship graph and with a contribution higher than the preset contribution threshold is taken as the starting point, and the upstream factors that may cause the abnormal feature change are traced upward, that is, the parent node of the starting point in the causal relationship graph is taken as the upstream factor; the downstream indicators that may be affected by the abnormality are traced downward, that is, the child node of the starting point in the causal relationship graph is taken as the downstream indicator.
[0014] Furthermore, the differentiated traceability strategy adopts different solutions for different types of uncertainties: For samples with high epistemic uncertainty: prioritize collecting more relevant data, introduce expert knowledge to assist in judgment, and increase model training on such samples; For samples with high random uncertainty: increase the sampling frequency or density, use ensemble methods to improve prediction robustness, and consider introducing additional sensors or measurement methods; For low-uncertainty samples: directly use the causal tracing results and implement a standardized problem-solving process.
[0015] The present invention provides a system for tracing the source of abnormalities in cultivated land quality assessment integrated with machine learning, which is used to execute the steps in the above-mentioned method for tracing the source of abnormalities in cultivated land quality assessment integrated with machine learning, including: The multi-scale feature extraction and fusion module is used to perform multi-scale decomposition of cultivated land quality evaluation data using Wavelet transform, obtain feature representations at different scales, and achieve adaptive weighted fusion of features through the attention mechanism; The Bayesian neural network module is used to build a Bayesian neural network model. By introducing parameter prior distribution and variational inference methods, it models the fusion features and outputs prediction results containing uncertainty information. The evidence deep learning module is used to implement the evidence deep learning method. By generating evidence vectors, it explicitly quantifies and decomposes uncertainty and distinguishes epistemic uncertainty from random uncertainty. The anomaly detection and traceability module is used to implement anomaly detection and traceability analysis based on gradient boosting trees and causal graph networks, locate and track detected anomalies, and identify potential causes and propagation paths; The adaptive traceability strategy execution module is used to formulate and execute adaptive traceability strategies based on uncertainty decomposition and traceability result credibility, adopt differentiated processing solutions according to different types of uncertainty, and improve resource utilization efficiency. The beneficial effects of the present invention are: The multi-scale feature extraction and fusion method can comprehensively capture the characteristic changes and abnormal performance of cultivated land quality evaluation indicators at different scales, and improve the accuracy and comprehensiveness of anomaly identification.
[0016] The method that combines Bayesian neural networks with evidence-based deep learning realizes the quantification and decomposition of uncertainty, can distinguish between epistemic uncertainty and random uncertainty, and provide a risk assessment basis for anomaly tracing decisions.
[0017] The anomaly detection and tracing method based on gradient boosting tree and causal graph network can not only accurately locate anomalies, but also trace the causes and propagation paths of anomalies, thereby achieving explainable tracing of anomalies.
[0018] Adaptive traceability strategies can adopt differentiated processing solutions based on different types of uncertainties, improve resource utilization efficiency, and achieve precision agricultural management.
[0019] Through the credibility assessment of traceability results and resource optimization allocation algorithm, risk level assessment and resource allocation recommendations are provided to decision makers, thereby improving the precise governance capabilities of cultivated land quality monitoring and management. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 This is a flow chart of a method for tracing the source of abnormalities in cultivated land quality evaluation that integrates machine learning according to the present invention; Figure 2 is a flow chart of the multi-scale feature extraction and fusion steps of the present invention; Figure 3 is a flow chart of the steps of constructing a Bayesian neural network of the present invention; Figure 4 is a flow chart of the steps for implementing deep learning of evidence of the present invention; Figure 5 is a flow chart of the anomaly detection and tracing steps of the present invention; Figure 6 It is a flow chart of the steps for executing the adaptive tracing strategy of the present invention. DETAILED DESCRIPTION
[0021] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that these embodiments are discussed solely to enable those skilled in the art to better understand and implement the subject matter described herein, and that the functions and arrangements of the elements discussed may be varied without departing from the scope of this specification. Various examples may omit, substitute, or add various processes or components as needed. Furthermore, features described in some examples may be combined in other examples.
[0022] At least one embodiment of the present invention discloses a method for tracing the source of abnormalities in cultivated land quality evaluation by integrating machine learning, such as Figures 1 to 6 As shown, the following steps are included: Step 1: Multi-scale feature extraction and fusion: Wavelet transform is used to perform multi-scale decomposition on the cultivated land quality evaluation data to obtain feature representations at different scales, and adaptive weighted fusion of features is achieved through the attention mechanism; Step 1.1, data preprocessing; The original cultivated land quality evaluation data are standardized and missing values are processed. Assume that the original cultivated land quality evaluation data matrix is ,in represents the number of samples, Indicates the number of indicators. The standardization process uses the following formula: ; in, is the mean vector of each indicator, is the standard deviation vector of each indicator, The data are standardized.
[0023] Step 1.2, Wavelet transform multi-scale decomposition; The standardized data Multi-scale decomposition is performed through Wavelet transform to extract feature representations in different frequency domains. To decompose the levels, , is the number of decomposition levels, the multi-scale decomposition can be expressed as: ; in, Indicates the Layer Wavelet transform operation, Indicates the The feature representation after layer decomposition. The specific decomposition process uses discrete Wavelet transform. For one-dimensional signals , its decomposition formula is: ; ; in, is the scaling function, is the wavelet function, is the initial scale, is the translation parameter, is the signal length, 、 They represent the approximate coefficient matrix and detail coefficient matrix of different scales respectively. For multi-dimensional data, tensor product is used for expansion.
[0024] Step 1.3, feature fusion of attention mechanism; The features of different scales are adaptively weighted and fused through the attention mechanism to obtain multi-granularity feature representation. The features obtained by layer decomposition are , then the fused feature representation It can be expressed as: ; in, is the attention weight, which is calculated as follows: ; ; in, 、 and is a learnable parameter, is the hyperbolic tangent activation function, Represents a transpose operation. 、 Respectively represent 、 The importance score of layer features is calculated. In this way, the model can adaptively focus on important features at different scales and improve the ability to identify multi-scale anomalies.
[0025] The multi-scale feature extraction and fusion method combining Wavelet transform and attention mechanism can effectively capture the characteristic changes of cultivated land quality evaluation indicators at different scales, and provide a more comprehensive feature representation for subsequent anomaly detection and traceability.
[0026] Step 2: Bayesian neural network construction: Build a Bayesian neural network model, introduce parameter prior distribution and variational inference methods to model the fusion features, and output prediction results containing uncertainty information; Step 2.1, Bayesian neural network architecture; The basic architecture of Bayesian neural network is constructed based on multi-layer perceptron. Assume that the network contains hidden layer, The weight matrix of the layer is , the bias vector is , the activation function is , the forward propagation process is as follows: ; ; ; in, is the input layer of the network, is the fusion feature obtained in step 1, For the The output of the hidden layer, For the The output of the hidden layer, is the network output, used to indicate normal or abnormal category prediction, 、 denote the weight matrix and bias vector of the output layer respectively.
[0027] Step 2.2, parameter prior distribution is introduced; The weight parameters in the neural network Introducing the prior distribution, using Gaussian prior, the calculation formula is: ; in, is the weight parameter The prior distribution of is a Gaussian distribution; is the mean vector, indicating that the expected value of the prior hypothesis weight is zero; is the variance; is the identity matrix; The complete expression of the Bayesian neural network is: ; in, Represents a given input and training dataset Under these conditions, the output The predicted distribution of represents the training dataset; is the posterior distribution of the parameter; For the given parameters and input The output distribution under the conditions.
[0028] Step 2.3, variational inference implementation; Since the posterior distribution It is usually difficult to calculate directly, so the variational inference method is used for approximation: ; in, For the parameterized approximate posterior distribution, you can choose the following form: ; in, is a Gaussian distribution, is the mean vector of weights, is the variance parameter, is the identity matrix.
[0029] The optimization goal is to minimize the KL divergence, which is calculated as follows: ; in, is the loss function for variational inference; Represents KL divergence, which is used to measure the difference between two probability distributions; Through the variational lower bound (ELBO) transformation, the optimization objective is equivalent to: ; in, represents the evidence lower bound loss function, Represents the distribution expectations, is the log-likelihood function, which measures how well the model fits the training data. It represents the KL divergence between the approximate posterior distribution and the prior distribution, and serves as a regularization term to prevent overfitting.
[0030] Step 2.4, predict distribution output; After the model training is completed, the new sample The predictive distribution of can be estimated by Monte Carlo sampling method: ; in, Represents a given new input sample and training dataset Under these conditions, the output The predicted distribution of is the number of Monte Carlo sampling; Express The summation operation of the sub-sampling results; Indicates that given the weight parameter and new input Output probability under the condition; express To approximate the posterior distribution Parameter samples sampled in .
[0031] Bayesian neural networks can not only give prediction results, but also quantify the uncertainty of the predictions, providing a basis for subsequent evidence deep learning and uncertainty decomposition.
[0032] Step 3: Implementing evidence deep learning: Implementing evidence deep learning methods to explicitly quantify and decompose uncertainty by generating evidence vectors, distinguishing epistemic uncertainty from stochastic uncertainty. Step 3.1, construction of evidence theory model; Based on the Dempster-Shafer evidence theory, an evidence deep learning model is constructed. For the binary classification problem (normal samples and abnormal samples), the Dirichlet distribution is introduced to model the confidence of the classification results. The calculation formula is: ; in, represents the class probability vector, represents the probability that the sample belongs to the normal class, represents the probability that the sample belongs to the abnormal class, represents the evidence vector, 、 Evidence parameters corresponding to normal and abnormal classes respectively, , is the parameter of Dirichlet distribution, which indicates that the model is The confidence level of the class, For the The amount of evidence for a class.
[0033] The probability density function of the Dirichlet distribution is: ; in, Indicates that the parameter is The Dirichlet distribution probability density function, is the gamma function, represents the sum of all evidence parameters, represents the product of the gamma function values corresponding to each evidence parameter, Represents the weighted product of the probabilities of each category, where Represents the normal class, Represents an exception class.
[0034] Step 3.2, evidence vector generation network; On the basis of the Bayesian neural network, an evidence generation layer is added to map the output of the last layer of the network into an evidence vector , and its calculation formula is: ; in, represents the output of the feature extraction network, As the evidence generation function, an exponential function can be used to ensure that the evidence is non-negative. The calculation formula is: ; in, is the evidence vector, is the activation value output by the network.
[0035] Step 3.3, uncertainty quantification and decomposition; The expected probability of the category is calculated by the evidence vector, and the calculation formula is: ; in, For the The expected value of the predicted probability of the class, Indicates the Evidence parameters of the class, is the sum of the evidence vectors.
[0036] The uncertainty of a class prediction can be decomposed into two parts: Epistemic uncertainty : Uncertainty caused by insufficient knowledge or data can be reduced by collecting more data. It is expressed using the variance of the Dirichlet distribution: ; Random uncertainty : Uncertainty caused by the inherent randomness of the data cannot be reduced by increasing the amount of data. It is expressed using the prediction entropy, which is calculated as follows: ; in, Indicates the The expected value of the predicted probability of the class; Total uncertainty , and its calculation formula is: ; Step 3.4, loss function of evidence deep learning; Defining the loss function for evidence-based deep learning consists of two parts: Negative log-likelihood loss , and its calculation formula is: ; in, For samples The true label, is the Dirichlet parameter corresponding to the label, For samples The sum of all parameters.
[0037] Evidence regularization term : Encourage the model to generate less evidence for uncertain samples. The calculation formula is: ; in, is the indicator function, For samples No. Class evidence parameters.
[0038] The comprehensive loss function is , and its calculation formula is: ; in, is the regularization coefficient.
[0039] Through the evidence-based deep learning method, the uncertainty of the anomaly detection results of cultivated land quality evaluation can be quantified and decomposed in a refined manner, providing a more reliable decision-making basis for subsequent anomaly tracing.
[0040] Step 4: Anomaly detection and traceability: Based on the gradient boosting tree and causal graph network, anomaly detection and traceability analysis are implemented to locate and track detected anomalies and identify potential causes and propagation paths. The detected anomalies are located by analyzing the contribution of abnormal features, and traceability paths are generated based on the cause-effect diagram to identify potential causes and propagation paths.
[0041] Step 4.1, gradient boosting tree anomaly detection; Based on the output of the above evidence deep learning, the gradient boosting tree algorithm is used to build an anomaly detection model. , using the following feature vector as input: ; in, is the characteristic vector of the sample, is the fusion feature of the sample, is the predicted probability, and They are epistemic uncertainty and aleatoric uncertainty respectively.
[0042] The training goal of the gradient boosting tree model is to minimize the following loss function, which is calculated as: ; in, is the loss function (such as cross entropy), is the true label of the sample, is the predicted value, For the The complexity penalty of a tree, is the total number of samples, is the number of trees.
[0043] After the model training is completed, anomaly detection is performed on new samples and anomaly probability scores are output. , and based on the set threshold Determine abnormality: ; Step 4.2, abnormal feature contribution analysis; For detected abnormal samples, SHAP values are used to analyze the contribution of each feature and identify the key features that cause the abnormality: ; in, Representation characteristics For samples Contribution to the prediction results, is the set of all features, Indicates that features are not included The feature subset of Indicates that only a subset of features is used Prediction Sample The result, Indicates the number of parameters in the collection. Indicates the use of feature subsets Add features For samples The model output for making predictions, Represents the factorial operation.
[0044] Sort features by contribution and generate a feature importance list , 、 、 Respectively represent 、 、 Important features, is the total number of features, where , Represents characteristics 、 、 For samples The contribution value of the prediction result.
[0045] Step 4.3, cause-effect diagram construction and analysis; Constructing a causal relationship diagram between cultivated land quality indicators , where the node set Represents each quality index, edge set Represents the causal relationship between indicators. The construction of a causal graph consists of two steps: Structural learning: Using PC algorithms or scoring-based methods to learn causal relationships between variables: ; in, is the optimal graph structure, For a given data set The structure below The scoring function of .
[0046] Parameter learning: determining the probability distribution of each condition ,in For nodes The parent node collection.
[0047] Step 4.4, abnormal tracing path generation; Based on the constructed causal graph and the abnormal feature contribution analysis results, the abnormal traceability path is generated: For detected abnormal samples , identify the feature subset with the highest contribution ,in, is the contribution threshold, Represents a sample A ranked list of feature importances.
[0048] In the causal diagram, Nodes in are used as starting points to trace potential causal paths: Tracing up: looking for upstream factors (causes) that may lead to abnormal characteristic changes; Tracking down: Evaluate downstream indicators (outcomes) that the anomaly may affect; The specific steps are: taking the feature identified in the causal relationship graph and with a contribution higher than the preset contribution threshold as the starting point, tracing upward the parent node of the starting point in the causal relationship graph as the upstream factor, and tracing downward the child node of the starting point in the causal relationship graph as the downstream indicator.
[0049] Calculate the influence strength of each node on the path , based on conditional probability and intervention evaluation, its calculation formula is: ; in, Indicates the intervention characteristics The value is Under the conditions, the characteristics The value is The probability of Represents a variable Intervene so that its value is , Representation characteristics The value is The marginal probability of .
[0050] Sort by impact intensity and generate the most likely anomaly traceability path set ,in, 、 、 Respectively represent 、 、 A set of paths, is the total number of paths, and each path is a sequence of nodes arranged in causal order.
[0051] By combining the gradient boosting tree and causal graph network method, we can not only accurately detect anomalies in cultivated land quality evaluation, but also deeply analyze the causes of anomalies and possible impact paths, providing a decision-making basis for subsequent adaptive traceability strategies.
[0052] Step 5: Adaptive traceability strategy execution: Based on uncertainty decomposition and the credibility of traceability results, an adaptive traceability strategy is formulated and executed. Differentiated processing solutions are adopted according to different types of uncertainty to improve resource utilization efficiency. Step 5.1, adaptive adjustment of uncertainty threshold; According to the historical traceability results and expert feedback, an adaptive adjustment mechanism for uncertainty thresholds is established. According to the historical traceability results and expert feedback, an adaptive adjustment mechanism for uncertainty thresholds is established. The uncertainty thresholds include cognitive uncertainty thresholds. , random uncertainty threshold and the total uncertainty threshold : ; ; ; in, and Respectively The thresholds of epistemic uncertainty and random uncertainty in the iteration, and Respectively The thresholds of epistemic uncertainty and random uncertainty in the iteration, is the learning rate, is the performance change, and Respectively represent In the round of iteration and The total uncertainty threshold in the round iteration.
[0053] Increase epistemic uncertainty to greater than the epistemic uncertainty threshold The samples with high epistemic uncertainty are identified as samples with high epistemic uncertainty, and the random uncertainty is greater than the random uncertainty threshold. The samples are identified as high random uncertainty samples, and the total uncertainty is lower than the total uncertainty threshold The samples are identified as low uncertainty samples.
[0054] Step 5.2, differentiated traceability strategy design; Based on the uncertainty decomposition results, design differentiated traceability strategies: Strategy for handling samples with high epistemic uncertainty: Prioritize collecting more relevant data; Introducing expert knowledge to assist in judgment; Increase the training of the model on samples of this type; For samples where epistemic uncertainty exceeds the threshold ( ), do the following: ; in, and Represents samples The epistemic uncertainty value and epistemic uncertainty threshold, For samples with high epistemic uncertainty The set of actions taken, Indicates data collection, Indicates expert consultation, Indicates targeted training; High random uncertainty sample processing strategy: Increase sampling frequency or density; Using ensemble methods to improve forecast robustness; Consider introducing additional sensors or measurements; For samples where the random uncertainty exceeds the threshold ( ),implement: ; in, and Represents samples The random uncertainty value and random uncertainty threshold of Indicates that for high random uncertainty samples The set of actions taken, Indicates increased sampling, represents the ensemble prediction, represents sensor fusion; Low uncertainty sample processing strategy: Directly adopt the causal tracing results; Implement standardized problem-solving processes; For samples with low total uncertainty, directly follow the traceability path Perform standardization.
[0055] Step 5.3, credibility assessment of traceability results; Combined total uncertainty and impact intensity , define the credibility score of traceability results : ; in, and is the weight coefficient, Indicates the The total uncertainty of the samples, Indicates the The set of all possible traceability paths for a sample, Indicates the path Middle slave node To Node A causal link, Indicates the path The product of the influence strengths of all causal links above.
[0056] Step 5.4, resource optimization allocation algorithm; Based on the credibility and uncertainty type of traceability results, we design a resource optimization allocation algorithm to maximize the traceability effect: ; in, Assigned to the sample The amount of resources, For an exceptional degree of urgency, is the total number of exceptions that need to be handled currently, is the total available resources, and Represents samples and The inverse of the credibility score of the traceability result.
[0057] Through the implementation of adaptive traceability strategies, resource allocation and processing plans can be flexibly adjusted according to different types of uncertainties and the credibility of traceability results, thereby improving the efficiency and accuracy of abnormal traceability in cultivated land quality evaluation and ultimately achieving the goal of precision agricultural management.
[0058] In one embodiment of the present invention, an application example of the aforementioned method for tracing the source of abnormalities in cultivated land quality assessment integrated with machine learning is provided: Application scenario description: Taking the cultivated land quality evaluation data of a certain agricultural demonstration area as an example, this dataset contains 1,000 cultivated land samples, each of which contains 38 indicators, covering soil physical and chemical properties (such as organic matter content, pH value, total nitrogen, total phosphorus, total potassium, etc.), environmental factors (such as precipitation, daylight duration, temperature, etc.), management measures (such as fertilizer application amount, irrigation frequency, etc.), and output indicators (such as yield per unit area).
[0059] The dataset is divided into two parts: 700 samples as a training set and 300 samples as a test set. Data preprocessing includes: filling missing values (using the mean or median), outlier handling (based on the 3σ principle), and feature normalization (Z-score normalization). The specific operations are as follows: A 4-layer Wavelet decomposition was used to perform multi-scale analysis on the original cultivated land quality evaluation data to obtain characteristic information in different frequency domains: The first layer (highest frequency): captures the short-term rapid changes in cultivated land quality evaluation indicators; The second layer: capturing short- to medium-term change characteristics; The third layer: capturing medium- and long-term change characteristics; The fourth layer (lowest frequency): captures long-term stable trend characteristics.
[0060] For the organic matter content index of a certain cultivated land sample, the decomposition results at different scales are shown in the following table:
[0061] An example of multi-scale feature fusion weights calculated by the attention mechanism is shown in the following table:
[0062] It can be seen that for mutation-type abnormal samples, the model automatically assigns higher weights to high-frequency features; while for trend-type abnormal samples, the model pays more attention to medium- and low-frequency features, reflecting the advantages of adaptive fusion.
[0063] The Bayesian neural network adopts a three-hidden layer structure, with the number of neurons in each layer being 64, 32, and 16 respectively. The network parameters are configured as follows: Activation function: ReLU; Prior distribution: zero-mean Gaussian distribution, precision parameter λ=1.0; Initial variational parameters in variational inference: μ_θ is initialized to the standard normal distribution sampling value, σ_θ is initialized to 0.1; Optimizer: Adam, learning rate set to 0.001; Batch size: 32; Number of training rounds: 100 rounds.
[0064] After training the Bayesian neural network, its prediction performance is shown in the following table:
[0065] By Monte Carlo sampling ( =30) shows a clear bimodal feature in the distribution of the estimated prediction uncertainty, reflecting the uncertainty differences of the model on different types of samples.
[0066] Based on the output of the Bayesian neural network, we build an evidence deep learning layer. The main configuration is as follows: Evidence generation function: Softplus function is used to ensure that the evidence is non-negative; Dirichlet distribution parameter initialization: α_1=α_2=1.01 (close to uninformative prior); Loss function weight: λ = 0.1 (weight of evidence regularization term); Number of training rounds: 50 rounds.
[0067] The evidence vectors and uncertainty decomposition results of some typical samples are shown in the following table:
[0068] The results show that typical normal samples and typical abnormal samples have higher evidence and lower total uncertainty; mildly abnormal samples have medium evidence and medium uncertainty; fuzzy samples at the classification boundary have lower evidence and the highest uncertainty.
[0069] Based on the output of evidence-based deep learning, anomaly detection is performed using a gradient boosting tree with the following configuration: Number of trees: 100; Maximum depth: 5; Learning rate: 0.1; Subsampling ratio: 0.8; Feature sampling ratio: 0.7.
[0070] For the abnormal sample with sample ID #156, the SHAP value analysis results show that the top five contributing features are: organic matter content (SHAP value: 0.37); pH value (SHAP value: 0.29); total phosphorus content (SHAP value: 0.21); irrigation frequency (SHAP value: 0.18); and groundwater level (SHAP value: 0.15). Through causal graph analysis, the main anomaly tracing paths generated are as follows: Path 1 (influence strength: 0.61): Irrigation frequency (too high) → groundwater level (increase) → pH value (increase) → organic matter content (decrease) → abnormal arable land quality; Path 2 (influence strength: 0.43): Fertilizer application (excessive) → total phosphorus content (too high) → microbial activity (reduced) → organic matter content (reduced) → abnormal arable land quality; Path 3 (influence strength: 0.28): Improper use of pesticides → increased soil heavy metal content → decreased organic matter content → abnormal arable land quality; For the abnormal sample #156 detected above, the system implements the following adaptive tracing strategy: Uncertainty analysis: epistemic uncertainty (0.14) is below the threshold (0.2), and aleatoric uncertainty (0.35) is above the threshold (0.3); Based on the uncertainty type, select a high random uncertainty handling strategy: Increase the number of sampling points in irrigation areas (from 3 to 7); Deploy additional real-time pH monitoring sensors; Adopting integrated prediction methods to enhance the monitoring accuracy of the region; Calculation of credibility score of traceability results: ; Resource allocation algorithm allocation results: Total available traceability resource units: 100; The current number of abnormal samples to be processed: 5; Urgency score for sample #156: 0.8 (based on abnormal organic matter content); Resources allocated to sample #156: 32 resource units; Specific traceability actions implemented: Adjust irrigation strategy and reduce irrigation frequency (consumes 15 resource units); Implement soil improvement measures to adjust pH (consume 10 resource units); Deploy a real-time soil organic matter monitoring system (consuming 7 resource units); After implementing the above-mentioned traceability strategy, a three-month follow-up monitoring was conducted in the area where sample #156 was located. The results showed that the organic matter content gradually recovered, the pH value tended to the normal range, and the cultivated land quality evaluation indicators returned to normal levels, proving the effectiveness of the proposed abnormality traceability method.
[0071] In one embodiment of the present invention, the machine learning-integrated farmland quality assessment anomaly tracing method can also be applied to the verification and review of the farmland quality rating database, effectively improving the accuracy and reliability of farmland quality assessments. This embodiment mainly includes two core steps: database review and AI-based report generation: 1. Review of the evaluation database; The anomaly tracing method of the present invention first reviews the county-level farmland quality grade evaluation spatial database, including two levels of formal review and compliance review: Formalities Examination: A formal review of the county-level spatial database for cultivated land quality evaluation was conducted, focusing on the three dimensions of database content: completeness, standardization, and accuracy. Completeness Review: Utilizing this invention's evidence-based deep learning model, the system automatically checks whether spatial data includes complete vector graphics such as a 1:10,000 cultivated land resource management unit map, a 1:10,000 cultivated land quality grading survey and evaluation point map, and a 1:10,000 quality change zone map (including quality development zones such as high-standard farmland construction and degraded cultivated land management, as well as areas of occupation-compensation balance). Furthermore, the system checks whether attribute data includes a table of cultivated land quality grading survey and evaluation point data, a table of cultivated land resource management unit attributes, and a table of cultivated land quality evaluation results. Multi-scale feature extraction allows for accurate identification of missing data and quantification of the potential impact of missing data on evaluation results.
[0072] Compliance Review: The anomaly detection module of this invention automatically verifies that the coordinate system, scale, storage format of spatial data, as well as the field name, field type, field length, and field content of attribute data, conform to data dictionary specifications. Using an uncertainty quantification mechanism constructed using a Bayesian neural network, risk ratings can be assigned to non-compliant data items, distinguishing varying degrees of compliance anomalies.
[0073] Accuracy Review: Combined with causal network analysis, this method verifies whether the spatial locations of survey and evaluation points are consistent with the administrative division information in the attribute data fields and whether they are located on the cultivated land map patches. An adaptive traceability strategy can be used to trace the root causes of any location inconsistencies, such as coordinate system conversion errors, data entry deviations, or equipment accuracy issues.
[0074] Compliance review: On the basis of formal review, the method of the present invention is further applied to the compliance review of spatial data reported at the county level, checking the logical relationship of fields in the attribute table and judging whether the attribute data is scientific and reasonable: At the regional scale, the multi-scale feature extraction and fusion technology of this invention is applied to examine whether the layout of survey sample points in the county space is based on the four types of areas: conventional utilization areas, quality improvement areas, cultivated land occupation and compensation areas, and damaged and destroyed areas. The density and number of sample points also meet the requirements of the "Technical Specification for Annual Cultivated Land Quality Grade Change Survey in Counties (NY / T4322-2023)." Feature fusion using the attention mechanism can adaptively focus on the feature differences between different types of areas, improving the accuracy of sample point layout anomaly detection.
[0075] At the regional scale, a Bayesian neural network model was used to examine whether the county-level assessed cultivated land area was consistent with the area published by the Guizhou Provincial Department of Natural Resources. The Bayesian model quantified the sources of uncertainty in area discrepancies and distinguished whether they were caused by random errors or systematic biases.
[0076] Using deep learning methods based on evidence, we review and analyze the annual comparative results of cultivated land quality ratings at a regional scale. By generating evidence vectors, we decompose the uncertainty of evaluation results that vary significantly from year to year, distinguishing between changes that fall within the normal range of fluctuation and those that represent potential anomalies.
[0077] Based on the spatial location of the plots, the anomaly detection and traceability module of this invention is used to compare the evaluation indicators and results with those of the previous year, marking plots with significant or illogical changes. SHAP value analysis is used to identify key indicators that cause abnormal changes, and cause-and-effect diagram analysis is used to examine the logical relationships between indicators within the same plot and year, such as the scientific correlation between organic matter and total nitrogen, texture and bulk density, and other indicators.
[0078] During the review process, the adaptive traceability strategy execution module of the present invention plays an important role, adopting differentiated processing solutions for different types of anomalies and uncertainties: For review results with high epistemic uncertainty (e.g., evaluation uncertainty caused by insufficient sample placement), priority should be given to collecting supplementary data and introducing expert knowledge to assist in judgment; For review results with high random uncertainty (e.g., fluctuations due to measurement error), increase sampling frequency and use ensemble methods to improve the robustness of judgments; For clear anomalies with low uncertainty, directly use the causal tracing results and execute a standardized problem-solving process.
[0079] 2. Generate reports based on AI; After completing the above review, the method of the present invention automatically generates two types of reports using the accumulated analysis results: Quality inspection report: After completing the space quality inspection, a quality inspection report is generated based on the credibility assessment mechanism of the abnormal traceability results of the present invention, which systematically analyzes the basis for unreasonable evaluation results. The report includes: Statistical and spatial distribution characteristics of various anomalies Abnormal traceability path analysis, including upward tracing of the cause chain and downward tracing of the impact chain Uncertainty quantification and decomposition results, distinguishing evaluation uncertainties from different sources Abnormal severity classification and priority treatment recommendations Evaluation Report: After quality inspection, a county-level farmland quality grade evaluation report is generated by combining the prediction results of deep evidence learning with causal graph analysis. This report not only includes traditional statistical analysis of quality grade areas and proportions, but also analyzes the reasons for improvements or declines in farmland quality grade by combining the previous year's evaluation results and applying the causal inference capabilities of this invention. Based on the confidence assessment of deep evidence learning, targeted countermeasures and recommendations are also provided.
[0080] By applying the proposed method for tracing the source of abnormalities in cultivated land quality assessments, which integrates machine learning, to the process of evaluating and verifying cultivated land quality grades, the reliability of evaluation data and the scientific nature of the results can be significantly improved. Compared with traditional manual review, this method has the following significant advantages: Multi-scale feature extraction and fusion capabilities can simultaneously focus on abnormal features at different spatial and temporal scales; The uncertainty quantification mechanism that combines Bayesian neural networks with deep learning of evidence provides a reliable confidence assessment for the evaluation results; Anomaly tracing methods based on gradient boosting trees and causal graph networks provide explainable analysis of evaluation bias sources; The adaptive traceability strategy can take differentiated verification and error correction measures based on different types of evaluation anomalies to improve resource utilization efficiency.
[0081] In the practice of arable land quality grade evaluation and verification in a certain province, compared with the traditional audit process, the method of the present invention increased the abnormality detection rate by 35%, the audit efficiency by 60%, and the consistency of the verification results by 28%, effectively reducing human judgment bias and providing scientific and efficient technical support for arable land quality evaluation.
[0082] The above describes an embodiment of the present invention, but this embodiment is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Ordinary technicians in this field can also make more forms of equivalent embodiments based on the inspiration of this embodiment, all of which are protected by this embodiment.
Claims
1. A method for tracing abnormalities in cultivated land quality assessment by integrating machine learning, characterized in that: The following steps are involved: Multi-scale feature extraction and fusion: Wavelet transform is used to perform multi-scale decomposition of cultivated land quality evaluation data to obtain feature representations at different scales, and adaptive weighted fusion of features is achieved through the attention mechanism; Bayesian neural network construction: Build a Bayesian neural network model, introduce parameter prior distribution and variational inference methods to model the fusion features, and output prediction results containing uncertainty information; Implementation of Evidence Deep Learning: Implementing evidence deep learning methods to explicitly quantify and decompose uncertainty by generating evidence vectors, distinguishing between epistemic uncertainty and random uncertainty; Anomaly detection and traceability: Based on gradient boosting trees and causal graph networks, anomaly detection and traceability analysis are implemented to locate and track detected anomalies and identify potential causes and propagation paths. Adaptive traceability strategy execution: Based on uncertainty decomposition and the credibility of traceability results, formulate and execute adaptive traceability strategies, adopt differentiated processing solutions according to different types of uncertainties, and improve resource utilization efficiency.
2. The method for tracing the abnormality of cultivated land quality evaluation by integrating machine learning according to claim 1 is characterized in that: The multi-scale feature extraction and fusion specifically include: Data preprocessing: original farmland quality evaluation data matrix Perform standardization, handle missing values, and obtain a normalized data matrix ; Wavelet transform multi-scale decomposition: normalized data matrix Apply multi-layer Wavelet transform to obtain approximate coefficient matrices of different scales and detail coefficient matrix ; Attention mechanism feature fusion: Adaptively weighted fusion of features of different scales is performed through the attention mechanism to obtain fusion features .
3. The method for tracing the abnormality of cultivated land quality evaluation by integrating machine learning according to claim 1 is characterized in that: The Bayesian neural network construction specifically includes: Bayesian neural network architecture: The basic architecture of Bayesian neural network based on multi-layer perceptron, including layer hidden layer; Introduction of parameter prior distribution: weight parameters in neural networks Introducing Gaussian prior distribution ; Variational inference implementation: Using variational inference method to posterior distribution Make an approximation and introduce the variational lower bound optimization objective; Prediction distribution output: Estimate the prediction distribution of new samples through the Monte Carlo sampling method and output the prediction results containing uncertainty information.
4. The method for tracing the abnormality of cultivated land quality evaluation by integrating machine learning according to claim 1 is characterized in that: The evidence deep learning implementation specifically includes: Evidence theory model construction: Based on the Dempster-Shafer evidence theory, the Dirichlet distribution is introduced to model the confidence of the classification results; Evidence vector generation network: Based on the Bayesian neural network, an evidence generation layer is added to map the network output into evidence vectors; Uncertainty quantification and decomposition: Calculate the expected probability of a category through the evidence vector and decompose uncertainty into epistemic uncertainty and random uncertainty; Loss function for evidence-based deep learning: A comprehensive loss function consisting of negative log-likelihood loss and evidence regularization term is defined for model training.
5. The method for tracing the abnormality of cultivated land quality evaluation by integrating machine learning according to claim 1 is characterized in that: The anomaly detection and tracing specifically include: Gradient boosted tree anomaly detection: Based on fusion features, epistemic uncertainty and random uncertainty, the gradient boosted tree algorithm is used to build an anomaly detection model; Abnormal feature contribution analysis: Use SHAP value to analyze the contribution of each feature and identify the key features that cause abnormalities; Causal graph construction and analysis: Constructing causal relationship graphs between farmland quality indicators, including structure learning and parameter learning; Anomaly traceability path generation: Based on anomaly feature contribution analysis and causal diagrams, anomaly traceability paths are generated to track the causes and impacts of anomalies.
6. The method for tracing the source of abnormalities in cultivated land quality assessment integrated with machine learning according to claim 1, characterized in that: The adaptive traceability strategy execution specifically includes: Adaptive adjustment of uncertainty thresholds: Establish an adaptive adjustment mechanism for uncertainty thresholds based on historical traceability results and expert feedback; Differentiated traceability strategy design: Based on the uncertainty decomposition results, differentiated traceability strategies are designed for samples with high epistemic uncertainty, high random uncertainty, and low uncertainty, where the epistemic uncertainty is greater than the epistemic uncertainty threshold. The samples with high epistemic uncertainty are identified as samples with high epistemic uncertainty, and the random uncertainty is greater than the random uncertainty threshold. The samples are identified as high random uncertainty samples, and the total uncertainty is lower than the total uncertainty threshold The samples are identified as low uncertainty samples; Credibility assessment of traceability results: Combine the total uncertainty and impact intensity to define the credibility score of the traceability results; Resource optimization allocation algorithm: Based on the credibility of traceability results and the urgency of exceptions, a resource optimization allocation algorithm is designed to maximize the traceability effect.
7. The method for tracing the abnormality of cultivated land quality evaluation by integrating machine learning according to claim 1 is characterized in that: The uncertainty decomposition formula is: Epistemic uncertainty , and its calculation formula is: ; in is the sum of the evidence vectors, For the Dirichlet parameter of the class; Random uncertainty , and its calculation formula is: ; in For category The predicted probability of .
8. The method for tracing the abnormality of cultivated land quality evaluation by integrating machine learning according to claim 5 is characterized in that: The generation of the abnormal tracing path includes two directions: upward tracing and downward tracing. The specific steps are as follows: The feature identified in the causal relationship graph and with a contribution higher than the preset contribution threshold is taken as the starting point, and the upstream factors that may cause the abnormal feature change are traced upward, that is, the parent node of the starting point in the causal relationship graph is taken as the upstream factor; the downstream indicators that may be affected by the abnormality are traced downward, that is, the child node of the starting point in the causal relationship graph is taken as the downstream indicator.
9. The method for tracing the abnormality of cultivated land quality evaluation by integrating machine learning according to claim 6 is characterized in that: The differentiated traceability strategy adopts different solutions to deal with different types of uncertainties: For samples with high epistemic uncertainty: prioritize collecting more relevant data, introduce expert knowledge to assist in judgment, and increase model training on such samples; For samples with high random uncertainty: increase the sampling frequency or density, use ensemble methods to improve prediction robustness, and consider introducing additional sensors or measurement methods; For low-uncertainty samples: directly use the causal tracing results and implement a standardized problem-solving process.
10. A system for tracing the source of abnormalities in cultivated land quality evaluation integrated with machine learning, for executing the steps in the method for tracing the source of abnormalities in cultivated land quality evaluation integrated with machine learning as claimed in any one of claims 1 to 9, characterized in that: include: The multi-scale feature extraction and fusion module is used to perform multi-scale decomposition of cultivated land quality evaluation data using Wavelet transform, obtain feature representations at different scales, and achieve adaptive weighted fusion of features through the attention mechanism; The Bayesian neural network module is used to build a Bayesian neural network model. By introducing parameter prior distribution and variational inference methods, it models the fusion features and outputs prediction results containing uncertainty information. The evidence deep learning module is used to implement the evidence deep learning method. By generating evidence vectors, it explicitly quantifies and decomposes uncertainty and distinguishes epistemic uncertainty from random uncertainty. The anomaly detection and traceability module is used to implement anomaly detection and traceability analysis based on gradient boosting trees and causal graph networks, locate and track detected anomalies, and identify potential causes and propagation paths; The adaptive traceability strategy execution module is used to formulate and execute adaptive traceability strategies based on uncertainty decomposition and the credibility of traceability results, adopt differentiated processing solutions according to different types of uncertainties, and improve resource utilization efficiency.
Citation Information
Patent Citations
Communication signal classification and recognition method based on multi-feature association and Bayesian network
CN109450834A
Method for quantitatively calibrating uncertainty in equipment fault diagnosis based on deep learning
CN115204227A
Data shortage estuary composite flood disaster research method considering climatic change
CN118760899A
Cultivated land quality evaluation method and system based on random forest
CN119089323A
Method for establishing fault diagnosis technique based on contingent Bayesian networks
US20190087294A1
Cited By
Data blood relationship construction and reasoning method and system for cultivated land quality monitoring
CN121705809A