An automated audit and analysis method for cultivated land soil data

By combining IoT sensors and drone multispectral imaging with extended Kalman filters, causal reasoning graph neural networks, and quantum computing-driven visualization technology, the problems of low soil data collection coverage and update frequency are solved, and efficient and secure soil data analysis and display are achieved, supporting precision agricultural management.

CN119151462BActive Publication Date: 2025-09-26HUBEI LIANGQING AGRI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411169066.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-23
Publication Date
2025-09-26
Estimated Expiration
2044-08-23

AI Technical Summary

Technical Problem

Existing soil data collection has limited coverage, low update frequency, low processing efficiency, single analysis method and insufficient data security, which makes it difficult to meet the fast, efficient and precise needs of modern agriculture.

Method used

IoT sensors and drone multispectral imaging are used for real-time multi-source data collection, combined with extended Kalman filter noise reduction and outlier removal, causal reasoning graph neural network is used for data analysis and fusion, quantum computing-driven visualization and natural language generation technology are used for display, and encryption and access control are implemented.

Benefits of technology

It improves the efficiency and accuracy of soil data review, ensures data security, provides a comprehensive soil data processing solution, and supports scientific decision-making in precision agriculture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119151462B_ABST
    Figure CN119151462B_ABST
Patent Text Reader

Abstract

This invention discloses an automated audit and analysis method for cultivated land soil data, comprising the following steps: S1. Collecting raw soil data using IoT sensors and drone multispectral imaging to form a dataset; S2. Using an extended Kalman filter to reduce noise and remove outliers, and using time series interpolation to complete missing data; S3. Auditing preprocessed data based on time series analysis to identify abnormal patterns; S4. Analyzing and fusing audit data using a graph neural network with causal reasoning; S5. Applying big data analysis techniques to perform multi-dimensional analysis and prediction on the fused data; S6. Implementing real-time data visualization and interactive display using a quantum computing-driven visualization method; S7. Generating reports on the soil dataset and analysis results using natural language generation technology; and S8. Implementing encryption and access control during data processing and transmission. This invention provides a comprehensive soil data processing solution that effectively improves the efficiency, accuracy, and security of data auditing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of agricultural information management, and in particular to an automated review and analysis method for cultivated land soil data. Background Art

[0002] In modern agriculture, cultivated land soil data has become an important production factor and strategic resource. With the development of agricultural information technology, the demand for soil data in various fields is increasing, including precision agriculture, environmental monitoring, and land management. Efficient data processing and analysis can promote the optimal allocation of resources, improve agricultural production efficiency, and promote scientific and technological innovation. However, soil data management and analysis also face many challenges, especially in data collection, processing efficiency, and data security. Traditional soil data processing methods include manual data collection, data cleaning, and simple statistical analysis. Although existing technologies have solved some problems in data processing to a certain extent, there are still many defects and challenges, which are specifically manifested in the following aspects:

[0003] First, there are limitations to data collection. Traditional soil data collection methods rely primarily on ground sensors and manual sampling, which have significant limitations in terms of coverage and data update frequency. Ground sensors can only cover a limited area, with few sampling points, making it difficult to fully reflect the soil conditions of large areas of farmland. Laboratory analysis processes are cumbersome, and data updates are infrequent, making it difficult to reflect dynamic changes in soil conditions in real time. These factors limit soil data collection and prevent comprehensive and accurate soil information from being obtained.

[0004] Secondly, data processing is inefficient. Existing soil data processing and review primarily relies on manual data cleaning, normalization, and outlier detection. Faced with massive amounts of soil data, manual processing is not only time-consuming and labor-intensive but also prone to human error, impacting the accuracy and reliability of data review results. This inefficient processing method fails to meet the demands of modern agriculture for fast and efficient data processing, resulting in low efficiency and accuracy in soil data review and analysis.

[0005] Third, the analytical approach is limited. Existing soil data analysis methods often rely on simple statistical analysis, making it difficult to deeply explore complex relationships within the data. For example, there is a lack of effective integration and analysis methods for multi-source, multi-temporal soil data, making it impossible to fully understand the soil's overall condition and development trends. This single analytical approach makes it difficult to support scientific decision-making in precision agriculture and cannot provide comprehensive and accurate soil condition assessment and prediction.

[0006] Furthermore, data security is insufficient. Existing technologies lack adequate encryption and access control measures during data transmission and processing, which can easily lead to data leakage and tampering, compromising data reliability and security. This lack of data security exposes soil data to significant risks during transmission and processing, impacting its credibility and usefulness.

[0007] Therefore, how to provide a method for automated review and analysis of cultivated land soil data is an urgent problem that needs to be solved by those skilled in the art. Summary of the Invention

[0008] One object of the present invention is to propose a method for automated audit and analysis of cultivated land soil data. The present invention achieves efficient collection, precise processing, multi-dimensional analysis and security management of soil data by combining Internet of Things sensors, drone multispectral imaging technology, extended Kalman filter, causal reasoning graph neural network, big data analysis, quantum computing-driven visualization technology and natural language generation technology. Specific methods include real-time multi-source data acquisition, noise reduction and outlier removal, time series analysis, deep analysis and data fusion, quantum computing visualization and automatic report generation. The method has the following advantages: it improves the efficiency, accuracy and security of data auditing, provides a comprehensive soil data processing solution, overcomes many shortcomings of existing technologies, and promotes precise management and scientific decision-making in modern agriculture.

[0009] According to an embodiment of the present invention, a method for automatically reviewing and analyzing cultivated land soil data includes the following steps:

[0010] S1. Collect raw soil data and conduct real-time multi-source data collection and transmission through IoT sensors and drone multispectral imaging to form a soil dataset.

[0011] S2. Preprocess the soil dataset, use the extended Kalman filter to reduce data noise and remove outliers, and use the time series interpolation method to fill in missing data;

[0012] S3, perform data audit on the preprocessed soil dataset based on time series analysis to identify potential abnormal data patterns;

[0013] S4. Use graph neural networks with causal reasoning to perform data analysis and fusion on the audited soil dataset;

[0014] S5. Use big data analysis technology to perform multidimensional data analysis and prediction on the fused soil dataset;

[0015] S6. Use quantum computing-driven visualization methods to perform real-time data visualization and interactive display of analysis and prediction results;

[0016] S7. Use natural language generation technology to generate reports based on the latest data and analysis results of the soil dataset, providing data summaries, anomaly detection results, analysis conclusions and recommendations, and supporting multiple formats for output;

[0017] S8. During the processing and transmission of soil datasets, encryption and access control policies are applied to each data block.

[0018] Optionally, the S2 specifically includes:

[0019] S21. Perform preliminary noise reduction on the soil dataset by using an improved adaptive extended Kalman filter, initialize the state vector and covariance matrix, and perform noise modeling in combination with the multivariate Student distribution, setting the state transfer equation and the Jacobian matrix of the state transfer function;

[0020] S22. Use Gaussian process regression to predict the state and estimate the noise, and optimize the state prediction through variational inference:

[0021]

[0022] Among them, P k|k-1 Forecast error covariance matrix, F k-1 is the Jacobian matrix of the state transfer function, P k-1 The error covariance matrix of the previous moment, ν is the degree of freedom parameter, w k-1 is the process noise vector, I is the identity matrix, α is the weight coefficient, is the Gaussian basis function, N is the number of basis functions, (x k-1 ) T variable x k-1 The transpose of

[0023] S23. Perform measurement updates through multimodal particle filtering, calculate measurement residuals and adaptive Kalman gains, and fuse the measurement results with the prediction results;

[0024] S24, using the adaptive Kalman gain to modify the state vector and covariance matrix to further reduce the impact of noise;

[0025] S25. Remove outliers from the denoised and corrected soil dataset by using an outlier detection method based on a deep generative model. The model is trained using a variational autoencoder, and a generative model and a discriminative model are constructed. The anomaly score for each data point is calculated, a threshold is set, and outliers are marked and removed.

[0026] S26. Complete the missing data using sparse Bayesian learning. Through sparse Bayesian matrix decomposition and deep generative adversarial network optimization, define the observation data matrix and dictionary matrix, establish a sparse Bayesian matrix decomposition model, and jointly optimize the generative adversarial network and sparse Bayesian learning:

[0027]

[0028] Among them, ln p(Y|X,A,E) is the log-likelihood of the observation data Y given the input X, matrix A and noise E, is the expected value of Y from the true data distribution, ln D(Y) is the log probability of the discriminator D for the true data Y, λ is the weight parameter, γ and δ are regularization parameters, is the expected value of X from the input data distribution, ln(1-D(G(X))) is the logarithmic probability of the discriminator D for the generated data G(X), ||·||1 represents the L1 norm, D and G represent the discriminator and generator of the generative adversarial network respectively, and the optimal sparse coefficient matrix X is obtained by jointly optimizing the above complex loss function. optimized And the dictionary matrix A, finally reconstruct the missing data.

[0029] Optionally, the S3 specifically includes:

[0030] S31, performing time series decomposition on the preprocessed soil data set, decomposing the time series into trend components, seasonal components, and residual components;

[0031] S32. Perform advanced trend analysis on trend components, using logarithmic transformation and nonlinear filtering methods to convert trend components into:

[0032]

[0033] Where T(t) is the original trend component, ∈ is a small constant used to avoid calculation errors in logarithmic transformation, and coeff n is the filter coefficient of the nth term, and α is the weight used to enhance the nonlinear processing capability of the model;

[0034] S33. Model the seasonal component S(t) using non-parametric Bayesian methods and use Gaussian process regression for prediction:

[0035]

[0036] Where m(t) is the mean function of the Gaussian process, is the function variance, l is the length scale, is the noise variance, p is the period length, corresponding to the seasonal variation period, and δ(t-t') is the Diracdelta function, which handles the autocorrelation problem;

[0037] S34. Perform nonlinear modeling on the residual component R(t), use variational autoencoder for residual generation and prediction, and optimize the objective function as follows:

[0038]

[0039] in, The loss function of the variational autoencoder, E represents the expected value, is the encoder, and its parameters are -lnp θ (x|z) is the negative of the log-likelihood, which represents the probability of generating data x given the latent variable z; p θ (x|z) is the decoder, p(z) is the prior distribution with parameter θ; z is the latent variable; N represents the Gaussian distribution; KL represents the Kullback-Leibler divergence;

[0040] S35. Detect abnormal data based on the decomposed components, mark the trend component and seasonal component as outliers, use the isolation forest method to determine the outlier threshold, and mark and remove abnormal data based on the variational autoencoder model of the residual component;

[0041] S36. Reconstruct the data after marking and removing outliers, combine the trend component, seasonal component and corrected residual component, and reconstruct the complete time series data.

[0042] Optionally, the S4 includes the following steps:

[0043] S41. Extract features from the audited soil dataset and use the graph neural network method to build a causal graph model of soil data. Generate the hidden state vector of each node by aggregating and updating node features:

[0044]

[0045] in, is the hidden state vector of node v in the k+1 layer, σ is the activation function, W (k) is the weight matrix of the kth layer, AGG is the aggregation function, is the hidden state vector of neighbor node u in layer k, e vu and e vv is the edge weight vector between nodes, b (k) is the bias vector of the kth layer, ⊙ represents element-wise multiplication;

[0046] S42. Perform causal reasoning on the constructed causal graph model and use the causal reasoning framework to evaluate the causal relationship in the soil dataset;

[0047] S43. Perform feature selection on the causal inference results, use the attention mechanism to calculate the weight of each feature, and select key features;

[0048] S44. Perform cluster analysis on the selected key features and use a deep embedding clustering method based on graph neural network. The optimization objective function is:

[0049]

[0050] Among them, z i is the representation of data point i in the embedding space, μ j is the representation of cluster center j, α is the degree of freedom parameter of Student’s t distribution, is the frequency of cluster center j, and j' is the index variable when summing the cluster centers;

[0051] S45. Based on the cluster analysis results, combined with feature selection and cluster analysis for causal reasoning, the soil dataset is updated, including the soil types of different clusters and the main characteristics of each soil type.

[0052] S46. Combine the updated soil dataset with the time series analysis results to perform data fusion.

[0053] Optionally, the S44 includes the following steps:

[0054] S441. Preprocess the selected key features and embed them into a high-dimensional space using a graph neural network method to generate embedded representations of nodes and initial cluster center representations.

[0055] S442, initialize cluster center μ j , select the initial cluster center in the embedding space;

[0056] S443, calculate the adaptive soft assignment q of each data point i to cluster center j ij :

[0057]

[0058] in, is the variance of cluster center j, z i is the embedding representation of data point i, μ j is the embedded representation of cluster center j;

[0059] S444. Calculate target distribution p ij , further optimize the soft allocation results:

[0060]

[0061] in, is the frequency of cluster center j;

[0062] S445. Define the regularized KL divergence loss function to measure the soft assignment q ij With the target distribution p ij The difference between , and add information entropy regularization:

[0063]

[0064] Among them, λ is the regularization parameter;

[0065] S446, based on the adaptive gradient descent method to minimize the regularized KL divergence loss function, optimize the objective function Update node embedding representations and cluster centers through backpropagation:

[0066]

[0067] Where η is the learning rate, and are the cluster center and node embedding representation at the t-th iteration, and are the cumulative sum of squared gradients;

[0068] S447, introduce adaptive neighborhood graph regularization and update the loss function to:

[0069]

[0070] Among them, β and δ are regularization parameters, is the edge set of the graph, cos(z i ,z j ) represents the cosine similarity between node embeddings;

[0071] S448. Combine contrastive learning to construct contrastive loss and update the loss function to:

[0072]

[0073] Among them, γ is the weight parameter of contrastive learning, 1 is the indicator function, and y i and y j are the category labels of data points i and j respectively;

[0074] S449, repeat steps S443 to S448 until the objective function Convergence is achieved and the final clustering result is obtained.

[0075] Optionally, the S6 includes the following steps:

[0076] S61. Build quantum visualization algorithms and design quantum circuits to process and map soil datasets.

[0077] S62. Define the encoding and decoding strategy of the quantum state, encode the soil dataset into the quantum state, and decode it after the quantum computation is completed;

[0078] S63. Execute quantum computing tasks, input analysis and prediction data into the quantum processor, and realize data processing through quantum state evolution:

[0079]

[0080] Among them, ψ evolved is the quantum state after evolution, β j is the evolution coefficient, is the quantum operator, M is the number of evolution steps, γ(x) is the weight function, and exp(iHx) is the quantum coherent operation involving the Hamiltonian H;

[0081] S64, converting the output data of the quantum calculation into classical data, obtaining the classical data through quantum state measurement, and reconstructing the classical data;

[0082] S65. Convert the classical data output by quantum computing into visual graphics, use classical computers to post-process the quantum results, and generate charts and visual interfaces;

[0083] S66, realize real-time interactive display, through which users can dynamically adjust and view analysis and prediction results;

[0084] S67, introduces visualization of quantum state entanglement and superposition to demonstrate the correlation between different quantum states of the soil dataset;

[0085] S68. Using quantum state coherence and decoherence analysis, we show how soil datasets change under different environments.

[0086]

[0087] in, is the Levee operator, ρ is the density matrix, L i is the environmental action operator, It's L i The conjugate transpose operator, γ jk is the decoherence rate, A j and A k is the coherence operator.

[0088] The beneficial effects of the present invention are:

[0089] (1) This invention significantly improves the efficiency and accuracy of arable land soil data review and analysis by combining efficient data preprocessing techniques, anomaly detection methods based on time series analysis, causal reasoning graph neural networks, and big data analysis techniques. In particular, in terms of data fusion and prediction, this invention overcomes the shortcomings of traditional methods, can quickly and accurately process large-scale, multi-source soil data, and provide a unified data representation.

[0090] (2) By introducing quantum computing-driven visualization methods and natural language generation technology, the present invention optimizes the presentation of data analysis results and report generation. Quantum computing-driven visualization methods enable real-time data visualization and interactive presentation, allowing users to intuitively understand and analyze data results. Natural language generation technology automatically generates detailed reports based on the latest data and analysis results, including data summaries, anomaly detection results, analysis conclusions, and recommendations, improving the efficiency and accuracy of users' understanding and application of data.

[0091] (3) This invention uses advanced data encryption and access control strategies to enhance data security and effectively prevent unauthorized access and data leakage. These measures can respond to ever-changing security threats and ensure the security of soil data during processing and transmission. BRIEF DESCRIPTION OF THE DRAWINGS

[0092] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0093] Figure 1 This is a general framework diagram of an automated audit and analysis method for cultivated land soil data proposed by the present invention;

[0094] Figure 2 This is a flow chart of the extended Kalman filter denoising and outlier removal for the automated audit and analysis method for cultivated land soil data proposed in the present invention;

[0095] Figure 3 This is a schematic diagram of the deep analysis and fusion module of the causal reasoning graph neural network of the automated audit and analysis method for cultivated land soil data proposed in the present invention;

[0096] Figure 4 This is a quantum computing-driven real-time data visualization and interactive display block diagram of the automated review and analysis method for cultivated land soil data proposed in the present invention. DETAILED DESCRIPTION

[0097] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.

[0098] refer to Figure 1-4 , a method for automatic review and analysis of cultivated land soil data, comprising the following steps:

[0099] S1. Collect raw soil data and conduct real-time multi-source data collection and transmission through IoT sensors and drone multispectral imaging to form a soil dataset.

[0100] S2. Preprocess the soil dataset, use the extended Kalman filter to reduce data noise and remove outliers, and use the time series interpolation method to fill in missing data;

[0101] S3, perform data audit on the preprocessed soil dataset based on time series analysis to identify potential abnormal data patterns;

[0102] S4. Use graph neural networks with causal reasoning to perform data analysis and fusion on the audited soil dataset;

[0103] S5. Use big data analysis technology to perform multidimensional data analysis and prediction on the fused soil dataset;

[0104] S6. Use quantum computing-driven visualization methods to perform real-time data visualization and interactive display of analysis and prediction results;

[0105] S7. Use natural language generation technology to generate reports based on the latest data and analysis results of the soil dataset, providing data summaries, anomaly detection results, analysis conclusions and recommendations, and supporting multiple formats for output;

[0106] S8. During the processing and transmission of soil datasets, encryption and access control policies are applied to each data block.

[0107] In this embodiment, S2 specifically includes:

[0108] S21. Perform preliminary noise reduction on the soil dataset by using an improved adaptive extended Kalman filter, initialize the state vector and covariance matrix, and perform noise modeling in combination with the multivariate Student distribution, setting the state transfer equation and the Jacobian matrix of the state transfer function;

[0109] S22. Use Gaussian process regression to predict the state and estimate the noise, and optimize the state prediction through variational inference:

[0110]

[0111] Among them, P k|k-1 Forecast error covariance matrix, F k-1 is the Jacobian matrix of the state transfer function, P k-1The error covariance matrix of the previous moment, ν is the degree of freedom parameter, w k-1 is the process noise vector, I is the identity matrix, α is the weight coefficient, is the Gaussian basis function, N is the number of basis functions, (x k-1 ) T variable x k-1 The transpose of

[0112] S23. Perform measurement updates through multimodal particle filtering, calculate measurement residuals and adaptive Kalman gains, and fuse the measurement results with the prediction results;

[0113] S24, using the adaptive Kalman gain to modify the state vector and covariance matrix to further reduce the impact of noise;

[0114] S25. Remove outliers from the denoised and corrected soil dataset by using an outlier detection method based on a deep generative model. The model is trained using a variational autoencoder, and a generative model and a discriminative model are constructed. The anomaly score for each data point is calculated, a threshold is set, and outliers are marked and removed.

[0115] S26. Complete the missing data using sparse Bayesian learning. Through sparse Bayesian matrix decomposition and deep generative adversarial network optimization, define the observation data matrix and dictionary matrix, establish a sparse Bayesian matrix decomposition model, and jointly optimize the generative adversarial network and sparse Bayesian learning:

[0116]

[0117] Among them, ln p(Y|X,A,E) is the log-likelihood of the observation data Y given the input X, matrix A and noise E, is the expected value of Y from the true data distribution, ln D(Y) is the log probability of the discriminator D for the true data Y, λ is the weight parameter, γ and δ are regularization parameters, is the expected value of X from the input data distribution, ln(1-D(G(X))) is the logarithmic probability of the discriminator D for the generated data G(X), ||·||1 represents the L1 norm, D and G represent the discriminator and generator of the generative adversarial network respectively, and the optimal sparse coefficient matrix X is obtained by jointly optimizing the above complex loss function. optimized And the dictionary matrix A, finally reconstruct the missing data.

[0118] In this embodiment, S3 specifically includes:

[0119] S31, performing time series decomposition on the preprocessed soil data set, decomposing the time series into trend components, seasonal components, and residual components;

[0120] S32. Perform advanced trend analysis on trend components, using logarithmic transformation and nonlinear filtering methods to convert trend components into:

[0121]

[0122] Where T(t) is the original trend component, ∈ is a small constant used to avoid calculation errors in logarithmic transformation, and coeff n is the filter coefficient of the nth term, and α is the weight used to enhance the nonlinear processing capability of the model;

[0123] S33. Model the seasonal component S(t) using non-parametric Bayesian methods and use Gaussian process regression for prediction:

[0124]

[0125] Where m(t) is the mean function of the Gaussian process, is the function variance, l is the length scale, is the noise variance, p is the period length, corresponding to the seasonal variation period, and δ(t-t') is the Diracdelta function, which handles the autocorrelation problem;

[0126] S34. Perform nonlinear modeling on the residual component R(t), use variational autoencoder for residual generation and prediction, and optimize the objective function as follows:

[0127]

[0128] in, The loss function of the variational autoencoder, E represents the expected value, is the encoder, and its parameters are -lnp θ (x|z) is the negative of the log-likelihood, which represents the probability of generating data x given the latent variable z; p θ (x|z) is the decoder, p(z) is the prior distribution with parameter θ; z is the latent variable; N represents the Gaussian distribution; KL represents the Kullback-Leibler divergence;

[0129] S35. Detect abnormal data based on the decomposed components, mark the trend component and seasonal component as outliers, use the isolation forest method to determine the outlier threshold, and mark and remove abnormal data based on the variational autoencoder model of the residual component;

[0130] S36. Reconstruct the data after marking and removing outliers, combine the trend component, seasonal component and corrected residual component, and reconstruct the complete time series data.

[0131] In this embodiment, the S4 specifically includes:

[0132] S41. Extract features from the audited soil dataset and use the graph neural network method to build a causal graph model of soil data. Generate the hidden state vector of each node by aggregating and updating node features:

[0133]

[0134] in, is the hidden state vector of node v in the k+1 layer, σ is the activation function, W (k) is the weight matrix of the kth layer, AGG is the aggregation function, is the hidden state vector of neighbor node u in layer k, e vu and e vv is the edge weight vector between nodes, b (k) is the bias vector of the kth layer, ⊙ represents element-wise multiplication;

[0135] S42. Perform causal reasoning on the constructed causal graph model and use the causal reasoning framework to evaluate the causal relationship in the soil dataset;

[0136] S43. Perform feature selection on the causal inference results, use the attention mechanism to calculate the weight of each feature, and select key features;

[0137] S44. Perform cluster analysis on the selected key features and use a deep embedding clustering method based on graph neural network. The optimization objective function is:

[0138]

[0139] Among them, z i is the representation of data point i in the embedding space, μ j is the representation of cluster center j, α is the degree of freedom parameter of Student’s t distribution, is the frequency of cluster center j, and j' is the index variable when summing cluster centers. This formula combines embedding representation, Student's t distribution, and frequency adjustment. By minimizing the loss function, it ultimately achieves cluster analysis of key features. It can not only process high-dimensional data, but also effectively identify outliers and abnormal values, ensuring the stability and accuracy of clustering results. This optimization objective function can effectively improve the effectiveness of data analysis and provide a reliable basis for subsequent data fusion and decision-making.

[0140] S45. Based on the cluster analysis results, combined with feature selection and cluster analysis for causal reasoning, the soil dataset is updated, including the soil types of different clusters and the main characteristics of each soil type.

[0141] S46. Combine the updated soil dataset with the time series analysis results to perform data fusion.

[0142] In this embodiment, the S44 specifically includes:

[0143] S441. Preprocess the selected key features and embed them into a high-dimensional space using a graph neural network method to generate embedded representations of nodes and initial cluster center representations.

[0144] S442, initialize cluster center μ j , select the initial cluster center in the embedding space;

[0145] S443, calculate the adaptive soft assignment q of each data point i to cluster center j ij :

[0146]

[0147] in, is the variance of cluster center j, z i is the embedding representation of data point i, μ j is the embedded representation of cluster center j;

[0148] S444. Calculate target distribution p ij , further optimize the soft allocation results:

[0149]

[0150] in, is the frequency of cluster center j;

[0151] S445. Define the regularized KL divergence loss function to measure the soft assignment q ij With the target distribution p ij The difference between , and add information entropy regularization:

[0152]

[0153] Among them, λ is the regularization parameter;

[0154] S446, based on the adaptive gradient descent method to minimize the regularized KL divergence loss function, optimize the objective function Update node embedding representations and cluster centers through backpropagation:

[0155]

[0156] Where η is the learning rate, and are the cluster center and node embedding representation at the t-th iteration, and are the cumulative sum of squared gradients;

[0157] S447, introduce adaptive neighborhood graph regularization and update the loss function to:

[0158]

[0159] Among them, β and δ are regularization parameters, is the edge set of the graph, cos(z i ,z j ) represents the cosine similarity between node embeddings;

[0160] S448. Combine contrastive learning to construct contrastive loss and update the loss function to:

[0161]

[0162] Among them, γ is the weight parameter of contrastive learning, 1 is the indicator function, and y i and y j are the category labels of data points i and j respectively;

[0163] S449, repeat steps S443 to S448 until the objective function Convergence is achieved and the final clustering result is obtained.

[0164] In this embodiment, S6 specifically includes:

[0165] S61. Build quantum visualization algorithms and design quantum circuits to process and map soil datasets.

[0166] S62. Define the encoding and decoding strategy of the quantum state, encode the soil dataset into the quantum state, and decode it after the quantum computation is completed;

[0167] S63. Execute quantum computing tasks, input analysis and prediction data into the quantum processor, and realize data processing through quantum state evolution:

[0168]

[0169] Among them, ψ evolved is the quantum state after evolution, β j is the evolution coefficient, is the quantum operator, M is the number of evolution steps, γ(x) is the weight function, and exp(iHx) is the quantum coherent operation involving the Hamiltonian H;

[0170] S64, converting the output data of the quantum calculation into classical data, obtaining the classical data through quantum state measurement, and reconstructing the classical data;

[0171] S65. Convert the classical data output by quantum computing into visual graphics, use classical computers to post-process the quantum results, and generate charts and visual interfaces;

[0172] S66, realize real-time interactive display, through which users can dynamically adjust and view analysis and prediction results;

[0173] S67, introduces visualization of quantum state entanglement and superposition to demonstrate the correlation between different quantum states of the soil dataset;

[0174] S68. Using quantum state coherence and decoherence analysis, we show how soil datasets change under different environments.

[0175]

[0176] in, is the Levee operator, ρ is the density matrix, L i is the environmental action operator, It's L i The conjugate transpose operator, γ jk is the decoherence rate, A j and A k is the coherence operator.

[0177] Example 1:

[0178] To verify the feasibility of this invention, a comprehensive test was conducted on a large farmland in the North China Plain. An automated soil data audit and analysis network was established. Each farmland area served as a node, and data was collected using IoT sensors and drone multispectral imaging technology. The data was then interconnected via a data processing and analysis network. The entire process included data collection, data preprocessing, in-depth analysis and fusion, data analysis and prediction, visualization, and report generation.

[0179] A network of IoT sensors has been deployed across various areas of the farmland, collecting real-time data on soil temperature, moisture, pH, and nutrient content. To obtain data over a wider area and at higher resolution, the team also uses drones for multispectral imaging, acquiring spectral information about the soil through remote sensing technology.

[0180] All collected data is first processed through an extended Kalman filter. This filter effectively reduces noise and removes outliers, ensuring data accuracy and reliability. The processed data is then supplemented using time series interpolation, addressing missing data caused by sensor failure or signal loss.

[0181] The pre-processed data is deeply analyzed and integrated using a causal inference graph neural network, identifying complex relationships between the data and revealing the causal relationship between soil health and the crop growth environment. For example, high soil pH in certain areas may be related to local fertilization practices, while insufficient soil moisture may be associated with a malfunction in the irrigation system.

[0182] The team then used big data analytics to conduct multi-dimensional analysis and predictions on the integrated data. By analyzing historical and current data, the system generated soil health reports and predicted future soil trends. For example, based on trends in soil nutrient content, the system predicted that certain areas might require additional fertilizer in the coming months, allowing for proactive preparation and preventing crop yield reductions due to nutrient deficiencies.

[0183] To make data analysis more intuitive, the research team introduced quantum computing-driven visualization technology. Leveraging the efficient processing power of quantum computing, the system generates real-time visualizations of soil health and displays them on an interactive display platform. Farmers can easily view soil conditions in different regions, historical trends, and future forecasts, enabling them to make informed agricultural management decisions.

[0184] Finally, the team used natural language generation technology to automatically generate a detailed soil health report based on the data analysis results. The report includes a data summary, anomaly detection results, analysis conclusions, and recommendations. All information is presented in plain language, making it easier for farmers to understand and apply.

[0185] Table 1 Test data and result analysis table of the present invention

[0186]

[0187]

[0188] During the six-month testing period, a large amount of soil data was processed, validating the effectiveness and superiority of this method in practical applications. All data remained encrypted during sharing and calculation, and no data leaks occurred. Regular security audits ensured the security of data storage and transmission at each node, achieving 100% privacy protection coverage.

[0189] By using the method of this invention, the average data processing time per transaction was reduced from 2 hours to 30 minutes, a 4-fold increase in efficiency. The result merging and verification time was reduced from 1 hour to 5 minutes, a 12-fold increase. The time for generating and broadcasting new blocks was reduced from 30 minutes to 3 minutes, a 10-fold increase. The total processing time was reduced from 3.5 hours to 38 minutes, an approximately 5.5-fold increase in efficiency. All data operations are recorded on-chain, with clear operation logs and transaction records. Smart contracts automatically enforce data authorization and sharing rules, eliminating manual intervention and ensuring the fairness and compliance of data operations.

[0190] In summary, the present invention provides a comprehensive soil data processing solution, which significantly improves the efficiency, accuracy and security of data review, overcomes many shortcomings of the existing technology, provides strong support for the precise management and scientific decision-making of modern agriculture, and has broad application prospects.

[0191] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A method for automated review and analysis of cultivated land soil data, characterized in that: The steps include: S1. Collect raw soil data and conduct real-time multi-source data collection and transmission through IoT sensors and drone multispectral imaging to form a soil dataset. S2. Preprocess the soil dataset, use the extended Kalman filter to reduce data noise and remove outliers, and use the time series interpolation method to fill in missing data; S3, perform data audit on the preprocessed soil dataset based on time series analysis to identify potential abnormal data patterns; S4. Use graph neural networks with causal reasoning to analyze and fuse the audited soil dataset, identify complex associations between the data, and reveal the causal relationship between soil health and the crop growth environment; S5. Use big data analysis technology to conduct multi-dimensional data analysis and prediction on the integrated soil data set. By analyzing historical and current data, a soil health report is generated and future soil change trends are predicted. S6. Use quantum computing-driven visualization methods to perform real-time data visualization and interactive display of analysis and prediction results, generate real-time visualization charts of soil health status, and display them on an interactive display platform; S7. Use natural language generation technology to automatically generate detailed soil health reports based on data analysis results. The reports include data summaries, anomaly detection results, analysis conclusions and recommendations, and support multiple output formats. S8. Apply encryption and access control policies to each data block during the processing and transmission of soil datasets; The S3 includes: S31, performing time series decomposition on the preprocessed soil data set, decomposing the time series into trend components, seasonal components, and residual components; S32. Perform advanced trend analysis on trend components, using logarithmic transformation and nonlinear filtering methods to extract long-term trend components: Where T(t) is the original trend component, ∈ is a small constant used to avoid calculation errors in logarithmic transformation, and coeff n is the filter coefficient of the nth term, and α is the weight used to enhance the nonlinear processing capability of the model; S33. Model the seasonal component S(t) using non-parametric Bayesian methods and use Gaussian process regression for prediction: Where m(t) is the mean function of the Gaussian process, is the function variance, l is the length scale, is the noise variance, p is the period length, corresponding to the seasonal variation period, and δ(t-t') is the Diracdelta function, which handles the autocorrelation problem; S34, performing nonlinear modeling on the residual component R(t), and using a variational autoencoder for residual generation and prediction; S35. Detect abnormal data based on the decomposed components, mark the trend component and seasonal component as outliers, use the isolation forest method to determine the outlier threshold, and mark and remove abnormal data based on the variational autoencoder model of the residual component; S36. Reconstruct the data after marking and removing outliers, combine the trend component, seasonal component and corrected residual component, and reconstruct the complete time series data.

2. The method for automatic review and analysis of cultivated land soil data according to claim 1, characterized in that: The S2 comprises the following steps: S21. Perform preliminary noise reduction on the soil dataset by using an improved adaptive extended Kalman filter, initialize the state vector and covariance matrix, and perform noise modeling in combination with the multivariate Student distribution, setting the state transfer equation and the Jacobian matrix of the state transfer function; S22, using Gaussian process regression to predict state and estimate noise, and optimizing state prediction through variational inference; S23. Perform measurement updates through multimodal particle filtering, calculate measurement residuals and adaptive Kalman gains, and fuse the measurement results with the prediction results; S24, using the adaptive Kalman gain to modify the state vector and covariance matrix to further reduce the impact of noise; S25. Remove outliers from the denoised and corrected soil dataset by using an outlier detection method based on a deep generative model. The model is trained using a variational autoencoder, and a generative model and a discriminative model are constructed. The anomaly score for each data point is calculated, a threshold is set, and outliers are marked and removed. S26. Complete the missing data using sparse Bayesian learning. Through sparse Bayesian matrix decomposition and deep generative adversarial network optimization, define the observation data matrix and dictionary matrix, establish a sparse Bayesian matrix decomposition model, and jointly optimize the generative adversarial network and sparse Bayesian learning: Among them, ln p(Y|X,A,E) is the log-likelihood of the observation data Y given the input X, dictionary matrix A and noise E, is the expected value of Y from the true data distribution, ln D(Y) is the log probability of the discriminator D for the true data Y, λ is the weight parameter, γ and δ are regularization parameters, is the expected value of X from the input data distribution, ln(1-D(G(X))) is the logarithmic probability of the discriminator D for the generated data G(X), ||·||1 represents the L1 norm, D and G represent the discriminator and generator of the generative adversarial network respectively, and the optimal sparse coefficient matrix X is obtained by jointly optimizing the above complex loss function. optimized And the dictionary matrix A, finally reconstruct the missing data.

3. The method for automatic review and analysis of cultivated land soil data according to claim 1, characterized in that: The S4 comprises the following steps: S41. Extract features from the audited soil dataset and use the graph neural network method to build a causal graph model of soil data. Generate the hidden state vector of each node by aggregating and updating node features: in, is the hidden state vector of node v in the k+1 layer, σ is the activation function, W (k) is the weight matrix of the kth layer, AGG is the aggregation function, is the hidden state vector of neighbor node u in layer k, e vu and e vv is the edge weight vector between nodes, b (k) is the bias vector of the kth layer, ⊙ represents element-wise multiplication; S42. Perform causal reasoning on the constructed causal graph model and use the causal reasoning framework to evaluate the causal relationship in the soil dataset; S43. Perform feature selection on the causal inference results, use the attention mechanism to calculate the weight of each feature, and select key features; S44. Perform cluster analysis on the selected key features and use a deep embedding clustering method based on graph neural network. The optimization objective function is: Among them, z i is the representation of data point i in the embedding space, μ j is the representation of cluster center j, α is the degree of freedom parameter of Student’s t distribution, f j is the frequency of cluster center j, j ' It is the index variable when summing the cluster centers; S45. Based on the cluster analysis results, combined with feature selection and cluster analysis for causal reasoning, the soil dataset is updated, including the soil types of different clusters and the main characteristics of each soil type. S46. Combine the updated soil dataset with the time series analysis results to perform data fusion.

4. The method for automatic review and analysis of cultivated land soil data according to claim 3, characterized in that: The S44 includes the following steps: S441. Preprocess the selected key features and embed them into a high-dimensional space using a graph neural network method to generate embedded representations of nodes and initial cluster center representations. S442, initialize cluster center μ j , select the initial cluster center in the embedding space; S443, calculate the adaptive soft assignment q of each data point i to cluster center j ij : in, is the variance of cluster center j, z i is the embedding representation of data point i, μ j is the embedded representation of cluster center j; S444. Calculate target distribution p ij , further optimize the soft allocation results: in, is the frequency of cluster center j; S445. Define the regularized KL divergence loss function to measure the soft assignment q ij With the target distribution p ij The difference between them is calculated and information entropy regularization is added: Among them, λ is the regularization parameter; S446, based on the adaptive gradient descent method to minimize the regularized KL divergence loss function, optimize the objective function Update node embedding representations and cluster centers through backpropagation: Where η is the learning rate, and are the cluster center and node embedding representation at the t-th iteration, and are the cumulative sum of squared gradients; S447, introduce adaptive neighborhood graph regularization and update the loss function to: Among them, β and δ are regularization parameters, is the edge set of the graph, cos(z i ,z j ) represents the cosine similarity between node embeddings; S448. Combine contrastive learning to construct contrastive loss and update the loss function to: Among them, γ is the weight parameter of contrastive learning, 1 is the indicator function, and y i and y j are the category labels of data points i and j respectively; S449, repeat steps S443 to S448 until the objective function Convergence is achieved and the final clustering result is obtained.

5. The method for automatic review and analysis of cultivated land soil data according to claim 1, characterized in that: The S6 comprises the following steps: S61. Build quantum visualization algorithms and design quantum circuits to process and map soil datasets. S62. Define the encoding and decoding strategy of the quantum state, encode the soil dataset into the quantum state, and decode it after the quantum computation is completed; S63. Execute quantum computing tasks, input analysis and prediction data into the quantum processor, and realize data processing through quantum state evolution: Among them, ψ evolved is the quantum state after evolution, β j is the evolution coefficient, is the quantum operator, M is the number of evolution steps, γ(x) is the weight function, and exp(iHx) is the quantum coherent operation involving the Hamiltonian H; S64, converting the output data of the quantum calculation into classical data, obtaining the classical data through quantum state measurement, and reconstructing the classical data; S65. Convert the classical data output by quantum computing into visual graphics, use classical computers to post-process the quantum results, and generate charts and visual interfaces; S66, realize real-time interactive display, through which users can dynamically adjust and view analysis and prediction results; S67, introduces visualization of quantum state entanglement and superposition to demonstrate the correlation between different quantum states of the soil dataset; S68. Using quantum state coherence and decoherence analysis, we show how soil datasets change under different environments. in, is the Levee operator, ρ is the density matrix, L i is the environmental action operator, It's L i The conjugate transpose operator, γ jk is the decoherence rate, A j and A k is the coherence operator.

Citation Information

Patent Citations

  • Visual obstetrical image examination processing method

    CN116825293A

  • Chemical safety automatic detection and monitoring system of cloud PLC (Programmable Logic Controller)

    CN117808166A

  • Ecological soil quality management method and system based on GIS and automation technology

    CN118350718A