Soil parameter data anomaly detection method and device based on diffusion model

Through the data anomaly detection and completion method based on diffusion model, the problem of data anomaly and missing in the field environment of soil sensors is solved, ensuring the completeness and accuracy of the data, and is suitable for precise agriculture and soil environment monitoring.

CN120372516APending Publication Date: 2025-07-25INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510550019.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Data collected by soil sensors in a field environment are susceptible to external interference and component performance limitations, resulting in abnormal and missing data, affecting the credibility of the data.

Method used

The data anomaly detection method based on the diffusion model is adopted, and the soil parameter data to be detected is gradually added and denoised, the data is reconstructed using the diffusion model, and the abnormal data is detected through the sliding window, and then the data is completed.

Benefits of technology

Effectively detect and compensate abnormal data, ensure the integrity and accuracy of high-spatial-time resolution data, and improve the credibility of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372516A_ABST
    Figure CN120372516A_ABST
Patent Text Reader

Abstract

The invention discloses a soil parameter data anomaly detection method and device based on a diffusion model, and the method comprises the steps: carrying out the noise addition of a first preset step number of to-be-detected soil parameter data, and obtaining first noise-added data; the first noise adding data are input into a denoising module in a diffusion model, first reconstruction data corresponding to the to-be-detected soil parameter data are obtained, and the first reconstruction data are denoising data obtained by conducting denoising of a first preset step number on the first noise adding data; first data of the to-be-detected soil parameter data in each sliding window and second data of the first reconstruction data in each sliding window are acquired by taking the first duration as the sliding window; for each sliding window, obtaining difference information of the first data and the second data under the current sliding window; and determining that the first data of the to-be-detected soil parameter data under any sliding window is abnormal data in response to the condition that the difference information of any sliding window meets a preset condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the field of data processing, and more specifically, to a method and apparatus for detecting anomalies in soil parameter data based on a diffusion model. Background Art

[0002] At present, soil sensors can be deployed in a large scale and with high density to form a grid monitoring network, and key areas can be densely deployed. It has the characteristic of supporting the acquisition of high spatio-temporal resolution data. However, the deployment environment is often in a greenhouse or a field. Especially in the field environment, the environment is relatively harsh, and it may be affected by various external interferences and limited by the performance of its own components. This may cause the measured data to be abnormal and missing, which leads to a decrease in the credibility of the subsequent analysis of the soil condition. Summary of the Invention

[0003] Embodiments of the present disclosure provide a method and apparatus for detecting anomalies in soil parameter data based on a diffusion model, which can effectively solve the problem of inaccurate data collected in the prior art.

[0004] In one general aspect, a method for detecting data anomalies is provided, including: adding noise to the soil parameter data to be detected for a first predetermined number of steps to obtain first noise-added data; inputting the first noise-added data into a denoising module of a diffusion model to obtain first reconstructed data corresponding to the soil parameter data to be detected, where the first reconstructed data is denoised data obtained by denoising the first noise-added data for a first predetermined number of steps; taking the first time period as a sliding window, and obtaining first data of the soil parameter data to be detected under each sliding window and second data of the first reconstructed data under each sliding window; for each sliding window, obtaining difference information between the first data and the second data in the current sliding window; in response to the difference information of any sliding window satisfying a predetermined condition, determining the first data of the soil parameter data to be detected under any sliding window as abnormal data.

[0005] Optionally, the above method further includes: removing all abnormal data from the soil parameter data to be detected to obtain third data; using the diffusion model to complete the third data.

[0006] Optionally, a diffusion model is used to complete the third data, including: adding random data to the data missing positions in the third data to obtain fourth data; adding noise to the fourth data for a second predetermined number of steps to obtain second noisy data; inputting the second noisy data into the denoising module of the diffusion model to obtain second reconstructed data corresponding to the fourth data, where the second reconstructed data is the denoised data obtained by denoising the second noisy data for the second predetermined number of steps; covering the data at the first predetermined positions in the third data to the corresponding positions in the second reconstructed data to obtain the third data after completion, where the first predetermined positions are the positions corresponding to the non-missing data positions in the third data.

[0007] Optionally, the second predetermined number of steps is determined as follows: in response to the data missing rate of the third data being less than or equal to the first threshold, determining the second predetermined number of steps as the first preset value; in response to the data missing rate of the third data being greater than the first threshold and less than the second threshold, determining the second predetermined number of steps as the second preset value; in response to the data missing rate of the third data being greater than or equal to the second threshold, determining the second predetermined number of steps as the third preset value; where the third preset value is greater than the second preset value, and the second preset value is greater than the first preset value.

[0008] Optionally, inputting the second noisy data into the denoising module of the diffusion model to obtain second reconstructed data corresponding to the fourth data includes: using the denoising module in the diffusion model to denoise the second noisy data for the second predetermined number of steps, where for each step of denoising in the second predetermined number of steps, the following processing is performed: in response to obtaining the denoised data of the previous step of the current step, covering the data at the corresponding positions in the denoised data with the data at the second predetermined positions in the predetermined noisy data to obtain intermediate data, where the predetermined noisy data is the noisy data obtained by adding noise to the fourth data for the current step, and the second predetermined positions are the positions corresponding to the non-missing data positions in the third data; performing denoising processing on the intermediate data for the current step to obtain the denoised data of the current step; determining the denoised data of the last step as the second reconstructed data.

[0009] Optionally, the diffusion model is trained as follows: constructing a training sample set, where each training sample includes noisy data, the number of processing steps, and the actual noise data added at each step in the number of processing steps; for each training sample in the training sample set, inputting the noisy data and the number of processing steps in the training sample into the denoising module of the initial diffusion model to obtain the estimated noise data added at each step in the number of processing steps of the noisy data; for each step in the number of processing steps, adjusting the parameters of the initial diffusion model according to the loss between the estimated noise data of the current step and the actual noise data, and training the initial diffusion model. In another general aspect, a data anomaly processing device is provided, comprising: a noise adding unit, configured to perform noise adding for a first predetermined number of steps on soil parameter data to be detected, to obtain first noisy data; a denoising unit, configured to input the first noisy data into a denoising module in a diffusion model, to obtain first reconstructed data corresponding to the soil parameter data to be detected, wherein the first reconstructed data is denoised data obtained by denoising the first noisy data for a first predetermined number of steps; a first acquisition unit, configured to obtain first data of the soil parameter data to be detected in each sliding window and second data of the first reconstructed data in each sliding window with a first time length as a sliding window; a second acquisition unit, configured to obtain, for each sliding window, difference information between the first data and the second data in the current sliding window; a detection unit, configured to determine that the first data of the soil parameter data to be detected in any sliding window is abnormal data in response to the difference information of any sliding window satisfying a predetermined condition.

[0010] Optionally, the data anomaly processing device further includes: a data completion unit configured to remove all abnormal data from the soil parameter data to be detected to obtain third data; and complete the third data using a diffusion model.

[0011] Optionally, the data completion unit is further configured to add random data to the data missing position in the third data to obtain fourth data; perform noise addition on the fourth data for a second predetermined number of steps to obtain second noisy data; input the second noisy data into a denoising module in the diffusion model to obtain second reconstructed data corresponding to the fourth data, wherein the second reconstructed data is denoised data obtained by denoising the second noisy data for a second predetermined number of steps; overwrite the data at the first predetermined position in the third data to the corresponding position in the second reconstructed data to obtain the completed third data, wherein the first predetermined position is a position corresponding to the position in the third data where data is not missing.

[0012] Optionally, the second predetermined number of steps is determined as follows: in response to the data missing rate of the third data being less than or equal to the first threshold, the second predetermined number of steps is determined to be the first preset value; in response to the data missing rate of the third data being greater than the first threshold and less than the second threshold, the second predetermined number of steps is determined to be the second preset value; in response to the data missing rate of the third data being greater than or equal to the second threshold, the second predetermined number of steps is determined to be the third preset value; wherein the third preset value is greater than the second preset value, and the second preset value is greater than the first preset value.

[0013] Optionally, the data completion unit is further configured to denoise the second noise-added data for a second predetermined number of steps using the denoising module in the diffusion model. For each step of denoising in the second predetermined number of steps, the following processing is performed: in response to obtaining the denoised data of the previous step of the current step, covering the corresponding position data in the denoised data with the data at the second predetermined position in the predetermined noise-added data to obtain intermediate data, where the predetermined noise-added data is the noise-added data obtained by adding noise to the fourth data for the current step, and the second predetermined position is the position corresponding to the non-missing position data in the third data; performing denoising processing on the intermediate data for the current step to obtain the denoised data of the current step; and determining the denoised data of the last step as the second reconstructed data.

[0014] Optionally, the diffusion model is trained in the following manner: constructing a training sample set, where each training sample includes noise-added data, the number of processing steps, and the actual noise data added in each step of the number of processing steps for the noise-added data; for each training sample in the training sample set, inputting the noise-added data and the number of processing steps in the training sample into the denoising module of the initial diffusion model to obtain the estimated noise data added in each step of the number of processing steps for the noise-added data; for each step of the number of processing steps, adjusting the parameters of the initial diffusion model according to the loss between the estimated noise data of the current step and the actual noise data, and training the initial diffusion model. In another general aspect, there is provided a computer-readable storage medium storing instructions, where when the instructions are run by at least one computing device, at least one computing device is caused to execute any one of the above data anomaly detection methods.

[0015] In another general aspect, there is provided a system including at least one computing device and at least one storage device storing instructions, where when the instructions are run by at least one computing device, at least one computing device is caused to execute any one of the above data anomaly detection methods.

[0016] In another general aspect, there is provided a computer program product including computer instructions, which when executed by a processor implement any one of the above data anomaly detection methods.

[0017] According to the soil parameter data anomaly detection method and device based on the diffusion model of the embodiments of the present disclosure, by gradually adding noise and gradually removing noise to the soil parameter data to be detected, the reconstructed data of the soil parameter data to be detected is obtained. According to the difference information between the reconstructed data and the soil parameter data to be detected, the anomaly of the soil parameter data to be detected can be effectively detected. Subsequently, without affecting the data collection process, data compensation and other processing can be performed on the abnormal data, thereby ensuring the integrity and accuracy of high spatio-temporal resolution data and improving the credibility of the data. Therefore, through the present disclosure, the problem of inaccurate data collected in the prior art can be effectively solved.

[0018] Additional aspects and / or advantages of the present disclosure will be set forth in part in the following description, and in part will be obvious from the description, or may be learned by practice of the present disclosure. Description of the Drawings

[0019] The above and other objects and features of the embodiments of the present disclosure will become more apparent from the following description taken in conjunction with the drawings showing embodiments thereof, wherein: Figure 1 is a flowchart of a data anomaly detection method according to an embodiment of the present disclosure; Figure 2 is a schematic structural diagram of a diffusion model according to an embodiment of the present disclosure; Figure 3 is a training flowchart of a diffusion model according to an embodiment of the present disclosure; Figure 4 is a schematic flowchart of a data optimization system according to an embodiment of the present disclosure; Figure 5 is a flowchart of an abnormal data detection according to an embodiment of the present disclosure; Figure 6 is a flowchart of a missing value completion according to an embodiment of the present disclosure; Figure 7 is a flowchart of a data anomaly processing device according to an embodiment of the present disclosure. Detailed Description of the Embodiments

[0020] The following detailed description is provided to assist the reader in obtaining a comprehensive understanding of the methods, apparatuses, and / or systems described herein. However, various changes, modifications, and equivalents of the methods, apparatuses, and / or systems described herein will be apparent after understanding the disclosure of this application. For example, the order of operations described herein is merely illustrative and is not limited to those set forth herein, but may be changed as will be apparent after understanding the disclosure of this application, except for operations that must occur in a particular order. Additionally, descriptions of features known in the art may be omitted for greater clarity and conciseness.

[0021] The features described herein may be implemented in different forms and should not be construed as limited to the examples described herein. Instead, the examples described herein are provided only to illustrate some of the many possible ways of implementing the methods, apparatuses, and / or systems described herein, which will be apparent after understanding the disclosure of this application.

[0022] As used herein, the term "and / or" includes any one of the associated listed items and any combination of any two or more thereof.

[0023] Although terms such as "first", "second", and "third" may be used herein to describe various components, elements, regions, layers, or parts, these components, elements, regions, layers, or parts should not be limited by these terms. Instead, these terms are only used to distinguish one component, element, region, layer, or part from another. Thus, without departing from the teachings of the examples, the first component, first element, first region, first layer, or first part described in the examples herein may also be referred to as the second component, second element, second region, second layer, or second part.

[0024] In the specification, when an element (such as a layer, region, or substrate) is described as being "on" another element, "connected to" or "coupled to" another element, the element can be directly "on" the other element, directly "connected to" or "coupled to" the other element, or there can be one or more other elements therebetween. In contrast, when an element is described as being "directly on" another element, "directly connected to" or "directly coupled to" another element, there can be no other elements therebetween.

[0025] The terms used herein are only for describing various examples and are not intended to limit the disclosure. Unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. The terms "comprising", "including", and "having" specify the presence of the stated features, quantities, operations, components, elements, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, quantities, operations, components, elements, and / or combinations thereof.

[0026] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains after understanding the disclosure. Unless clearly defined herein, terms (such as those defined in a general dictionary) should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and should not be interpreted in an idealized or overly formal manner.

[0027] Furthermore, in the description of the examples, when a detailed description of related structures or functions that are considered well-known would cause an unclear interpretation of the disclosure, such detailed descriptions will be omitted.

[0028] The method and apparatus for abnormal detection of soil parameter data based on a diffusion model according to the present disclosure will be described in detail below with reference to the accompanying drawings.

[0029] The present disclosure provides a method for data anomaly detection. Figure 1 is a flowchart showing the data anomaly detection method according to an embodiment of the present disclosure. Refer to Figure 1, the data anomaly detection method includes the following steps: In step S101, add noise to the soil parameter data to be detected for a first predetermined number of steps to obtain first noisy data.

[0030] As an example, the soil parameter data to be detected may include data of various soil parameters obtained by a soil sensor over a period of time. The soil parameters include but are not limited to: soil moisture, soil conductivity, soil temperature, soil water potential, soil pH, etc. The noise added in the above noise addition process can be Gaussian noise or other noise, and the present disclosure does not limit this.

[0031] As an example, the above first predetermined number of steps is the number of noise addition steps when adding noise to the soil parameter data to be detected, which can be set as needed. It should be noted that when adding noise to the soil parameter data to be detected for outlier detection, the number of noise addition steps (i.e., the first predetermined number of steps) should not be too large or too small, as long as it can achieve the purpose of destroying the dependence on the soil parameter data to be detected when the diffusion model denoises but not completely destroying the key features of the soil parameter data to be detected. For example, the number of noise addition steps can be selected as 50. After observation, the key features of the soil parameter data to be detected can still be observed in the data after adding noise at this number of steps, and the dependence on the soil parameter data to be detected is eliminated to a certain extent.

[0032] As an example, the process of adding noise to the soil parameter data to be detected can be implemented in the following manner: First, gradually add Gaussian noise to the soil parameter data to be detected, as shown in the following formula: (1) where t represents the current step in the first predetermined number of steps, N(0, I) is the standard Gaussian noise, represents the degree of noise addition at the current step of noise addition, which can be set as needed, represents the noisy data after adding noise at the current step.

[0033] Secondly, let , then formula (1) is converted into the following formula (2): (2) Thirdly, the noise addition process can be simplified, that is, formula (2) is simplified into the following formula (3): (3) where, .

[0034] It should be noted that the degree of noise addition of the diffusion process can choose a cosine schedule, as shown below: (4) Among them, It is determined by the cosine function, specifically as follows: (5) Among them, t represents the current step (normalized to [0, T]), and T s is the maximum number of steps of the first predetermined number of steps, and s is a smoothing parameter, which can take 0.008 to prevent premature decay to 0.

[0035] Return Figure 1 , in step S102, the first noisy data is input into the denoising module of the diffusion model to obtain the first reconstructed data corresponding to the soil parameter data to be detected, where the first reconstructed data is the denoised data obtained by denoising the first noisy data for the first predetermined number of steps.

[0036] As an example, the trained diffusion model can predict the noise data added at each step in the first predetermined number of steps for the first noisy data, and perform denoising processing on the first noisy data for the first predetermined number of steps through the predicted noise data added at each step. Specifically, the denoising processing of the first noisy data can be completed by iterative denoising. For example, it can be as shown in the following formula: (17) Among them, represents the denoised data after the t step, t represents the current step in the first predetermined number of steps, represents the degree of noise addition when adding noise to the current step, , is the noise data predicted by the diffusion model to be added at the current step in the first predetermined number of steps t , is the coefficient for controlling the smoothness of the denoising process, is the new random noise, that is, a small amount of noise is added after each step of denoising to prevent the model from overfitting, but no noise is added when t = 0. It should be noted that and are set according to needs, so they can be directly obtained.

[0037] As an example, before performing the above step S102, it is necessary to construct a diffusion model and train the constructed diffusion model in advance. The following will be described separately from two aspects: constructing the diffusion model and training the diffusion model.

[0038] 1. Construct the diffusion model: First, obtain the position information of the data, where the data is the data at which step among the first predetermined number of steps of adding noise. The position information can be represented by PE, and the calculation formula of PE can be as follows: (6) (7) where t represents the number of noise-adding steps for obtaining the current data, d represents the coding dimension, 2i represents the dimension of the even steps, that is, the position information of the even steps adopts ; 2i + 1 represents the dimension of the odd steps, that is, the position information of the odd steps adopts , it should be noted that 2i ≤ d, 2i + 1 ≤ d. The above position information PE can be used as a conditional input to guide the diffusion model to adjust the weights at different diffusion stages. For example, the position information can be explicitly embedded in the diffusion model, so that the diffusion model can effectively distinguish the noise levels and achieve controllable generation from pure noise to clean data.

[0039] Secondly, in the diffusion model, structured long convolution (SLConv) is used to extract the temporal features of the data. Through multi-scale convolutional kernels and learnable skip connections, long-range dependencies in the signal can be efficiently captured. The structured long convolution can be constructed in the following way: 1) Set the required scale of the convolutional kernel according to needs. For each scale, the initial convolutional kernel kernel is interpolated to the size of the corresponding scale: (8) where i represents the serial number of the current scale, represents the convolutional kernel after kernel is interpolated to the size of the i th scale.

[0040] The interpolated convolutional kernel also needs to be multiplied by a scale decay factor, as shown below: (9) where num - scales represents i the maximum value of, i represents the serial number of the current scale.

[0041] Splice the convolutional kernels of each scale together to form a "long convolutional kernel", as shown below: (10) where represents the convolutional kernel spliced in the length direction.

[0042] The length of the spliced convolutional kernel, that is, the length of the "long convolutional kernel" is: (11) Normalize the concatenated convolution kernel, that is, the "long convolution kernel", as follows: (12) where kernel_norm is the L2 norm of the convolution kernel, calculated as follows: (13) Normalizing the convolution kernel ensures the stability of the numerical range of the convolution kernel, thereby improving the training stability and generalization ability of the diffusion model. It should be noted that the diffusion model can perform convolution operations in a multi-head manner, that is, initialize convolution kernels corresponding to the number of heads for each channel and perform convolution operations, which is not limited in this disclosure.

[0043] 2) Add a learnable skip matrix D, that is, map the input signal of the diffusion model to the same format (shape) as the result of each step of convolution operation through linear transformation and perform addition processing. In this way, it can help the diffusion model better learn long-range dependence information. It should be noted that in each step of the denoising process of the diffusion model, the input signal of the diffusion model needs to be mapped to the same format (shape) as the result of the current step of convolution operation and perform addition processing to avoid losing the original features of the data after each step of denoising.

[0044] As an example, Figure 2 the structure of the denoising module in a diffusion model is given, as Figure 2 shown. First, calculate the position information of the data after denoising in the previous step, specifically referring to formulas (6) and (7); then, perform preliminary feature extraction on the data after denoising in the previous step (i.e., Figure 2 the input data) and the corresponding position information through Conv1D, then use structured long convolution (SLConv) for local and global feature extraction, and then use 1×1 convolution to reduce the dimension of the output of SLConv. The feature extraction and dimension reduction operations of structured convolution are performed three times in a loop. Finally, the data is scaled to the original size through a Conv1D and output. This output is the noise data of the current step predicted by the denoising module in the diffusion model based on the input data. Then, subtract the predicted noise data of the current step from the data after denoising in the previous step to obtain the denoised data of the current step. The denoising process of the next step continues the above process.

[0045] 2. Train the diffusion model: According to an embodiment of the present disclosure, the diffusion model can be trained in the following manner: constructing a training sample set, where each training sample includes noisy data, the number of processing steps, and the actual noise data added at each step in the number of processing steps; for each training sample in the training sample set, input the noisy data and the number of processing steps in the training sample into the denoising module of the initial diffusion model to obtain the corresponding estimated noise data added at each step in the number of processing steps for the noisy data; for each step in the number of processing steps, adjust the parameters of the initial diffusion model according to the loss between the estimated noise data and the actual noise data at the current step, and train the initial diffusion model. Through the training method of this embodiment, an effective diffusion model can be obtained.

[0046] As an example, Figure 3 shows the training process of the diffusion model. As Figure 3 shown, noisy data, the number of steps t of the noisy data (i.e., the above-mentioned number of processing steps), and the actual noise data added at each step of the noisy data can be obtained. The first two can be used as training samples, and the latter can be used as the label of the training sample. When multiple training samples are obtained, a training sample set can be formed. It should be noted that the number of processing steps can be randomly generated during the training process, that is, the noisy data is obtained by adding noise through randomly generated processing steps, and the added noise can be obtained by adding noise to the collected data through the corresponding number of processing steps. The collected data can include various parameters, and the present disclosure does not limit this.

[0047] After starting the training, input the noisy data and the number of processing steps into the denoising module of the diffusion model, so that the denoising module can denoise the noisy data for the number of steps t, and output the predicted noise data added at each step of the noisy data; then for each step in the number of processing steps, calculate the loss using the predicted noise data and the actual noise data, which can be specifically shown as follows: First, calculate the residual between the actual noise data and the predicted noise data diff : (14) where noise represents the actual noise data, and res represents the predicted noise data.

[0048] Then, the Mahalanobis distance of the residual can be calculated through the following formula: (15) where S represents the covariance matrix.

[0049] Finally, the loss function of the diffusion model can be defined as the square of the Mahalanobis distance, that is: (16) After obtaining the corresponding loss, adjust the parameters of the initial diffusion model according to the obtained loss, and train the initial diffusion model.

[0050] During the actual training process, the number of epochs for training the diffusion model can be 1000, the maximum number of noise addition steps can be 1000, the batch size can be 32, and the Adam optimizer can be used. The parameters of the trained diffusion model can be saved in the.pt format, which is not limited in this disclosure. If it is applied in the field of soil detection, the dataset used for training needs to be soil data without anomalies. For example, the sampling interval can be 1 hour, and the sampled data can be selected for a relatively long time, such as 720, that is, soil data for 1 month can be collected at 1-hour intervals as an input sample. It should be noted that the diffusion model requires a large amount of data for training, so the input sample can include various parameters. For example, in actual use, the input sample includes but is not limited to: soil parameters such as soil water content, temperature, water potential, conductivity, and pH5.

[0051] In step S103, using the first time period as a sliding window, obtain the first data of the soil parameter data to be detected under each sliding window and the second data of the first reconstructed data under each sliding window.

[0052] As an example, the above-mentioned first time period can be set as needed, such as it can be set to 48 hours, which is not limited in this disclosure. The interval of the start time of each sliding window can also be set as needed, such as it can be set to 6 hours, which is also not limited in this disclosure.

[0053] As an example, after obtaining the first reconstructed data, a sliding window with a duration of 48 hours can be selected, that is, corresponding to 2 days of actual data, and the interval of the start time of each sliding window can be selected as 6 hours; then, for each sliding window, obtain the data of the soil parameter data to be detected and the first reconstructed data under this sliding window, that is, the above-mentioned first data and second data.

[0054] In step S104, for each sliding window, obtain the difference information between the first data and the second data under the current sliding window.

[0055] As an example, after obtaining the first data and the second data of each sliding window, the difference information between the first data and the second data under each sliding window can be calculated. Here, the mean square error (MSE) can be used to calculate the difference information, as shown in the following formula: (18) where n represents the length of the sliding window, i represents the i th hour of the sliding window, yRepresents the first data of the soil parameter data to be detected under this sliding window, which is the second data of the first reconstructed data under this sliding window.

[0056] In step S105, in response to the difference information of any sliding window satisfying a predetermined condition, it is determined that the first data of the soil parameter data to be detected under any sliding window is abnormal data. As an example, the above-mentioned predetermined condition may be that the difference information of any sliding window is greater than a preset threshold, such as a preset value of 0.05, and the present disclosure is not limited thereto.

[0057] As an example, in actual use, when the MSE value (i.e., the difference information) calculated from the soil parameter data to be detected and the first reconstructed data in a sliding window is greater than 0.05, some data of the soil parameter data to be detected in this sliding window are determined to be abnormal data, and further optimization and processing are required in the follow-up. For example, this part of abnormal data can be selected to be removed from the soil parameter data to be detected and the diffusion model can be used to complete the missing values.

[0058] According to an embodiment of the present disclosure, all abnormal data can be removed from the soil parameter data to be detected to obtain the third data; the diffusion model is used to complete the data of the third data. Through this embodiment, removing the abnormal data of the soil parameter data to be detected and using the diffusion model to complete the data of the third data can effectively recover the missing data in the third data.

[0059] As an example, the missing of the third data may be caused by removing all abnormal data from the soil parameter data to be detected after detecting abnormal data, or may be caused by abnormalities in the actual acquisition process of the soil sensor. The present disclosure is not limited thereto.

[0060] According to an embodiment of the present disclosure, using the diffusion model to complete the data of the third data may include: adding random data to the data missing positions in the third data to obtain the fourth data; performing noise addition for a second predetermined number of steps on the fourth data to obtain the second noise-added data; inputting the second noise-added data into the denoising module of the diffusion model to obtain the second reconstructed data corresponding to the fourth data, where the second reconstructed data is the denoised data obtained by denoising the second noise-added data for the second predetermined number of steps; covering the data at the first predetermined position in the third data to the corresponding position in the second reconstructed data to obtain the completed third data, where the first predetermined position is the position corresponding to the non-missing position of the data in the third data. Through this embodiment, by gradually adding noise and gradually removing noise, a relatively accurate second reconstructed data can be obtained, and then the missing data can be effectively recovered by using the second reconstructed data, thereby ensuring the integrity and accuracy of the high spatio-temporal resolution data.

[0061] As an example, random numbers can be generated from a Gaussian distribution and added to the data missing positions in the third data to obtain the fourth data. Then, the fourth data is gradually noise-added and gradually noise-removed to obtain the reconstructed data corresponding to the fourth data. The data at the first predetermined position in the reconstructed data is added to the corresponding position in the third data to obtain the third data after completion, that is, the data with the missing data part completed.

[0062] According to an embodiment of the present disclosure, the second predetermined number of steps can be determined in the following manner: in response to the data missing rate of the third data being less than or equal to the first threshold, determining the second predetermined number of steps as the first preset value; in response to the data missing rate of the third data being greater than the first threshold and less than the second threshold, determining the second predetermined number of steps as the second preset value; in response to the data missing rate of the third data being greater than or equal to the second threshold, determining the second predetermined number of steps as the third preset value; where the third preset value is greater than the second preset value, and the second preset value is greater than the first preset value. Through this embodiment, different numbers of noise-adding and noise-removing steps are set for different degrees of data missing situations, which can be more adapted to the actual situation and avoid being out of touch with the actual situation.

[0063] As an example, the above first threshold and second threshold can be set as needed. For example, the first threshold can be set to 5%, and the second threshold can be set to 20%. The present disclosure does not limit this.

[0064] For example, to cope with the actual situation, different numbers of noise-adding and noise-removing steps, that is, the above second predetermined number of steps, can be set for different degrees of data missing situations. For the situation where the data missing rate is less than 5%, the second predetermined number of steps is set to 200; for the situation where the data missing rate is greater than 5% and less than 20%, the second predetermined number of steps is set to 500; for the situation where the data missing rate is greater than 20%, the second predetermined number of steps is set to 1000.

[0065] According to an embodiment of the present disclosure, inputting the second noisy data into the denoising module of the diffusion model to obtain the second reconstructed data corresponding to the fourth data may include: denoising the second noisy data by the denoising module of the diffusion model for a second predetermined number of steps. For each step of denoising among the second predetermined number of steps, the following processing is performed: in response to obtaining the denoised data of the previous step of the current step, covering the data at the corresponding position in the denoised data with the data at the second predetermined position in the predetermined noisy data, where the predetermined noisy data is the noisy data obtained by adding noise to the fourth data for the current step, and the second predetermined position is the position corresponding to the non-missing data position in the third data; performing denoising processing on the intermediate data for the current step to obtain the denoised data of the current step; determining the denoised data of the last step as the second reconstructed data. Through this embodiment, for the position corresponding to the non-missing data position of the third data in the denoised data used in each step of the denoising process, the original data in the noisy data added to the current step is used, avoiding data distortion caused by the denoising process and ensuring the accuracy of the second reconstructed data.

[0066] As an example, a mask vector corresponding to the third data may be obtained first, where the positions corresponding to the missing data positions of the third data in the mask vector are set to 0, and the positions corresponding to the non-missing data positions of the third data in the mask vector are set to 1; for the denoised data obtained in each step of denoising, the mask vector can be used to cover all the non-missing data positions in the denoised data with the actual data in the noisy data added to the current step, that is, the data at the second predetermined position in the above-mentioned predetermined noisy data covers the data at the corresponding position in the denoised data. Specifically, it can be expressed as follows: (19) where mask represents the mask vector, is the noisy data obtained by adding noise to the fourth data up to the current step t and is the denoised data obtained by denoising the second noisy data up to the current step t of the second noisy data.

[0067] To facilitate understanding of the present disclosure, the following is a systematic description in conjunction with Figure 4 、 Figure 5 and Figure 6 for a systematic description.

[0068] Figure 4 shows the data optimization system based on the diffusion model of the present disclosure, as shown in Figure 4As shown, the data optimization system includes two major parts: model training and model inference. For the model training part, an input vector is first constructed, that is, training samples are constructed, and then the training samples are noise-added through forward noise addition. Then, the diffusion model is trained with the training samples to obtain the trained diffusion model. For the model inference part, the data to be optimized is first obtained, and then the data to be optimized is noise-added through forward noise addition. The noise-added data is input into the denoising module of the diffusion model for anomaly data detection or missing value filling, etc., which is not limited in this disclosure.

[0069] Figure 5 shows the process of anomaly data detection. As Figure 5 shown, Gaussian noise is added to the soil parameter data to be detected (i.e., the data to be detected in Figure 5 ) to obtain the noise-added data. Then, the noise-added data is denoised for t steps to obtain the reconstructed data corresponding to the soil parameter data to be detected. Specifically, after the t-th step of denoising the noise-added data, it is judged whether t is equal to 0. If t is equal to 0, it means the denoising process ends, and the denoised data at this time is the reconstructed data. If t is not equal to 0, the (t - 1)-th step of denoising is performed until t is equal to 0 to obtain the reconstructed data. After obtaining the reconstructed data, with a predetermined time length as the sliding window, the first data of the soil parameter data to be detected under each sliding window and the second data of the reconstructed data under each sliding window are obtained, and the difference information between the first data and the second data under each sliding window is calculated. If the difference information of a certain sliding window is greater than the threshold, the first data of the soil parameter data to be detected in this sliding window is determined as anomaly data, otherwise it is determined that there is no anomaly data.

[0070] Figure 6 shows the process of missing value filling. As Figure 6 shown, the data to be filled for the missing data is obtained, random numbers are added to the missing parts of the data to be filled, the filled data is noise-added to obtain the noise-added data, and the noise-added data is denoised to obtain the filled data. Specifically, in each step of denoising, the data of the non-missing part in the denoised data of the current step is replaced with the actual data in the noise-added data noise-added to the current step until the last step of denoising, and the data of the non-missing part in the denoised data of the last step is replaced with the actual data in the data to be filled.

[0071] Hereinafter, the present disclosure will be introduced by taking soil parameter data as an example.

[0072] First, based on the data of multiple soil parameters (such as moisture, conductivity, temperature, water potential, pH, etc.) collected by soil sensors, these data are used to construct an input vector, which is also the subsequent training sample or data to be optimized. For constructing the training sample, considering that the data of soil sensors exhibit certain temporal characteristics due to regular watering and diurnal temperature changes, therefore, the temporal data of C soil parameters can be selected and aligned by time as the training sample and applied to the training of the diffusion model. Moreover, considering the time dependence and environmental impact of soil data, a time series with a longer duration can be used as the training sample. By introducing training samples with a longer duration, the diffusion model can be made to have stronger generalization ability, so as to be applicable to the optimization processing of soil anomaly data for a longer time. Assume that the duration of the temporal data of a soil parameter in the input vector is L, and the dimension of the input vector is C. Secondly, in the forward diffusion process, Gaussian noise is gradually added to the input vector to simulate data degradation. Then, the diffusion model is trained with the noise-added input vector to denoise in the reverse direction to recover the original data from the noise-added input vector. Specifically, the diffusion model can be trained to learn to predict the noise added at the corresponding steps and minimize the difference between the predicted noise and the actual noise at each step to achieve the training of the diffusion model. After obtaining the trained diffusion model, anomaly data detection and missing value filling, etc. can be performed on the soil parameter data. The specific process has been introduced in detail above and will not be elaborated here.

[0073] In summary, in view of the advantages of the diffusion model in dealing with complex data distributions and generating high-quality samples, the present disclosure optimizes the temporal data collected by soil sensors with reference to the method of generating data by the diffusion model, such as anomaly data detection and missing value filling. Specifically, noise is gradually added to and gradually removed from the soil parameter data to be detected to obtain the reconstructed data of the soil parameter data to be detected. According to the difference information between the reconstructed data and the soil parameter data to be detected, the anomalies of the soil parameter data to be detected can be effectively detected. Thus, subsequently, on the basis of not affecting the data collection process, data recovery and other processing can be performed on the anomaly data, so as to ensure the integrity and accuracy of high spatio-temporal resolution data and provide more high-quality data for research related to soil sensors.

[0074] The present disclosure can optimize the soil sensor data at the software algorithm level without affecting the data collection process of the sensor, solve the data anomalies and data missing problems that occur in the soil sensor data collection, and improve the credibility of the soil sensor data. Moreover, it has the characteristics of simple operation and strong adaptability, provides high-quality data support for precision agriculture, soil environment monitoring and related research, and helps the intelligent development of agricultural production. It should be noted that a network service platform for sensor data optimization can also be built based on the soil sensor data optimization model, which simplifies the usage process of sensor data optimization.

[0075] Figure 7 is a flow chart showing a data exception processing device according to an embodiment of the present disclosure, such as Figure 7 As shown, the device includes a noise adding unit 70 , a noise removing unit 71 , a first acquiring unit 72 , a second acquiring unit 73 and a detecting unit 74 .

[0076] The denoising unit 70 is configured to perform denoising for a first predetermined number of steps on the soil parameter data to be detected to obtain first denoised data; the denoising unit 71 is configured to input the first denoised data into a denoising module in a diffusion model to obtain first reconstructed data corresponding to the soil parameter data to be detected, wherein the first reconstructed data is denoised data obtained by denoising the first denoised data for a first predetermined number of steps; the first acquisition unit 72 is configured to obtain the first data of the soil parameter data to be detected in each sliding window and the second data of the first reconstructed data in each sliding window with a first time length as a sliding window; the second acquisition unit 73 is configured to obtain, for each sliding window, the difference information between the first data and the second data in the current sliding window; the detection unit 74 is configured to determine that the first data of the soil parameter data to be detected in any sliding window is abnormal data in response to the difference information of any sliding window satisfying a predetermined condition.

[0077] According to an embodiment of the present disclosure, the data anomaly processing device further includes: a data completion unit configured to remove all abnormal data from the soil parameter data to be detected to obtain third data; and complete the third data using a diffusion model.

[0078] According to an embodiment of the present disclosure, the data completion unit is further configured to add random data to a data missing position in the third data to obtain fourth data; perform noise addition on the fourth data for a second predetermined number of steps to obtain second noisy data; input the second noisy data into a denoising module in a diffusion model to obtain second reconstructed data corresponding to the fourth data, wherein the second reconstructed data is denoised data obtained by denoising the second noisy data for a second predetermined number of steps; overwrite the data at a first predetermined position in the third data to a corresponding position in the second reconstructed data to obtain the completed third data, wherein the first predetermined position is a position corresponding to a position in the third data where data is not missing.

[0079] According to an embodiment of the present disclosure, the second predetermined number of steps is determined as follows: in response to the data missing rate of the third data being less than or equal to the first threshold, the second predetermined number of steps is determined to be the first preset value; in response to the data missing rate of the third data being greater than the first threshold and less than the second threshold, the second predetermined number of steps is determined to be the second preset value; in response to the data missing rate of the third data being greater than or equal to the second threshold, the second predetermined number of steps is determined to be the third preset value; wherein the third preset value is greater than the second preset value, and the second preset value is greater than the first preset value.

[0080] According to an embodiment of the present disclosure, the data completion unit is further configured to use the denoising module in the diffusion model to perform denoising on the second noisy data for a second predetermined number of steps. For each step of denoising in the second predetermined number of steps, the following processing is performed: in response to obtaining the denoised data of the previous step of the current step, the data at the second predetermined position in the predetermined noisy data is used to overwrite the corresponding position data in the denoised data to obtain intermediate data, where the predetermined noisy data is the noisy data obtained by adding noise to the fourth data for the current step, and the second predetermined position is the position corresponding to the non-missing position data in the third data; perform denoising processing on the intermediate data for the current step to obtain the denoised data of the current step; and determine the denoised data of the last step as the second reconstructed data.

[0081] According to an embodiment of the present disclosure, the diffusion model is trained in the following manner: a training sample set is constructed, where each training sample includes noisy data, the number of processing steps, and the actual noise data added at each step in the number of processing steps for the noisy data; for each training sample in the training sample set, the noisy data and the number of processing steps in the training sample are input into the denoising module of the initial diffusion model to obtain the estimated noise data added at each step in the number of processing steps for the noisy data; for each step in the number of processing steps, the parameters of the initial diffusion model are adjusted according to the loss between the estimated noise data of the current step and the actual noise data, and the initial diffusion model is trained. According to an embodiment of the present disclosure, there is provided a computer-readable storage medium storing instructions, where when the instructions are run by at least one computing device, at least one computing device is caused to execute the data anomaly detection method according to any one of the above embodiments.

[0082] According to an embodiment of the present disclosure, there is provided a system including at least one computing device and at least one storage device storing instructions, where when the instructions are run by at least one computing device, at least one computing device is caused to execute the data anomaly detection method according to any one of the above embodiments.

[0083] According to an embodiment of the present disclosure, there is provided a computer program product including computer instructions, where when the computer instructions are executed by a processor, any one of the above data anomaly detection methods is implemented.

[0084] Although some embodiments of the present disclosure have been shown and described, those skilled in the art should understand that these embodiments can be modified without departing from the principles and spirit of the present disclosure defined by the claims and their equivalents.

Claims

1. A method for detecting anomalies in soil parameter data based on a diffusion model, characterized in that, Including: Adding noise to the soil parameter data to be detected for a first predetermined number of steps to obtain first noisy data; Inputting the first noisy data into the denoising module of the diffusion model to obtain first reconstructed data corresponding to the soil parameter data to be detected, where the first reconstructed data is the denoised data obtained by denoising the first noisy data for the first predetermined number of steps; Taking the first duration as a sliding window to obtain first data of the soil parameter data to be detected under each sliding window and second data of the first reconstructed data under each sliding window; For each sliding window, obtaining the difference information between the first data and the second data under the current sliding window; In response to the difference information of any sliding window satisfying a predetermined condition, determining the first data of the soil parameter data to be detected under the any sliding window as abnormal data.

2. The data anomaly detection method according to claim 1, wherein Also including: Removing all abnormal data from the soil parameter data to be detected to obtain third data; Using the diffusion model to complete the data of the third data.

3. The data anomaly detection method according to claim 2, wherein The using the diffusion model to complete the data of the third data includes: Adding random data to the data missing positions in the third data to obtain fourth data; Adding noise to the fourth data for a second predetermined number of steps to obtain second noisy data; Inputting the second noisy data into the denoising module of the diffusion model to obtain second reconstructed data corresponding to the fourth data, where the second reconstructed data is the denoised data obtained by denoising the second noisy data for the second predetermined number of steps; Covering the data at a first predetermined position in the third data to the corresponding position in the second reconstructed data to obtain the third data after completion, where the first predetermined position is the position corresponding to the non-missing position of the data in the third data.

4. The data anomaly detection method according to claim 3, characterized in that The second predetermined number of steps is determined by the following method: In response to the data missing rate of the third data being less than or equal to a first threshold, determining the second predetermined number of steps as a first preset value; In response to the data missing rate of the third data being greater than the first threshold and less than a second threshold, determining the second predetermined number of steps as a second preset value; In response to the data missing rate of the third data being greater than or equal to the second threshold, determining the second predetermined number of steps as a third preset value; Wherein, the third preset value is greater than the second preset value, and the second preset value is greater than the first preset value.

5. The data anomaly detection method according to claim 3, wherein The inputting the second noisy data into the denoising module of the diffusion model to obtain second reconstructed data corresponding to the fourth data includes: Using the denoising module in the diffusion model to denoise the second noisy data for the second predetermined number of steps, where for each step of denoising in the second predetermined number of steps of denoising, the following processing is performed: In response to obtaining the denoised data of the previous step of the current step, the data at the second predetermined position in the predetermined noise-added data is used to overwrite the data at the corresponding position in the denoised data to obtain intermediate data, where the predetermined noise-added data is the noise-added data obtained by adding noise in the current step to the fourth data, and the second predetermined position is the position corresponding to the non-missing data position in the third data; Perform denoising processing on the intermediate data in the current step to obtain the denoised data of the current step; Determine the denoised data of the last step as the second reconstructed data.

6. The data anomaly detection method according to claim 1, wherein The diffusion model is trained in the following manner: Construct a training sample set, where each training sample includes noise-added data, the number of processing steps, and the actual noise data added in each step of the noise-added data during the processing steps; For each training sample in the training sample set, input the noise-added data and the number of processing steps in the training sample into the denoising module of the initial diffusion model to obtain the estimated noise data added in each step of the noise-added data during the processing steps; For each step in the processing steps, adjust the parameters of the initial diffusion model according to the loss between the estimated noise data of the current step and the actual noise data, and train the initial diffusion model.

7. A data anomaly handling device, characterized in that, Including: A noise-adding unit configured to add noise to the soil parameter data to be detected for a first predetermined number of steps to obtain first noise-added data; A denoising unit configured to input the first noise-added data into the denoising module of the diffusion model to obtain a first reconstructed data corresponding to the soil parameter data to be detected, where the first reconstructed data is the denoised data obtained by denoising the first noise-added data for the first predetermined number of steps; A first obtaining unit configured to use the first time period as a sliding window to obtain the first data of the soil parameter data to be detected under each sliding window and the second data of the first reconstructed data under each sliding window; A second obtaining unit configured to obtain the difference information between the first data and the second data under the current sliding window for each sliding window; A detecting unit configured to, in response to the difference information of any sliding window satisfying a predetermined condition, determine the first data of the soil parameter data to be detected under the any sliding window as abnormal data.

8. A computer-readable storage medium for storing instructions, characterized in that, When the instruction is run by at least one computing device, it causes the at least one computing device to execute the data anomaly detection method according to any one of claims 1 to 6.

9. A system comprising at least one computing device and at least one storage device storing instructions, characterized in that, When the instruction is run by the at least one computing device, it causes the at least one computing device to execute the data anomaly detection method according to any one of claims 1 to 6.

10. A computer program product comprising computer instructions, characterized in that, When the computer instruction is executed by a processor, it implements the data anomaly detection method according to any one of claims 1 to 6.