Pesticide residue data analysis method and system based on big data

By adopting a blocking processing method based on big data in the analysis of pesticide residue data, combining the European distance and data sequence correlation to calculate the similarity, and applying the LOF algorithm in each block, the problem of inaccurate analysis results caused by the traditional LOF algorithm ignoring regional concentration differences is solved, and the accuracy and reliability of abnormal data recognition are improved.

CN119939479AActive Publication Date: 2025-05-06SHANDONG RUNDA TESTING TECH CO LTD

Patent Information

Application Number
CN202510421684.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-06
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

The traditional LOF algorithm ignores the differences in pesticide spraying concentrations in different regions during pesticide spraying, resulting in inaccurate results of pesticide residue data analysis.

Method used

The pesticide residue data analysis method based on big data is used to calculate the correlation between the European distance between monitoring points and the pesticide residue data sequence, calculate the similarity, and use the seed points as the clustering center for block processing, and finally use the LOF algorithm to perform abnormal detection in each block.

Benefits of technology

This method can adaptively block monitoring points, comprehensively consider spatial location and data characteristics, improve the accuracy and reliability of abnormal data identification, and reduce false alarms and missed reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939479A_ABST
    Figure CN119939479A_ABST
Patent Text Reader

Abstract

The invention relates to the field of data processing, in particular to a pesticide residue data analysis method and system based on big data, and the method comprises the steps: carrying out the partitioning processing of all monitoring points disposed at crops; and monitoring the pesticide residue data of the monitoring points in each block by using a local abnormal factor algorithm, and identifying the data of which the abnormal degree is greater than a preset threshold value as abnormal data. The method can effectively capture the local abnormal condition of the pesticide residue data, and has the effect of improving the accuracy and reliability of abnormal data recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and in particular to a pesticide residue data analysis method and system based on big data. Background Art

[0002] With the rapid development of agricultural modernization and food industry, the application of pesticides in crop production has become increasingly common. However, the problem of pesticide residues not only affects the quality and yield of crops, but also poses potential hazards to human health and the environment. Therefore, pesticide residue data analysis plays a vital role in agricultural product quality control. By monitoring and analyzing pesticide residue data, potential sources of pollution can be discovered in a timely manner, the risk of pesticide abuse can be reduced, and food safety can be guaranteed.

[0003] In the process of pesticide residue monitoring, data is easily affected by many factors, among which climate conditions may cause abnormal changes in pesticide residues. Therefore, it is necessary to monitor and remind staff to pay attention to these changes in a timely manner. The traditional LOF (Local Outlier Factor) algorithm calculates the Euclidean distance between all data points and obtains the local density of each data point, which can realize the analysis of pesticide residue data.

[0004] However, during the pesticide spraying process, as the sprayer system continues to operate, the spraying area continues to increase, and the load of the sprayer will also change, resulting in unstable operating pressure. This instability will cause differences in pesticide spraying concentrations in different areas. The traditional LOF algorithm only considers the Euclidean distance between data points, but ignores the impact of differences in pesticide spraying concentrations in different areas. Therefore, if the traditional LOF algorithm is directly used to analyze pesticide residue data, the analysis results may be inaccurate. Summary of the invention

[0005] In order to solve the technical problem that the traditional LOF algorithm has large errors in the analysis results of pesticide residue data, the present application provides a pesticide residue data analysis method and system based on big data.

[0006] In the first aspect, the present application provides a pesticide residue data analysis method based on big data, which adopts the following technical solution: A pesticide residue data analysis method based on big data comprises the following steps: performing block processing on all monitoring points arranged at crops; monitoring the pesticide residue data of the monitoring points in each block by using a local anomaly factor algorithm, and identifying the data with an abnormality greater than a preset threshold as abnormal data; the block division method is: calculating the Euclidean distance between any two monitoring points and the correlation of the pesticide residue data sequences between the monitoring points, and calculating the similarity between any two monitoring points according to the correlation and the Euclidean distance; for any monitoring point, calculating the standard deviation of the similarity between the monitoring point and other monitoring points, normalizing the standard deviation to obtain a normalized result, and taking the monitoring point with a normalized result greater than the preset threshold as a seed point; taking the seed point as the clustering center, clustering the monitoring points to obtain a plurality of cluster clusters, and the monitoring points in each cluster cluster are taken as the monitoring points in the same block.

[0007] The beneficial effects are: adaptively divide the monitoring points into blocks, comprehensively consider the spatial location and data characteristics of the monitoring points, divide the monitoring points with similar pesticide residue characteristics into the same block, and provide a more targeted data subset for the subsequent local anomaly factor algorithm detection. Compared with the block method based solely on spatial distance or data characteristics, it can more effectively capture the local anomalies of pesticide residue data, improve the accuracy and reliability of abnormal data identification, and reduce false positives and false negatives.

[0008] Optionally, the similarity calculation formula between any two monitoring points is: ; In the formula, Indicates Monitoring points and The similarity of the monitoring points; Indicates Monitoring points and Correlation of pesticide residue sequences at each monitoring point; Indicates Monitoring points and The Euclidean distance between monitoring points; is a hyperparameter, Represents the standard normalization function.

[0009] The beneficial effects are: combining Euclidean distance and data sequence correlation to quantify the similarity between monitoring points, taking into account both the physical distance between monitoring points and the similarity of data features, and being able to more comprehensively reflect the actual relationship between monitoring points, while considering the spatial position of the monitoring points and the similarity of the pesticide residue data sequence. Hyperparameters are used to adjust the weights of Euclidean distance and correlation, which is more flexible and more adaptable to different practical applications.

[0010] Optionally, the similarity calculation formula between any two monitoring points is: ; In the formula, Indicates Monitoring points and The similarity of the monitoring points; Indicates Monitoring points and Correlation of pesticide residue sequences at each monitoring point; Indicates Monitoring points and The Euclidean distance between monitoring points; Represents the standard normalization function.

[0011] The beneficial effects are: it takes into account both the physical distance between monitoring points and the similarity of data features, can more comprehensively reflect the actual relationship between monitoring points, and at the same time considers the spatial position of the monitoring points and the similarity of the pesticide residue data sequence, and the calculation process is simple and efficient.

[0012] Optionally, the similarity calculation formula between any two monitoring points is: ; In the formula, Indicates Monitoring points and The similarity of the monitoring points; Indicates Monitoring points and Correlation of pesticide residue sequences at each monitoring point; Indicates Monitoring points and The Euclidean distance between monitoring points; represents the standard normalization function; Indicates The initial concentration at each monitoring point, Indicates The initial concentration at each monitoring point.

[0013] The beneficial effects are: the exponential decay term of the initial concentration difference is introduced, the similarity calculation is further refined, and by considering the difference in initial concentration, the difference in pesticide residue characteristics between monitoring points can be more accurately reflected, avoiding misjudgment caused by large differences in initial concentrations. It can better adapt to the complex relationship between different monitoring points, improve the accuracy of block segmentation and the sensitivity of abnormal data detection.

[0014] Optionally, the method for calculating the correlation of pesticide residue data sequences between monitoring points includes: for any monitoring point, arranging the pesticide degradation rates obtained at the monitoring point in chronological order to obtain the pesticide residue concentration sequence of the monitoring point; for the pesticide residue sequences of any two monitoring points, calculating the Pearson correlation coefficient of the pesticide residue sequences of the two monitoring points; and normalizing the Pearson correlation coefficient as the correlation.

[0015] The beneficial effect is that by calculating the Pearson correlation coefficient of the pesticide residue concentration sequence of the monitoring point and normalizing it as the correlation, the linear correlation degree of the pesticide residue data sequence between the monitoring points can be quantified. This method is simple and easy to implement, and can quickly evaluate the similarity between the monitoring points, providing effective data features for subsequent segmentation and anomaly detection. The normalization of the Pearson correlation coefficient makes the correlation value have a uniform range, which is convenient for combining with other factors (such as Euclidean distance) for comprehensive similarity calculation, thereby improving the accuracy of segmentation and the reliability of abnormal data detection, helping to timely discover abnormal changes in pesticide residues and provide strong support for risk warning and prevention and control of pesticide residues.

[0016] Optionally, the method for calculating the correlation of pesticide residue data sequences between monitoring points includes: for any monitoring point, arranging the pesticide degradation rates obtained at the monitoring point in chronological order to obtain the pesticide residue concentration sequence of the monitoring point; for the pesticide residue sequences of any two monitoring points, calculating the DTW values ​​of the pesticide residue sequences of the two monitoring points; and normalizing the DTW values ​​as the correlation.

[0017] The beneficial effect is that compared with the Pearson correlation coefficient, DTW can better capture the temporal similarity of the pesticide residue concentration sequences between monitoring points, even if these sequences are misaligned or have different rates of change in time. This calculation method can more accurately reflect the actual similarity between monitoring points, thus providing a more reliable basis for block processing.

[0018] Optionally, the normalization operation is standard normalization or maximum and minimum value normalization.

[0019] Optionally, the monitoring points are clustered to obtain a plurality of clusters, and the clustering is performed using a K-means clustering algorithm.

[0020] Optionally, pesticide residue data is collected by a multispectral sensor.

[0021] In the second aspect, the present application provides a pesticide residue data analysis system based on big data, which adopts the following technical solutions: A pesticide residue data analysis system based on big data comprises: a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the pesticide residue data analysis method based on big data is implemented.

[0022] The beneficial effect is: the above-mentioned pesticide residue data analysis method based on big data is generated into a computer program and stored in a memory so as to be loaded and executed by a processor, thereby making a system based on the memory and the processor for easy use.

[0023] The present application has the following technical effects: adaptively divide the monitoring points into blocks, comprehensively consider the spatial location and data characteristics of the monitoring points, divide the monitoring points with similar pesticide residue characteristics into the same block, and provide a more targeted data subset for subsequent local anomaly factor algorithm detection. Compared with the block division method based solely on spatial distance or data characteristics, it can more effectively capture the local anomalies of pesticide residue data, improve the accuracy and reliability of abnormal data identification, and reduce false positives and false negatives. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 This is a method flow chart of a pesticide residue data analysis method based on big data in an embodiment of the present application.

[0025] Figure 2 It is a flowchart of the block method in a pesticide residue data analysis method based on big data in an embodiment of the present application. DETAILED DESCRIPTION

[0026] The embodiment of the present application discloses a pesticide residue data analysis method based on big data, which adaptively divides the monitoring points used to collect pesticide residue data into blocks, and performs independent anomaly detection based on the divided data. Specifically, the monitoring points are first divided into a number of blocks with similar characteristics based on the geographical characteristics of the monitoring area and the pesticide spraying conditions, and then the LOF algorithm is applied to each block for anomaly detection. This block processing method can effectively take into account regional differences, making the detection results more consistent with the actual distribution characteristics, thereby improving the accuracy of anomaly detection. Reference Figure 1 The pesticide residue data analysis method based on big data includes steps S1 and S2, which are as follows: S1: Divide all monitoring points placed on crops into blocks. Figure 2 The block division method includes steps S10 to S12, which are as follows: S10: Calculate the Euclidean distance between any two monitoring points and the correlation between the pesticide residue data sequences between the monitoring points, and calculate the similarity between any two monitoring points based on the correlation and the Euclidean distance.

[0027] The method for collecting pesticide residue data is to collect pesticide residue data of crops by deploying multispectral sensors. In the area where the crops are located, a multispectral sensor is deployed at intervals of 2 meters to collect pesticide residues on the surface of crop leaves. The spectrum probe is placed 1 cm on both sides of the midrib of the leaf, and the data is collected three times in a row, and the average value is taken as the pesticide residue concentration data of that time. Repeat the above steps to complete the collection of pesticide residue concentration data of all monitoring points. The collection cycle can be 1 hour.

[0028] In order to ensure that the collected pesticide residue data is as accurate as possible, before using a multispectral sensor to collect pesticide residue data of crops, ultrapure water is needed to be used as a blank sample to calibrate the multispectral sensor.

[0029] In one embodiment, the calculation method of the correlation of the pesticide residue data sequences between the monitoring points is as follows: for any monitoring point, the pesticide degradation rates obtained at the monitoring point are arranged in chronological order to obtain the pesticide residue concentration sequence of the monitoring point; for the pesticide residue sequences of any two monitoring points, the Pearson correlation coefficient of the pesticide residue sequences of the two monitoring points is calculated; the value range of the Pearson correlation coefficient is In order to facilitate subsequent calculations, the Pearson correlation coefficient is normalized and used as the correlation. The normalization operation is standard normalization or maximum and minimum value normalization, and the prior art will not be described in detail here.

[0030] Among them, for any monitoring point, monitoring starts after the crop at the monitoring point is applied with pesticides, and the pesticide residue concentration data of the monitoring point is obtained every hour. The calculation formula for the pesticide degradation rate starts from the third collection time. ; In the formula, Indicates Pesticide degradation rate at each collection moment; Indicates Pesticide residue concentration data at each moment; Indicates Pesticide residue concentration data at each moment; Respectively represent The pesticide residue concentration data collected at the moment. The ratio of the concentration difference between two adjacent moments and the previous moment is taken as the pesticide degradation rate at the current moment. The larger the value, the faster the pesticide degradation rate is as the collection time continues. On the contrary, the slower the pesticide degradation rate is as the collection time continues.

[0031] In other embodiments, the DTW (Dynamic Time Warping) values ​​of the pesticide residue sequences of the two monitoring points can also be calculated; the DTW values ​​are normalized and used as correlation. DTW is an algorithm for measuring the similarity of time series, which can effectively handle situations where the lengths of time series are inconsistent or the time axes are not completely aligned. The prior art will not be described here. Even if these sequences have a certain misalignment in time or different rates of change. This calculation method can more accurately reflect the actual similarity between the monitoring points, thereby providing a more reliable basis for block processing.

[0032] S11: For any monitoring point, calculate the standard deviation of the similarity between the monitoring point and other monitoring points, normalize the standard deviation to obtain a normalized result, and use the monitoring point whose normalized result is greater than a preset threshold as a seed point.

[0033] In one embodiment, the calculation formula for the similarity between any two monitoring points is: ; In the formula, Indicates Monitoring points and The similarity of the monitoring points; Indicates Monitoring points and Correlation of pesticide residue sequences at each monitoring point; Indicates Monitoring points and The Euclidean distance between monitoring points; Represents the standard normalization function.

[0034] For any two monitoring points, if the correlation between the pesticide residue sequences corresponding to the two monitoring points is higher and the Euclidean distance between the two monitoring points is closer, the similarity between the two monitoring points is higher. Otherwise, the similarity is lower.

[0035] In one embodiment, the calculation formula for the similarity between any two monitoring points can be: ; In the formula, Indicates Monitoring points and The similarity of the monitoring points; Indicates Monitoring points and Correlation of pesticide residue sequences at each monitoring point; Indicates Monitoring points and The Euclidean distance between monitoring points; Represents the standard normalization function.

[0036] is a hyperparameter used to adjust the weight of Euclidean distance and correlation, for example During the pesticide spraying process, the closer the Euclidean distance between the monitoring points is, the closer the pesticide spraying concentration and soil The closer the values ​​are, the closer the pesticide residue data of the two monitoring points are. Therefore, in this application, the weight corresponding to the Euclidean distance is higher than the weight corresponding to the similarity.

[0037] In one embodiment, the calculation formula for the similarity between any two monitoring points is: ; In the formula, Indicates Monitoring points and The similarity of the monitoring points; Indicates Monitoring points and Correlation of pesticide residue sequences at each monitoring point; Indicates Monitoring points and The Euclidean distance between monitoring points; Represents the standard normalization function.

[0038] Indicates The initial concentration at each monitoring point, Indicates The initial concentration of each monitoring point. The exponential decay term of the initial concentration difference is introduced. The larger the initial concentration difference, The smaller the value of , the greater the impact of the initial concentration difference on the similarity, and vice versa.

[0039] For any monitoring point, if the value of the normalized standard deviation of the similarity between the monitoring point and other monitoring points is smaller, the similarity between the monitoring point and other monitoring points is higher or the similarity between the monitoring point and other monitoring points is lower; conversely, the larger the value, the more inconsistent the similarity difference between the monitoring point and other monitoring points is, and there are both monitoring points with higher similarity to the monitoring point and monitoring points with lower similarity to the monitoring point.

[0040] If you want to assign similar monitoring points to the same block, you need to first select seed points. If you use the monitoring points with smaller normalized results as seed points, all monitoring points will be in the same block. Therefore, you need to use monitoring points with different similarities to different blocks as seed points.

[0041] S12: Taking the seed point as the cluster center, the monitoring points are clustered to obtain multiple clusters, and the monitoring points in each cluster are used as monitoring points in the same block.

[0042] In one embodiment, an empirical threshold of 0.7 is preset, and all monitoring points are clustered using the K-means clustering algorithm to obtain clustering results, wherein the clustering parameter K is the number of seed points, and the initial cluster center is the seed point. Each cluster is regarded as a block.

[0043] S2: Use the local anomaly factor algorithm to monitor the pesticide residue data of the monitoring points in each block, and identify the data with an abnormal degree greater than the preset threshold as abnormal data.

[0044] The LOF algorithm is used to monitor the pesticide residue data of the monitoring points (collected at the same time) in each block, and the data with an abnormality greater than a preset threshold (for example, 1) is regarded as abnormal data, and the abnormal data is stored separately from the normal data to facilitate analysis by the staff. The LOF algorithm is an existing technology and will not be described in detail here.

[0045] An embodiment of the present application also discloses a pesticide residue data analysis system based on big data, including a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the pesticide residue data analysis method based on big data according to the present application is implemented.

[0046] The above system also includes other components well known to those skilled in the art, such as a communication bus and a communication interface, and their configuration and functions are known in the art, so they will not be described in detail here.

[0047] In the present application, the aforementioned memory may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, a computer-readable storage medium may be any suitable magnetic storage medium or magneto-optical storage medium, such as a resistive random access memory, a dynamic random access memory, a static random access memory, etc., or any other medium that can be used to store the required information and can be accessed by an application, a module, or both. Any such computer storage medium may be part of a device or accessible or connectable to a device.

[0048] The above are all preferred embodiments of the present application, and the protection scope of the present application is not limited thereto. Therefore, any equivalent changes made according to the structure, shape, and principle of the present application should be included in the protection scope of the present application.

Claims

1. A pesticide residue data analysis method based on big data, characterized in that: Includes steps: All monitoring points arranged at crops are processed in blocks; The local anomaly factor algorithm is used to monitor the pesticide residue data of the monitoring points in each block, and the data with an abnormal degree greater than the preset threshold is identified as abnormal data; The block division method is: calculate the Euclidean distance between any two monitoring points and the correlation of the pesticide residue data series between the monitoring points, and calculate the similarity between any two monitoring points based on the correlation and the Euclidean distance; For any monitoring point, calculate the standard deviation of the similarity between the monitoring point and other monitoring points, normalize the standard deviation to obtain a normalized result, and use the monitoring point whose normalized result is greater than a preset threshold as a seed point; Taking the seed point as the cluster center, the monitoring points are clustered to obtain multiple clusters, and the monitoring points in each cluster are used as monitoring points in the same block.

2. The pesticide residue data analysis method based on big data according to claim 1, characterized in that: The calculation formula for the similarity between any two monitoring points is: ; In the formula, Indicates Monitoring points and The similarity of the monitoring points; Indicates Monitoring points and Correlation of pesticide residue sequences at each monitoring point; Indicates Monitoring points and The Euclidean distance between monitoring points; is a hyperparameter, Represents the standard normalization function.

3. The pesticide residue data analysis method based on big data according to claim 1, characterized in that: The calculation formula for the similarity between any two monitoring points is: ; In the formula, Indicates Monitoring points and The similarity of the monitoring points; Indicates Monitoring points and Correlation of pesticide residue sequences at each monitoring point; Indicates Monitoring points and The Euclidean distance between monitoring points; Represents the standard normalization function.

4. The pesticide residue data analysis method based on big data according to claim 1, characterized in that: The calculation formula for the similarity between any two monitoring points is: ; In the formula, Indicates Monitoring points and The similarity of the monitoring points; Indicates Monitoring points and Correlation of pesticide residue sequences at each monitoring point; Indicates Monitoring points and The Euclidean distance between monitoring points; represents the standard normalization function; Indicates The initial concentration at each monitoring point, Indicates The initial concentration at each monitoring point.

5. The pesticide residue data analysis method based on big data according to claim 1, characterized in that: The calculation method of the correlation of pesticide residue data series between monitoring points includes: For any monitoring point, the pesticide degradation rates obtained at the monitoring point are arranged in chronological order to obtain the pesticide residue concentration sequence of the monitoring point; For the pesticide residue sequences of any two monitoring points, the Pearson correlation coefficient of the pesticide residue sequences of the two monitoring points is calculated; the Pearson correlation coefficient is normalized and used as the correlation.

6. The pesticide residue data analysis method based on big data according to claim 1, characterized in that: The calculation method of the correlation of pesticide residue data series between monitoring points includes: For any monitoring point, the pesticide degradation rates obtained at the monitoring point are arranged in chronological order to obtain the pesticide residue concentration sequence of the monitoring point; For the pesticide residue sequences of any two monitoring points, the DTW values ​​of the pesticide residue sequences of the two monitoring points are calculated; the DTW values ​​are normalized and used as correlation.

7. The pesticide residue data analysis method based on big data according to claim 5 or 6, characterized in that: The normalization operation is standard normalization or maximum and minimum normalization.

8. The pesticide residue data analysis method based on big data according to claim 1, characterized in that: The monitoring points are clustered into multiple clusters, and the K-means clustering algorithm is used for clustering.

9. The pesticide residue data analysis method based on big data according to claim 1, characterized in that: Pesticide residue data are collected by multispectral sensors.

10. A pesticide residue data analysis system based on big data, characterized in that: include: A processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the pesticide residue data analysis method based on big data according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Method and system for visual analysis of high dimensional data of pesticide residues based on subspace clustering

    CN109344194A

  • Edge identification method, system and equipment for agricultural abnormal data and medium

    CN116668473A

  • Agricultural non-point source pollution monitoring system

    CN117517609A

  • Optimization method of pesticide residue monitoring process based on time sequence predictive analysis

    CN117787510A

  • Spectral data intelligent processing method for vegetable pesticide residue detection

    CN117809070A

Cited By

  • Pesticide residue anomaly detection method and system based on spatial-temporal characteristic clustering

    CN121682008A

  • A pesticide residue anomaly detection method and system based on spatiotemporal feature clustering

    CN121682008B