A method and system for analyzing pesticide residue data based on big data
Through adaptive block processing and similarity calculation, combined with European distance and data sequence correlation, the problem of the impact of concentration differences in pesticide residue data analysis is solved, and more accurate abnormal data recognition is achieved.
Patent Information
- Application Number
- CN202510421684.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-07
AI Technical Summary
The traditional LOF algorithm failed to effectively consider the spray concentration differences in the analysis of pesticide residue data, resulting in inaccurate analysis results.
Adaptive blocking processing method is adopted, combining the European-style distance and pesticide residue data sequence correlation, and abnormal data is identified through similarity calculation and clustering analysis.
It improves the accuracy and reliability of pesticide residue data analysis, reduces false alarms and missed reports, and can better capture local abnormalities.
Smart Images

Figure CN119939479B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and in particular, to a method and system for analyzing pesticide residue data based on big data. Background Art
[0002] With the rapid development of agricultural modernization and the food industry, the application of pesticides in crop production has become increasingly common. However, the problem of pesticide residues not only affects the quality and yield of crops, but also poses potential hazards to human health and the environment. Therefore, the analysis of pesticide residue data plays a crucial role in the quality control of agricultural products. By monitoring and analyzing pesticide residue data, potential pollution sources can be detected in a timely manner, and the risk of pesticide abuse can be reduced, thus ensuring food safety.
[0003] During the process of pesticide residue monitoring, the data is vulnerable to various factors, and among them, climatic conditions may cause abnormal changes in pesticide residues. Therefore, it is necessary to monitor in a timely manner and remind the staff to pay attention to these changes. The traditional LOF (Local Outlier Factor) algorithm can analyze pesticide residue data by calculating the Euclidean distance between all data points and obtaining the local density of each data point.
[0004] However, during the pesticide spraying process, as the sprayer system continues to operate, the spraying area continuously increases, and the load of the sprayer will also change accordingly, which will lead to unstable operating pressure. This instability will cause differences in the pesticide spraying concentration in different areas. The traditional LOF algorithm only considers the Euclidean distance between data points and ignores the influence of the difference in pesticide spraying concentration in different areas. Therefore, if the traditional LOF algorithm is directly used to analyze pesticide residue data, the analysis result may be inaccurate. Summary of the Invention
[0005] In order to solve the technical problem of large errors in the analysis results of the traditional LOF algorithm for pesticide residue data, this application provides a method and system for analyzing pesticide residue data based on big data.
[0006] In the first aspect, this application provides a method for analyzing pesticide residue data based on big data, adopting the following technical solution:
[0007] A method for analyzing pesticide residue data based on big data, comprising the steps of: performing block processing on all monitoring points arranged at crop locations; using the Local Outlier Factor algorithm to monitor the pesticide residue data of the monitoring points within each block, and identifying the data with an abnormal degree greater than a preset threshold as abnormal data; the method of block division is: calculating the Euclidean distance between any two monitoring points and the correlation of the pesticide residue data sequences between the monitoring points, and calculating the similarity between any two monitoring points according to the correlation and the Euclidean distance; for any one monitoring point, calculating the standard deviation of the similarity between this monitoring point and other monitoring points, normalizing the standard deviation to obtain a normalization result, and taking the monitoring points with a normalization result greater than the preset threshold as seed points; using the seed points as clustering centers, clustering the monitoring points to obtain multiple clustering clusters, and taking the monitoring points in each clustering cluster as the monitoring points in the same block.
[0008] The beneficial effects are as follows: adaptively dividing the monitoring points into blocks, comprehensively considering the spatial positions and data characteristics of the monitoring points, dividing the monitoring points with similar pesticide residue characteristics into the same block, and providing a more targeted data subset for the subsequent detection by the Local Outlier Factor algorithm. Compared with the block division method based solely on spatial distance or data characteristics, it can more effectively capture the local anomalies in the pesticide residue data, improve the accuracy and reliability of abnormal data identification, and reduce false alarms and missed reports.
[0009] Optionally, the calculation formula for the similarity between any two monitoring points is:
[0010] ; in the formula, represents the similarity between the th monitoring point and the th monitoring point; represents the correlation of the pesticide residue sequences between the th monitoring point and the th monitoring point; represents the Euclidean distance between the th monitoring point and the th monitoring point; is a hyperparameter, represents the standard normalization function.
[0011] The beneficial effects are as follows: combining the Euclidean distance and the correlation of data sequences to quantify the similarity between monitoring points, considering both the physical distance between monitoring points and the similarity of data characteristics, being able to more comprehensively reflect the actual relationship between monitoring points, simultaneously considering the spatial position of monitoring points and the similarity of pesticide residue data sequences, and the hyperparameter is used to adjust the weights of the Euclidean distance and the correlation, with higher flexibility and better adaptability to different practical applications.
[0012] Optionally, the calculation formula for the similarity between any two monitoring points is:
[0013] ; where, represents the similarity between the th monitoring point and the th monitoring point; represents the correlation of the pesticide residue sequences between the th monitoring point and the th monitoring point; represents the Euclidean distance between the th monitoring point and the th monitoring point; represents the standard normalization function.
[0014] The beneficial effects are as follows: It takes into account both the physical distance between the monitoring points and the similarity of data features, can more comprehensively reflect the actual relationship between the monitoring points, simultaneously considers the spatial position of the monitoring points and the similarity of the pesticide residue data sequences, and the calculation process is simple and efficient.
[0015] Optionally, the calculation formula for the similarity between any two monitoring points is:
[0016] ; where, represents the similarity between the th monitoring point and the th monitoring point; represents the correlation of the pesticide residue sequences between the th monitoring point and the th monitoring point; represents the Euclidean distance between the th monitoring point and the th monitoring point; represents the standard normalization function; represents the initial concentration of the th monitoring point, represents the initial concentration of the th monitoring point.
[0017] The beneficial effects are as follows: An exponential decay term of the initial concentration difference is introduced, further refining the similarity calculation. By considering the difference in the initial concentration, it can more accurately reflect the difference in the pesticide residue characteristics between the monitoring points, avoiding misjudgment caused by a large difference in the initial concentration. It better adapts to the complex relationships between different monitoring points, improving the accuracy of block division and the sensitivity of abnormal data detection.
[0018] Optionally, the calculation method for the correlation of pesticide residue data sequences between monitoring points includes: for any one monitoring point, arranging the pesticide degradation rates obtained at this monitoring point in chronological order to obtain the pesticide residue concentration sequence of this monitoring point; for the pesticide residue sequences of any two monitoring points, calculating the Pearson correlation coefficient of the pesticide residue sequences of these two monitoring points; performing a normalization operation on the Pearson correlation coefficient and using it as the correlation.
[0019] The beneficial effects are as follows: By calculating the Pearson correlation coefficient of the pesticide residue concentration sequences of the monitoring points and normalizing it as the correlation, the linear correlation degree of the pesticide residue data sequences between the monitoring points can be quantified. This method is simple and easy to implement, can quickly evaluate the similarity between the monitoring points, and provides effective data features for subsequent chunking and anomaly detection. The normalization process of the Pearson correlation coefficient makes the correlation values have a unified range, which is convenient to combine with other factors (such as Euclidean distance) for comprehensive similarity calculation, thereby improving the accuracy of chunking and the reliability of anomaly data detection, helping to timely discover abnormal change trends of pesticide residues, and providing strong support for the risk warning and prevention and control of pesticide residues.
[0020] Optionally, the calculation method for the correlation of pesticide residue data sequences between monitoring points includes: for any one monitoring point, arranging the pesticide degradation rates obtained at this monitoring point in chronological order to obtain the pesticide residue concentration sequence of this monitoring point; for the pesticide residue sequences of any two monitoring points, calculating the DTW value of the pesticide residue sequences of these two monitoring points; performing a normalization operation on the DTW value and using it as the correlation.
[0021] The beneficial effects are as follows: Compared with the Pearson correlation coefficient, DTW can better capture the similarity in time of the pesticide residue concentration sequences between the monitoring points, even if there is a certain time dislocation or different change rates in these sequences. This calculation method can more accurately reflect the actual similarity degree between the monitoring points, thus providing a more reliable basis for chunking processing.
[0022] Optionally, the normalization operation is standard normalization or maximum-minimum normalization.
[0023] Optionally, when clustering the monitoring points to obtain multiple clusters, the K-means clustering algorithm is used for clustering.
[0024] Optionally, the pesticide residue data is collected by a multispectral sensor.
[0025] In a second aspect, the present application provides a big data-based pesticide residue data analysis system, adopting the following technical solution:
[0026] A big data-based pesticide residue data analysis system includes: a processor and a memory. The memory stores computer program instructions. When the computer program instructions are executed by the processor, the above-mentioned big data-based pesticide residue data analysis method is implemented.
[0027] The beneficial effects are as follows: Generate a computer program for the above-mentioned big data-based pesticide residue data analysis method and store it in the memory to be loaded and executed by the processor. Thus, make a system according to the memory and the processor for convenient use.
[0028] The present application has the following technical effects: Adaptively divide the monitoring points, comprehensively consider the spatial positions and data characteristics of the monitoring points, and divide the monitoring points with similar pesticide residue characteristics into the same block, providing a more targeted data subset for the subsequent local outlier factor algorithm detection. Compared with the block division methods based solely on spatial distance or data characteristics, it can more effectively capture the local anomalies in the pesticide residue data, improve the accuracy and reliability of anomaly data recognition, and reduce false alarms and missed reports. Description of the Drawings
[0029] Figure 1 It is a method flow chart of a big data-based pesticide residue data analysis method according to an embodiment of the present application.
[0030] Figure 2 It is a method flow chart of block division in a big data-based pesticide residue data analysis method according to an embodiment of the present application. Detailed Embodiments
[0031] An embodiment of the present application discloses a big data-based pesticide residue data analysis method, which adaptively divides the monitoring points for collecting pesticide residue data and performs independent anomaly detection according to the divided data. Specifically, first, divide the monitoring points into several blocks with similar characteristics according to the geographical characteristics of the monitoring area and the pesticide spraying situation, and then apply the LOF algorithm for anomaly detection in each block respectively. This block processing method can effectively consider the regional differences, make the detection results more in line with the actual distribution characteristics, and thus improve the accuracy of anomaly detection. Refer to Figure 1 Based on this, the big data-based pesticide residue data analysis method includes steps S1 - step S2, which are specifically as follows:
[0032] S1: Perform block processing on all monitoring points arranged at the crops. Refer to Figure 2 Based on this, the method of block division includes steps S10 - step S12, which are specifically as follows:
[0033] S10: Calculate the Euclidean distance between any two monitoring points and the correlation of the pesticide residue data sequences between the monitoring points, and calculate the similarity between any two monitoring points according to the correlation and the Euclidean distance.
[0034] The method for collecting pesticide residue data is as follows: Deploy multi-spectral sensors to collect pesticide residue data of crops. Among them, a multi-spectral sensor is deployed every 2 meters in the area where the crops are located to collect the pesticide residues on the surface of the crop leaves. Place the spectral probe 1 cm on both sides of the midrib of the leaf and collect continuously for 3 times, and take the average value as the pesticide residue concentration data for this time. Repeat the above steps to complete the collection of pesticide residue concentration data for all monitoring points. The collection period can be 1 hour.
[0035] In order to ensure that the collected pesticide residue data is as accurate as possible, before using the multi-spectral sensor to collect the pesticide residue data of crops, it is necessary to calibrate the multi-spectral sensor using ultrapure water as a blank sample.
[0036] In one embodiment, the calculation method for the correlation of the pesticide residue data sequence between monitoring points is as follows: For any monitoring point, arrange the obtained pesticide degradation rate of this monitoring point in chronological order to obtain the pesticide residue concentration sequence of this monitoring point; for the pesticide residue sequences of any two monitoring points, calculate the Pearson correlation coefficient of the pesticide residue sequences of these two monitoring points; since the value range of the Pearson correlation coefficient is , for the convenience of subsequent calculation, the Pearson correlation coefficient is normalized and used as the correlation. Among them, the normalization operation is standard normalization or maximum-minimum normalization, and the prior art will not be elaborated here.
[0037] Among them, for any monitoring point, start monitoring from the time when pesticides are applied to the crops at this monitoring point, and obtain the pesticide residue concentration data of this monitoring point every hour. Starting from the third collection moment, the calculation formula for the pesticide degradation rate can be: ; in the formula, represents the pesticide degradation rate at the th collection moment; represents the pesticide residue concentration data at the moment; represents the pesticide residue concentration data at the moment; respectively represent the pesticide residue concentration data collected at the moment. Take the ratio of the concentration difference between two adjacent moments and its previous moment as the pesticide degradation rate at the current moment. The larger its value, the faster the pesticide degradation rate is continuously accelerating as the collection time continues. On the contrary, as the collection time continues, the pesticide degradation rate is continuously slowing down.
[0038] In other embodiments, the DTW (Dynamic Time Warping) value of the pesticide residue sequences at these two monitoring points can also be calculated; after normalizing the DTW value, it is used as the correlation. DTW is an algorithm for measuring the similarity of time series, which can effectively handle the situation where the lengths of time series are inconsistent or the time axes are not completely aligned. The prior art will not be elaborated here. Even if there are certain misalignments or different change rates in time for these sequences. This calculation method can more accurately reflect the actual similarity degree between the monitoring points, thus providing a more reliable basis for block processing.
[0039] S11: For any one monitoring point, calculate the standard deviation of the similarity between this monitoring point and other monitoring points, normalize the standard deviation to obtain a normalization result, and use the monitoring points with the normalization result greater than the preset threshold as seed points.
[0040] In one embodiment, the calculation formula for the similarity between any two monitoring points is:
[0041] ; where represents the similarity between the th monitoring point and the th monitoring point; represents the correlation of the pesticide residue sequences between the th monitoring point and the th monitoring point; represents the Euclidean distance between the th monitoring point and the th monitoring point; represents the standard normalization function.
[0042] For any two monitoring points, if the correlation of the pesticide residue sequences corresponding to these two monitoring points is higher and the Euclidean distance between these two monitoring points is closer, then the similarity between these two monitoring points is higher. Conversely, the similarity is lower.
[0043] In one embodiment, the calculation formula for the similarity between any two monitoring points can be:
[0044] ; where represents the similarity between the th monitoring point and the th monitoring point; represents the correlation of the pesticide residue sequences between the th monitoring point and the th monitoring point; represents the Euclidean distance between the th monitoring point and the th monitoring point; Represents the standard normalization function.
[0045] Is a hyperparameter used to adjust the weights of the Euclidean distance and correlation. For example ; During the process of pesticide spraying, the closer the Euclidean distance between monitoring points, the closer the pesticide spraying concentrations at these two monitoring points and the The closer the values, the closer the pesticide residue data at these two monitoring points. Therefore, in this application, the weight corresponding to the Euclidean distance is higher than the weight corresponding to similarity.
[0046] In one embodiment, the calculation formula for the similarity between any two monitoring points is:
[0047] ; In the formula, Represents the similarity between the th monitoring point and the th monitoring point; Represents the correlation between the pesticide residue sequences of the th monitoring point and the th monitoring point; Represents the Euclidean distance between the th monitoring point and the th monitoring point; Represents the standard normalization function.
[0048] Represents the initial concentration of the th monitoring point, Represents the initial concentration of the th monitoring point. By introducing the exponential decay term of the initial concentration difference, the greater the initial concentration difference, The smaller the value of, indicating that the greater the impact of the initial concentration difference on similarity, and vice versa, the smaller the impact.
[0049] For any monitoring point, if the normalized result of the standard deviation of the similarity between this monitoring point and other monitoring points has a smaller value, then the similarity between this monitoring point and other monitoring points is either relatively high or relatively low; conversely, the larger its value, the more inconsistent the similarity between this monitoring point and other monitoring points, that is, there are monitoring points with relatively high similarity to this monitoring point and monitoring points with relatively low similarity to this monitoring point.
[0050] If we want to assign similar monitoring points to the same block, we need to first select seed points. If we use the monitoring points with smaller normalized results as seed points, then all monitoring points will be in the same block. Therefore, we need to use the monitoring points with different similarities to different blocks as seed points.
[0051] S12: Using the seed points as the clustering centers, cluster the monitoring points to obtain multiple clustering clusters, and use the monitoring points in each clustering cluster as the monitoring points in the same block.
[0052] In one embodiment, a preset empirical threshold is 0.7. Use the K-means clustering algorithm to cluster all the monitoring points to obtain the clustering result. Among them, the clustering parameter K is the number of seed points, and the initial clustering centers are the seed points. Each clustering cluster is used as a block.
[0053] S2: Use the Local Outlier Factor (LOF) algorithm to monitor the pesticide residue data of the monitoring points in each block, and identify the data with an abnormal degree greater than the preset threshold as abnormal data.
[0054] Use the LOF algorithm to monitor the pesticide residue data of the monitoring points (collected at the same moment) in each block respectively. The data with an abnormal degree greater than the preset threshold (exemplarily, it can be 1) is used as abnormal data, and the abnormal data and the normal data are stored separately for the convenience of staff analysis. The LOF algorithm is a prior art and will not be elaborated here.
[0055] The embodiment of the present application also discloses a big data-based pesticide residue data analysis system, including a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, the big data-based pesticide residue data analysis method according to the present application is implemented.
[0056] The above system also includes other components well known to those skilled in the art such as a communication bus and a communication interface. Their settings and functions are known in the art, so they will not be elaborated here.
[0057] In the present application, the aforementioned memory can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or component. For example, the computer-readable storage medium can be any suitable magnetic storage medium or magneto-optical storage medium, such as a resistive random access memory, a dynamic random access memory, a static random access memory, etc., or any other medium that can be used to store the required information and can be accessed by an application program, module, or both. Any such computer storage medium can be a part of the device or accessible or connectable to the device.
[0058] The above are all the preferred embodiments of the present application. The protection scope of the present application is not limited by this. Therefore, all equivalent changes made according to the structure, shape, and principle of the present application should be covered within the protection scope of the present application.
Claims
1. A method for analyzing pesticide residue data based on big data, characterized in that, Including the steps: Perform block processing on all monitoring points arranged at the crops; For any one monitoring point, arrange the pesticide degradation rates obtained at this monitoring point in chronological order to obtain the pesticide residue concentration sequence of this monitoring point; Use the Local Outlier Factor algorithm to monitor the pesticide residue data of the monitoring points within each block, and identify the data with an anomaly degree greater than the preset threshold as abnormal data; The method of block division is: calculate the Euclidean distance between any two monitoring points and the correlation of the pesticide residue data sequences between the monitoring points, and calculate the similarity between any two monitoring points according to the correlation and the Euclidean distance; For any one monitoring point, calculate the standard deviation of the similarity between this monitoring point and other monitoring points, normalize the standard deviation to obtain the normalization result, and use the monitoring points with the normalization result greater than the preset threshold as seed points; Taking the seed points as the clustering centers, cluster the monitoring points to obtain multiple clustering clusters, and use the monitoring points in each clustering cluster as the monitoring points in the same block.
2. The method for analyzing pesticide residue data based on big data according to claim 1, wherein The calculation formula for the similarity between any two monitoring points is: ; wherein, represents the similarity between the -th monitoring point and the -th monitoring point; represents the correlation of the pesticide residue sequences between the -th monitoring point and the -th monitoring point; represents the Euclidean distance between the -th monitoring point and the -th monitoring point; is a hyperparameter, represents the standard normalization function.
3. The method for analyzing pesticide residue data based on big data according to claim 1, wherein The calculation formula for the similarity between any two monitoring points is: ; In the formula, represents the similarity between the th monitoring point and the th monitoring point; represents the correlation of the pesticide residue sequences between the th monitoring point and the th monitoring point; represents the Euclidean distance between the th monitoring point and the th monitoring point; represents the standard normalization function.
4. The method for analyzing pesticide residue data based on big data according to claim 1, wherein The calculation formula for the similarity between any two monitoring points is: ; where, represents the similarity between the -th monitoring point and the -th monitoring point; represents the correlation between the pesticide residue sequences of the -th monitoring point and the -th monitoring point; represents the Euclidean distance between the -th monitoring point and the -th monitoring point; represents the standard normalization function; represents the initial concentration of the -th monitoring point, and represents the initial concentration of the 5. The method for analyzing pesticide residue data based on big data according to claim 1, wherein The calculation method for the correlation of the pesticide residue data sequences between the monitoring points includes: For the pesticide residue sequences of any two monitoring points, calculate the Pearson correlation coefficient of the pesticide residue sequences of these two monitoring points; after normalizing the Pearson correlation coefficient, use it as the correlation.
6. The method for analyzing pesticide residue data based on big data according to claim 1, wherein The calculation method for the correlation of the pesticide residue data sequences between the monitoring points includes: For any one monitoring point, arrange the pesticide degradation rates obtained at this monitoring point in chronological order to obtain the pesticide residue concentration sequence of this monitoring point; For the pesticide residue sequences of any two monitoring points, calculate the DTW value of the pesticide residue sequences of these two monitoring points; after normalizing the DTW value, use it as the correlation.
7. The method for analyzing pesticide residue data based on big data according to claim 5 or 6, characterized in that The normalization operation is standard normalization or min-max normalization.
8. The method for analyzing pesticide residue data based on big data according to claim 1, characterized in that Among the multiple clustering clusters obtained by clustering the monitoring points, use the K-means clustering algorithm for clustering.
9. The method for analyzing pesticide residue data based on big data according to claim 1, characterized in that, The pesticide residue data is collected by a multispectral sensor.
10. A pesticide residue data analysis system based on big data, characterized in that, Including: A processor and a memory, the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the big data-based pesticide residue data analysis method according to any one of claims 1-9 is implemented.
Citation Information
Patent Citations
Agricultural non-point source pollution monitoring system
CN117517609A
Root cause analysis method, device and equipment and computer readable storage medium
CN117950892A