A filtering method and device for real-time measurement data of vehicle-mounted sensors
Through the combination of manual filtering and automatic filtering, the problem of removing outliers in vehicle-mounted sensor data is solved, data quality and prediction accuracy are improved, and high-quality data support is provided for precision agriculture.
Patent Information
- Application Number
- CN202510442556.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-04-10
AI Technical Summary
There are errors and redundant data in the process of data collection by vehicle sensors, and it is difficult for the existing technology to effectively filter outliers, affecting the data quality and the implementation of precision agriculture.
Combining manual filtering and automatic filtering methods, the outliers are initially cleaned up by manual filtering, and the clustering algorithm is used to further identify and remove abnormal data to ensure the robustness of the data filtering process.
It realizes phased and efficient cleaning of outliers, improves data accuracy and consistency, and improves the robustness and prediction accuracy of the data filtering process.
Smart Images

Figure CN119961633B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data filtering, and particularly to a method and device for filtering real-time measurement data of vehicle-mounted sensors. Background Art
[0002] Precision agriculture, as the basis of sustainable agriculture, is the main trend of modern agricultural development. Therefore, obtaining real-time farmland conditions and spatial distribution information of various parameters such as soil and crops is of great significance for guiding the implementation of precise farmland management, farmland soil fertility evaluation, etc. In recent years, with the application of position location technologies such as the Global Positioning System, vehicle-mounted sensing technologies have developed by leaps and bounds. These sensors are installed on agricultural machinery and can obtain soil, crop traits, and environmental parameters in real time during movement, providing data with higher spatial resolution and larger scale range than traditional sampling methods, and providing strong support for precision agriculture and soil difference analysis.
[0003] Vehicle-mounted sensing technologies can be combined with machine position information to measure agricultural indicators with high precision at close range and accurately relocate these measurement values to specific areas. Although the range of vehicle-mounted sensors is relatively small compared to spaceborne or airborne remote sensing systems, it can provide dynamic real-time information with improved accuracy within a sufficiently large range, such as soil conductivity, vegetation coverage, yield information, etc. Such sensing systems can collect a large amount of data on a large scale in a short time, greatly improving the data collection efficiency. Although rich data is important for on-site management decisions, the data collection process may also contain certain incorrect data that needs to be preprocessed such as filtered before further processing and analysis. Among them, some errors come from sensor performance, such as repeatability, accuracy, and resolution, etc.; more incorrect or redundant data comes from the influence of machine movement posture, movement speed, operator habits, and special conditions of specific fields, etc.
[0004] The sensor-based data set still belongs to a spatial data set. Different from remote sensing data, it is not obtained as a whole in pixel form, but is obtained sequentially as the machine moves, thus adding a time dimension, and the data position accuracy depends on the machine GPS accuracy. This requires relevant filtering algorithms to improve the data set quality. Summary of the Invention
[0005] The purpose of the present application is to provide a method and device for filtering real-time measurement data of vehicle-mounted sensors, which not only realizes the efficient cleaning of outliers in stages by combining manual filtering and automatic filtering, but also ensures the robustness of the data filtering process.
[0006] To achieve the above object, the present application provides the following solutions.
[0007] In a first aspect, the present application provides a method for filtering real-time measurement data of vehicle-mounted sensors, including the following steps.
[0008] Obtain the measurement data of the vehicle-mounted sensor within the target field plot.
[0009] Manually filter the outliers from the measurement data to obtain the data after manual filtering.
[0010] Apply a clustering method to cluster the data after manual filtering to obtain a clustering result.
[0011] Automatically filter the abnormal data from the data after manual filtering according to the clustering result to obtain the data after automatic filtering.
[0012] In a second aspect, the present application provides a filtering device for real-time measurement data of a vehicle-mounted sensor, including the following modules.
[0013] A data acquisition module for obtaining the measurement data of the vehicle-mounted sensor within the target field plot.
[0014] A manual filtering module for manually filtering the outliers from the measurement data to obtain the data after manual filtering.
[0015] A clustering module for applying a clustering method to cluster the data after manual filtering to obtain a clustering result.
[0016] An automatic filtering module for automatically filtering the abnormal data from the data after manual filtering according to the clustering result to obtain the data after automatic filtering.
[0017] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned filtering method for real-time measurement data of a vehicle-mounted sensor.
[0018] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the above-mentioned filtering method for real-time measurement data of a vehicle-mounted sensor.
[0019] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the above-mentioned filtering method for real-time measurement data of a vehicle-mounted sensor.
[0020] According to the specific embodiments provided by the present application, the following technical effects are disclosed in the present application.
[0021] The present application provides a method and device for filtering real-time measurement data of vehicle-mounted sensors, which includes obtaining measurement data collected by vehicle-mounted sensors within a target field plot; manually filtering out data outliers from the measurement data to obtain manually filtered data; applying a clustering method to the manually filtered data to obtain a clustering result; and automatically filtering out abnormal data from the manually filtered data according to the clustering result to obtain automatically filtered data. During the data filtering process, outliers in the data are first manually filtered to achieve preliminary filtering. On the basis of the preliminary filtering, a clustering algorithm is applied for automatic fine filtering. The manual filtering step uses a series of rules to perform necessary preliminary cleaning of the outliers, and at the same time avoids interference of these outliers with the clustering parameters of the automatic filtering, thereby ensuring the effectiveness of the automatic filtering. The combination of manual filtering and automatic filtering in the present application not only realizes the efficient cleaning of outliers in stages, but also ensures the robustness of the data filtering process. Description of the Drawings
[0022] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0023] Figure 1 It is an application environment diagram of a method for filtering real-time measurement data of vehicle-mounted sensors in an embodiment of the present application.
[0024] Figure 2 It is a schematic flowchart of a method for filtering real-time measurement data of vehicle-mounted sensors provided in an embodiment of the present application.
[0025] Figure 3 It is a schematic diagram of the original data collected by four deep soil apparent electrical conductivities provided in an embodiment of the present application.
[0026] Figure 4 It is a schematic diagram of the soil apparent electrical conductivity data at four depths after manual filtering provided in an embodiment of the present application.
[0027] Figure 5 It is a schematic diagram of the spatial neighborhood and temporal neighborhood of an example sample point provided in an embodiment of the present application.
[0028] Figure 6 It is a schematic diagram of the soil apparent electrical conductivity data at four depths after automatic filtering provided in an embodiment of the present application.
[0029] Figure 7 It is the data filtering effect at a depth of 0.54 m (PRP1) provided in an embodiment of the present application.
[0030] Figure 8 Data filtering effect at a depth of 1.03 m (PRP2) provided by an embodiment of the present application.
[0031] Figure 9 Data filtering effect at a depth of 1.55 m (HCP1) provided by an embodiment of the present application.
[0032] Figure 10 Data filtering effect at a depth of 3.18 m (HCP2) provided by an embodiment of the present application.
[0033] Figure 11 Schematic diagram of Kriging interpolation of the original data collected for soil apparent electrical conductivity at four depths provided by an embodiment of the present application.
[0034] Figure 12 Schematic diagram of Kriging interpolation of soil apparent electrical conductivity data at four depths after manual filtering provided by an embodiment of the present application.
[0035] Figure 13 Schematic diagram of Kriging interpolation of soil apparent electrical conductivity data at four depths after automatic filtering provided by an embodiment of the present application.
[0036] Figure 14 Distribution of sampling points for modeling set and validation set for predicting soil sand content provided by an embodiment of the present application.
[0037] Figure 15 Correlation distribution diagram between soil sand content and data after different data filtering steps provided by an embodiment of the present application.
[0038] Figure 16 Schematic diagram of functional modules of a filtering device for real-time measurement data of vehicle-mounted sensors provided by another embodiment of the present application.
[0039] Figure 17 Schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0040] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0041] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0042] The filtering method for real-time measurement data of in-vehicle sensors provided by the embodiments of the present application can be applied to an application environment as Figure 1 shown. Among them, the terminal communicates with the server through a network. The data storage system can store the data that the server needs to process. The data storage system can be set up separately, integrated on the server, or placed on the cloud or other servers. The terminal can send the measurement data collected by the in-vehicle sensors in the target field block to be processed to the server. After receiving the measurement data of the in-vehicle sensors in the target field block, the server manually filters the data outliers of the measurement data to obtain the manually filtered data; applies a clustering method to the manually filtered data to obtain a clustering result; automatically filters the abnormal data in the manually filtered data according to the clustering result to obtain the automatically filtered data. The server can feedback the obtained automatically filtered data to the terminal. In addition, in some embodiments, the filtering method for real-time measurement data of in-vehicle sensors can also be implemented by the server or the terminal alone. For example, the terminal can directly filter the measurement data collected by the in-vehicle sensors in the target field block to be processed, or the server can obtain the measurement data collected by the in-vehicle sensors in the target field block to be processed from the data storage system and perform data filtering.
[0043] Among them, the terminal can be, but is not limited to, various desktop computers, laptop computers, smartphones, tablet computers, Internet of Things devices, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.
[0044] In an exemplary embodiment, neither the manual filtering method based on experience alone nor the automatic filtering method based on the clustering algorithm can filter all outliers. For the manual filtering method based on experience, on the one hand, it requires a large amount of work; on the other hand, it cannot filter outliers in the case of local mutation outliers caused by the instrument being affected by the environment. For the automatic filtering method based on the clustering algorithm, it is easy to have difficulty in identifying outliers generated due to instrument posture and vehicle operation but may not be significantly intuitively reflected in the readings themselves, and at the same time, the parameters of the algorithm itself and the filtering effect are easily affected by these outliers. Therefore, it is necessary to combine the manual filtering method based on prior experience and error sources and the automatic filtering method based on the clustering algorithm to form a general abnormal data filtering framework. For this, as Figure 2 shown, a filtering method for real-time measurement data of in-vehicle sensors is provided. This method is executed by a computer device, and can be specifically executed by a computer device such as a terminal or a server alone, or jointly executed by a terminal and a server. In the embodiments of the present application, this method is applied to Figure 1Taking the server in [description] as an example, the following steps 101 to 104 are included.
[0045] Step 101: Obtain the measurement data at different depths collected by on-vehicle sensors within the target field block.
[0046] In this embodiment, the filtering step is described by taking the soil apparent conductivity data measured at different depths as an example. The measurement data of the on-vehicle sensors within the target field block may be data of different categories, measurement data at a certain depth of the same category, or measurement data at different depths of the same category. Among them, the measurement data includes but is not limited to soil apparent conductivity (ECa), temperature, humidity, crop yield, etc.
[0047] The presence of a drainage pipe passing through underground within the target field block will affect the measurement of soil ECa at the corresponding position, which is also one of the difficulties in data filtering. Based on the DUALEM-21S sensor, the ECa of the soil at multiple depths within the target field block is collected, including the ECa of the soil at four depths: 0.54 m (denoted as PRP1), 1.03 m (denoted as PRP2), 1.55 m (denoted as HCP1), and 3.18 m (denoted as HCP2), as Figure 3 shown, which are the original data of the soil apparent conductivity collected at four depths. In addition, the measurement data also includes the attitude data of the DUALEM-21S during movement, which is divided into pitch angle and roll angle. The global navigation satellite system (GNSS) receiver Trimble AgGPS 542 based on the real-time kinematic (RTK) technology collects the longitude and latitude of each sampling point during movement. A total of 9454 sampling point data are collected.
[0048] Step 102: Manually filter the data outliers in the measurement data at each depth to obtain the data after manual filtering.
[0049] Step 103: Apply a clustering method to the data after manual filtering at each depth for clustering to obtain a clustering result.
[0050] Step 104: For the data after manual filtering at each depth, automatically filter the abnormal data in the data after manual filtering according to the clustering result to obtain the data after automatic filtering.
[0051] Implementing the above-mentioned steps 101 to 104, the combination of manual filtering and automatic filtering not only realizes the phased and efficient cleaning of outliers, but also ensures the robustness of the data filtering process. The regularized operation of manual filtering accurately locates specific types of outliers and solves some problems that are difficult to directly handle by automatic filtering (such as local data anomalies). Automatic filtering identifies and filters the remaining outliers through a clustering algorithm, further improving the overall quality of the data on the basis of manual filtering, and forming an effective synergy. The present invention realizes a significant improvement in the quality of measurement data by combining manual filtering and automatic filtering strategies. Among them, manual filtering not only directly removes some outliers, but also provides good initial conditions for the parameter determination and clustering analysis of automatic filtering. The synergy between the two effectively improves the consistency and spatial correlation of the data, greatly improving the prediction accuracy after data filtering, and providing important support for data applications in fields such as precision agriculture and environmental monitoring.
[0052] In another exemplary embodiment of the present application, in step 102, manually filtering data outliers from the measurement data at each depth to obtain data after manual filtering specifically includes: for the measurement data at each depth, filtering outliers when the vehicle-mounted sensor starts measurement, data exceeding the threshold range defined by prior knowledge, data whose change exceeds the preset change degree, and data whose attitude change exceeds the preset attitude change degree during the measurement by the vehicle-mounted sensor.
[0053] As an example, filter startup (warm-up) uncertainty values. Specifically, filter data points that are less than 1 m or more than 6 m apart from the previous data point in space, and filter data points that are more than 1.5 s apart from the previous data point in time. A total of 37 data points are filtered.
[0054] As an example, filter soil ECa values outside the minimum and maximum absolute threshold values defined by the researcher based on prior knowledge. Specifically, filter data points with soil ECa values less than 0. A total of 23 data points are filtered.
[0055] As an example, remove data points with extremely large changes in soil ECa. Specifically, filter data points with soil ECa values that change by more than 20 mS / m between the previous and subsequent data points and between different depths of the same data point. A total of 24 data points are filtered.
[0056] As an example, remove data points with pitch changes outside the acceptable range. Specifically, filter points with pitch changes exceeding 10°. A total of 20 data points are filtered.
[0057] As an example, remove data points with roll changes outside the acceptable range. Specifically, filter data points with roll changes exceeding 10°. A total of 13 data points are filtered. Figure 4They are the soil apparent conductivity data at four depths after manual filtering.
[0058] In another exemplary embodiment of the present application, a clustering method is applied to the manually filtered data at each depth for clustering, and a clustering result is obtained, which specifically includes the following steps.
[0059] (1) For each data point in the manually filtered data at each depth, calculate the median in the temporal neighborhood and the median in the spatial neighborhood of each data point.
[0060] For any data point (referred to as an example sample point here), the data points before and after it are adjacent to it in time, and the data points on its left and right paths are adjacent to it in space. Take four points before and after each data point as its "temporal neighborhood", and take the points within a certain distance range on the left and right paths of each data point as its "spatial neighborhood", and calculate the median of the measured values of the data points included in the temporal neighborhood and the spatial neighborhood respectively. Figure 5 It is a schematic diagram of the spatial neighborhood and temporal neighborhood of the example sample point.
[0061] (2) Calculate the clustering coordinates of each data point according to the median in the temporal neighborhood and the median in the spatial neighborhood of each data point.
[0062] Calculate the difference (x) between the measured value of the example sample point and the median of the measured values of the data points in the temporal neighborhood and the difference (y) between the measured value of the example sample point and the median of the measured values of the data points in the spatial neighborhood, and use them as the clustering coordinates (x, y) of the example sample point. Since the distances of each data point from the left and right paths are not necessarily equal. Therefore, when defining the spatial neighborhood, it is divided into three ranges of 10m, 15m, and 20m respectively and the differences are calculated respectively, and finally the average value of the differences is taken as y.
[0063] (3) According to the clustering coordinates of each data point in the manually filtered data at each depth, calculate the Euclidean distance between the data points in the manually filtered data and estimate the kernel density of the data points, and determine the distance threshold parameter epsilon and the minimum number of points parameter minPoints in the Density-Based Spatial Clustering of Applications with Noise (DBSCAN).
[0064] (4) Based on the distance threshold parameter epsilon and the minimum number of points parameter minPoints, apply the DBSCAN algorithm for clustering to obtain the clustering result corresponding to the manually filtered data at each depth.
[0065] The DBSCAN algorithm was used for clustering, and the main categories were retained as the final filtered data at that depth (a total of 3556 data points were filtered). Figure 6 It is the soil apparent conductivity data at four depths after automatic filtering.
[0066] Unlike traditional distance-based methods, the DBSCAN algorithm identifies clusters based on the density of data points. It can identify areas of relatively high density and treat them as clusters, while also identifying noise points. The DBSCAN algorithm defines clusters using two main parameters: a distance threshold parameter, epsilon, and a minimum number of points, minPoints. The epsilon distance threshold parameter defines a neighborhood centered on the data point, while the minPoints minimum number of points within that neighborhood indicates the minimum number of data points required to be considered a core point. The DBSCAN algorithm works by randomly selecting a point in the dataset and checking whether the number of points within its epsilon neighborhood meets the criteria for being a core point. If it is a core point, the surrounding points are assigned to the same cluster. If it is a border point, it is assigned to the corresponding core point's cluster. This process is repeated until all points have been assigned to their respective clusters or marked as noise points.
[0067] In another exemplary embodiment of the present application, when each data point includes different types of measurement data of the target field, after executing the step of "automatically filtering abnormal data in the manually filtered data according to the clustering results to obtain automatically filtered data", the filtering method of real-time measurement data of the on-board sensor also includes: for each data point, determining whether each type of measurement data corresponding to the data point is normal data; if so, retaining each type of measurement data of the data point; if not, filtering all types of measurement data corresponding to the data point.
[0068] For example, if the same sensor can simultaneously measure soil ECa and magnetic susceptibility, meaning each data point contains two or more types of data from the same sensor, then if one type of data is problematic, that data point will be filtered. For example, if the soil ECa does not filter out a data point, but the magnetic susceptibility measurement at that point is abnormal, the sensor status is considered problematic, and the soil ECa at that sampling point is considered a potential outlier. This means that the intersection of the filtering results for different types of data is taken.
[0069] In another exemplary embodiment of the present application, after executing the step of "automatically filtering abnormal data in the manually filtered data at each depth according to the clustering results to obtain automatically filtered data", the filtering method of the real-time measurement data of the on-board sensor also includes: evaluating the filtering effect of the automatically filtered data.
[0070] fromFigure 3 - Figure 4 and Figure 6 、 Figure 11 - Figure 13 It can be seen that the filtering method of this application removes the influence of the drainage pipe. Further analysis is carried out from two aspects: geostatistics and soil sand content prediction performance.
[0071] 1. Analysis by geostatistics (detailed results are shown in Table 1, Figure 7 - Figure 10 . Figure 7 - Figure 10 The "model" in [[ ]] refers to the fitting model of the semivariogram function. "Averaged" means that when calculating the semivariogram, the data is grouped by distance intervals, and the average value of the semivariogram values in each interval is taken to reduce noise and more clearly display the spatial structure; the "1:1 line" represents the situation where the predicted value and the measured value are the same, that is, the prediction error is 0).
[0072] (1) Cross-validation: The root mean square error (RMSE) in cross-validation is used to further measure the effect of data filtering. The specific calculation method is shown in formula (1).
[0073]
[0074] where y obs is the observed value, y pred is the predicted value, and n is the number of samples.
[0075] The RMSE values in the original data are relatively high, indicating that there is a large error between the observed value and the predicted value during cross-validation of the original data. After manual filtering, the RMSE values generally decrease, indicating that manually removing outliers or noise improves the fitting degree. After further automatic filtering, the RMSE values further decrease, indicating that automatic filtering further reduces the prediction error and improves the prediction accuracy by effectively processing the inconsistencies and noise in the data.
[0076] (2) Nugget effect: The nugget effect reflects the short-range variability of the data, usually caused by noise, outliers, or local errors. There is a large short-range variability in the original data, and the nugget effect decreases significantly after manual filtering. This indicates that the outliers or noise in the data have been greatly reduced through manual filtering. After manual and automatic filtering, the nugget effect further decreases overall, especially for HCP2. This indicates that automatic filtering further removes the outliers in the data. However, the nugget effect of PRP1 slightly increases after automatic filtering, which may be due to the data in some regions with local fluctuations.
[0077] (3) Range: The range reflects the spatial autocorrelation of the data at a certain scale. Outliers and noise can introduce additional randomness in the estimation process, potentially masking or distorting the true spatial structure of the data. After filtering out these outliers, the spatial autocorrelation of the data itself can be more accurately reflected, and its impact on the performance of the range is not fixed. If the outliers mainly mask the spatial correlation at a larger scale, the estimated range may increase after filtering (PRP2, HCP1), that is, the spatial dependence decays at a greater distance. If the outliers mainly affect local details, and the data becomes smoother after removal, with more obvious local correlation, the range may decrease (PRP1, HCP2), because the true spatial structure shows a high degree of correlation at a shorter distance.
[0078] (4) Sill: The sill represents the overall variability of the data. After data filtering, if noise and outliers mainly increase the random fluctuations of the data, then after removing them, the data becomes smoother, and the estimated overall variability will decrease, resulting in a lower sill (PRP1, HCP2). Conversely, if the filtering process makes the regions with larger variability originally account for a higher proportion in the overall data, then the estimated variability may increase instead, leading to an increase in the sill (PRP2, HCP1).
[0079] Table 1 Parameter situations after semi-variance function prediction with different data filtering methods
[0080]
[0081] 2. Analyze through the correlation with soil sand content and the accuracy of predicting soil sand content
[0082] Surface soil samples were collected on the target field plot and the soil sand content was measured. Predictions were made using the original data, manually filtered data, and data after manual filtering + automatic filtering respectively. The coefficient of determination (R 2 ) and RMSE were used to evaluate the prediction accuracy and further measure the effect of data filtering. The specific calculation method of R 2 is shown in formula (2).
[0083]
[0084] where yobs is the observed value, ypred is the predicted value, and n is the number of samples.
[0085] Figure 14 The sample point distribution in Figure 15The correlation distribution map in can visually display the effect of data filtering. The detailed results are shown in Table 2. It can be seen that the correlation and accuracy are both improved after filtering. Overall, manual filtering improves the correlation between the soil ECa dataset and the soil sand content, while the manual filtering + automatic filtering method can further improve the correlation and even the prediction accuracy, obtaining a better prediction effect.
[0086] Table 2 Accuracy of predicting soil sand content after different data filtering methods
[0087]
[0088] Processing tool: The data filtering process is completed by programming in Python language; the geostatistics-related part is completed by using ArcGISpro software.
[0089] By defining a detection and filtering mechanism for data anomalies, the present invention aims to solve the problem of outliers caused by sensor attitude changes, environmental interference or equipment errors, and achieve data quality optimization. The method includes three main steps: data acquisition, manual filtering, and automatic filtering. First, target data and auxiliary information during the movement, such as attitude angles, geographical locations, and timestamps, are collected based on in-vehicle sensors. Secondly, through the manual filtering step, uncertain values at the start of the vehicle or instrument, outliers beyond the threshold range defined by prior knowledge, values with extremely large data changes, and points where the attitude changes of the instrument or vehicle exceed the acceptable range are filtered to complete the manual filtering operation. Subsequently, the median of the measured values is calculated based on the data in the neighborhood of each data point. Further, according to the deviation between the data point and the neighborhood median, a clustering coordinate is constructed, and the parameters of the clustering algorithm are determined through Euclidean distance and kernel density estimation. Cluster analysis is performed to identify and remove outliers, and the data of the main categories are retained as the filtered result. The present invention provides a general and modular data filtering method, which is applicable to the outlier processing of various in-vehicle sensor data, can effectively improve the accuracy and reliability of the measured data of in-vehicle sensors, and provide high-quality data support for precision agriculture and environmental monitoring. Compared with the prior art, it has the advantages of significantly improved data quality, precise filtering of outliers, significantly improved prediction accuracy, and strong applicability. These advantages come from the organic combination of the dual data filtering strategies of manual filtering and automatic filtering. In particular, manual filtering lays the foundation for the effectiveness of automatic filtering, which is specifically described as follows.
[0090] (1) Significantly improved data quality, with manual filtering as the primary key step
[0091] The manual filtering step conducts necessary preliminary cleaning of outliers through a series of rules, filtering outliers determined based on prior knowledge, and avoiding these outliers from interfering with the clustering parameters of automatic filtering.
[0092] Manual filtering clears some high measurement error and noise points in the data by filtering out sampling points with abnormal filtering distances or time intervals, measurement values outside the absolute threshold range, and points with abnormal changes (such as abnormal changes in the amplitude of data values and the angle of sensor attitude). Manual filtering reduces the root mean square error value of cross-validation, indicating that it provides a relatively clean initial data set for subsequent automatic filtering, avoiding the influence of outliers on the determination of clustering parameters in the automatic filtering process, and thus ensuring the effectiveness of automatic filtering.
[0093] (2)Combination of manual and automatic filtering effectively improves the data filtering effect
[0094] The combination of manual and automatic filtering significantly improves the convenience of data filtering and achieves obvious effects. The root mean square error value of cross-validation for the four depth data is further significantly reduced, and the prediction accuracy for the sand grain content with high correlation is significantly improved, indicating that after combining automatic filtering, the noise in the data is greatly reduced and the remaining data is more usable.
[0095] (3)Synergistic effect of the dual filtering strategy
[0096] The combination of manual filtering and automatic filtering not only realizes the efficient cleaning of outliers in stages but also ensures the robustness of the data filtering process: the regularized operation of manual filtering accurately locates specific types of outliers and solves some problems that are difficult for automatic filtering to directly handle (such as extreme values and local data anomalies). Automatic filtering further improves the overall quality of the data by identifying and filtering the remaining outliers based on the clustering algorithm, forming an effective synergistic effect.
[0097] In summary, the present invention significantly improves the quality of measurement data by combining manual filtering and automatic filtering strategies. Among them, manual filtering not only directly clears outliers but also provides good initial conditions for the parameter determination and clustering analysis of automatic filtering. The two work together effectively to improve the consistency and spatial correlation of the data, greatly improving the prediction accuracy after data filtering, and providing important support for data applications in fields such as precision agriculture and environmental monitoring.
[0098] This application also provides an application scenario that applies the above method for filtering real-time measurement data of vehicle-mounted sensors. Specifically: The method for filtering real-time measurement data of vehicle-mounted sensors provided in this embodiment can be applied in the scenario of precision agriculture development. This scenario includes a data acquisition link (for collecting measurement data using vehicle-mounted sensors), a data filtering link (for filtering the collected measurement data), and an agricultural guidance and development link (for precisely guiding agricultural development based on the filtered data). The method for filtering real-time measurement data of vehicle-mounted sensors provided in this embodiment belongs to the data filtering link.
[0099] Based on the same inventive concept, an embodiment of the present application further provides a filtering device for real-time measurement data of in-vehicle sensors for implementing the filtering method for real-time measurement data of in-vehicle sensors involved above. The implementation solutions provided by this device to solve problems are similar to those recorded in the above method. Therefore, the specific limitations in one or more embodiments of the filtering device for real-time measurement data of in-vehicle sensors provided below can refer to the limitations on the filtering method for real-time measurement data of in-vehicle sensors in the above text, and will not be elaborated here.
[0100] In an exemplary embodiment, as Figure 16 shown, a filtering device for real-time measurement data of in-vehicle sensors is provided, including the following modules.
[0101] A data acquisition module M1, configured to acquire measurement data at different depths collected by in-vehicle sensors within a target field block.
[0102] A manual filtering module M2, configured to manually filter data outliers from the measurement data at each depth to obtain manually filtered data.
[0103] A clustering module M3, configured to apply a clustering method to the manually filtered data at each depth for clustering to obtain a clustering result.
[0104] An automatic filtering module M4, configured to automatically filter outlier data from the manually filtered data at each depth according to the clustering result to obtain automatically filtered data.
[0105] In an exemplary embodiment, a computer device is provided. This computer device can be a server or a terminal, and its internal structure diagram can be as Figure 17 shown. This computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of this computer device is used to provide computing and control capabilities. The memory of this computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of this computer device is used to store filtering data of real-time measurement data of in-vehicle sensors. The input / output interface of this computer device is used for the processor to exchange information with external devices. The communication interface of this computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a filtering method for real-time measurement data of in-vehicle sensors.
[0106] Those skilled in the art can understand that Figure 17 The structure shown in Figure 17 is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0107] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0108] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0109] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0110] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0111] The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0112] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0113] Specific examples are used in this article to elaborate on the principles and implementation methods of the present application. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be construed as a limitation of the present application.
Claims
1. A filtering method for real-time measurement data of vehicle-mounted sensors, characterized in that The method for filtering real-time measurement data of in-vehicle sensors includes: Obtaining the measurement data of in-vehicle sensors within a target field; Manually filtering outlier data from the measurement data to obtain manually filtered data; Applying a clustering method to cluster the manually filtered data to obtain a clustering result; Automatically filtering outlier data from the manually filtered data according to the clustering result to obtain automatically filtered data; Among them, manually filtering outlier data from the measurement data to obtain manually filtered data specifically includes: Filtering the measurement data for outliers when the in-vehicle sensors start measuring, data beyond the threshold range defined by prior knowledge, data whose change exceeds a preset change degree, and data whose attitude change during in-vehicle sensor measurement exceeds a preset attitude change degree; Among them, applying a clustering method to cluster the manually filtered data to obtain a clustering result specifically includes: (1) For each data point in the manually filtered data, calculate the median within the time neighborhood and the median within the spatial neighborhood of each data point; (2) Calculate the clustering coordinates of each data point according to the median within the time neighborhood and the median within the spatial neighborhood of each data point; each data point includes multiple spatial neighborhoods; Among them, calculating the clustering coordinates of each data point according to the median within the time neighborhood and the median within the spatial neighborhood of each data point specifically includes: For each data point, calculate the difference between the measurement value of this data point and the median of the measurement values of each data point within the time neighborhood to obtain a first difference; Calculate the differences between this data point and the medians of the measurement values of each data point within the corresponding each spatial neighborhood respectively to obtain second differences; Calculate the mean value of the corresponding second differences of this data point to obtain a difference mean value; Take the corresponding first difference and difference mean value of this data point as the clustering coordinates of this data point; (3) According to the clustering coordinates of each data point in the manually filtered data, calculate the Euclidean distance between data points in the manually filtered data and estimate the kernel density of data points to determine the distance threshold parameter and the minimum number of points parameter in the DBSCAN algorithm; (4) Apply the DBSCAN algorithm for clustering to obtain the clustering result corresponding to the manually filtered data.
2. The filtering method for real-time measurement data of in-vehicle sensors according to claim 1, wherein After executing the step "Automatically filtering outlier data from the manually filtered data according to the clustering result to obtain automatically filtered data", the method for filtering real-time measurement data of in-vehicle sensors further includes: Evaluating the filtering effect of the automatically filtered data.
3. The filtering method for real-time measurement data of vehicle-mounted sensors according to claim 1, wherein When each data point includes different types of measurement data of the target field, after executing the step "Automatically filtering outlier data from the manually filtered data according to the clustering result to obtain automatically filtered data", the method for filtering real-time measurement data of in-vehicle sensors further includes: For each data point, determine whether each type of measurement data corresponding to the data point is normal data; if so, retain the measurement data of each type of the data point; if not, filter all types of measurement data corresponding to the data point.
4. A filtering device for real-time measurement data of vehicle-mounted sensors, characterized in that, The filtering device for real-time measurement data of in-vehicle sensors includes: A data acquisition module for obtaining the measurement data of in-vehicle sensors within a target field; A manual filtering module for manually filtering data outliers from the measurement data to obtain manually filtered data; A clustering module for calculating clustering coordinates and applying a clustering method to cluster the manually filtered data to obtain a clustering result; Among them, manually filtering data outliers from the measurement data to obtain manually filtered data specifically includes: Filtering the measurement data for outliers when the vehicle-mounted sensor starts measurement, data exceeding the threshold range defined by prior knowledge, data whose change exceeds a preset change degree, and data whose attitude change exceeds a preset attitude change degree during vehicle-mounted sensor measurement; Among them, applying a clustering method to cluster the manually filtered data to obtain a clustering result specifically includes: (1) For each data point in the manually filtered data, calculate the median in the time neighborhood and the median in the spatial neighborhood of each data point; (2) Calculate the clustering coordinates of each data point according to the median in the time neighborhood and the median in the spatial neighborhood of each data point; each data point includes multiple spatial neighborhoods; Among them, calculating the clustering coordinates of each data point according to the median in the time neighborhood and the median in the spatial neighborhood of each data point specifically includes: For each data point, calculate the difference between the measurement value of this data point and the median of the measurement values of each data point in the time neighborhood to obtain a first difference; Calculate the differences between this data point and the medians of the measurement values of each data point in the corresponding each spatial neighborhood respectively to obtain second differences; Calculate the mean value of the corresponding differences of this data point to obtain a difference mean value; Use the first difference and the difference mean value corresponding to this data point as the clustering coordinates of this data point; (3) According to the clustering coordinates of each data point in the manually filtered data, calculate the Euclidean distance between the data points in the manually filtered data and estimate the kernel density of the data points to determine the distance threshold parameter and the minimum number of points parameter in the DBSCAN algorithm; (4) Apply the DBSCAN algorithm for clustering to obtain the clustering result corresponding to the manually filtered data; An automatic filtering module for automatically filtering abnormal data in the manually filtered data according to the clustering result to obtain automatically filtered data.
5. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the filtering method for real-time measurement data of a vehicle-mounted sensor according to any one of claims 1-3.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the filtering method for real-time measurement data of a vehicle-mounted sensor according to any one of claims 1-3.
7. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the filtering method for real-time measurement data of a vehicle-mounted sensor according to any one of claims 1-3.
Citation Information
Patent Citations
Method and device for cleaning abnormal data of wind turbine generator
CN115438030A
Multi-modal model optimization retrieval training method and storage medium
CN118094216A