Method and device for filtering real-time measurement data of vehicle-mounted sensor

By combining manual filtering and automatic filtering, the problems of outliers and redundant data in vehicle-mounted sensor measurement data are solved, and the data quality is significantly improved and the filtration process is robust.

CN119961633AActive Publication Date: 2025-05-09CHINA AGRI UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510442556.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-09
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

During the real-time measurement of data by on-board sensors, outliers and redundant data exist, which affects the data quality and requires effective filtering methods to improve the quality of the data set.

Method used

A combination of manual filtering and automatic filtering is adopted. First, the outliers are cleaned up first by manually filtering, and then the clustering method is used to automatically filter to ensure the robustness of the data filtering process.

Benefits of technology

It realizes phased and efficient cleaning of outliers, improves the quality and accuracy of data filtering, and ensures the robustness of the data filtering process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961633A_ABST
    Figure CN119961633A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for filtering real-time measurement data of a vehicle-mounted sensor, and relates to the field of data filtering, and the method comprises the steps: obtaining the measurement data of the vehicle-mounted sensor in a target field; manually filtering data abnormal values of the measurement data to obtain manually filtered data; applying a clustering method to the manually filtered data to obtain a clustering result; and automatically filtering abnormal data in the manually filtered data according to a clustering result to obtain filtered data. By means of combination of manual filtering and automatic filtering, staging efficient cleaning of abnormal values is achieved, and robustness of the data filtering process is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data filtering, and in particular to a method and device for filtering real-time measurement data of a vehicle-mounted sensor. Background Art

[0002] Precision agriculture, as the foundation of sustainable agriculture, is the main trend of modern agricultural development. Therefore, obtaining the real-time status of farmland and the spatial distribution information of various parameters such as soil and crops is of great significance for guiding the implementation of precise farmland management and farmland soil fertility evaluation. In recent years, with the application of location positioning technologies such as the global positioning system, vehicle-mounted sensing technology has made great progress. Such sensors are installed on agricultural machinery and can obtain soil, crop properties and environmental parameters in real time during movement. They provide data with higher spatial resolution and larger scale range than traditional sampling methods, providing strong support for precision agriculture and soil difference analysis.

[0003] Vehicle-mounted sensing technology can be combined with machine location information to measure agricultural indicators at close range with high precision, and accurately relocate these measurements to specific areas. Although the range of vehicle-mounted sensors is relatively small compared to satellite- or airborne-based remote sensing systems, it can provide dynamic, real-time information with improved accuracy within a sufficiently large range, such as soil conductivity, vegetation coverage, yield information, and so on. This type of sensing system can collect a large amount of data on a large scale in a short period of time, greatly improving the efficiency of data collection. Although rich data is important for on-site management decisions, the collection process may also contain certain erroneous data, which needs to be filtered and pre-processed before further processing and analysis. Among them, some errors come from sensor performance, such as repeatability, accuracy, and resolution; more errors or redundant data come from the machine's motion posture, motion speed, operator habits, and the impact of special conditions on specific fields.

[0004] Sensor-based data sets are still spatial data sets. Unlike remote sensing data, they are not acquired as a whole pixel by pixel, but acquired sequentially as the machine moves, thus adding a time dimension. The accuracy of data location depends on the accuracy of the machine's GPS. This requires relevant filtering algorithms to improve the quality of the data set. Summary of the invention

[0005] The purpose of this application is to provide a method and device for filtering real-time measurement data of vehicle-mounted sensors, which combines manual filtering with automatic filtering to not only achieve efficient staged cleaning of outliers, but also ensure the robustness of the data filtering process.

[0006] To achieve the above objectives, this application provides the following solutions.

[0007] In a first aspect, the present application provides a method for filtering real-time measurement data of a vehicle-mounted sensor, comprising the following steps.

[0008] Obtain measurement data from vehicle-mounted sensors within the target field.

[0009] Manually filter the measured data for outliers to obtain manually filtered data.

[0010] A clustering method is applied to cluster the manually filtered data to obtain a clustering result.

[0011] The abnormal data in the manually filtered data is automatically filtered according to the clustering result to obtain the automatically filtered data.

[0012] In a second aspect, the present application provides a device for filtering real-time measurement data of a vehicle-mounted sensor, comprising the following modules.

[0013] The data acquisition module is used to obtain the measurement data of the vehicle-mounted sensor in the target field.

[0014] The manual filtering module is used to manually filter out abnormal values ​​of the measured data to obtain manually filtered data.

[0015] The clustering module is used to cluster the manually filtered data using a clustering method to obtain a clustering result.

[0016] The automatic filtering module is used to automatically filter abnormal data in the manually filtered data according to the clustering results to obtain the automatically filtered data.

[0017] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned method for filtering real-time measurement data of vehicle-mounted sensors.

[0018] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned method for filtering real-time measurement data of a vehicle-mounted sensor.

[0019] In a fifth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned method for filtering real-time measurement data of a vehicle-mounted sensor.

[0020] According to the specific embodiments provided in this application, this application discloses the following technical effects.

[0021] The present application provides a method and device for filtering real-time measurement data of vehicle-mounted sensors, which obtains measurement data collected by vehicle-mounted sensors in a target field; manually filters data outliers from the measurement data to obtain manually filtered data; applies a clustering method to the manually filtered data to obtain clustering results; and automatically filters outliers in the manually filtered data based on the clustering results to obtain automatically filtered data. In the process of data filtering, the outliers in the data are first manually filtered to achieve preliminary filtering, and on the basis of the preliminary filtering, a clustering algorithm is applied to perform automatic fine filtering. The manual filtering step uses a series of rules to perform necessary preliminary cleaning of outliers, while avoiding interference of these outliers with the clustering parameters of automatic filtering, thereby ensuring the effectiveness of automatic filtering. The combination of manual filtering and automatic filtering in the present application not only achieves efficient stage-by-stage cleaning of outliers, but also ensures the robustness of the data filtering process. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0023] Figure 1 This is a diagram of the application environment of a method for filtering real-time measurement data of a vehicle-mounted sensor in one embodiment of the present application.

[0024] Figure 2 A flowchart of a method for filtering real-time measurement data of a vehicle-mounted sensor provided in one embodiment of the present application.

[0025] Figure 3 A schematic diagram of raw data collected from soil apparent conductivity at four depths provided in an embodiment of the present application.

[0026] Figure 4 A schematic diagram of soil apparent conductivity data at four depths after manual filtering provided in an embodiment of the present application.

[0027] Figure 5 A schematic diagram of the spatial neighborhood and temporal neighborhood of an example sample point provided in an embodiment of the present application.

[0028] Figure 6 A schematic diagram of soil apparent conductivity data at four depths after automatic filtering provided in an embodiment of the present application.

[0029] Figure 7 This is the data filtering effect at a depth of 0.54 m (PRP1) provided in an embodiment of the present application.

[0030] Figure 8 This is the data filtering effect at a depth of 1.03 m (PRP2) provided in one embodiment of the present application.

[0031] Fig. 9 This is the data filtering effect at a depth of 1.55 m (HCP1) provided in one embodiment of the present application.

[0032] Fig.10 This is the data filtering effect at a depth of 3.18 m (HCP2) provided in one embodiment of the present application.

[0033] Fig.11 A schematic diagram of Kriging interpolation of raw data collected from soil apparent conductivity at four depths is provided for one embodiment of the present application.

[0034] Fig.12 A schematic diagram of Kriging interpolation of soil apparent conductivity data at four depths after manual filtering is provided for an embodiment of the present application.

[0035] Fig.13 A schematic diagram of Kriging interpolation of soil apparent conductivity data at four depths after automatic filtering is provided for an embodiment of the present application.

[0036] Fig.14 The distribution of modeling set sample points and verification set sample points for soil sand content prediction provided in one embodiment of the present application.

[0037] Fig.15 A correlation distribution diagram between the soil sand content and the data after different data filtering steps provided in an embodiment of the present application.

[0038] Fig.16 A schematic diagram of functional modules of a device for filtering real-time measurement data of a vehicle-mounted sensor provided in another embodiment of the present application.

[0039] Fig.17 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0040] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0041] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0042] The filtering method for real-time measurement data of the vehicle-mounted sensor provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal communicates with the server through the network. The data storage system can store the data that the server needs to process. The data storage system can be set up separately, integrated on the server, or placed on the cloud or other servers. The terminal can send the measurement data collected by the vehicle-mounted sensor in the target field to be processed to the server. After the server receives the measurement data of the vehicle-mounted sensor in the target field, the server manually filters the data abnormal values ​​of the measurement data to obtain the manually filtered data; applies the clustering method to the manually filtered data to obtain the clustering results; automatically filters the abnormal data in the manually filtered data according to the clustering results to obtain the automatically filtered data. The server can feed back the obtained automatically filtered data to the terminal. In addition, in some embodiments, the filtering method of the real-time measurement data of the vehicle-mounted sensor can also be implemented by the server or the terminal alone, such as the terminal can directly filter the measurement data collected by the vehicle-mounted sensor in the target field to be processed, or the server can obtain the measurement data collected by the vehicle-mounted sensor in the target field to be processed from the data storage system and filter the data.

[0043] The terminal may be, but is not limited to, various desktop computers, laptops, smart phones, tablet computers, IoT devices, and portable wearable devices. The server may be implemented as an independent server or a server cluster consisting of multiple servers, or may be a cloud server.

[0044] In an exemplary embodiment, neither the manual filtering method based on experience nor the automatic filtering method based on clustering algorithm can filter out all outliers. For the manual filtering method based on experience, on the one hand, it requires a large workload; on the other hand, it is impossible to filter outliers when the instrument is affected by the environment and produces local mutation outliers. For the automatic filtering method based on clustering algorithm, it is difficult to identify outliers that are caused by instrument posture and vehicle operation but may be difficult to reflect significantly and intuitively from the readings themselves. At the same time, the parameters of the algorithm itself and even the filtering effect are easily affected by these outliers. Therefore, it is necessary to combine the manual filtering method based on prior experience and error sources and the automatic filtering method based on clustering algorithm to form a universal abnormal data filtering framework. In this regard, Figure 2 As shown, a method for filtering real-time measurement data of a vehicle-mounted sensor is provided. The method is executed by a computer device, and specifically can be executed by a computer device such as a terminal or a server alone, or can be executed by a terminal and a server together. In the embodiment of the present application, the method is applied to Figure 1The server in is taken as an example to illustrate, including the following steps 101 to 104.

[0045] Step 101, obtaining measurement data at different depths collected by a vehicle-mounted sensor in a target field.

[0046] This embodiment uses the soil apparent conductivity data measured at different depths as an example to illustrate the filtering steps. The measurement data of the vehicle-mounted sensor in the target field may be data of different categories, measurement data at a certain depth of the same category, or measurement data at different depths of the same category. The measurement data includes but is not limited to soil apparent conductivity (ECa), temperature, humidity, crop yield, etc.

[0047] There are drainage pipes passing through the target field, which will affect the soil ECa measurement at the corresponding location. This is also one of the difficulties in data filtering. Based on the DUALEM-21S sensor, the ECa of the soil at multiple depths of the target field is collected, including the soil ECa at four depths of 0.54m (recorded as PRP1), 1.03m (recorded as PRP2), 1.55m (recorded as HCP1) and 3.18m (recorded as HCP2). Figure 3 The following are the raw data collected from the apparent conductivity of soil at four depths. In addition, the measured data also includes the attitude data of DUALEM-21S during the movement, which is divided into pitch angle and roll angle. The TrimbleAgGPS 542 global navigation satellite system (GNSS) receiver based on real-time kinematic (RTK) technology collects the longitude and latitude of each sampling point during the movement. A total of 9454 sampling points were collected.

[0048] Step 102: Manually filter the measured data at each depth to remove abnormal values, and obtain manually filtered data.

[0049] Step 103: cluster the manually filtered data at each depth using a clustering method to obtain a clustering result.

[0050] Step 104 , for the manually filtered data at each depth, abnormal data in the manually filtered data is automatically filtered according to the clustering result to obtain automatically filtered data.

[0051] By implementing the above-mentioned steps 101 to 104, the combination of manual filtering and automatic filtering not only realizes the efficient cleaning of outliers in stages, but also ensures the robustness of the data filtering process. The regularized operation of manual filtering accurately locates specific types of outliers and solves some problems that are difficult to directly handle by automatic filtering (such as local data anomalies). Automatic filtering identifies and filters the remaining outliers through a clustering algorithm, further improves the overall quality of the data on the basis of manual filtering, and forms an effective synergistic effect. The present invention achieves a significant improvement in the quality of measurement data by combining manual filtering with automatic filtering strategies. Among them, manual filtering not only directly removes some outliers, but also provides good initial conditions for parameter determination and cluster analysis of automatic filtering. The synergistic effect of the two effectively improves the consistency and spatial correlation of the data, greatly improves the prediction accuracy after data filtering, and provides important support for data applications in fields such as precision agriculture and environmental monitoring.

[0052] In another exemplary embodiment of the present application, in step 102, the measurement data at each depth is manually filtered for data outliers to obtain manually filtered data, specifically including: for the measurement data at each depth, filtering outliers when the vehicle-mounted sensor starts measuring, data that exceeds a threshold range defined by prior knowledge, data whose changes exceed a preset degree of change, and data whose posture changes exceed a preset degree of posture change when the vehicle-mounted sensor measures.

[0053] As an example, the startup (warm-up) uncertainty value is filtered. Specifically, the data points that are less than 1m or more than 6m away from the previous data point in space are filtered, and the data points that are more than 1.5s away from the previous data point in time are filtered, and a total of 37 data points are filtered.

[0054] As an example, the soil ECa values ​​outside the minimum and maximum absolute thresholds defined by the researcher based on prior knowledge are filtered. Specifically, the data points with soil ECa values ​​less than 0 are filtered, and a total of 23 data points are filtered.

[0055] As an example, data points with abnormally large changes in soil ECa are removed. Specifically, data points with soil ECa values ​​that vary by more than 20 mS / m relative to the previous and next data points and between different depths of the same data point are filtered out, and a total of 24 data points are filtered out.

[0056] As an example, data points whose pitch changes exceed the acceptable range are removed. Specifically, points whose pitch changes exceed 10° are filtered, and a total of 20 data points are filtered.

[0057] As an example, data points with a roll variation exceeding an acceptable range are removed. Specifically, data points with a roll variation exceeding 10° are filtered, and a total of 13 data points are filtered. Figure 4These are the soil apparent conductivity data at four depths after manual filtering.

[0058] In another exemplary embodiment of the present application, a clustering method is applied to the manually filtered data at each depth to perform clustering to obtain a clustering result, which specifically includes the following steps.

[0059] (1) For each data point in the manually filtered data at each depth, calculate the median of each data point in its temporal neighborhood and the median of its spatial neighborhood.

[0060] For any data point (here called a sample point), the data points before and after it are adjacent to it in time, and the data points on its left and right paths are adjacent to it in space. The four points before and after each data point are taken as its "time neighborhood", and the points within a certain distance range on the left and right paths of each data point are taken as its "spatial neighborhood". The median of the measured values ​​of the data points contained in the time neighborhood and the spatial neighborhood are calculated respectively. Figure 5 It is a schematic diagram of the spatial and temporal neighborhood of the example sample point.

[0061] (2) Calculate the cluster coordinates of each data point based on the median of its temporal neighborhood and the median of its spatial neighborhood.

[0062] Calculate the difference (x) between the sample point measurement value and the median of the measurement value of each data point in the time neighborhood and the difference (y) between the sample point measurement value and the median of the measurement value of each data point in the spatial neighborhood as the cluster coordinates (x, y) of the sample point. Since the distance between each data point and the left and right paths is not necessarily equal, when defining the spatial neighborhood, three ranges of 10m, 15m and 20m are defined and the difference is calculated respectively, and the average value of the difference is taken as y.

[0063] (3) According to the cluster coordinates of each data point in the manually filtered data at each depth, the Euclidean distance between the data points in the manually filtered data and the kernel density of the estimated data points are calculated, and the distance threshold parameter epsilon and the minimum point number parameter minPoints in the density-based spatial clustering of applications with noise (DBSCAN) algorithm are determined.

[0064] (4) Based on the distance threshold parameter epsilon and the minimum point number parameter minPoints, the DBSCAN algorithm is applied to perform clustering to obtain the clustering results corresponding to the manually filtered data at each depth.

[0065] The DBSCAN algorithm was executed for clustering, and the main categories were retained as the final filtered data at this depth (a total of 3556 data points were filtered). Figure 6 It is the soil apparent conductivity data at four depths after automatic filtering.

[0066] The DBSCAN algorithm is different from traditional distance-based methods in that it identifies clusters based on the density of data points. It is able to find areas with relatively high density and treat them as a cluster, while identifying noise points. The DBSCAN algorithm defines clusters through two main parameters: the distance threshold parameter epsilon and the minimum number of points parameter minPoints. The distance threshold parameter epsilon determines a neighborhood range centered on the data point, and the minimum number of points parameter minPoints indicates the minimum number of data points in the neighborhood that are considered core points. The DBSCAN algorithm works by randomly selecting a point from the data set and checking whether the number of points in its epsilon neighborhood meets the conditions for becoming a core point. If it is a core point, the points around it are assigned to the same cluster. If it is a boundary point, it is assigned to the cluster of the corresponding core point. This process is repeated until all points are assigned to the corresponding cluster or marked as noise points.

[0067] In another exemplary embodiment of the present application, when each data point includes different types of measurement data of the target field, after executing the step of "automatically filtering abnormal data in the manually filtered data according to the clustering results to obtain automatically filtered data", the filtering method of real-time measurement data of the vehicle-mounted sensor also includes: for each data point, determining whether each type of measurement data corresponding to the data point is normal data; if so, retaining each type of measurement data of the data point; if not, filtering all types of measurement data corresponding to the data point.

[0068] For example, for the same sensor, if soil ECa and magnetic susceptibility can be measured simultaneously, that is, each data point has two or more types of data from the same sensor and measured simultaneously, then if one type of data is problematic, then this data point will be filtered. For example, if soil ECa does not filter out a data point, but the magnetic susceptibility measurement of this point is abnormal, then it is considered that the sensor state at this time is problematic, and the soil ECa of this sampling point is considered a potential outlier, that is, the intersection of different types of data filtering results is taken.

[0069] In another exemplary embodiment of the present application, after executing the step of "for the manually filtered data at each depth, automatically filtering abnormal data in the manually filtered data according to the clustering results to obtain automatically filtered data", the filtering method for real-time measurement data of the vehicle-mounted sensor also includes: evaluating the filtering effect of the automatically filtered data.

[0070] from Figure 3-Figure 4 and Figure 6 , Figure 11-13 It can be seen that the filtering method of the present application removes the influence of the drainage pipe. Further analysis is conducted through geostatistics and soil sand content prediction performance.

[0071] 1. Analysis by geostatistics (detailed results are shown in Table 1, Figure 7-10 . Figure 7-10 The "model" in the figure refers to the fitting model of the semivariance function, and "averaged" means that when calculating the semivariance, the data are grouped by distance intervals and the semivariance values ​​of each interval are averaged to reduce noise and show the spatial structure more clearly; the "1:1 line" represents the case where the predicted value and the measured value are the same, that is, the prediction error is 0).

[0072] (1) Cross-validation: The root mean square error (RMSE) in cross-validation is used to further measure the effect of data filtering. The specific calculation method is shown in formula (1).

[0073]

[0074] Among them, y obs is the observed value, y pred is the predicted value and n is the number of samples.

[0075] The RMSE values ​​in the original data were relatively high, indicating that there was a large error between the observed and predicted values ​​when cross-validating the original data. After manual filtering, the RMSE values ​​generally decreased, indicating that manually removing outliers or noise improved the fit. After further automatic filtering, the RMSE values ​​were further reduced, indicating that automatic filtering further reduced the prediction error and improved the accuracy of the prediction by effectively dealing with inconsistencies and noise in the data.

[0076] (2) Nugget value: The nugget value reflects the short-range variability of the data, which is usually caused by noise, outliers, or local errors. There is a large short-range variability in the original data. After manual filtering, the nugget value dropped significantly. This shows that the outliers or noise in the data have been greatly reduced through manual filtering. After manual and automatic filtering, the nugget value is further reduced overall, especially for HCP2. This shows that automatic filtering further removes outliers in the data. However, the nugget value of PRP1 is slightly larger after automatic filtering, which may be due to the data in some areas with local fluctuations.

[0077] (3) Range: The range reflects the spatial autocorrelation of the data at a certain scale. Outliers and noise will introduce additional randomness in the estimation process, which may mask or distort the true spatial structure of the data. After filtering out these outliers, the spatial autocorrelation of the data itself can be more accurately reflected, and their impact on the range is not fixed. If the outliers mainly mask the spatial correlation at a larger scale, filtering may increase the estimated range (PRP2, HCP1), that is, the spatial dependence decays at a larger distance. If the outliers mainly affect local details, the data becomes smoother after removal, and the local correlation is more obvious, then the range may decrease (PRP1, HCP2), because the true spatial structure shows high correlation within a shorter distance.

[0078] (4) Sill value: The sill value represents the overall variability of the data. After data filtering, if noise and outliers mainly increase the random fluctuations of the data, then after removing them, the data becomes smoother and the estimated overall variability will decrease, thereby reducing the sill value (PRP1, HCP2). Conversely, if the filtering process causes those areas with large variability to account for a higher proportion of the overall data, then the estimated variability may increase, resulting in an increase in the sill value (PRP2, HCP1).

[0079] Table 1 Parameters of semivariogram prediction after different data filtering methods

[0080] 2. Analysis by correlation with soil sand content and accuracy of predicting soil sand content Surface soil samples were collected from the target fields and the sand content of the soil was measured. The original data, the data after manual filtration, and the data after manual filtration + automatic filtration were used for prediction. The coefficient of determination (R 2 ) and RMSE are used to evaluate the prediction accuracy and further measure the effect of data filtering. 2 The specific calculation method of is shown in formula (2).

[0081]

[0082] Among them, yobs is the observed value, ypred is the predicted value, and n is the number of samples.

[0083] Fig.14 The sample point distribution in is to illustrate the sample point distribution used for soil sand content prediction modeling and verification. Fig.15The correlation distribution diagram in the figure can intuitively show the effect of data filtering. The detailed results are shown in Table 2. It can be seen that the correlation and accuracy are improved after filtering. Overall, manual filtering improves the correlation between the soil ECa dataset and the soil sand content, and the manual filtering + automatic filtering method can further improve the correlation and even the prediction accuracy, and obtain better prediction results.

[0084] Table 2 Prediction accuracy of soil sand content after different data filtering methods

[0085] Processing tools: The data filtering process is completed using Python programming language; the geostatistical related parts are completed using ArcGIS Pro software.

[0086] The present invention aims to solve the problem of outliers caused by sensor posture changes, environmental interference or equipment errors by defining a detection and filtering mechanism for data anomalies, and realize data quality optimization. The method includes three main steps: data collection, manual filtering and automatic filtering. First, target data and auxiliary information during motion, such as posture angle, geographic location and timestamp, are collected based on vehicle-mounted sensors. Secondly, the uncertain values ​​when the vehicle or instrument is started, the outliers beyond the threshold range defined by prior knowledge, the values ​​with abnormally large data changes, and the points where the instrument or vehicle posture changes beyond the acceptable range are filtered out through the manual filtering step to complete the manual filtering operation. Subsequently, the median of the measured value is calculated based on the data in the neighborhood of each data point, and the clustering coordinates are constructed based on the deviation between the data point and the neighborhood median. The parameters of the clustering algorithm are determined by Euclidean distance and kernel density estimation, and cluster analysis is performed to identify and remove outliers, and the main category of data is retained as the filtered result. The present invention provides a universal and modular data filtering method, which is suitable for outlier processing of various vehicle-mounted sensor data, can effectively improve the accuracy and reliability of vehicle-mounted sensor measurement data, and provide high-quality data support for precision agriculture and environmental monitoring. Compared with the existing technology, it has the advantages of significantly improved data quality, accurate filtering of outliers, significantly improved prediction accuracy and strong applicability. These advantages come from the organic combination of manual filtering and automatic filtering dual data filtering strategies, especially manual filtering lays the foundation for the effectiveness of automatic filtering, as described below.

[0087] (1) Data quality is significantly improved, with manual filtering being the primary key step The manual filtering step performs necessary preliminary cleaning of outliers through a series of rules, filtering outliers determined based on prior knowledge to prevent these outliers from interfering with the clustering parameters of automatic filtering.

[0088] Manual filtering removes some high measurement errors and noise points in the data by filtering sampling points with abnormal distance or time interval, measurement values ​​beyond the absolute threshold range, and points with abnormal changes (such as abnormal changes in data value amplitude and sensor attitude angle). Manual filtering reduces the root mean square error value of cross-validation, indicating that it provides a relatively clean initial data set for subsequent automatic filtering, avoiding outliers from affecting the determination of clustering parameters in the automatic filtering process, thereby ensuring the effectiveness of automatic filtering.

[0089] (2) Combining manual and automatic filtering effectively improves data filtering effects The combination of manual and automatic filtering significantly improved the convenience of data filtering and achieved obvious results. The root mean square error value of the cross-validation of the four depth data was further significantly reduced, and the prediction accuracy of the highly correlated sand content was significantly improved, indicating that after combining with automatic filtering, the noise of the data was greatly reduced and the remaining data was more usable.

[0090] (3) Synergy of dual filtering strategies The combination of manual filtering and automatic filtering not only achieves efficient phased cleaning of outliers, but also ensures the robustness of the data filtering process: the regularized operation of manual filtering accurately locates specific types of outliers and solves some problems that are difficult to directly handle with automatic filtering (such as extreme values ​​and local data anomalies). Automatic filtering uses clustering algorithms to identify and filter remaining outliers, further improving the overall quality of the data on the basis of manual filtering, forming an effective synergy.

[0091] In summary, the present invention achieves a significant improvement in the quality of measurement data by combining manual filtering with automatic filtering strategies. Among them, manual filtering not only directly removes outliers, but also provides good initial conditions for parameter determination and cluster analysis of automatic filtering. The synergistic effect of the two effectively improves the consistency and spatial correlation of the data, greatly improves the prediction accuracy after data filtering, and provides important support for data applications in fields such as precision agriculture and environmental monitoring.

[0092] The present application also provides an application scenario, which applies the above-mentioned method for filtering real-time measurement data of vehicle-mounted sensors. Specifically: The method for filtering real-time measurement data of vehicle-mounted sensors provided in this embodiment can be applied in the precision agriculture development scenario. The scenario includes a data collection link (used to collect measurement data using vehicle-mounted sensors), a data filtering link (used to filter the collected measurement data) and an agricultural guidance development link (used to accurately guide agricultural development based on the filtered data). The method for filtering real-time measurement data of vehicle-mounted sensors provided in this embodiment belongs to the data filtering link.

[0093] Based on the same inventive concept, the embodiment of the present application also provides a filtering device for real-time measurement data of vehicle-mounted sensors for implementing the filtering method for real-time measurement data of vehicle-mounted sensors involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in the embodiments of the filtering device for real-time measurement data of one or more vehicle-mounted sensors provided below can refer to the limitations of the filtering method for real-time measurement data of vehicle-mounted sensors above, and will not be repeated here.

[0094] In an exemplary embodiment, Fig.16 As shown, a device for filtering real-time measurement data of a vehicle-mounted sensor is provided, comprising the following modules.

[0095] The data acquisition module M1 is used to obtain measurement data at different depths collected by the vehicle-mounted sensor in the target field.

[0096] The manual filtering module M2 is used to manually filter outliers in the measurement data at each depth to obtain manually filtered data.

[0097] The clustering module M3 is used to apply a clustering method to the manually filtered data at each depth to cluster and obtain a clustering result.

[0098] The automatic filtering module M4 is used to automatically filter abnormal data in the manually filtered data at each depth according to the clustering result to obtain automatically filtered data.

[0099] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Fig.17 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store filtering data of real-time measurement data of vehicle-mounted sensors. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for filtering real-time measurement data of a vehicle-mounted sensor is implemented.

[0100] Those skilled in the art will understand that Fig.17 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components. In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.

[0101] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0102] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0103] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0104] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0105] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.

[0106] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0107] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A method for filtering real-time measurement data of a vehicle-mounted sensor, characterized in that: The method for filtering the real-time measurement data of the vehicle-mounted sensor comprises: Obtain measurement data from vehicle-mounted sensors within the target field; Manually filtering the measured data for outliers to obtain manually filtered data; Applying a clustering method to cluster the manually filtered data to obtain a clustering result; The abnormal data in the manually filtered data is automatically filtered according to the clustering result to obtain the automatically filtered data.

2. The method for filtering real-time measurement data of a vehicle-mounted sensor according to claim 1, characterized in that: Manually filtering the measured data for outliers to obtain manually filtered data, specifically including: For the measurement data, abnormal values ​​when the vehicle-mounted sensor starts measuring, data exceeding a threshold range defined by prior knowledge, data whose data changes exceed a preset degree of change, and data whose posture changes when the vehicle-mounted sensor measures exceed a preset degree of posture change are filtered.

3. The method for filtering real-time measurement data of a vehicle-mounted sensor according to claim 1, characterized in that: Applying a clustering method to cluster the manually filtered data to obtain a clustering result, specifically including: For each data point in the manually filtered data, calculating the median in the time neighborhood and the median in the spatial neighborhood of each data point; The cluster coordinates of each data point are calculated based on the median of its temporal neighborhood and the median of its spatial neighborhood; According to the cluster coordinates of each data point in the manually filtered data, the Euclidean distance between the data points in the manually filtered data is calculated and the kernel density of the data points is estimated, and the distance threshold parameter and the minimum number of points parameter in the DBSCAN algorithm are determined; The DBSCAN algorithm is used for clustering to obtain the clustering results corresponding to the manually filtered data.

4. The method for filtering real-time measurement data of a vehicle-mounted sensor according to claim 1, characterized in that: After executing the step of "automatically filtering abnormal data in the manually filtered data according to the clustering result to obtain automatically filtered data", the method for filtering the real-time measurement data of the vehicle-mounted sensor also includes: Evaluate the filtering effect of the data after automatic filtering.

5. The method for filtering real-time measurement data of a vehicle-mounted sensor according to claim 1, characterized in that: When each data point includes different types of measurement data of the target field, after executing the step of "automatically filtering abnormal data in the manually filtered data according to the clustering result to obtain automatically filtered data", the filtering method for real-time measurement data of the vehicle-mounted sensor further includes: For each data point, determine whether each type of measurement data corresponding to the data point is normal data; if so, retain each type of measurement data of the data point; if not, filter all types of measurement data corresponding to the data point.

6. A filtering device for real-time measurement data of a vehicle-mounted sensor, characterized in that: The filtering device for real-time measurement data of the vehicle-mounted sensor comprises: A data acquisition module is used to obtain measurement data from vehicle-mounted sensors in the target field; A manual filtering module, used for manually filtering outliers in the measurement data to obtain manually filtered data; A clustering module, used for clustering the manually filtered data using a clustering method to obtain a clustering result; The automatic filtering module is used to automatically filter abnormal data in the manually filtered data according to the clustering results to obtain the automatically filtered data.

7. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for filtering real-time measurement data of a vehicle-mounted sensor according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for filtering real-time measurement data of a vehicle-mounted sensor according to any one of claims 1 to 5 is implemented.

9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for filtering real-time measurement data of a vehicle-mounted sensor according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Abnormal value elimination method for lake water level long-time sequence monitoring data

    CN114817228A

  • Expressway traffic accident black spot road section identification method and computer device

    CN115424430A

  • Method and device for cleaning abnormal data of wind turbine generator

    CN115438030A

  • Multi-modal model optimization retrieval training method and storage medium

    CN118094216A

  • Data abnormality detection method and apparatus, electronic device and storage medium

    WO2021184727A1