A track data cleaning method based on fuzzy clustering

Aircraft track data cleaning is performed using fuzzy clustering methods, which solves the insufficient application of fuzzy clustering in track data processing in the existing technology, achieves more efficient data cleaning, maintains information integrity and reduces bias.

CN116842315BActive Publication Date: 2025-09-09HUNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310703567.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-14
Publication Date
2025-09-09
Estimated Expiration
2043-06-14

AI Technical Summary

Technical Problem

The existing technology does not use fuzzy clustering methods in the cleaning and processing of aircraft track data, which affects the processing efficiency.

Method used

A track data cleaning method based on fuzzy clustering is adopted, including data deduplication, missing item deletion, density clustering and two-dimensional interpolation processing, to identify and mark outliers without deletion or replacement.

Benefits of technology

It improves the accuracy and stability of data, reduces information loss and bias, and can better identify fuzzy outliers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116842315B_ABST
    Figure CN116842315B_ABST
Patent Text Reader

Abstract

The present invention discloses a track data cleaning method based on fuzzy clustering, comprising the following steps: Step 1: Acquire ADS‑B raw data; Step 2: Deduplication of data; Step 3: Deletion of missing items; Step 4: Detection and analysis of outliers using fuzzy clustering; Step 5: Extraction of data with appropriate spacing using Newton interpolation as the final result of data cleaning. The present invention is capable of identifying fuzzy outliers: Since fuzzy clustering can assign data points to multiple fuzzy groups, it can better identify fuzzy outliers, i.e., situations where data points do not completely belong to any one group; outliers will not be deleted or replaced, so useful information will not be lost. Instead, they can mark outliers by assigning metrics for reference in further analysis, while better handling uncertainty and complexity, thereby reducing the possibility of bias.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to a track data cleaning method based on fuzzy clustering. Background Art

[0002] In data mining and machine learning, outliers are data points that differ significantly from the majority of data points, usually due to errors, noise, or unusual circumstances. These outliers may interfere with the performance and results of the model, so outlier handling is necessary.

[0003] Traditional outlier handling methods include removing, replacing, or labeling outliers. However, these methods may lose useful information or introduce bias. Therefore, some researchers have explored the use of clustering methods for outlier handling.

[0004] Fuzzy clustering is a clustering method that takes into account that data points may belong to multiple groups rather than just one. Therefore, fuzzy clustering can be used to assign data points to multiple fuzzy groups rather than just one clear group.

[0005] Outlier handling methods based on fuzzy clustering assign data points to fuzzy groups and use their assignment metric as an outlier score. If a data point is assigned to a small, distinct fuzzy group, it is likely an outlier. Therefore, this method can identify and handle outliers using the assignment metric.

[0006] The performance of the fuzzy clustering method depends on the selection of some parameters, such as the number of fuzzy groups and the fuzziness parameter. Therefore, when implementing the outlier processing method based on fuzzy clustering, parameter optimization and adjustment are required to obtain the best results.

[0007] However, in the prior art, there has been no application of fuzzy clustering methods in aircraft track data cleaning and processing, which affects the efficiency of aircraft track data cleaning and processing. Summary of the Invention

[0008] 1. Technical problems to be solved

[0009] The purpose of the present invention is to solve the problem that the fuzzy clustering method has not yet been applied in the aircraft track data cleaning process in the prior art, and to propose a track data cleaning method based on fuzzy clustering.

[0010] 2. Technical solution

[0011] In order to achieve the above object, the present invention adopts the following technical solutions:

[0012] A track data cleaning method based on fuzzy clustering includes the following steps:

[0013] Step 1: Get ADS-B raw data. Get ADS-B data of several historical tracks from VariFlight global flight real-time tracking radar platform. Each track contains several point tracks P i , each point contains some feature information Q i The relationship between the track L, the point track Pi and the characteristic items can be expressed as follows:

[0014]

[0015]

[0016] Among them, P i is the i-th point in the track, Q i P i The i-th characteristic information in the ADS-B data is read, and the trace P after the above steps is processed i It can be expressed as:

[0017] P i ={timestamp, altitude, speed, azimuth, longitude, latitude};

[0018] Step 2: Data deduplication. To facilitate subsequent outlier search and analysis, data deduplication must be performed in advance.

[0019] Step 3: Deletion of missing items. To prevent incomplete data, data with missing items should be deleted. First, search for data groups with missing items and delete the data groups found to prevent them from affecting the processing of other data.

[0020] Step 4: Use fuzzy clustering to detect and analyze outliers. The density clustering method is used for outlier processing. The specific implementation process and ideas are as follows: Assume that there are several points in a two-dimensional plane and cluster these points. The density clustering process is as follows: For each point, calculate the two variables of "local density" and "local distance" of this point;

[0021] Step 5: Use the two-dimensional interpolation method to extract data with appropriate spacing as the final result of data cleaning, use the two-dimensional function formula for fitting, and use interpolation to convert the data into points with the same spacing to facilitate subsequent operations; linearize the data and take points with the same time distance on a line as the points needed for subsequent extraction.

[0022] Preferably, each point track in each track in the raw data obtained from VariFlight in step 1 contains 9 feature items, namely: Time (time stamp), UTC (UTC time), Anum (aircraft registration number), Fnum (flight number), Height (altitude), Speed ​​(speed), Angle (azimuth), Longitude (longitude), and Latitude (latitude). Invalid data including UTC (time), Anum, and Fnum can be deleted after processing.

[0023] Preferably, the data deduplication in step 2 specifically includes the following steps:

[0024] S2.1 sorts the data according to the time sequence, making the order of the data clearer and easier to process the data and retrieve the output data after processing.

[0025] Compare and sort each set of data according to the time series, so that each set of data in the data set can be accurately sorted according to time.

[0026] S2.2 Delete data with the same time. During the flight of the aircraft, points with the same time in the row set must be caused by errors, so the track points with the same timestamp should be deleted to avoid unnecessary impact on the actual data and make the data inaccurate.

[0027] Search for data with the same time point according to the time series, and delete them from the original data set after finding them to ensure the uniqueness and accuracy of the timestamp.

[0028] S2.3 Delete data with the same longitude and latitude. Because aircraft flight is unidirectional, over time, there will no longer be track points with the same longitude and latitude. Therefore, it is necessary to delete track points with the same longitude and latitude. Extract the longitude and latitude of each data set separately, find data with the same longitude and latitude, record their sequence numbers, and delete all data sets with the same longitude and latitude to avoid errors.

[0029] Preferably, the local density in step 4 depends on the dc value, which is a hyperparameter manually specified before density clustering. It is a range distance. Then the local density of a point is the number of neighboring points within the dc range around the point, that is, the number of neighboring points of the circle with this point as the center and dc as the radius.

[0030] Preferably, the dc value is generally selected to be a value that makes the average number of neighbors of all data points be 1-2% of the total data points.

[0031] Preferably, in step 4, the local distance is: assuming that the local density of all points has been obtained, for each point's local distance: for the point with the highest local density, its local distance is the maximum value of the distances between it and all other points, that is, the distance to the point farthest from it; for a point that does not have the highest local density, its local distance is: the distance to the nearest point among all points with higher local density than it.

[0032] Preferably, the total distance in step 4 is: the total distance of each point = local density + local distance; if the data is ultimately to be divided into 3 categories, then the first three points with the largest total distance are the three center points, and the remaining points are classified into the category to whichever point they are closer to; the group with the most data around the center is taken as the required data, and the remaining data is discarded; in the actual operation process, fuzzy clustering processing is performed on latitude and longitude, altitude, and speed with time as the independent variable to eliminate outliers.

[0033] Preferably, other data at the same time in step 5 are also modified in the same way as the final result of data cleaning to facilitate subsequent data calls.

[0034] 3. Beneficial effects

[0035] Compared with the prior art, the advantages of the present invention are:

[0036] (1) In the present invention, the track data cleaning method based on fuzzy clustering is able to identify fuzzy outliers: since fuzzy clustering can assign data points to multiple fuzzy groups, it can better identify fuzzy outliers, that is, situations where data points do not completely belong to any group.

[0037] (2) In the present invention, the fuzzy clustering-based track data cleaning method does not lose information: Compared with traditional outlier processing methods, the fuzzy clustering-based methods do not delete or replace outliers, so they do not lose useful information. Instead, they can mark outliers by assigning metrics for reference in further analysis.

[0038] (3) In the present invention, the fuzzy clustering track data cleaning method can reduce bias: compared with the method of deleting or replacing outliers, the fuzzy clustering-based method can better handle uncertainty and complexity, thereby reducing the possibility of bias. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 This is a flow chart of a track data cleaning method based on fuzzy clustering proposed by the present invention;

[0040] Figure 2 Acquisition and display of original track data proposed by the present invention;

[0041] Figure 3 This is the data fitting diagram under the fuzzy clustering proposed by the present invention;

[0042] Figure 4 This is a data distribution diagram of longitude interpolation using the Newton interpolation method proposed in the present invention;

[0043] Figure 5 This is a data distribution diagram of longitude interpolation using the two-dimensional interpolation method proposed in the present invention;

[0044] Figure 6 This is a diagram of the specific steps of outlier processing based on fuzzy clustering proposed by the present invention. DETAILED DESCRIPTION

[0045] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0046] Example 1:

[0047] Reference Figure 1 and 6 ,A track data cleaning method based on fuzzy clustering, including the following steps:

[0048] Step 1: Obtain ADS-B raw data

[0049] Get ADS-B raw data. Get ADS-B data of several historical tracks from VariFlight global flight real-time tracking radar platform. Each track contains several point tracks P i , each point contains some feature information Q i The relationship between the track L, the point track Pi and the characteristic items can be expressed as follows:

[0050]

[0051]

[0052] Among them, P i is the i-th point in the track, Q i P i The i-th feature information in .

[0053] Read valid data from ADS-B data. Each point in each track in the raw data obtained from VariFlight contains 9 feature items: Time (timestamp), UTC TIME (UTC time), Anum (aircraft registration number), Fnum (flight number), Height (altitude), Speed ​​(speed), Angle (azimuth), Longitude (longitude), Latitude (latitude). After processing, the invalid data (UTC TIME, Anum, Fnum) can be deleted. After the above steps, the track P i It can be expressed as:

[0054] P i ={timestamp, altitude, speed, azimuth, longitude, latitude}, see Figure 2 Step 2: Data deduplication. To facilitate subsequent outlier search and analysis, the data must be deduplicated in advance. This includes the following three steps:

[0055] S2.1 Sort the data according to chronological order.

[0056] This makes the order of data clearer, making it easier to perform subsequent data processing and retrieve and apply the processed output data.

[0057] Compare and sort each set of data according to the time series, so that each set of data in the data set can be accurately sorted according to time.

[0058] S2.2 Delete data of the same time

[0059] During the flight of the aircraft, points with the same time in the row set must be caused by errors, so the track points with the same timestamp should be deleted to avoid unnecessary impact on the actual data and make the data inaccurate.

[0060] Search for data with the same time point according to the time series, and delete them from the original data set after finding them to ensure the uniqueness and accuracy of the timestamp.

[0061] S2.3 Delete data with the same latitude and longitude

[0062] Because aircraft flight is unidirectional, as time goes by, there will no longer be a situation where the longitude and latitude of the track are the same. Therefore, it is necessary to delete track points with the same longitude and latitude. The longitude and latitude of each set of data are extracted separately, and the data with the same longitude and latitude are found and recorded. The entire set of data with the same longitude and latitude is deleted to avoid errors.

[0063] Step 3: Deletion of missing items

[0064] To prevent incomplete data, data with missing items should be deleted. First, search for data groups with missing items and delete the found data groups to prevent them from affecting the processing of other data.

[0065] Table 1 Examples of duplicate and missing track points

[0066]

[0067] As shown in Table 1, the first and second data sets have the same characteristic item: speed, which is a speed duplication problem; the fourth and fifth data sets have the same characteristic item: longitude, which is a location duplication problem; the sixth data set is missing azimuth characteristic information, which is a missing information problem; the third and seventh data sets are normal. Therefore, the second, fifth, and sixth data sets are deleted, and the results after deduplication and deletion are as follows:

[0068] Table 2 Example of track points after deduplication and deletion

[0069]

[0070] Step 4: Detect and analyze outliers using fuzzy clustering

[0071] The outlier processing here uses the density clustering method. The specific implementation process and ideas are summarized as follows: Suppose there are several points in a two-dimensional plane, and we want to cluster these points. Then the density clustering process is as follows: For each point, we need to calculate the two variables of "local density" and "local distance" of this point.

[0072] The definition is as follows: Local density: Local density depends on the dc value. The dc value is a hyperparameter manually specified before density clustering. It is a range distance. Then the local density of a point is the number of neighboring points within the dc range around this point (that is, the circle with this point as the center and dc as the radius).

[0073] Regarding the selection of dc value, it is generally a value that makes the average number of neighbors of all data points 1-2% of the total data points. The relevant definitions are as follows: Local distance: Assuming that we have now obtained the local density of all points, for the local distance of each point:

[0074] (1) For the point with the highest local density (just that one point), its local distance is the maximum distance between that point and all other points, that is, the distance to the point farthest from it.

[0075] (2) For the point with the largest non-local density (all the remaining points), its local distance is: the distance to the nearest point among all points with higher local density than it.

[0076] Total distance: At this point, the total distance of each point = local density + local distance. If the data is ultimately divided into three clusters, the first three points with the largest total distances will be the three centers, and the remaining points will be assigned to the cluster to whichever point they are closest to. (So, "total distance" can also be thought of as "total weight": the greater the total weight, the more qualified it is to become a cluster center.)

[0077] Take the group with the most data around the center as the data we need and discard the rest.

[0078] In the actual operation process, the latitude and longitude, altitude and speed are fuzzy clustered with time as the independent variable, and the abnormal values ​​are eliminated. Figure 3 ;

[0079] Step 5: Use the two-dimensional interpolation method to extract data with appropriate spacing as the final result of data cleaning: use the two-dimensional function formula for fitting, and use interpolation to convert the data into points with the same spacing to facilitate subsequent operations; linearize the data, and take points with the same time distance on a line as the points needed for subsequent extraction. Other data at the same time are also modified in the same way as the final result of data cleaning for subsequent data calls. Figure 4 and 5 .

[0080] In the present invention, it can be seen that the outlier processing based on fuzzy clustering has high accuracy, can greatly improve the accuracy and stability of data, thereby making the data more valuable and improving various performances of data preprocessing.

[0081] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A track data cleaning method based on fuzzy clustering, characterized in that: The following steps are involved: Step 1: Get ADS-B raw data. Get ADS-B raw data from VariFlight global flight real-time tracking radar platform to obtain ADS-B data of several historical tracks. Each track contains several points P i , each point contains some feature information Q i , where the relationship between the track L, point track Pi and the characteristic items can be expressed as: Among them, P i is the i-th point in the track, Q i P i The i-th characteristic information in the ADS-B data is read, and the trace P after the above steps is processed i It can be expressed as: P i ={timestamp, altitude, speed, azimuth, longitude, latitude}; Step 2: Data deduplication. To facilitate subsequent outlier search and analysis, data deduplication must be performed in advance. Step 3: Deletion of missing items. To prevent incomplete data, data with missing items should be deleted. First, search for data groups with missing items and delete the found data groups to prevent them from affecting the processing of other data. Step 4: Use fuzzy clustering to detect and analyze outliers. The density clustering method is used for outlier processing. The specific implementation process and ideas are as follows: Assume that there are several points in a two-dimensional plane and cluster these points. The density clustering process is as follows: For each point, calculate the two variables of "local density" and "local distance" of this point; Step 5: Use the two-dimensional interpolation method to extract data with appropriate spacing as the final result of data cleaning: use the difference table, difference quotient table and squat interpolation formula for fitting, and use interpolation to convert the data into points with the same spacing to facilitate subsequent operations; linearize the data and take points with the same time distance on a line as the points needed for subsequent extraction.

2. The track data cleaning method based on fuzzy clustering according to claim 1, characterized in that: In the raw data obtained from VariFlight in step 1, each point in each track contains nine feature items: Time (timestamp), UTC (UTC time), Anum (aircraft registration number), Fnum (flight number), Height (altitude), Speed ​​(speed), Angle (azimuth), Longitude (longitude), and Latitude (latitude). Invalid data (UTC time, Anum, and Fnum) can be deleted after processing.

3. The track data cleaning method based on fuzzy clustering according to claim 1, characterized in that: The data deduplication in step 2 specifically includes the following steps: S2.1 Sort the data in chronological order, making the order of the data clearer and easier to process and retrieve the output data after processing. Compare and sort each set of data according to the time series, so that each set of data in the data set can be accurately sorted according to time; S2.2 Delete data with the same time. During the flight of an aircraft, points with the same time in the row set must be caused by errors. Therefore, track points with the same timestamp should be deleted to avoid unnecessary impact on the actual data and make the data inaccurate. Search for data with the same time point according to the time series, and delete them from the original data set after finding them to ensure the uniqueness and accuracy of the timestamp; S2.3 Delete data with the same longitude and latitude. Because aircraft flight is unidirectional, as time goes by, there will no longer be a situation where the longitude and longitude of the track are the same. Therefore, it is necessary to delete track points with the same longitude and latitude. Extract the longitude and latitude of each set of data separately, find the data with the same longitude and longitude and record its serial number, and delete the entire set of data with the same longitude and longitude to avoid errors.

4. The track data cleaning method based on fuzzy clustering according to claim 1, characterized in that: The local density in step 4 depends on the dc value. The dc value is a hyperparameter manually specified before density clustering. It is a range distance. Then the local density of a point is the number of neighboring points within the dc range around this point, that is, the number of neighboring points of the circle with this point as the center and dc as the radius.

5. The track data cleaning method based on fuzzy clustering according to claim 4 is characterized in that: The dc value is generally selected so that the average number of neighbors of all data points is 1-2% of the total data points.

6. The track data cleaning method based on fuzzy clustering according to claim 1, characterized in that: The local distance in step 4: Assuming that the local density of all points has been obtained, for the local distance of each point: for the point with the largest local density, its local distance is the maximum value of the distance between this point and all other points, that is, the distance to the point farthest from it; for the point with the largest non-local density, its local distance is: the distance to the point closest to it among all points with higher local density than it.

7. The track data cleaning method based on fuzzy clustering according to claim 1, characterized in that: The total distance in step 4 is: the total distance of each point = local density + local distance; if the data is ultimately to be divided into three categories, then the first three points with the largest total distance are the three center points, and the remaining points are classified into the category to whichever point they are closer to; the group with the most data around the center is taken as the required data, and the rest of the data is discarded; in the actual operation process, fuzzy clustering is performed on latitude and longitude, altitude, and speed with time as the independent variable to eliminate outliers.

8. The track data cleaning method based on fuzzy clustering according to claim 1, characterized in that: The other data at the same time in step 5 are also modified in the same way as the final result of data cleaning to facilitate subsequent data calls.

Citation Information

Patent Citations

  • ADS-B track cleaning and calibration method based on local traversal density clustering

    CN110362559A

  • ADS-B flight path data cleaning method based on fuzzy clustering

    CN113254432A