An aircraft trajectory clustering method in the terminal area based on area division and cosine distance
By combining the DBSCAN algorithm based on area division and cosine distance, the accuracy and efficiency problems of the existing track clustering methods when dealing with speed differences and flight direction differences are solved, and more efficient and accurate track clustering results are achieved.
Patent Information
- Application Number
- CN202210356450.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-06
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-04-06
AI Technical Summary
The existing track clustering method has reduced accuracy when processing aircraft speed differences, and failed to effectively consider flight direction differences, has high time complexity, and is inefficient in processing big data.
The terminal area aircraft trajectory clustering method based on area division and cosine distance is adopted, and the area and cosine distance between tracks are calculated by triangular segmentation method, and clustering is combined with the DBSCAN algorithm to reduce the time complexity and improve the rationality of the clustering results.
It improves the accuracy of track clustering, can effectively distinguish tracks with significantly different flight directions, reduces time complexity, improves the efficiency of processing big data, and makes the results more objective through comprehensive clustering performance index.
Smart Images

Figure CN114742150B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of aircraft flight track data research, and particularly to a method for clustering aircraft trajectories in the terminal area based on area division and cosine distance. Background Art
[0002] With the rapid development of civil aviation in China, the air traffic flow continues to grow, and the operating environment in the terminal area is becoming increasingly complex. The existing arrival and departure route design and sector division can no longer meet the requirements, and it is necessary to make optimization adjustments according to the actual operating conditions of aircraft. Therefore, mining representative traffic flows from flight track data has become a research hotspot in the current air traffic field. And flight track clustering is an important technical means to grasp the operating characteristics of aircraft and mine prevalent traffic flows. Flight track clustering divides flight tracks with high similarity by measuring the similarity between flight tracks, and its results can provide technical support for abnormal flight track detection and flight track prediction, provide a basis for arrival and departure route design and sector division, and have certain reference value for reasonably allocating airspace resources in the terminal area and improving the operating efficiency of the terminal area.
[0003] In the clustering process, how to define the similarity between flight tracks has a significant impact on the clustering results. According to the requirements for the number of track points of each flight track, it can be divided into two categories. The first category requires the same number of track points. The flight tracks are pre-processed such as resampling, and then methods such as Euclidean distance are used to calculate the similarity between flight tracks; the other is to use methods that can directly measure flight tracks of different lengths as similarity models, such as Hausdorff distance, dynamic time warping, etc.
[0004] However, the deficiencies of the above methods are reflected in: when the speed differences of aircraft are large, resulting in significant differences in the number of track points, the accuracy will decrease; only the spatial distance is considered, and the difference in flight direction is not considered; the time complexity is relatively high, and it is inefficient when dealing with big data. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for clustering aircraft trajectories in the terminal area based on area division and cosine distance, which analyzes aircraft trajectories considering both spatial distance and flight direction, improves the rationality of the trajectory analysis results, and reduces the time complexity. The technical solution adopted by the present invention is as follows.
[0006] On the one hand, the present invention provides a method for clustering aircraft trajectories in the terminal area based on area division and cosine distance, including:
[0007] Obtain the flight track data to be clustered;
[0008] Pre-process the flight track data to be clustered;
[0009] Based on the preprocessed track data, use the triangular segmentation method based on area division to calculate the area between different tracks, and determine the first track-to-track distance considering space according to the area between tracks;
[0010] Based on the preprocessed track data, calculate the cosine distance between different tracks, and determine the second track-to-track distance considering space and direction according to the first track-to-track distance and the cosine distance, and obtain a distance matrix including the second track-to-track distances between all different tracks;
[0011] Based on the distance matrix, use the DBSCAN algorithm to cluster the track data to obtain a track clustering result.
[0012] Optionally, the preprocessing of the track data to be clustered includes: deleting the data in the track that is not within the terminal area, deleting the taxiing data in the track, and deleting the track with irregular changes between track points.
[0013] Optionally, the preprocessing of the track data to be clustered further includes: converting the longitude and latitude coordinate data in the radar track data into plane coordinate data, and the conversion formula is:
[0014]
[0015]
[0016]
[0017] where R is the radius of the earth, lon and lat are the longitude and latitude coordinates of the track point respectively, Lon and Lat are the longitude and latitude coordinates of the coordinate reference point respectively, (x, y) is the plane coordinate after converting the longitude and latitude coordinates (lon, lat), and a is the intermediate parameter.
[0018] Optionally, the use of the triangular segmentation method based on area division to calculate the area between different tracks and determine the track-to-track distance includes:
[0019] For any two tracks p i ={p i1 , p i2 ,..., p im} and p j ={p j1 , p j2 ,..., p jm}, starting from the track start point, sequentially select the connection lines of the track points between the two tracks until all the track points of at least one track are connected, obtaining a plurality of non-overlapping triangles, and iteratively calculate the area SA between the two tracks through the following formula:
[0020]
[0021] area(·) represents the area of a triangle with the three points in the parentheses as the three corner points, dist(·) represents the length of the line connecting the two track points in the parentheses, SA m represents the area of all the connected triangles obtained from the m-th iterative calculation. The initial value of SA is 0, m is a non-zero positive integer, p ki and p gj respectively represent the k-th track point in p i and the g-th track point in p j ;
[0022] Use the following formula to determine the track-to-track distance APL based on the area between tracks:
[0023] APL = SA / Length
[0024] where Length is the average length of the parts of the two tracks involved in the area calculation.
[0025] Optionally, calculating the cosine distance between different tracks based on the preprocessed track data includes:
[0026] For any two tracks, calculate their cosine distance according to the following formula:
[0027]
[0028] In the formula, V i and V j are respectively the vectors formed by the starting and ending points of the two tracks p i , p j respectively, and θ ij is the included angle between V i and V j .
[0029] Optionally, determining the second track-to-track distance r considering space and direction based on the first track-to-track distance and the cosine distance ij , the formula is:
[0030]
[0031] In the formula, M represents an infinite number;
[0032] The distance matrix of the second track-to-track distances between all different tracks is:
[0033]
[0034] Optionally, based on the distance matrix, using the DBSCAN algorithm to cluster the track data to obtain the track clustering result, including:
[0035] Set the initial values of the input parameters of the DBSCAN algorithm. The output parameters include the neighborhood range parameter Eps and the minimum number of tracks parameter MinPts in the neighborhood.
[0036] Based on the currently set input parameters, use the DBSCAN algorithm to traverse the terminal area track set P = {p 1 , p 2 ,..., p n}, and find all core tracks according to the distance matrix R.
[0037] Connect and aggregate all core tracks into track clusters according to the neighborhood range parameter Eps, and add the line segments connected to the density of each track cluster in the non-core tracks to the corresponding track clusters to obtain the clustering result.
[0038] Optionally, using the DBSCAN algorithm to cluster the track data to obtain the track clustering result further includes:
[0039] Take the track clustering result obtained by clustering with the initial values of the input parameters of the DBSCAN algorithm as the initial clustering result. According to the initial clustering result, calculate the silhouette coefficient SC and the Davies-Bouldin index DBI. The formulas are as follows:
[0040]
[0041]
[0042] In the formula, a(p i ) is the average distance from track p i to other tracks in its belonging cluster; b(p i ) is the average distance from track p i to all tracks in the cluster closest to its current cluster; S i is the average distance from all tracks in the i-th cluster to the cluster center; M ij is the distance between the centers of two clusters; n is the total number of tracks.
[0043] Calculate the comprehensive clustering performance index CCPM according to the silhouette coefficient SC and the Davies-Bouldin index DBI. The formula is:
[0044] CCPM = SC + 1 / DBI
[0045] Judge whether CCPM meets the preset threshold range and whether there are tracks that are significantly different but classified into the same category. If CCPM meets the preset threshold range and there are no tracks that are significantly different but classified into the same category, then take the corresponding clustering result as the final clustering result. Otherwise, adjust the input parameters of the DBSCAN algorithm and re-cluster the tracks according to the distance matrix R.
[0046] Optionally, the use of the DBSCAN algorithm to cluster the track data to obtain a track clustering result further includes:
[0047] Set initial values of input parameters of multiple groups of DBSCAN algorithms to obtain corresponding initial clustering results respectively;
[0048] For each initial clustering result, calculate the silhouette coefficient SC and the Davies-Bouldin index DBI respectively, and calculate the comprehensive clustering performance index CCPM according to the silhouette coefficient SC and the Davies-Bouldin index DBI;
[0049] Select the clustering result obtained by the DBSCAN algorithm with the largest comprehensive clustering performance index CCPM as the final clustering result.
[0050] Optionally, the method further includes: for clusters that are difficult to distinguish in the clustering result, adjust the parameters of the DBSCAN algorithm and perform clustering again.
[0051] In a second aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for clustering aircraft trajectories in the terminal area based on area division and cosine distance as introduced in the first aspect.
[0052] Beneficial effects
[0053] Compared with the prior art, the present invention has the following advantages and improvements:
[0054] 1. For tracks that are relatively similar near the runway but have significantly different flight directions, traditional similarity measurement methods may classify them into one category. However, the present invention not only considers the spatial distance of the tracks but also considers the flight direction for cosine distance calculation, which can distinguish the significant differences in direction between the two tracks, avoid classifying tracks with significantly different directions into one category, and improve the clustering accuracy;
[0055] 2. Compared with traditional similarity measurement methods based on track points, the present invention can reduce the time complexity from T(n) = O(n×m) to T(n) = O(n + m), and the triangular segmentation method based on area division has better robustness in the case of unequal numbers of track points, without the need to perform operations such as resampling the tracks in advance;
[0056] 3. Using the DBSCAN algorithm for clustering, different clustering results can be obtained by adjusting parameters, with strong flexibility and controllability, and it can identify radar-guided tracks with a small number of tracks;
[0057] 4. Introducing the comprehensive clustering performance index, combining the silhouette coefficient and the Davies-Bouldin index, facilitates the selection of appropriate input parameters and clustering results, making the selection of results more objective;
[0058] 5. Re - clustering the difficult - to - divide tracks can optimize the results, making the clustering results more refined and interpretable. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 The figure shows a schematic flow chart of an embodiment of the method for clustering aircraft trajectories in the terminal area of the present invention;
[0060] Figure 2 (a), (b), and (c) are respectively the clustering result diagrams for three combinations of input parameters;
[0061] Figure 3 (a), (b), (c), (d), and (e) are respectively the decomposition diagrams for each class of the clustering results;
[0062] Figure 4 (a) is a difficult - to - divide track cluster, Figure 4 (b) and (c) are respectively the result diagrams of re - clustering. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0063] The following is further described in conjunction with the drawings and specific embodiments.
[0064] Embodiment 1
[0065] This embodiment introduces a method for clustering aircraft trajectories in the terminal area based on area division and cosine distance, including:
[0066] Obtaining the track data to be clustered;
[0067] Pre - processing the track data to be clustered;
[0068] Based on the pre - processed track data, using the triangular segmentation method based on area division to calculate the area between different tracks, and determining the first track - to - track distance considering space according to the area between tracks;
[0069] Based on the pre - processed track data, calculating the cosine distance between different tracks, and determining the second track - to - track distance considering space and direction according to the first track - to - track distance and the cosine distance, obtaining a distance matrix including the second track - to - track distances between all different tracks;
[0070] Based on the distance matrix, using the DBSCAN algorithm to cluster the track data to obtain the track clustering result.
[0071] Refer to Figure 1 , and the following specifically introduces each part.
[0072] I. Track Data Pre - processing
[0073] In this embodiment, the terminal area track data includes approach, departure and overflight track data, and the preprocessing of the track data includes:
[0074] (1.1) Delete track point data that is not in the terminal area;
[0075] (1.2) Delete any two track points with irregular changes that are obviously different from the normal track in the adjacent time period;
[0076] (1.3) In order to calculate the area between tracks and visualize the clustering results, the longitude and latitude coordinates in the radar track data are converted into plane coordinates x and y with the airport reference point as the coordinate origin using the Mercator projection:
[0077]
[0078]
[0079]
[0080] Among them, R is the radius of the earth, lon and lat are the longitude and latitude coordinates of the track point, and Lon and Lat are the longitude and latitude coordinates of the airport reference point;
[0081] (1.4) Separate the arrival and departure data to facilitate separate processing of arrival and departure data;
[0082] (1.5) Since the complete track data is collected, including the data of taxiing to the parking position after landing, if it is used directly, it will greatly increase the amount of calculation data and affect the accuracy of the results. Therefore, it is necessary to eliminate redundant data according to the altitude. The aircraft in the data landed on runways 16L, 16R and 17R. The initial approach point altitude of the landing procedure for runways 16L and 16R is 600m, and the initial approach point altitude of the landing procedure for runway 17R is 900m. In order to eliminate the taxiing data and ensure that too much data is not lost due to excessive interception altitude, the track data above 600m is retained.
[0083] 2. Calculate the first track distance in the considered space
[0084] The present invention calculates the areas between different tracks based on a triangulation method of area division, and then determines the first inter-track distance of the considered space according to the inter-track area.
[0085] Assuming the track p i It is composed of several track point data, represented by p i ={p i1 ,p i2 ,...,p im}, track p j ={p j1, p j2 ,..., p jm}, points a and b respectively belong to p i and p j and respectively loop through each track point in p i and p j ;
[0086] At the beginning of the loop, a and b are the starting points of two tracks, and the line connecting a and b is side L 1 ,;
[0087] Compare the lengths of the line connecting a and point (b + 1) and the line connecting b and point (a + 1), and take the shorter one as L 2 ;
[0088] Determine L 1 and L 2 After determining L 3 and L
[0089] the third side L of the triangle can be determined, and then the area of the triangle can be calculated
[0090] Then a or b traverses to the next track point (a + 1) or (b + 1), and repeat the previous method for determining the triangle side. When a or b traverses to the end of the track, the loop ends
[0091]
[0092] The calculation formula for the area SA between two tracks is: k , a g , a k+1 ) represents the area of the triangle enclosed by a k , b g and a k+1 ;
[0093] The above formula can also be written as:
[0094]
[0095] area(·) represents the area of the triangle with the three points in the parentheses as the three corner points, dist(·) represents the length of the line connecting the two track points in the parentheses, SA m represents the area of all the connected triangles obtained in the m-th iteration calculation. The initial value of SA is 0, m is a non-zero positive integer, p ki and p gj respectively represent the k-th track point in p i and the g-th track point in p j ;
[0096] Finally, the area per length (APL) is obtained, which is the track distance:
[0097] APL = SA / Length
[0098] where Length is the average length of the parts of the two tracks involved in the calculation.
[0099] III. Calculate the cosine distance between tracks
[0100] The cosine distance, also known as cosine similarity, measures the similarity between two vectors by calculating the cosine value of the angle between them. Define the vector V as formed by connecting the starting and ending points of the track. The cosine distance calculation formula for two tracks is:
[0101]
[0102] where V i and V j are the vectors formed by the starting and ending points of track p i and track p j respectively, and θ ij is the angle formed by V i and V j respectively.
[0103] When the angle between the two tracks is within the range of π / 2, the APL representing the spatial distance between the tracks and the cosine distance measuring the difference in flight directions between the tracks are combined, and the ratio of APL to the cosine distance is used to represent the distance between the two tracks, making the distance between tracks with a larger difference in flight directions larger. When θ ij ≥ π / 2, that is, cosθ ij ≤ 0, it is considered that there is an obvious deviation in the flight directions of track p i and p j , and the distance can be set to M, where M is infinity. The distance between track p i and p j is defined as:
[0104]
[0105] At this time, the distance matrix between the tracks can be obtained by combining the aforementioned first track distance:
[0106]
[0107] IV. Track clustering
[0108] Use the distance matrix as the input of the DBSCAN algorithm to cluster the tracks.
[0109] The two input parameters of DBSCAN are the neighborhood range parameter Eps and the minimum number of tracks in the neighborhood parameter MinPts. Under the allowable experimental conditions, the range of the input parameters should be as large as possible. After determining the input parameters and the distance matrix, the use of the DBSCAN algorithm for clustering can refer to the existing technology and will not be elaborated here.
[0110] After multiple experiments, it is found that when Eps is in the range of 4000 to 5000, the clustering effect is better. There is a guiding principle for the selection of MinPts, that is, MinPts ≥ dim + 1, where dim represents the dimension of the data to be clustered.
[0111] Set the hyperparameter grid of the DBSCAN algorithm as shown in Table 1.
[0112] Table 1 Hyperparameter grid of the DBSCAN algorithm
[0113]
[0114] The definitions used in the DBSCAN algorithm combined with track data are as follows:
[0115] Eps neighborhood: The Eps neighborhood of track p i refers to the area within the terminal area traffic flow set P with p i as the center and Eps as the radius, that is:
[0116] N Eps (p i ) = {p j ∈P|r ij ≤ Eps} (6)
[0117] Core track: Given the neighborhood range parameter Eps and the minimum number of tracks in the neighborhood parameter MinPts, if the number of tracks within N Eps (p i ) ≥ MinPts, then p i is considered a core track.
[0118] Density direct reach: If there exists a core track p i , p j is included in N Eps (p i ), then p j is considered to be density directly reachable from p i .
[0119] Density reachable: If there exists density direct reach from p i to p k , and density direct reach from p k to p j is also density direct reach, then it is considered that from p i to p jDensity can reach.
[0120] Density connected: If there is p j and p i At the same time from p k The density can be reached, then it is considered that p j and p i Density connected.
[0121] Noise point: A track that does not belong to any category is a noise point.
[0122] In this embodiment, based on the distance matrix, the DBSCAN algorithm is used to cluster the track data to obtain the track clustering results, including:
[0123] Set the initial values of the input parameters of the DBSCAN algorithm, wherein the output parameters include the neighborhood range parameter Eps and the neighborhood minimum track parameter MinPts;
[0124] Based on the currently set input parameters, the DBSCAN algorithm is used to traverse the terminal area track set P = {p 1 ,p 2 ,...,p n}, find all core tracks according to the distance matrix R;
[0125] According to the neighborhood range parameter Eps, all core tracks are connected and aggregated into track clusters, and the line segments in non-core tracks that are connected to various track cluster densities are added to the corresponding track clusters to obtain the clustering results.
[0126] In order to ensure the reliability of the clustering results of the DBSCAN algorithm, this embodiment can pre-set multiple groups of DBSCAN algorithm input parameter combinations to obtain corresponding clustering results respectively, and select the better clustering results from them. It can also adjust the input parameters of the DBSCAN algorithm after one clustering and cluster again until the clustering result meets the requirements.
[0127] For clusters that are difficult to distinguish in the clustering results, the parameters of the DBSCAN algorithm can be adjusted to perform clustering again separately.
[0128] V. Evaluation of clustering results
[0129] In this embodiment, the evaluation of the clustering results is achieved by calculating the following indicators.
[0130] Calculate the silhouette coefficient SC, the formula is as follows:
[0131]
[0132] Among them, a(p i ) is the track p i The average distance to other tracks in the cluster to which it belongs; b(p i) is the track p i The average distance to all tracks within the cluster closest to the cluster where it is located.
[0133] Calculate the Davies-Bouldin index DBI, and its formula is as follows:
[0134]
[0135] Among them, S i is the average distance from all tracks in the i-th cluster to the cluster center; M ij is the distance between the centers of two clusters; n is the total number of tracks.
[0136] Introduce the comprehensive clustering performance index CCPM, and its formula is as follows:
[0137] CCPM = SC + 1 / DBI
[0138] Theoretically, the larger the SC value, the smaller the DBI value, and the larger the CCPM value, the better the clustering effect.
[0139] In the experiment, the three groups of parameter combinations with better performance are shown in Table 2.
[0140] Table 2 Results of clustering evaluation indicators
[0141]
[0142] Figure 2 is the clustering result diagram of the three groups of parameter combinations. The coordinate system in the figure takes the reference point of Shanghai Pudong International Airport as the origin of coordinates, the due east direction as the positive direction of the x-axis, and the due north direction as the positive direction of the y-axis. Figure 3 is Figure 2 (a) The decomposition diagram of each category, and the track with a triangle represents the standard approach procedure. Thus, the c-th group of input parameters with the largest CCPM value can be selected to obtain a better clustering result.
[0143] For Figure 3 (e), for the clusters that are difficult to distinguish, in order to obtain a more significant partitioning effect, the re-clustering method is used to optimize the result, and this part of the tracks is re-clustered separately. Figure 4 is the clustering result when Eps = 1600 and MinPts = 3.
[0144] In summary of the above examples, it can be seen that the present invention accurately divides the tracks along different approach procedures, and can identify abnormal tracks, which helps to optimize the arrival and departure procedures, analyze the reasons for the generation of abnormal tracks, rationally allocate the airspace resources in the terminal area, and improve the operation efficiency of the terminal area.
[0145] Example 2
[0146] Based on the same inventive concept as in Embodiment 1, this embodiment introduces a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for clustering aircraft trajectories in the terminal area based on area division and cosine distance as introduced in Embodiment 1.
[0147] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0148] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0149] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0150] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0151] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims. All of these fall within the protection scope of the present invention.
Claims
1. A method for clustering aircraft trajectories in the terminal area based on area division and cosine distance, characterized in that, it includes: Obtain the trajectory data to be clustered; Preprocess the trajectory data to be clustered; Based on the preprocessed trajectory data, use the triangular segmentation method based on area division to calculate the area between different trajectories, and determine the first distance between trajectories considering space according to the area between trajectories; Based on the preprocessed trajectory data, calculate the cosine distance between different trajectories, and determine the second distance between trajectories considering space and direction according to the first distance between trajectories and the cosine distance, and obtain a distance matrix including the second distances between all different trajectories; Based on the distance matrix, use the DBSCAN algorithm to cluster the trajectory data to obtain a trajectory clustering result; Among them, the step of using the DBSCAN algorithm to cluster the trajectory data based on the distance matrix to obtain a trajectory clustering result includes: Set the initial values of the input parameters of the DBSCAN algorithm, where the input parameters include the neighborhood range parameter and the minimum number of tracks in the neighborhood parameter ; Based on the currently set input parameters, use the DBSCAN algorithm to traverse the track set in the terminal area , according to the distance matrix find all core tracks; According to the neighborhood range parameter All core tracks are connected and aggregated into track clusters, and the line segments in the non-core tracks that are density-connected to various track clusters are added to the corresponding track clusters to obtain the clustering result; Take the trajectory clustering result obtained by clustering using the initial value of the input parameters of the DBSCAN algorithm as the initial clustering result, and according to the initial clustering result, calculate the silhouette coefficient SC and the Davies-Bouldin index DBI, and the formulas are respectively: , , Wherein, is the average distance from the track to other tracks within its affiliated cluster; is the average distance from the track to all tracks within the cluster closest to its current cluster; is the average distance from all tracks in the th cluster to the cluster center; is the distance between the centers of two clusters; is the total number of tracks; Calculate the comprehensive clustering performance index CCPM according to the silhouette coefficient SC and the Davies-Bouldin index DBI, and the formula is: , Judge whether the CCPM meets the preset threshold range and whether there are tracks that are significantly different but classified into the same category. If the CCPM meets the preset threshold range and there are no tracks that are significantly different but classified into the same category, then use the corresponding clustering result as the final clustering result; otherwise, adjust the input parameters of the DBSCAN algorithm and then re-perform track clustering based on the distance matrix for track clustering; Or, Set multiple groups of initial values of the input parameters of the DBSCAN algorithm, and obtain the corresponding initial clustering results respectively. Among them, for the case where the clusters in the initial clustering results are difficult to distinguish, adjust the parameters of the DBSCAN algorithm and cluster again; For each initial clustering result, calculate the silhouette coefficient SC and the Davies-Bouldin index DBI respectively, and calculate the comprehensive clustering performance index CCPM according to the silhouette coefficient SC and the Davies-Bouldin index DBI; Select the clustering result obtained by the DBSCAN algorithm with the largest comprehensive clustering performance index CCPM as the final clustering result.
2. The method according to claim 1, characterized in that, The preprocessing of the trajectory data to be clustered includes: deleting the data not in the terminal area in the trajectory, deleting the taxiing data in the trajectory, and deleting the trajectories with irregular changes between trajectory points.
3. The method according to claim 1, characterized in that, The preprocessing of the trajectory data to be clustered further includes: converting the longitude and latitude coordinate data in the radar trajectory data into plane coordinate data, and the conversion formula is: , , , In the formula, is the radius of the earth, are the longitude and latitude coordinates of the track point respectively, are the longitude and latitude coordinates of the coordinate reference point respectively, is for the longitude and latitude coordinates after conversion to the plane coordinates, is an intermediate parameter.
4. The method according to claim 1, characterized in that, The step of using the triangular segmentation method based on area division to calculate the area between different trajectories and determining the distance between trajectories according to the area between trajectories includes: For any two tracks and , starting from the track starting points, successively select the connecting lines of the track points between the two tracks until all the track points of at least one track are connected, obtaining multiple non-overlapping triangles, and iteratively calculate the area between the two tracks through the following formula : , represents the area of a triangle with the three points in parentheses as the three corner points, represents the length of the line segment connecting the two track points in parentheses, represents the total area of all the connected triangles obtained from the th iterative calculation, and the initial value of is 0, and respectively represent the th track point in and the th track point in ; Determine the distance between tracks based on the area between tracks using the following formula : , Among them, is the average length of the part where two tracks participate in the area calculation.
5. The method according to claim 1, characterized in that, The step of calculating the cosine distance between different trajectories based on the preprocessed trajectory data includes: For any two trajectories, calculate the cosine distance between them according to the following formula: , In the formula, and are respectively vectors formed by the starting and ending points of two respective tracks of and , is and 's included angle.
6. The method according to claim 5, characterized in that, According to the distance between the first tracks and the cosine distance, determine the distance between the second tracks considering space and direction , and the formula is: , In the formula, represents an infinite number; The distance matrix of the second distances between all different trajectories is: 。 7. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, it implements the method for clustering aircraft trajectories in the terminal area based on area division and cosine distance as described in any one of claims 1-6.
Citation Information
Patent Citations
Indoor people clustering method based on multiple labels
CN107944498A
Hurricane movement track trend calculation method based on MDL and speed direction
CN112633389A