Method for determining the track running state of a mobile target based on clustering
By performing multidimensional feature clustering and intra-cluster distribution parameter analysis on the tracks, the nearest neighbor set and deviation probability of the tracks are determined, which solves the problem of low accuracy in abnormal track identification in existing technologies and improves the accuracy and security of abnormal track identification.
Patent Information
- Application Number
- CN202511650838.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-12
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2045-11-12
AI Technical Summary
The accuracy of identifying abnormal behavior states in existing technologies is low, leading to safety hazards for maritime and aerial targets.
The clustering-based method clusters tracks using multidimensional features, determines the first nearest neighbor track set using the intra-cluster distribution parameters of the track clusters, and calculates the deviation probability to identify abnormal operating states.
It improves the accuracy of abnormal flight path identification, reduces false positives and false negatives, and enhances the safety management capabilities of maritime and aerial targets.
Smart Images

Figure CN121117899B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of anomaly detection technology, and more specifically, to a method for determining the trajectory and operational status of a moving target based on clustering. Background Technology
[0002] Analyzing the behavioral patterns and operational rules of mobile targets at sea and in the air can provide some auxiliary decision-making support for ensuring the safety management of the corresponding maritime and air areas.
[0003] In the process of realizing the concept of this invention, it was found through research that the accuracy of abnormal behavior state track identification is low in related technologies. Summary of the Invention
[0004] In view of this, the present invention provides a method for determining the trajectory and operational status of a moving target based on clustering.
[0005] One aspect of the present invention provides a method for determining the trajectory operation status of a moving target based on clustering, comprising: clustering multiple trajectories according to their respective multidimensional features to obtain multiple trajectory clusters; determining a first nearest neighbor trajectory set for each of the multiple trajectory clusters from the trajectory clusters to which each of the multiple trajectory clusters belongs or from the trajectory clusters closest to each of the multiple trajectory clusters, wherein the intra-cluster distribution parameters include at least one of the following: number of trajectories, speed information, and trajectory cluster density; obtaining the deviation probability of each of the multiple trajectory clusters based on the first nearest neighbor trajectory set of each of the multiple trajectory clusters; and determining the operation status of each of the multiple trajectory clusters based on a predetermined deviation threshold and the deviation probability of each of the multiple trajectory clusters.
[0006] According to an embodiment of the present invention, the track is a first track or a second track, and the track cluster is a first track cluster or a second track cluster. The first track cluster includes at least one first track, and the second track cluster includes at least one second track. The first track cluster and the second track cluster satisfy at least one of the following conditions: the track cluster density of the first track cluster is greater than the track cluster density of the second track cluster; the number of tracks in the first track cluster is greater than the number of tracks in the second track cluster; and / or the operational stability of the first track cluster is higher than the operational stability of the second track cluster. Speed information indicates the operational stability of the track cluster. Specifically, based on the intra-cluster distribution parameters of each of the multiple track clusters, the data is obtained from the track cluster to which each of the multiple tracks belongs or from the track cluster to which each of the multiple tracks belongs. Determining the first nearest neighbor set of multiple tracks within the nearest track cluster includes: dividing the multiple track clusters into a first track cluster set and a second track cluster set based on the intra-cluster distribution parameters and a preset distribution threshold, wherein the first track cluster set includes at least one first track cluster, and the second track cluster set includes at least one second track cluster; determining the first nearest neighbor set of each of the at least one first track included in the first track cluster set based on the first track cluster set and a first predetermined distance threshold; and determining the first nearest neighbor set of each of the at least one second track included in the second track cluster set based on the first track cluster set, the second track cluster set, and the second predetermined distance threshold.
[0007] According to an embodiment of the present invention, determining the first nearest neighbor set of each of at least one first track included in the first track cluster set, based on a first track cluster set and a first predetermined distance threshold, includes: for any first track cluster in the at least one first track cluster included in the first track cluster set, and for any first track in the plurality of first tracks included in any first track cluster, determining the first nearest neighbor set of any first track based on the first predetermined distance threshold and the distance between any first track in any first track cluster and other first tracks; determining the first nearest neighbor set of each of at least one second track included in the second track cluster set, based on the first track cluster set, the second track cluster set, and the second predetermined distance threshold, includes: for any second track cluster in the at least one second track cluster included in the second track cluster set, and for any second track in the plurality of second tracks included in any second track cluster, determining the first nearest neighbor set of any second track based on the second predetermined distance threshold and the distance between each first track in the first track cluster closest to any second track cluster and any second track.
[0008] According to an embodiment of the present invention, the velocity information includes average velocity and average acceleration, and the preset distribution threshold includes a preset track cluster density threshold, a preset ratio threshold, a preset acceleration threshold, and a preset velocity threshold. Based on the intra-cluster distribution parameters of each track cluster and the preset distribution threshold, the multiple track clusters are divided into a first track cluster set and a second track cluster set, including: determining track clusters with a track cluster density greater than or equal to the preset track cluster density threshold as first track clusters, and assigning track clusters with a track cluster density less than the preset track cluster density threshold to a first undetermined cluster set; sorting the multiple first undetermined clusters according to the number of tracks in each of the multiple first undetermined clusters in the first undetermined cluster set to obtain a first undetermined cluster sequence, so as to obtain the current first undetermined cluster and the next first undetermined cluster from the first undetermined cluster sequence, wherein the ratio between the number of tracks in the current first undetermined cluster and the number of tracks in the next first undetermined cluster is greater than or equal to a preset ratio threshold. In the case of a first undetermined cluster, the current first undetermined cluster is designated as the boundary track cluster corresponding to the number of tracks. If the ratio between the number of tracks in the current first undetermined cluster and the number of tracks in the next first undetermined cluster is less than a preset ratio threshold, the operation of designating the next first undetermined cluster in the first undetermined cluster sequence as the new current first undetermined cluster is repeated until the ratio between the number of tracks in the new current first undetermined cluster and the number of tracks in the new next first undetermined cluster is greater than or equal to the preset ratio threshold. The new current first undetermined cluster is then designated as the boundary track cluster corresponding to the number of tracks. Based on the boundary track clusters corresponding to the number of tracks, the first undetermined cluster sequence is divided into a first track cluster and a second undetermined cluster set. The second undetermined cluster in the second undetermined cluster set whose average speed is greater than or equal to a preset speed threshold and whose average acceleration is less than or equal to a preset acceleration threshold is designated as the first track cluster. The second undetermined clusters in the second undetermined cluster set other than the first track cluster are designated as the second track clusters.
[0009] According to an embodiment of the present invention, obtaining the deviation probability of each of the multiple tracks based on their respective first nearest neighbor track sets includes: for any track among the multiple tracks, obtaining a first probability set distance of the track based on the distances between the track and each of the multiple first nearest neighbor tracks in the track's first nearest neighbor track set, wherein the first probability set distance represents the absolute deviation of the track relative to the first nearest neighbor track set; and for any first nearest neighbor track among the multiple first nearest neighbor tracks, obtaining a second probability set of the first nearest neighbor track based on the distances between the first nearest neighbor tracks and the first nearest neighbor track in the second nearest neighbor track set of the first nearest neighbor track. The distance is defined as follows: the second nearest neighbor track set includes at least one second nearest neighbor track of the first nearest neighbor track; the second probability set distance represents the absolute deviation of the first nearest neighbor track relative to the second nearest neighbor track set; the average probability set distance of the first nearest neighbor track set is determined based on the second probability set distances of the multiple first nearest neighbor tracks; the average probability set distance represents the average deviation of the multiple first nearest neighbor tracks in the first nearest neighbor track set relative to the second nearest neighbor track sets of the multiple first nearest neighbor tracks; and the deviation probability of the track is obtained based on the first probability set distance of the track and the average probability set distance of the first nearest neighbor track set of the first track.
[0010] According to an embodiment of the present invention, obtaining the deviation probability of a track based on the first probability set distance of the track and the average probability set distance of the first nearest neighbor track set includes: obtaining a local deviation factor of the track based on the first probability set distance of the track and the average probability set distance of the first nearest neighbor track set, wherein the local deviation factor represents the relative deviation of the track relative to the first nearest neighbor track set; obtaining a global outlier factor of the track cluster to which the track belongs based on the local deviation factor of the track, the number of first nearest neighbor tracks included in the first nearest neighbor track set, and a predetermined normalization factor, wherein the global outlier factor represents the overall deviation of the track; and obtaining the deviation probability of the track based on the local deviation factor of the track and the global outlier factor of the track cluster to which the track belongs.
[0011] According to an embodiment of the present invention, multiple tracks are clustered based on their respective multidimensional features to obtain multiple track clusters, including: inputting the multidimensional features of each of the multiple tracks into a trained encoder, which is a trained autoencoder, to obtain low-dimensional latent representations of each of the multiple tracks; for any track among the multiple tracks, determining the track cluster to which the track belongs based on the distance between the low-dimensional latent representation of the track and each of the multiple target cluster centers; wherein the multiple target cluster centers and the trained autoencoder are obtained by training multiple initial cluster centers and a pre-trained autoencoder based on multiple second sample tracks using a joint loss function, the joint loss function being determined based on a non-clustering loss function and a clustering loss function, and the pre-trained autoencoder being obtained by training an autoencoder based on a non-clustering loss function using multiple first sample tracks. The clustering loss function is determined based on the weighting function and the reconstruction loss function. The weighting function calculates the weighting coefficients of the sample track based on the sample track and its reconstructed sample track. The reconstruction loss function calculates the difference between the sample track and its reconstructed sample track, where the sample track includes either a first sample track or a second sample track, and the reconstructed sample track includes either a first reconstructed sample track of the first sample track or a second reconstructed sample track of the second sample track. The clustering loss function evaluates the difference between the soft-assignment probability distribution matrix and the auxiliary probability distribution matrix. The soft-assignment probability distribution matrix includes the first element values of M×J first elements, and the auxiliary probability distribution matrix includes the second element values of M×J second elements, where M represents the number of second sample tracks, J represents the number of cluster centers, and q... mj q represents the value of the first element in the m-th row and j-th column. mj Used to indicate z m Belonging to the j-th cluster center u j The original probability of the corresponding cluster, p mj p represents the value of the second element in the m-th row and j-th column. mj Used to indicate z m Belonging to the j-th cluster center u j The corrected probability of the corresponding cluster, z m Let M and J represent the low-dimensional latent representation of the m-th second sample track, where M and J are integers greater than 1, m∈{1,……,M}, j∈{1,……,J}, and the initial value of the cluster center is the initial cluster center.
[0012] According to an embodiment of the present invention, the weighting function is determined based on a first weighting function and a second weighting function; wherein, when the absolute value is less than a predetermined threshold, the weighting coefficient of the sample track is determined based on the first weighting function, which is the product of a predetermined coefficient and an absolute value; when the absolute value is greater than or equal to the predetermined threshold, the weighting coefficient of the sample track is determined based on the second weighting function, which is an absolute value representing the absolute value of the difference between the sample track and the sample reconstructed track of the sample track, and the predetermined coefficient is greater than 0 and less than 1.
[0013] According to an embodiment of the present invention, q mj It is based on z m with u j The distance between and z m and The distance between them is determined. Let p represent the j-th cluster center; where p mj It is based on q mj , and Definitely. This represents the sum of the first element values of the M first elements in the j-th column. Let J' represent the sum of the first element values of the M first elements in the j'-th column; where j' ∈ {1, ..., J}.
[0014] According to an embodiment of the present invention, the pre-trained autoencoder is trained as follows: multiple first sample trajectories are input into the encoder and decoder of the initial autoencoder to obtain the first sample reconstructed trajectories of each of the multiple first sample trajectories; based on the non-clustering loss function, the non-clustering loss function value is obtained according to the multiple first sample trajectories and the first sample reconstructed trajectories of each of the multiple first sample trajectories; the model parameters of the autoencoder are adjusted according to the non-clustering loss function value to obtain the pre-trained autoencoder; the pre-trained autoencoder includes a pre-trained encoder and a pre-trained decoder, and the multiple target cluster centers and the trained autoencoder are trained as follows. The results are as follows: Multiple second sample tracks are input into a pre-trained autoencoder to obtain the second sample reconstructed tracks of each of the multiple second sample tracks; multiple second sample tracks are input into a pre-trained encoder to obtain the sample low-dimensional latent representations of each of the multiple second sample tracks; based on the joint loss function, the joint loss function value is obtained according to the multiple second sample tracks, the second sample reconstructed tracks of each of the multiple second sample tracks, the sample low-dimensional latent representations of each of the multiple second sample tracks, and multiple initial cluster centers; the model parameters of the pre-trained autoencoder and multiple initial cluster centers are adjusted according to the joint loss function value to obtain multiple target cluster centers and the trained autoencoder.
[0015] Another aspect of the present invention provides an electronic device comprising: one or more processors; and a memory for storing one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to perform the method as described above.
[0016] Another aspect of the present invention provides a computer-readable storage medium storing computer-executable instructions, which, when executed, are used to implement the method described above.
[0017] Another aspect of the present invention provides a computer program product comprising computer-executable instructions which, when executed, are used to implement the method described above.
[0018] According to embodiments of the present invention, multiple tracks are clustered based on their respective multidimensional features to obtain multiple track clusters with similar behavioral patterns, thus solving the technical problem that clustering multiple tracks using univariate features is difficult to fully reflect track behavioral patterns. Based on the number of tracks, speed information, or track cluster density of each track's respective track cluster, a first nearest neighbor track set is determined from the track cluster to which each track belongs or from the nearest track cluster, thereby solving the technical problem of ignoring the behavioral correlation between tracks and their respective track clusters when determining the first nearest neighbor track set. Then, based on the first nearest neighbor tracks of each track... The set of tracks is used to calculate the deviation probability, which serves as the judgment condition for the operating state of each track. In other words, the deviation degree of a track from the first nearest neighbor track set is quantified into a probability. The operating state of the track is determined based on the deviation probability and a predetermined deviation threshold, thereby identifying tracks with abnormal operating states. Compared with selecting nearest neighbor tracks based solely on distance, this method flexibly determines whether the first nearest neighbor track set is selected across clusters or within clusters based on the distribution parameters of the track cluster. This ensures that the determined first nearest neighbor track set takes into account both distance and track behavior patterns, thereby improving the reliability of the deviation probability, reducing misjudgments and omissions of tracks with abnormal behavior states, and improving the accuracy of identifying tracks with abnormal behavior states. Attached Figure Description
[0019] The above and other objects, features and advantages of the present invention will become more apparent from the following description of embodiments of the invention with reference to the accompanying drawings, in which:
[0020] Figure 1 A flowchart is shown for a method for determining the trajectory status of a moving target based on clustering according to an embodiment of the present invention.
[0021] Figure 2 A schematic diagram illustrating clustering of multiple tracks according to an embodiment of the present invention is shown.
[0022] Figure 3A schematic diagram of a reconstructed track by a trained autoencoder according to an embodiment of the present invention is shown.
[0023] Figure 4 A schematic diagram of the intra-cluster distribution of the first track cluster and the second track cluster according to an embodiment of the present invention is shown.
[0024] Figure 5 A block diagram of an electronic device suitable for implementing the clustering-based method for determining the trajectory status of a moving target as described above is shown according to an embodiment of the present invention. Detailed Implementation
[0025] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the invention. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the invention for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concept of the invention.
[0026] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0027] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0028] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0029] With the development of machine learning technology, more and more machine learning algorithms are being applied to the field of anomaly detection. This involves building models to learn normal patterns in data and identifying anomalous data that does not conform to these patterns. Deep learning, as a branch of machine learning, can effectively capture and learn deep-level representations of data through multi-layer neural network architectures. However, neural networks can only output attributes of a limited number of target variables and cannot fully analyze anomalous tracks. Furthermore, due to the multivariate and multi-scenario characteristics of time-series data for maritime and aerial targets, spatiotemporal anomaly detection algorithms based on univariate analysis often face technical challenges in practical applications, such as insufficient feature extraction and weak cross-scenario adaptability.
[0030] Therefore, related technologies determine abnormal states based on behavioral differences between the track to be determined and its nearest neighbor tracks, along with a predetermined threshold. However, because the determination of nearest neighbor tracks is based solely on Euclidean distance or Hausdorff distance, the correlation between multiple tracks within the track cluster to which the track belongs is ignored, leading to misjudgments of nearest neighbor tracks, and consequently, misjudgments and missed detections of abnormal tracks. Therefore, the accuracy of abnormal state detection in these technologies is relatively low, posing certain safety risks to maritime or aerial targets.
[0031] In view of this, embodiments of the present invention provide a method for determining the trajectory running state of a moving target based on clustering. The method involves clustering multiple trajectories according to their multidimensional features to obtain multiple trajectory clusters; determining the first nearest neighbor trajectory set for each trajectory from its own trajectory cluster or the trajectory cluster closest to each trajectory based on the intra-cluster distribution parameters of each trajectory cluster; wherein the intra-cluster distribution parameters include at least one of the following: number of trajectories, velocity information, or trajectory cluster density; obtaining the deviation probability of each trajectory based on its first nearest neighbor trajectory set; and determining the running state of each trajectory based on a predetermined deviation threshold and the deviation probability of each trajectory. Thus, in the process of determining the running state, not only are the multidimensional features of the trajectories considered, but the intra-cluster distribution parameters of the clustered trajectory clusters are also flexibly selected to select the nearest neighbor trajectory set corresponding to each trajectory, improving the accuracy of abnormal trajectory identification.
[0032] Figure 1 A flowchart is shown for a method for determining the trajectory status of a moving target based on clustering according to an embodiment of the present invention.
[0033] like Figure 1 As shown, the method 100 for determining the track operation status includes operations S110 to S140.
[0034] In operation S110, multiple tracks are clustered based on their multidimensional features to obtain multiple track clusters.
[0035] According to an embodiment of the present invention, a set of tracks T consisting of multiple tracks can be represented as T = , where t1, t2, t i t I Let represent the 1st track, the 2nd track, the i-th track, and the I-th track, respectively (where i and I are both positive integers and i ≤ I). A track can include multiple track points; therefore, track t... i It can be represented as a sequence of trackpoints consisting of multiple trackpoint vectors arranged in chronological order, i.e., t i = ,in, Representing the trajectory t respectively i The first, second, y-th, and Y-th trackpoint vectors in the vector (where y and Y are both positive integers and y≤Y) include multi-dimensional features corresponding to the current trackpoint.
[0036] According to an embodiment of the present invention, the track is generated by the movement of maritime targets and aerial targets. The multidimensional features of the track may include at least two of the following features: target number, attribute, type, time, longitude, latitude, altitude, speed, or heading. Among them, the target number serves as a unique identifier for the track and can be the number of the target corresponding to the track (e.g., XX ship, XX aircraft, XX UAV, etc.). The attribute can be the attribute of the target corresponding to the track (e.g., civilian). The type can be the type of the target corresponding to the track (e.g., aircraft, speedboat, ship).
[0037] According to an embodiment of the present invention, multiple tracks can be clustered based on the K-means clustering algorithm according to multidimensional features in space and time, so that the clustering of multiple tracks not only considers the spatial distribution of each track point in the track, but also takes into account the temporal dynamics and physical motion attributes.
[0038] In operation S120, based on the intra-cluster distribution parameters of each of the multiple track clusters, the first nearest neighbor track set of each of the multiple tracks is determined from the track cluster to which each of the multiple tracks belongs or from the track cluster that is closest to each of the multiple tracks.
[0039] According to embodiments of the present invention, the intra-cluster distribution parameters of the multiple track clusters obtained by clustering are different. For example, the intra-cluster distribution parameters may include track cluster density, the number of tracks in the track cluster, velocity information, etc. The intra-cluster distribution parameters can reflect the quality and reliability of the track cluster to a certain extent. Specifically, the track cluster density is the ratio of the track cluster area to the number of tracks.
[0040] Therefore, the source of the first nearest neighbor set of a track can be determined by the intra-cluster distribution parameters of the track cluster to which the track belongs. This avoids the technical problem of misjudging or missing judgments in the subsequent determination of abnormal states for some tracks within track clusters with low reliability, since the first nearest neighbor set is still determined from the track cluster to which the current track belongs.
[0041] In operation S130, the deviation probability of each of the multiple tracks is obtained based on the first nearest neighbor track set of each track.
[0042] According to embodiments of the present invention, the deviation probability can be used to quantify the degree to which the behavior pattern of a flight path deviates from the normal flight path. Generally, a higher deviation probability indicates a greater degree of deviation from the normal flight path.
[0043] In operation S140, the operating status of each of the multiple tracks is determined based on the predetermined deviation threshold and the deviation probability of each track.
[0044] According to embodiments of the present invention, the predetermined deviation threshold can be adjusted based on the type of target corresponding to the track or the area where the target is located. For example, for a research vessel navigating in sparsely populated sea areas, the predetermined deviation threshold can be 0.7. Thus, when the deviation probability is higher than or equal to the predetermined deviation threshold, the target state of the current track can be determined to be an abnormal state, while when the deviation probability is lower than the predetermined deviation threshold, the target state of the current track can be determined to be a normal state.
[0045] It should be noted that the distinction between abnormal and normal states is not absolute. Depending on the specific application scenario, the determination of the target state of the trajectory may change as the predetermined deviation threshold is adjusted. The target state can be transmitted back to the host computer in coded form, and the host computer will determine the warning level based on the difference between the deviation probability and the predetermined probability threshold.
[0046] According to embodiments of the present invention, multiple tracks are clustered based on their respective multidimensional features to obtain multiple track clusters with similar behavioral patterns, thus solving the technical problem that clustering multiple tracks using univariate features is difficult to fully reflect track behavioral patterns. Based on the number of tracks, speed information, or track cluster density of each track's respective track cluster, a first nearest neighbor track set is determined from the track cluster to which each track belongs or from the nearest track cluster. This solves the problem of ignoring the behavioral correlation between tracks and their respective track clusters when determining the first nearest neighbor track set. Then, based on the first nearest neighbor track sets of each track... The deviation probability, which serves as the judgment condition for the operating state of multiple tracks, is calculated. That is, the degree of deviation of a track from the first nearest neighbor track set is quantified into a probability. The operating state of the track is determined based on the deviation probability and a predetermined deviation threshold, thereby identifying tracks with abnormal operating states. Compared with selecting nearest neighbor tracks based solely on distance, the first nearest neighbor track set is flexibly determined based on the distribution parameters of the track cluster. This allows the determined first nearest neighbor track set to take into account both distance and track behavior patterns, thereby improving the reliability of the deviation probability, reducing misjudgments and omissions of tracks with abnormal behavior states, and improving the accuracy of identifying tracks with abnormal behavior states.
[0047] According to an embodiment of the present invention, multiple tracks are clustered based on their respective multidimensional features to obtain multiple track clusters, including: inputting multiple tracks into a trained encoder, which is a trained autoencoder, to obtain low-dimensional latent representations of each track; for any track among the multiple tracks, determining the track cluster to which the track belongs based on the distance between the low-dimensional latent representation of the track and each of the multiple target cluster centers.
[0048] With track t i For example, in the encoding stage of the autoencoder, a nonlinear mapping as shown in equation (1) can be used to transform the track point vector. Mapping to the target dimension compresses multidimensional features into low-dimensional latent representations, which can preserve key features of the track (such as spatiotemporal features related to track trends, speed change patterns, etc.).
[0049] (1);
[0050] Where D is the waypoint vector Feature dimension, z y for The dimension-reduced vector, D´ is z y The feature dimension (i.e., the target dimension). R D R represents a space of dimension D. D ´ represents a space of dimension D´.
[0051] By performing the dimensionality reduction process shown in formula (1) on each track point vector, the track t shown in formula (2) can be obtained. i Low-dimensional latent representation .
[0052] (2).
[0053] The trained autoencoder is obtained by training a pre-trained autoencoder based on a joint loss function and multiple second sample tracks, while the pre-trained autoencoder is obtained by training an autoencoder based on a non-clustering loss function and multiple first sample tracks. To better understand the training process of the autoencoder in this embodiment of the invention, the pre-training process of the autoencoder will be explained in detail below using formulas (3) to (5).
[0054] According to an embodiment of the present invention, the pre-trained autoencoder includes a pre-trained encoder and a pre-trained decoder, and the specific pre-training process of the autoencoder includes the following steps:
[0055] Multiple first sample tracks are input into the encoder and decoder of the initial autoencoder to obtain the first sample reconstructed tracks of each of the multiple first sample tracks, such as the first sample track. The corresponding first sample reconstructed track is .
[0056] The non-clustering loss function based on formula (3) The non-clustering loss function value is obtained by reconstructing the trajectory based on multiple first sample trajectories and the first samples of each of the multiple first sample trajectories.
[0057] (3);
[0058] in, This represents the first sample track. This indicates that the first sample reconstructs the trajectory. This represents the weighting function. Let β represent the reconstruction loss function, and let β represent the hyperparameter that controls the magnitude of the learning error change for each sample track.
[0059] According to an embodiment of the present invention, the weighting function It is determined based on the first weighting function and the second weighting function, and the first weighting function and the second weighting function respectively correspond to different situations of the magnitude relationship between the absolute value of the difference between the sample track and the sample reconstruction track of the sample track and the predetermined threshold. Here, the sample track may include the first sample track mentioned above, or it may include the second sample track used in the cluster center training process.
[0060] Specifically, the first weighting function is used to obtain the weighting coefficient of the sample track based on the product between the predetermined coefficient and the absolute value when the absolute value is less than the predetermined threshold. The second weighting function is used to obtain the weighting coefficient of the sample track based on the absolute value when the absolute value is greater than or equal to the predetermined threshold. The predetermined coefficient is usually a value greater than 0 and less than 1. Taking the predetermined coefficient as 1 / 2 as an example, the weighting function can be expressed as the following formula (4).
[0061] (4);
[0062] Where h represents the predetermined threshold, Represents absolute value. This represents the first weighting function.
[0063] Correspondingly, the reconstruction loss function is also determined based on the first reconstruction loss function and the second reconstruction loss function, and the first and second reconstruction loss functions respectively correspond to different situations regarding the magnitude relationship between the absolute value of the difference between the sample track and the sample reconstructed track of the sample track and a predetermined reconstruction error threshold. Reconstruction Loss Function It can be expressed as the following formula (5).
[0064] (5);
[0065] in, Denotes the first reconstruction loss function. Let δ represent the second reconstruction loss function, and let δ represent the predetermined reconstruction error threshold, which is usually a positive real number.
[0066] The model parameters of the autoencoder are iteratively adjusted based on the non-clustering loss function value determined by the non-clustering function to obtain a pre-trained autoencoder.
[0067] According to an embodiment of the present invention, a weighting function is added to the reconstruction loss function. The non-clustering loss function is weighted differently based on the difference between the sample track and the reconstructed track. This makes the pre-training process focus more on the sample track with larger errors. Therefore, the autoencoder is pre-trained based on the non-clustering loss function determined by the weighting function and the reconstruction loss function. This results in the pre-trained autoencoder having a better dimensionality reduction effect for multi-dimensional features and a smaller error between the reconstructed track obtained from the low-dimensional latent representation obtained in the encoding stage and the input sample track. Thus, in practical applications, the autoencoder removes noise and redundant information while reducing dimensionality, highlighting the key features of the track. Therefore, when using low-dimensional latent representation for clustering, the clustering results can more accurately reflect the similarity and differences between multiple tracks.
[0068] According to an embodiment of the present invention, multiple target cluster centers and a trained autoencoder are obtained based on a joint loss function, using multiple initial cluster centers and a pre-trained autoencoder trained from multiple second sample trajectories. The specific training process of the multiple target cluster centers and the trained autoencoder includes: inputting multiple second sample trajectories into the pre-trained autoencoder to obtain the second sample reconstructed trajectories of each of the multiple second sample trajectories; inputting multiple second sample trajectories into the pre-trained encoder to obtain the sample low-dimensional latent representations of each of the multiple second sample trajectories; obtaining a joint loss function value based on the joint loss function, using the multiple second sample trajectories, the second sample reconstructed trajectories of each of the multiple second sample trajectories, the sample low-dimensional latent representations of each of the multiple second sample trajectories, and the multiple initial cluster centers; and adjusting the model parameters of the pre-trained autoencoder and the multiple initial cluster centers based on the joint loss function value to obtain multiple target cluster centers and the trained autoencoder.
[0069] The joint loss function is determined based on the non-clustering loss function and the clustering loss function. The clustering loss function is used to evaluate the difference between the soft-assignment probability distribution matrix Q and the auxiliary probability distribution matrix P. The soft-assignment probability distribution matrix includes the first element values of M×J first elements, and the auxiliary probability distribution matrix includes the second element values of M×J second elements. M represents the number of second sample tracks, J represents the number of cluster centers, and q... mj q represents the value of the first element in the m-th row and j-th column. mj It is based on z m with u j The distance between and z m and The distance between them is fixed, q mj Used to indicate z m Belongs to u j The original probability can be expressed as the following formula (6).
[0070] (6);
[0071] Where M and J are integers greater than 1, m∈{1, ..., M}, j∈{1, ..., J}, Indicate z m and The distance between them (e.g., Euclidean distance), z m This represents the low-dimensional latent representation of the m-th second sample track. Let j' be the j-th cluster center, where j' ∈ {1, ..., J}. Indicate z m and The distance between them (e.g., Euclidean distance).
[0072] Formula (6) uses the student's t-distribution to measure z. m The similarity between cluster centers and clusters is considered; the smaller the distance, the higher the similarity. The initial cluster centers are set to initial values. The clustering loss function is calculated using the soft-assignment probability distribution matrix and the auxiliary probability distribution matrix to reduce the cluster size through training. m and The distance between them is used to obtain the target cluster center. The auxiliary probability distribution matrix can be used to highlight the contribution of the first element with high confidence, suppress the interference of the first element with low confidence, and guide the clustering to optimize towards "higher intra-cluster similarity". The auxiliary probability distribution matrix will be updated after a predetermined number of iterations.
[0073] The second element value p of the second element in the m-th row and j-th column of the auxiliary probability distribution matrix P mj Used to indicate z m Belongs to u j The corrected probability, p mj It is based on q mj , and Definitely. This represents the sum of the first element values of the M first elements in the j-th column. p represents the sum of the first element values of the M first elements in the j'-th column. mj It can be expressed as the following formula (7).
[0074] (7);
[0075] in, ; Let z represent the value of the first element in the m-th row and j'-th column, used to indicate z. m Belongs to the j´th cluster center The original probability of the corresponding cluster.
[0076] The degree of difference between the soft-assignment probability distribution matrix Q and the auxiliary probability distribution matrix P can be represented by the KL divergence, and the clustering loss function can be expressed as the following formula (8).
[0077] (8).
[0078] Joint loss function , This represents a predetermined weighting coefficient used to adjust the proportion of the clustering loss function in the joint loss function, so as to achieve a reasonable balance between the clustering loss function and the non-clustering loss function.
[0079] According to an embodiment of the present invention, after determining the joint loss function, multiple target cluster centers and a trained autoencoder are obtained by training in the following manner: inputting multiple second sample trajectories into the pre-trained autoencoder to obtain the second sample reconstructed trajectories of each of the multiple second sample trajectories; inputting the multiple second sample trajectories into the pre-trained encoder to obtain the sample low-dimensional latent representations of each of the multiple second sample trajectories; based on the joint loss function, obtaining the joint loss function value according to the multiple second sample trajectories, the second sample reconstructed trajectories of each of the multiple second sample trajectories, the sample low-dimensional latent representations of each of the multiple second sample trajectories, and the multiple initial cluster centers; adjusting the model parameters of the pre-trained autoencoder and the multiple initial cluster centers according to the joint loss function value to obtain multiple target cluster centers and the trained autoencoder.
[0080] Figure 2 A schematic diagram illustrating clustering of multiple tracks according to an embodiment of the present invention is shown.
[0081] like Figure 2 As shown, the multidimensional features of multiple tracks are sequentially input into a trained autoencoder. Dimensionality reduction is performed during the encoding stage to obtain low-dimensional latent representations. Using the target cluster centers determined by the joint loss function, multiple tracks are clustered based on the low-dimensional latent representations to obtain multiple track clusters. Meanwhile, the reconstructed tracks involved in the weighting function and reconstruction loss function in the non-clustering loss function corresponding to the pre-training of the autoencoder are obtained during the decoding stage of the autoencoder, reconstructed based on the low-dimensional latent representations obtained during the encoding stage.
[0082] Figure 3 A schematic diagram of a reconstructed track by a trained autoencoder according to an embodiment of the present invention is shown.
[0083] like Figure 3 As shown, the autoencoder trained by the joint loss function and the autoencoder trained by the mean square error loss in the related technology simultaneously reconstruct the same verification track. The comparison shows that the reconstructed track obtained by the autoencoder based on the embodiment of the present invention is closer to the input track than the reconstructed track obtained by the autoencoder based on the related technology, that is, the error is smaller.
[0084] According to an embodiment of the present invention, by constructing a soft-assignment probability distribution matrix and an auxiliary probability distribution matrix, the weight of the first element with higher confidence in the soft-assignment probability distribution matrix is amplified, thereby pushing the soft-assignment probability distribution matrix to approximate the auxiliary probability distribution matrix. The clustering loss function is determined, and the autoencoder feature extraction and clustering are optimized in a coordinated manner based on the joint loss function composed of the clustering loss function and the non-clustering loss function. This provides accurate track cluster division for the determination of subsequent operating states and reduces false alarms and missed alarms of abnormal tracks.
[0085] According to an embodiment of the present invention, determining a first nearest neighbor set of multiple track clusters from the track clusters to which each track belongs or the nearest track cluster based on the intra-cluster distribution parameters of each track cluster includes: dividing the multiple track clusters into a first track cluster set and a second track cluster set based on the intra-cluster distribution parameters of each track cluster and a preset distribution threshold, wherein the first track cluster set includes at least one first track cluster and the second track cluster set includes at least one second track cluster; determining a first nearest neighbor set of each of the at least one first track included in the first track cluster set based on the first track cluster set and a first predetermined distance threshold; and determining a first nearest neighbor set of each of the at least one second track included in the second track cluster set based on the first track cluster set, the second track cluster set, and the second predetermined distance threshold.
[0086] Wherein, the track is either the first track or the second track, and the track cluster is either the first track cluster or the second track cluster. The first track cluster includes at least one first track, and the second track cluster includes at least one second track. The first track cluster and the second track cluster satisfy at least one of the following conditions: the track cluster density of the first track cluster is greater than the track cluster density of the second track cluster, the number of tracks in the first track cluster is greater than the number of tracks in the second track cluster, or the operational stability of the first track cluster is higher than the operational stability of the second track cluster. Speed information indicates the operational stability of the track cluster.
[0087] To better illustrate the differences in distribution parameters between the first and second track clusters, the following will be conducted through... Figure 4 Please illustrate each point separately.
[0088] Figure 4 A schematic diagram of the intra-cluster distribution of the first track cluster and the second track cluster according to an embodiment of the present invention is shown.
[0089] Therefore, it can be seen that the first track cluster is denser and contains more tracks, while the second track cluster is relatively sparse and contains fewer tracks.
[0090] Determining the first nearest neighbor set of each of the at least one first track included in the first track cluster set, based on the first track cluster set and the first predetermined distance threshold, includes: for any first track cluster in the at least one first track cluster included in the first track cluster set, and for any first track in the plurality of first tracks included in the first track cluster, determining the first nearest neighbor set of the first track based on the first predetermined distance threshold and the distance between the first track in the first track cluster and other first tracks.
[0091] Determining the first nearest neighbor set of each of the at least one second track included in the second track cluster set, based on the first track cluster set, the second track cluster set, and the second predetermined distance threshold, includes: for any second track cluster among the at least one second track cluster included in the second track cluster, and for any second track among the plurality of second tracks included in the second track cluster, determining the first nearest neighbor set of the second track based on the second predetermined distance threshold and the distance between each first track in the first track cluster closest to the second track cluster and the second track.
[0092] For a first track within a first track cluster, the selection of its first nearest neighbor track set is limited to the first track cluster to which the first track belongs. Other first tracks that are relatively close are filtered out using a first predetermined distance threshold. Since multiple tracks within a cluster obtained through clustering have similar operating patterns, for the first track in a first track cluster with high cluster density, a large number of tracks, or high operating stability, the first nearest neighbor track set is preferentially determined from this cluster. This avoids the deviation in nearest neighbor track determination caused by cross-cluster selection and focuses on the group association characteristics between the first track and other first tracks within its first track cluster.
[0093] For the second track within the second track cluster, its first nearest neighbor track set is not determined based on its own cluster, but rather selected from the nearest first track cluster and filtered using a second predetermined distance threshold. Therefore, considering that the second track cluster to which the second track belongs has a small number of tracks, low cluster density, and poor operational stability, making it difficult to form an effective first nearest neighbor track set within the cluster, selecting nearest neighbor tracks based on the nearest first track cluster effectively improves the reliability of nearest neighbor track determination.
[0094] According to an embodiment of the present invention, the first track cluster represents a track cluster that satisfies at least one of the following conditions: the track cluster density is high, the number of tracks is large, or the average speed is fast and the average acceleration is small, wherein the speed information includes average speed and average acceleration, which are used to represent the operational stability of the track cluster.
[0095] Therefore, based on the intra-cluster distribution parameters and preset distribution thresholds of each of the multiple track clusters, the multiple track clusters are divided into a first track cluster set and a second track cluster set. Firstly, track clusters with a density greater than or equal to the preset track cluster density threshold are identified as the first track cluster, while track clusters with a density less than the preset track cluster density threshold are assigned to the first set of undetermined clusters. Thus, the multiple track clusters are first-stage filtered based on their density, and track clusters with a density greater than or equal to the preset track cluster density threshold are identified as the first track cluster. The remaining track clusters are then used as the first set of undetermined clusters for the next round of filtering based on the number of tracks.
[0096] Subsequently, during the second screening of the first set of clusters to be determined, the first sequence of clusters to be determined is obtained by first sorting the multiple clusters in the first set of clusters to be determined. Then, based on the sorting of the first track clusters to be determined in the first sequence of clusters to be determined, the current cluster to be determined and the next cluster to be determined are obtained from the sequence. If the ratio between the number of tracks of the current cluster to the number of tracks of the next cluster to be determined is greater than or equal to a preset ratio threshold, the current cluster to be determined is designated as the boundary track cluster corresponding to the number of tracks. If the ratio between the number of tracks is less than a preset ratio threshold, the following operations are repeated: according to the sorting of the first clusters to be determined in the first clusters to be determined sequence, the next first cluster to be determined is taken as the new current first cluster to be determined; the next first cluster to be determined is obtained from the first clusters to be determined sequence; and the operation of determining whether the ratio between the number of tracks of the new current first cluster to be determined and the number of tracks of the next first cluster to be determined is greater than or equal to the preset ratio threshold is repeated, until the new current first cluster to be determined is taken as the boundary track cluster corresponding to the number of tracks. According to the boundary track cluster corresponding to the number of tracks, the first clusters to be determined sequence is divided into the first track cluster and the second clusters to be determined set.
[0097] For example, multiple clusters in the set of clusters to be determined can be sorted from largest to smallest according to the number of tracks to obtain a sequence of clusters to be determined. c1, c2, c K All are in the first undetermined cluster, and the number of tracks Nc1 < c1. K The number of tracks Nc K .
[0098] The calculation starts from the ratio of the number of tracks in c1 to the number of tracks in c2 (i.e., Nc1 / Nc2), and continues until the new current first cluster to be determined c. b The next first cluster to be determined with the new current first cluster to be determined c b+1 (2≤b≤k-1) The ratio of the number of each track (D is a preset multiple threshold), the new current first cluster to be determined c b As a boundary track cluster corresponding to the number of tracks, c1~c b It was determined to be the second track cluster, and c b+1 ~c K They were assigned to the second set of undetermined clusters for the next round of screening based on track operational stability.
[0099] The selection based on the stability of the flight path is to identify the second set of undetermined clusters whose average speed is greater than or equal to a preset speed threshold and whose average acceleration is less than or equal to a preset acceleration threshold as the first flight path cluster, and to identify the second set of undetermined clusters other than the first flight path cluster as the second flight path cluster.
[0100] According to an embodiment of the present invention, track clusters with high density are preferentially classified into a first track cluster set through track cluster density screening. Secondly, boundary clusters corresponding to the number of tracks are determined based on the ratio of track counts. Significant differences in track counts allow for the identification of a second batch of track clusters to be classified into the first track cluster set. Finally, further screening is performed using average speed and average acceleration. Track clusters with high operational stability (speed not lower than a threshold, acceleration not exceeding a threshold) within a cluster are identified as the first track cluster. This ensures that the first track clusters in the first track cluster set are not only representative in spatial distribution and scale, but their motion patterns also conform to conventional patterns. Therefore, based on multiple constraints, the first track clusters in the obtained first track cluster set more accurately represent normal operating conditions, and the second track clusters in the second track cluster set are more likely to contain tracks in abnormal operating conditions. This provides a reliable classification basis for subsequent calculation of deviation probabilities based on the first nearest neighbor track set of track clusters, thereby reducing misjudgments and improving detection accuracy.
[0101] According to an embodiment of the present invention, obtaining the deviation probability of each of the multiple tracks based on their respective first nearest neighbor track sets includes: for any track among the multiple tracks, obtaining a first probability set distance of the track based on the distances between the track and each of the multiple first nearest neighbor tracks in the track's first nearest neighbor track set, wherein the first probability set distance represents the absolute deviation of the track relative to the first nearest neighbor track set; for any first nearest neighbor track among the multiple first nearest neighbor tracks, obtaining the deviation probability based on the distances between the first nearest neighbor tracks and the multiple second nearest neighbor tracks in the second nearest neighbor track set of the first nearest neighbor track. Obtain the second probability set distance of the first nearest neighbor track, where the second probability set distance represents the absolute deviation of the first nearest neighbor track relative to the set of second nearest neighbor tracks; determine the average probability set distance of the set of first nearest neighbor tracks based on the second probability set distances of each of the multiple first nearest neighbor tracks, where the average probability set distance represents the average deviation of the multiple first nearest neighbor tracks in the set of first nearest neighbor tracks relative to the sets of second nearest neighbor tracks of each of the multiple first nearest neighbor tracks; obtain the deviation probability of the track based on the first probability set distance of the track and the average probability set distance of the first nearest neighbor track set of the first track.
[0102] Each of the first nearest neighbor tracks in the first nearest neighbor track set also has its own second nearest neighbor track set. The method for determining the second nearest neighbor track set is similar to that for determining the first nearest neighbor track set. It is also necessary to consider whether the first nearest neighbor track is the first track or the second track, and to flexibly determine whether to determine it from within the cluster or the nearest cluster based on the intra-cluster distribution parameters of the track cluster to which the first nearest neighbor track belongs. This will not be elaborated here.
[0103] Taking the first track o as an example, the distance of the first probability set It can be expressed as the following formula (9).
[0104] (9);
[0105] Where S0 represents the set of the first nearest neighbor tracks of o, and λ represents the predetermined normalization factor. The value of λ comes from the empirical "three sigma" rule, which states that in a normal distribution, 68%, 95%, or 99.7% of the data are within one, two, or three standard deviations of the mean. Therefore, the value of λ is: (λ=1) 68.0%, λ=2 95.0%, λ=3 99.7%). The standard distance between the first track o and its first nearest neighbor track set S0 can be expressed as follows (10).
[0106] (10);
[0107] Where i represents the number of nearest neighbor tracks to the first target in the first nearest neighbor track set. This represents the distance between the first track o and a neighboring track s of a first target.
[0108] Similarly, each of the first nearest neighbor tracks in the first nearest neighbor track set also has its own second nearest neighbor track set. The second probability set distance of the first nearest neighbor track can be obtained by considering the distances between the second nearest neighbor tracks in the second nearest neighbor track set of the first nearest neighbor track and the first nearest neighbor track.
[0109] The distances between multiple second-target neighbor tracks and the first-target neighbor tracks in the set of second-target neighbor tracks of the first-target neighbor tracks are used to obtain the probability set distances of the first-target neighbor tracks. Then, the average probability set distance of the set of first-target neighbor tracks is calculated. This average probability set distance represents the average deviation of the multiple first-target neighbor tracks in the set from their respective sets of second-target neighbor tracks. After determining the second probability set distances of the multiple first-target neighbor tracks, the average probability set distance of the set of first-target neighbor tracks is determined based on these individual second probability set distances.
[0110] According to an embodiment of the present invention, obtaining the deviation probability of a track based on the first probability set distance of the track and the average probability set distance of the first nearest neighbor track set of the first track includes: obtaining a local deviation factor of the track based on the first probability set distance of the track and the average probability set distance of the first nearest neighbor track set of the track; obtaining a global outlier factor of the track cluster to which the track belongs based on the local deviation factor of the track, the number of first nearest neighbor tracks included in the first nearest neighbor track set, and a predetermined normalization factor; and obtaining the deviation probability of the track based on the local deviation factor of the track and the global outlier factor of the track cluster to which the track belongs.
[0111] Taking the first track o in the above embodiment as an example, the local deviation factor of o It can be expressed as the following formula (11).
[0112] (11);
[0113] Among them, S s Let E[ ] represent the set of the second nearest neighbor tracks of the first nearest neighbor track s, E[ ] represent the expectation, and the local deviation factor represents the relative deviation of the track from the set of the first nearest neighbor tracks.
[0114] According to an embodiment of the present invention, the global outlier factor of the first track cluster to which the first track o belongs. It can be expressed as the following formula (12).
[0115] (12);
[0116] in, This indicates the first track cluster to which the first track o belongs. express The total number of first tracks, λ represents the predetermined normalization factor, and the global outlier factor represents the overall deviation of the tracks.
[0117] According to an embodiment of the present invention, the Gaussian error function can be used to transform the local group factor obtained above from the deviation degree form into the deviation probability in the probabilistic form, so as to be better applicable to the determination of the target state.
[0118] Taking the first track o as an example, the deviation probability of o It can be expressed as the following formula (13).
[0119] (13);
[0120] in, This represents the Gaussian error function.
[0121] According to an embodiment of the present invention, the local deviation factor, by comparing the distance of the track's own first probability set with the average probability set distance of its first nearest neighbor track set, intuitively reflects the relative deviation of the track from its first nearest neighbor track set, and is used to capture the subtle differences between the track itself and the first nearest neighbor track set. At the same time, the global outlier factor, combined with the local deviation factor, the number of first nearest neighbor tracks included in the first nearest neighbor track set, and a predetermined normalization factor, considers the local deviation within the overall track cluster, reducing the bias that may be caused by isolated analysis of the local deviation factor. Finally, the deviation probability obtained by combining the local deviation factor and the global outlier factor, while retaining the subtle differences between the track itself and the first nearest neighbor track set, also takes into account the overall performance of the track in the entire track cluster, making the deviation probability more consistent with the actual operating state of the track. This effectively distinguishes whether the track is in a normal fluctuation state or an abnormal state, reduces false alarms and false negatives, and improves the reliability of anomaly detection.
[0122] Figure 5 A block diagram of an electronic device suitable for implementing the clustering-based method for determining the trajectory status of a moving target as described above is shown according to an embodiment of the present invention. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.
[0123] like Figure 5As shown, an electronic device 500 according to an embodiment of the present invention includes a processor 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage portion 508 into a random access memory (RAM) 503. The processor 501 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 501 may also include onboard memory for caching purposes. The processor 501 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0124] RAM 503 stores various programs and data required for the operation of electronic device 500. Processor 501, ROM 502, and RAM 503 are interconnected via bus 504. Processor 501 executes various operations of the method flow according to embodiments of the present invention by executing programs in ROM 502 and / or RAM 503. It should be noted that the programs may also be stored in one or more memories other than ROM 502 and RAM 503. Processor 501 may also execute various operations of the method flow according to embodiments of the present invention by executing programs stored in said one or more memories.
[0125] According to an embodiment of the present invention, the electronic device 500 may further include an input / output (I / O) interface 505, which is also connected to a bus 504. The electronic device 500 may also include one or more of the following components connected to the input / output (I / O) interface 505: an input section 506 including a keyboard, mouse, etc.; an output section 507 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the input / output (I / O) interface 505 as needed. A removable medium 511, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 510 as needed so that computer programs read from it can be installed into the storage section 508 as needed.
[0126] According to embodiments of the present invention, the method flow according to embodiments of the present invention can be implemented as a computer software program. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing program code for performing the method shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via communication section 509, and / or installed from removable medium 511. When the computer program is executed by processor 501, it performs the functions defined in the system of the embodiments of the present invention. According to embodiments of the present invention, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0127] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of the present invention.
[0128] According to embodiments of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium. Examples include, but are not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In the present invention, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0129] For example, according to embodiments of the present invention, a computer-readable storage medium may include the ROM 502 and / or RAM 503 described above and / or one or more memories other than ROM 502 and RAM 503.
[0130] Embodiments of the present invention also include a computer program product comprising a computer program containing program code for executing the methods provided in the embodiments of the present invention. When the computer program product is run on an electronic device, the program code is used to enable the electronic device to implement the method for determining the track operation state provided in the embodiments of the present invention.
[0131] When the computer program is executed by the processor 501, it performs the functions defined in the system / apparatus of this embodiment of the invention. According to embodiments of the invention, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0132] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 509, and / or installed from a removable medium 511. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0133] According to embodiments of the present invention, program code for executing the computer programs provided in the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0134] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions. Those skilled in the art will understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, the features described in the various embodiments of the present invention can be combined and / or combined in various ways without departing from the spirit and teachings of the present invention. All such combinations and / or pairings fall within the scope of this invention.
[0135] The embodiments of the present invention have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of the invention. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of the invention, and all such substitutions and modifications should fall within the scope of the invention.
Claims
1. A method for determining the trajectory and operational status of a moving target based on clustering, characterized in that, include: Based on the multidimensional features of each of the multiple tracks, the multiple tracks are clustered to obtain multiple track clusters; Based on the intra-cluster distribution parameters of each of the multiple track clusters, the first nearest neighbor track set of each of the multiple tracks is determined from the track cluster to which each of the multiple tracks belongs or from the track cluster closest to each of the multiple tracks, including: Based on the intra-cluster distribution parameters and preset distribution thresholds of the multiple track clusters, the multiple track clusters are divided into a first track cluster set and a second track cluster set, wherein the first track cluster set includes at least one first track cluster, and the second track cluster set includes at least one second track cluster; Based on the first track cluster set and the first predetermined distance threshold, determine the first nearest neighbor track set for each of the at least one first track included in the first track cluster set; Based on the first track cluster set, the second track cluster set, and a second predetermined distance threshold, a first nearest neighbor track set is determined for each of at least one second track included in the second track cluster set; wherein, the cluster distribution parameters include at least one of the following: number of tracks, speed information, and track cluster density, wherein the track is a first track or a second track, the track cluster is a first track cluster or a second track cluster, the first track cluster includes at least one first track, the second track cluster includes at least one second track, and the first track cluster and the second track cluster satisfy at least one of the following conditions: the track cluster density of the first track cluster is greater than the track cluster density of the second track cluster, the number of tracks in the first track cluster is greater than the number of tracks in the second track cluster, and / or the operational stability of the first track cluster is higher than the operational stability of the second track cluster, and the speed information indicates the operational stability of the track cluster; Based on the first nearest neighbor track set of each of the multiple tracks, the deviation probability of each of the multiple tracks is obtained; The operating status of each of the multiple tracks is determined based on a predetermined deviation threshold and the deviation probability of each track.
2. The method according to claim 1, characterized in that, The step of determining the first nearest neighbor set of each of the first track clusters, based on the first track cluster set and a first predetermined distance threshold, includes: For any one of the first track clusters included in the first track cluster set, For any first track among the multiple first tracks included in any first track cluster, a first nearest neighbor track set for any first track is determined based on the first predetermined distance threshold and the distance between any first track in any first track cluster and other first tracks. The step of determining the first nearest neighbor set of each of the at least one second track included in the second track cluster set based on the first track cluster set, the second track cluster set, and the second predetermined distance threshold includes: For any second track cluster in at least one of the second track clusters included in the second track cluster set, For any second track among the plurality of second tracks included in any second track cluster, a first nearest neighbor track set for any second track is determined based on the second predetermined distance threshold and the distance between each first track in the first track cluster closest to the second track cluster and the second track.
3. The method according to claim 2, characterized in that, The velocity information includes average velocity and average acceleration. The preset distribution thresholds include a preset track cluster density threshold, a preset ratio threshold, a preset acceleration threshold, and a preset velocity threshold. The step of dividing the multiple track clusters into a first track cluster set and a second track cluster set based on their respective intra-cluster distribution parameters and preset distribution thresholds includes: The track clusters whose density is greater than or equal to the preset track cluster density threshold are identified as the first track cluster, and the track clusters whose density is less than the preset track cluster density threshold are assigned to the first set of clusters to be determined. Based on the number of tracks of each of the multiple first-to-determine clusters in the first set of clusters to be determined, the multiple first-to-determine clusters are sorted to obtain a sequence of first-to-determine clusters. The current first-to-determine cluster and the next first-to-determine cluster of the current first-to-determine cluster are obtained from the first sequence of first-to-determine clusters. If the ratio between the number of tracks of the current first-to-determine cluster and the number of tracks of the next first-to-determine cluster is greater than or equal to the preset ratio threshold, the current first-to-determine cluster is regarded as the boundary track cluster corresponding to the number of tracks. If the ratio between the number of tracks of the current first cluster to be determined and the number of tracks of the next first cluster to be determined is less than the preset ratio threshold, the operation of taking the next first cluster to be determined in the first cluster to be determined sequence as the new current first cluster to be determined is repeated until the ratio between the number of tracks of the new current first cluster to be determined and the number of tracks of the new next first cluster to be determined is greater than or equal to the preset ratio threshold. The new current first cluster to be determined is then taken as the boundary track cluster corresponding to the number of tracks. Based on the boundary track cluster corresponding to the number of tracks, the first cluster to be determined sequence is divided into the first track cluster and the second cluster to be determined set. The second cluster in the second set of clusters to be determined, whose average speed is greater than or equal to the preset speed threshold and whose average acceleration is less than or equal to the preset acceleration threshold, is determined as the first track cluster, and the second cluster in the second set of clusters to be determined, other than the first track cluster, is determined as the second track cluster.
4. The method according to claim 2, characterized in that, Based on the first nearest neighbor set of each of the multiple tracks, the deviation probability of each of the multiple tracks is obtained, including: For any one of the multiple flight paths The first probability set distance of the trajectory is obtained based on the distances between the trajectory and each of the multiple first nearest neighbor trajectories in the first nearest neighbor trajectories set, wherein the first probability set distance represents the absolute deviation of the trajectory relative to the first nearest neighbor trajectories set. For any one of the plurality of first nearest neighbor tracks Based on the distances between multiple second neighbor tracks and the first neighbor track in the second neighbor track set of the first neighbor track, the second probability set distance of the first neighbor track is obtained, wherein the second neighbor track set includes at least one second neighbor track of the first neighbor track, and the second probability set distance represents the absolute deviation of the first neighbor track from the second neighbor track set; Based on the second probability set distance of each of the plurality of first nearest neighbor tracks, the average probability set distance of the first nearest neighbor track set is determined, wherein the average probability set distance represents the average deviation of the plurality of first nearest neighbor tracks in the first nearest neighbor track set relative to the second nearest neighbor track set of each of the plurality of first nearest neighbor tracks; The deviation probability of the trajectory is obtained based on the first probability set distance of the trajectory and the average probability set distance of the first nearest neighbor trajectory set of the first trajectory.
5. The method according to claim 4, characterized in that, The step of obtaining the deviation probability of the trajectory based on the first probability set distance of the trajectory and the average probability set distance of the first nearest neighbor trajectory set of the first trajectory includes: The local deviation factor of the trajectory is obtained based on the first probability set distance of the trajectory and the average probability set distance of the first nearest neighbor trajectory set, wherein the local deviation factor represents the relative deviation of the trajectory relative to the first nearest neighbor trajectory set; Based on the local deviation factor of the track, the number of first neighbor tracks included in the first neighbor track set, and a predetermined normalization factor, the global outlier factor of the track cluster to which the track belongs is obtained, wherein the global outlier factor represents the overall deviation of the track. The deviation probability of the trajectory is obtained based on the local deviation factor of the trajectory and the global outlier factor of the trajectory cluster to which the trajectory belongs.
6. The method according to claim 2, characterized in that, The method involves clustering the multiple flight paths based on their respective multidimensional features to obtain multiple flight path clusters, including: The multidimensional features of each of the multiple tracks are input into the trained encoder included in the trained autoencoder to obtain the low-dimensional latent representations of each of the multiple tracks. For any one of the multiple tracks, the track cluster to which the track belongs is determined from the track clusters to which each of the multiple target cluster centers belongs, based on the distance between the low-dimensional potential representation of the track and each of the multiple target cluster centers. The plurality of target cluster centers and the trained autoencoder are obtained by training a plurality of initial cluster centers and a pre-trained autoencoder based on a plurality of second sample trajectories using a joint loss function. The joint loss function is determined based on a non-clustering loss function and a clustering loss function. The pre-trained autoencoder is obtained by training an autoencoder based on a plurality of first sample trajectories using the non-clustering loss function. The non-clustering loss function is determined based on a weighting function and a reconstruction loss function. Wherein, the weighting function is used to obtain the weighting coefficient of the sample track based on the sample track and the sample reconstructed track of the sample track, and the reconstruction loss function is used to obtain the degree of difference between the sample track and the sample reconstructed track of the sample track based on the sample track and the sample reconstructed track of the sample track, wherein the sample track includes the first sample track or the second sample track, and the sample reconstructed track of the sample track includes the first sample reconstructed track of the first sample track or the second sample reconstructed track of the second sample track; The clustering loss function is used to evaluate the difference between the soft-assignment probability distribution matrix and the auxiliary probability distribution matrix. The soft-assignment probability distribution matrix includes the first element values of M×J first elements, and the auxiliary probability distribution matrix includes the second element values of M×J second elements. M represents the number of second sample tracks, J represents the number of cluster centers, and q mj q represents the value of the first element in the m-th row and j-th column. mj Used to indicate z m Belonging to the j-th cluster center u j The original probability of the corresponding cluster, p mj p represents the value of the second element in the m-th row and j-th column. mj Used to indicate z m Belonging to the j-th cluster center u j The corrected probability of the corresponding cluster, z m Let M and J represent the low-dimensional latent representation of the m-th second sample track, where M and J are integers greater than 1, m∈{1,……,M}, j∈{1,……,J}, and the initial value of the cluster center is the initial cluster center.
7. The method according to claim 6, characterized in that, The weighting function is determined based on the first weighting function and the second weighting function; Specifically, when the absolute value is less than a predetermined threshold, the weighting coefficient of the sample track is determined according to the first weighting function, which is the product of the predetermined coefficient and the absolute value; when the absolute value is greater than or equal to the predetermined threshold, the weighting coefficient of the sample track is determined according to the second weighting function, which is the absolute value, representing the absolute value of the difference between the sample track and the sample reconstructed track of the sample track, and the predetermined coefficient is greater than 0 and less than 1.
8. The method according to claim 6, characterized in that, q mj It is based on z m with u j The distance between and z m and The distance between them is determined. This represents the j'-th cluster center; Where, p mj It is based on q mj , and Definitely. This represents the sum of the first element values of the M first elements in the j-th column. Let represent the sum of the first element values of the M first elements in the j'-th column, where j'∈{1,……,J}.
9. The method according to claim 6, characterized in that, The pre-trained autoencoder was trained in the following manner: The multiple first sample tracks are input into the encoder and decoder of the initial autoencoder to obtain the first sample reconstructed track of each of the multiple first sample tracks; Based on the non-clustering loss function, the trajectory is reconstructed according to the plurality of first sample trajectories and the first sample of each of the plurality of first sample trajectories, and the non-clustering loss function value is obtained; The model parameters of the autoencoder are adjusted based on the non-clustering loss function value to obtain the pre-trained autoencoder, wherein the pre-trained autoencoder includes a pre-trained encoder and a pre-trained decoder; the plurality of target cluster centers and the trained autoencoder are obtained through training in the following manner: The multiple second sample tracks are input into the pre-trained autoencoder to obtain the second sample reconstructed tracks of each of the multiple second sample tracks; The plurality of second sample tracks are input into the pre-trained encoder to obtain the sample low-dimensional latent representation of each of the plurality of second sample tracks; Based on the joint loss function, the joint loss function value is obtained according to the plurality of second sample tracks, the second sample reconstructed tracks of each of the plurality of second sample tracks, the low-dimensional latent representations of each of the plurality of second sample tracks, and the plurality of initial cluster centers; The model parameters of the pre-trained autoencoder and the plurality of initial cluster centers are adjusted according to the joint loss function value to obtain the plurality of target cluster centers and the trained autoencoder.
Citation Information
Patent Citations
A method for ship route extraction and track deviation detection
CN109543715A
Four-dimensional track online abnormity detection method based on unsupervised learning
CN109977546A
Flight trajectory yaw detection method, electronic equipment and storage medium
CN116363908A