Abnormal trajectory detection method and device, electronic equipment and storage medium
By performing two clustering operations on historical trajectory data and utilizing a distributed computing engine, the problems of low accuracy and high cost in existing abnormal trajectory detection technologies have been solved, achieving efficient and accurate abnormal trajectory detection.
Patent Information
- Application Number
- CN202310307791.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-24
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2043-03-24
AI Technical Summary
Existing abnormal trajectory detection methods suffer from high labor costs and difficulties in mesh generation in unconstrained location spaces, resulting in low detection accuracy.
By performing first and second clustering on historical trajectory data, and using a distributed computing engine to process the trajectory data, the characteristics of the target trajectory model are obtained, and it is determined whether the trajectory to be detected is abnormal.
It improves the accuracy of abnormal trajectory detection, reduces computational complexity and labor costs, and enables real-time detection.
Smart Images

Figure CN116244356B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of data processing, and particularly relates to an abnormal trajectory detection method and device, an electronic device, and a storage medium. BACKGROUND
[0002] Abnormal trajectory detection is a hot issue in the field of trajectory data mining. The purpose of abnormal trajectory detection is mainly to find out abnormal trajectories that are uncommon and different from normal trajectories from a large amount of trajectory data, such as trajectories corresponding to behaviors that have significant differences in shape or travel time from other trajectories, or deviations in geographic location, sudden changes in motion direction and speed, etc.
[0003] Analyzing the found abnormal trajectories can obtain the causes and rules of the occurrence of abnormal trajectories, can more comprehensively reflect the user's daily behavior habits, and can effectively avoid the problem of trajectory abnormality, providing new ideas and directions for improving and updating positioning products.
[0004] At present, abnormal trajectory detection is usually performed by classification or grid division. When abnormal trajectory detection is performed by classification, the trajectory needs to be divided based on labels. In order to obtain labels that can accurately represent all abnormal classification situations, manual labeling is required, which consumes a high labor cost. When abnormal trajectory detection is performed by grid division, it is difficult to divide the grid when facing unconstrained and non-forced road network positions, and it is difficult to perform abnormal trajectory detection. SUMMARY
[0005] The present application aims to at least solve one of the technical problems existing in the prior art. To this end, the present application provides an abnormal trajectory detection method, device, electronic device and storage medium, which improves the accuracy of abnormal trajectory detection.
[0006] In a first aspect, the present application provides an abnormal trajectory detection method, which comprises: obtaining a to-be-detected trajectory;
[0007] Based on the to-be-detected trajectory and the target trajectory model feature, a target trajectory similarity is determined;
[0008] Based on the target trajectory similarity, trajectory state information of the to-be-detected trajectory is determined, the trajectory state information being a trajectory abnormal state or a trajectory normal state;
[0009] The target trajectory model feature is obtained by the following steps:
[0010] Based on the trajectory position information of the historical trajectory data set, the historical trajectory data set is first clustered to obtain at least one first clustering cluster;
[0011] Based on the trajectory similarity of the first cluster, the first cluster is further clustered to obtain at least one second cluster;
[0012] Based on the second cluster center corresponding to the second cluster, the features of the target trajectory model are obtained.
[0013] According to the abnormal trajectory detection method of this application, historical trajectory data is divided into first clusters of different regions through first clustering, and then second clustering is performed based on trajectory similarity to obtain the target trajectory model features corresponding to normal trajectory data in the same region. The target trajectory similarity is calculated to determine whether the trajectory to be detected is abnormal, which can improve the accuracy of abnormal trajectory detection.
[0014] According to one embodiment of this application, after performing a first clustering on the historical trajectory dataset based on trajectory location information to obtain at least one first cluster, and before performing a second clustering on the first cluster based on trajectory similarity to obtain at least one second cluster, the method further includes:
[0015] The historical trajectory data in the first cluster is processed by a distributed computing engine to obtain the elastic distributed dataset corresponding to the first cluster.
[0016] The step of performing a second clustering on the first cluster based on the trajectory similarity of the first cluster to obtain at least one second cluster includes:
[0017] Based on the trajectory similarity of the first cluster, the elastic distributed dataset corresponding to the first cluster is subjected to the second clustering to obtain at least one second cluster.
[0018] According to one embodiment of this application, the step of performing a second clustering on the first cluster based on the trajectory similarity of the first cluster to obtain at least one second cluster includes:
[0019] Based on the trajectory similarity of the first cluster, the number of target clusters of the first cluster is determined;
[0020] Based on the number of target clusters and the trajectory similarity of the first cluster, the historical trajectory data in the first cluster are clustered to obtain the second cluster.
[0021] According to one embodiment of this application, the step of clustering historical trajectory data in the first cluster based on the number of target clusters and the trajectory similarity of the first cluster to obtain the second cluster includes:
[0022] Based on the target number of clusters, an initial cluster center trajectory set is obtained, which includes the target number of initial cluster centers;
[0023] Based on the trajectory similarity between the initial cluster center trajectory set and the first cluster, K-means clustering is performed on the historical trajectory data in the first cluster to obtain the second cluster.
[0024] According to one embodiment of this application, determining the number of target clusters for the first cluster based on the trajectory similarity of the first cluster includes:
[0025] Based on the trajectory similarity of the first cluster, the first cluster is clustered according to the Canopy algorithm to determine the number of target clusters.
[0026] According to one embodiment of this application, after acquiring the trajectory to be detected and before determining the similarity of the target trajectory based on the features of the trajectory to be detected and the target trajectory model, the method further includes:
[0027] The target trajectory model features similar to the trajectory to be detected are found by indexing in the trajectory model feature database. The trajectory model feature database includes multiple trajectory model features.
[0028] Secondly, this application provides an abnormal trajectory detection device, the device comprising:
[0029] The acquisition module is used to acquire the trajectory to be detected;
[0030] The first processing module is used to determine the similarity of the target trajectory based on the features of the target trajectory model and the trajectory to be detected;
[0031] The second processing module is used to determine the trajectory status information of the trajectory to be detected based on the target trajectory similarity, wherein the trajectory status information is either an abnormal trajectory status or a normal trajectory status.
[0032] The target trajectory model features are obtained through the following steps:
[0033] Based on the trajectory location information of the historical trajectory dataset, the historical trajectory dataset is subjected to a first clustering to obtain at least one first cluster.
[0034] Based on the trajectory similarity of the first cluster, the first cluster is further clustered to obtain at least one second cluster;
[0035] Based on the second cluster center corresponding to the second cluster, the features of the target trajectory model are obtained.
[0036] According to the abnormal trajectory detection device of this application, historical trajectory data is divided into first clusters of different regions through first clustering, and then second clustering is performed based on trajectory similarity to obtain the target trajectory model features corresponding to normal trajectory data in the same region. The target trajectory similarity is calculated to determine whether the trajectory to be detected is abnormal, which can improve the accuracy of abnormal trajectory detection.
[0037] Thirdly, this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the abnormal trajectory detection method as described in the first aspect above.
[0038] Fourthly, this application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the abnormal trajectory detection method as described in the first aspect above.
[0039] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the abnormal trajectory detection method as described in the first aspect above.
[0040] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0041] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:
[0042] Figure 1 This is one of the flowcharts illustrating the abnormal trajectory detection method provided in the embodiments of this application;
[0043] Figure 2 This is a schematic diagram of the process for obtaining target trajectory model features provided in an embodiment of this application;
[0044] Figure 3 This is a histogram of the trajectory size versus running time for serial clustering and distributed parallel clustering provided in the embodiments of this application;
[0045] Figure 4 This is a second schematic flowchart of the abnormal trajectory detection method provided in the embodiments of this application;
[0046] Figure 5 This is the third flowchart illustrating the abnormal trajectory detection method provided in the embodiments of this application;
[0047] Figure 6 This is a flowchart illustrating the second clustering process provided in an embodiment of this application;
[0048] Figure 7 This is a schematic diagram of the abnormal trajectory detection device provided in the embodiments of this application;
[0049] Figure 8 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0050] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0051] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0052] The following is combined with Figures 1-8 The present application provides a detailed description of the abnormal trajectory detection method, abnormal trajectory detection device, electronic device, and readable storage medium through specific embodiments and application scenarios.
[0053] Among them, the abnormal trajectory detection method can be applied to the terminal, and can be executed by the hardware or software in the terminal.
[0054] The terminal includes, but is not limited to, portable communication devices such as mobile phones or tablets with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads). It should also be understood that, in some embodiments, the terminal may not be a portable communication device, but rather a desktop computer with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads).
[0055] The following embodiments describe a terminal including a display and a touch-sensitive surface. However, it should be understood that the terminal may include one or more other physical user interface devices such as a physical keyboard, mouse, and joystick.
[0056] The abnormal trajectory detection method provided in this application embodiment can be executed by an electronic device or a functional module or entity in an electronic device that can implement the abnormal trajectory detection method. The electronic devices mentioned in this application embodiment include, but are not limited to, mobile phones, tablets, computers, cameras, and wearable devices. The abnormal trajectory detection method provided in this application embodiment will be described below using an electronic device as the execution subject as an example.
[0057] like Figure 1 As shown, the abnormal trajectory detection method includes steps 110 to 130.
[0058] Step 110: Obtain the trajectory to be detected.
[0059] Among them, the trajectory to be detected is the trajectory that needs to be analyzed for abnormal trajectories.
[0060] In practice, the trajectory to be detected can be the movement trajectory of a user holding a mobile device in the city, or the movement trajectory of a vehicle in the city.
[0061] In this embodiment, the positioning system in the mobile device can record the user's location information at a certain moment. By recording the user's location information at multiple moments, the user's movement trajectory can be obtained. The mobile device will save the user's movement trajectory, and the dashcam in the vehicle can also save the vehicle's movement trajectory.
[0062] In practice, the movement trajectory of the corresponding vehicle in the dashcam can be read as the trajectory to be detected, or the trajectory to be detected can be obtained from the user's personal trajectory database.
[0063] Step 120: Determine the similarity of the target trajectory based on the model features of the trajectory to be detected and the target trajectory.
[0064] Among them, the target trajectory model features can be obtained by analyzing historical trajectory data to identify the model features of a normal trajectory.
[0065] In this step, the similarity between the target trajectory and the model features of the trajectory to be detected is calculated. The similarity between the target trajectory and the model features of the target trajectory can characterize the degree of similarity between the model features of the target trajectory and the trajectory to be detected, thereby determining whether the trajectory to be detected is an abnormal trajectory.
[0066] For example, the similarity between the target trajectory and the average difference between the distances between points on the model features of the trajectory to be detected and the target trajectory can be used to determine whether the trajectory to be detected is an abnormal trajectory.
[0067] In this embodiment, the difference between each point in the trajectory to be detected and each point in the feature of the target trajectory model can be calculated based on the distance between each point. After error analysis and normalization, the average difference between the distances between each point in the trajectory to be detected and the feature of the target trajectory model can be used as the similarity of the target trajectory.
[0068] For example, the similarity between the target trajectory and the longest common sub-trajectory of the model features of the trajectory to be detected can be used to determine whether the trajectory to be detected is an abnormal trajectory.
[0069] In this embodiment, the trajectory sequence of the trajectory to be detected can be obtained first. The coordinates of each point in the trajectory sequence are (y1, h1, t1), where y1 represents the difference in latitude and longitude from the previous point, h1 represents the difference in altitude from the previous point, and t1 represents the timestamp from the previous point.
[0070] To calculate the positional information corresponding to the trajectory sequence of the target trajectory, in actual execution, each point is converted into a multi-dimensional feature point x(x1, x2, x3, x4, x5), where x1 represents the longitude of the point, x2 represents the latitude of the point, x3 represents the velocity of the point, x4 represents the curvature of the point, and x5 represents the distance of the point from the previous point, thus obtaining a trajectory sequence including N multi-dimensional feature points. .
[0071] Based on the latitude and longitude of the multidimensional feature points in the trajectory to be detected, the LCSS algorithm can be used to calculate the distance between the trajectory to be detected and the target trajectory model features. Then, according to the formula The similarity score of the target trajectory is determined, with a value ranging from 0 to 1. Here, A represents the trajectory to be detected, B represents the target trajectory model feature, N represents the length of the trajectory to be detected, and M represents the length of the target trajectory model feature. This represents the minimum value among the features of the trajectory to be detected and the target trajectory model.
[0072] For example, The value of can be 2, the value of N is 6, and the value of M is 8, according to the formula. The target trajectory similarity is 0.3.
[0073] In practice, the LCSS algorithm is highly resistant to interference. Using the LCSS algorithm to calculate trajectory similarity can accurately obtain the trajectory similarity between two trajectories in different environments.
[0074] Among them, such as Figure 2 As shown, the target trajectory model features can be obtained through the following steps:
[0075] Step 210: Based on the trajectory location information of the historical trajectory dataset, perform the first clustering on the historical trajectory dataset to obtain at least one first cluster.
[0076] The historical trajectory dataset includes multiple historical trajectory data sets, which can be trajectory data corresponding to the user's normal trajectory.
[0077] In this embodiment, the trajectory location information represents the geographical location information corresponding to the historical trajectory data in the historical trajectory dataset, which may include information such as latitude and longitude.
[0078] In this embodiment, the average latitude and longitude of all points on the historical trajectory data can be determined based on the coordinate sequence of points on the historical trajectory data in the historical trajectory dataset, and the geographical location information of the points on the historical trajectory data can be determined using a geographic segmentation algorithm.
[0079] Based on the geographical location information of historical trajectory data, the historical trajectory data in the historical trajectory dataset is first clustered, dividing multiple historical trajectory data in the historical trajectory dataset into first clusters corresponding to different geographical regions. Then, the historical trajectory data corresponding to the first cluster of the same region can be further clustered, which improves the efficiency of clustering.
[0080] For example, based on geographical location information, historical trajectory datasets can be divided into Shanghai trajectory sets, Hefei trajectory sets, and Beijing trajectory sets.
[0081] It is understandable that when there is little historical trajectory data corresponding to certain geographical location information, the number of data objects in the first cluster is small. The underlying influencing factors cannot be reflected from the small amount of data, indicating that the historical trajectory data corresponding to the geographical location information does not have data mining value. After performing the first clustering, the historical trajectory data in this partition is filtered out and no further processing steps are performed.
[0082] In practice, a trajectory filtering threshold can be set. After the first clustering is completed, multiple first clusters are obtained. When the number of historical trajectory data in a first cluster is less than the trajectory filtering threshold, all historical trajectory data in that first cluster is filtered out, and the historical trajectory data in that first cluster is not processed in subsequent steps.
[0083] By filtering historical trajectory data in the first cluster where the number of historical trajectory data is less than the trajectory filtering threshold, the amount of historical trajectory data that needs to be clustered in subsequent clustering can be reduced, improving the efficiency of clustering. It also ensures that the data obtained from subsequent clustering has data mining value, avoids invalid data processing, and thus improves the detection efficiency of abnormal trajectories.
[0084] Step 220: Based on the trajectory similarity of the first cluster, perform a second clustering on the first cluster to obtain at least one second cluster.
[0085] The trajectory similarity can be the trajectory similarity between pairwise historical trajectory data in the first cluster, and the trajectory similarity is the clustering criterion of the second cluster.
[0086] In this embodiment, based on trajectory similarity, historical trajectory data with similar trajectories are grouped into the same second cluster, resulting in one or more second clusters.
[0087] Understandably, the historical trajectory data of the second cluster can characterize the user's normal trajectory behavior habits.
[0088] Taking a user's Shanghai trajectory cluster as the first cluster as an example.
[0089] Based on the trajectory similarity of the user's Shanghai trajectory cluster, a second clustering is performed on the Shanghai trajectory cluster, resulting in three second clusters. These three second clusters correspond to the user's movement trajectories during the three time periods of commuting to and from get off work and weekends, respectively.
[0090] Step 230: Based on the second cluster center corresponding to the second cluster, obtain the target trajectory model features.
[0091] The second cluster center can be the center point of the historical trajectory data in the second cluster.
[0092] In this embodiment, the second cluster center can characterize the model features of the normal trajectory corresponding to the second cluster, and then the target trajectory model features characterizing the normal trajectory can be determined based on the second cluster center corresponding to the second cluster.
[0093] Step 130: Determine the trajectory status information of the trajectory to be detected based on the similarity of the target trajectory.
[0094] The trajectory status information includes either an abnormal trajectory status or a normal trajectory status.
[0095] In this step, if the trajectory to be detected is not similar to the target trajectory model features based on the target trajectory similarity, the trajectory to be detected can be determined to be an abnormal trajectory. If the trajectory to be detected is similar to the target trajectory model features based on the target trajectory similarity, the trajectory to be detected can be determined to be a normal trajectory.
[0096] In practice, a similarity detection threshold can be set. When the similarity of the target trajectory is less than the similarity detection threshold, the trajectory to be detected is determined to be an abnormal trajectory. When the similarity of the target trajectory is greater than the similarity detection threshold, the trajectory to be detected is determined to be a normal trajectory.
[0097] For example, the similarity detection threshold can be set to 0.85. When the similarity of the target trajectory is less than 0.85, the trajectory to be detected is determined to be an abnormal trajectory. When the similarity of the target trajectory is greater than 0.85, the trajectory to be detected is determined to be a normal trajectory.
[0098] Among related technologies, some techniques have emerged that use clustering to identify abnormal trajectories. These techniques typically cluster the trajectory to be identified with historical trajectories. When the trajectory to be identified is not in the cluster obtained from the clustering, it is considered an abnormal trajectory. However, these methods have low detection accuracy.
[0099] In this embodiment, by performing a first clustering on the historical trajectory dataset based on the trajectory location information corresponding to different historical trajectory data, the historical trajectory data is divided into first clusters of different regions. This allows for subsequent clustering of historical trajectory data in the same region, improving the efficiency of clustering. Then, a second clustering based on trajectory similarity is performed on the first clusters, dividing the historical trajectory data into multiple second clusters corresponding to different trajectory features. Subsequently, based on the target trajectory similarity between the target trajectory model features and the trajectory to be detected, it is determined whether the trajectory to be detected is an abnormal trajectory. By accurately extracting normal trajectories in the same region through two clustering operations to construct target trajectory model features, and using the target trajectory model features to determine whether the trajectory to be detected is abnormal, the accuracy of abnormal trajectory detection can be effectively improved.
[0100] According to the abnormal trajectory detection method provided in the embodiments of this application, historical trajectory data is divided into first clusters of different regions through first clustering, and then second clustering is performed based on trajectory similarity to obtain target trajectory model features corresponding to normal trajectory data in the same region. Target trajectory similarity is calculated to determine whether the trajectory to be detected is abnormal, which can improve the accuracy of abnormal trajectory detection.
[0101] In some embodiments, after step 210, performing a first clustering on the historical trajectory dataset based on the trajectory location information of the historical trajectory dataset to obtain at least one first cluster, and before step 220, performing a second clustering on the first cluster based on the trajectory similarity of the first cluster to obtain at least one second cluster, the abnormal trajectory detection method further includes:
[0102] The historical trajectory data in the first cluster is processed by the distributed computing engine to obtain the elastic distributed dataset corresponding to the first cluster.
[0103] Step 230: Based on the trajectory similarity of the first cluster, perform a second clustering on the first cluster to obtain at least one second cluster, including:
[0104] Based on the trajectory similarity of the first cluster, a second cluster is performed on the elastic distributed dataset corresponding to the first cluster to obtain at least one second cluster.
[0105] The distributed computing engine can be Apache Spark, and the elastic distributed dataset is a computing container in the distributed computing engine. All data processing in the distributed computing engine is based on the elastic distributed dataset. Converting the first cluster into the corresponding elastic distributed dataset can facilitate data processing by the distributed computing engine.
[0106] In this embodiment, when performing the second clustering, the first cluster is deployed to the distributed computing engine. The distributed computing engine can transform the first cluster into a corresponding elastic distributed dataset and perform the second clustering based on the elastic distributed dataset to obtain the second cluster corresponding to the elastic distributed dataset.
[0107] In related technologies, conventional clustering methods are computationally inefficient when faced with a large amount of historical trajectory data. This is because the algorithm needs to perform N×(N-1) trajectory similarity calculations (where N is the number of trajectories), making it impossible to meet the requirements for real-time detection.
[0108] In this embodiment, by handing over the historical trajectory data in at least one first cluster to a distributed computing engine and adopting a distributed programming model, the historical trajectory data in multiple first clusters are calculated and then summarized simultaneously. When faced with a large amount of historical trajectory data, the clustering efficiency and accuracy of the second cluster are greatly improved, thereby improving the efficiency of trajectory anomaly detection and enabling the real-time detection requirement.
[0109] like Figure 3 As shown, historical trajectory data from at least one first cluster is handed over to a distributed computing engine. Using a distributed programming model, historical trajectory data from multiple first clusters are calculated simultaneously and then summarized and compared with conventional serial clustering methods. By comparing historical trajectory data from multiple first clusters and performing serial calculations, the required running time is significantly reduced when faced with a large amount of historical trajectory data.
[0110] In some embodiments, step 220, performing a second clustering on the first cluster based on the trajectory similarity of the first cluster to obtain at least one second cluster, may include:
[0111] Based on the trajectory similarity of the first cluster, the number of target clusters for the first cluster is determined;
[0112] Based on the number of target clusters and the trajectory similarity of the first cluster, the historical trajectory data in the first cluster are clustered to obtain the second cluster.
[0113] The target cluster number is the number of second clusters to be obtained when performing the second clustering.
[0114] In this embodiment, the trajectory similarity between historical trajectory data in the first cluster can be determined, and the target number of the first cluster can be determined based on multiple trajectory similarities. Based on the target number of clusters and the trajectory similarity of the first cluster, the historical trajectory data in the first cluster can be clustered to obtain the target number of second clusters.
[0115] It should be noted that by pre-determining the number of target clusters based on the trajectory similarity of the first cluster, the subsequent clustering process can avoid falling into the "local optimum trap" and improve the efficiency of the clustering process.
[0116] In some embodiments, based on the number of target clusters and the trajectory similarity of the first cluster, historical trajectory data in the first cluster are clustered to obtain a second cluster, including:
[0117] Based on the target number of clusters, an initial cluster center trajectory set is obtained, which includes the target number of initial cluster centers.
[0118] Based on the trajectory similarity between the initial cluster center trajectory set and the first cluster, K-means clustering is performed on the historical trajectory data in the first cluster to obtain the second cluster.
[0119] The initial cluster center trajectory set includes historical trajectory data corresponding to multiple predetermined initial cluster centers.
[0120] In this embodiment, the initial cluster center trajectory set can be obtained using the K-means++ method.
[0121] In practice, a random sampling algorithm can be used to randomly select a historical trajectory data point from the elastic distributed dataset as the second sample historical trajectory data point. This second sample historical trajectory data point is then used as an initial cluster center, and the clustering is performed according to the formula...
[0122]
[0123] Calculate the probability that all historical trajectory data in the elastic distributed dataset, excluding the second sample historical trajectory data, will be selected as the next initial cluster center.
[0124] in, , This represents the trajectory similarity between the historical trajectory data (excluding the second sample historical trajectory data) and the historical trajectory data of the second sample in the elastic distributed dataset.
[0125] The historical trajectory data with the highest probability is used as the next initial cluster center, and this historical trajectory data is used as the new second sample historical trajectory data. The above iterative operation of calculating the probability and determining the next initial cluster center is repeated until the target cluster number of initial cluster centers is obtained. The target cluster number of initial cluster centers are summarized to obtain the initial cluster center trajectory set, and the first cluster is initialized using the initial cluster center trajectory set.
[0126] In related technologies, when using K-means for clustering, it is usually necessary to select K initial cluster centers in the sample set. If they are simply selected randomly from the sample set, the initial cluster centers obtained will differ significantly from the actual cluster centers, resulting in slow convergence of the K-means clustering algorithm and low efficiency in clustering.
[0127] In this embodiment of the application, by determining the initial cluster centers based on the initial cluster center trajectory set before performing the second clustering on the first cluster cluster, the problem of randomly selecting the initial cluster centers during K-means clustering can be avoided, thereby accelerating the convergence speed of the K-means algorithm and improving the efficiency of clustering processing.
[0128] like Figure 6 As shown, after completing the initialization operation of the K-means clustering algorithm, the distributed computing engine can save the initial cluster center trajectory set to broadcast. Then, it uses mapPartitions to slice the elastic distributed dataset to obtain multiple local elastic distributed datasets. The initial cluster centers in broadcast are distributed to the corresponding local elastic distributed datasets, and the computing cores in multiple distributed computing engines are started in parallel to process multiple local elastic distributed datasets.
[0129] For each locally resilient distributed dataset, the computational core vectorizes the locally resilient distributed dataset. By calculating the trajectory similarity between the vector elements in the locally resilient distributed dataset and the corresponding initial cluster centers, it determines which initial cluster center has the greatest trajectory similarity between the vector elements in the locally resilient distributed dataset and the first cluster center. The vector element is then added to the cluster corresponding to the initial cluster center with the greatest trajectory similarity to the first cluster center, thus obtaining the second cluster.
[0130] In actual implementation, within each locally elastic distributed dataset, the computational core utilizes the ReduceByKey operator for re-aggregation calculation, calculating the average similarity of the second clusters in the locally elastic distributed dataset. The historical trajectory data with the highest average similarity among the historical trajectory data in each second cluster is taken as the second cluster center of that second cluster. All locally elastic distributed datasets are aggregated to obtain multiple second clusters. It is determined whether the second cluster centers in each second cluster have converged and remained unchanged, or whether the distance between them and the initial cluster centers is less than the cluster center threshold. If neither of these conditions is met, the second clusters that meet these conditions are removed. The remaining second clusters are then iteratively calculated again until convergent iteration results are obtained. The trajectories corresponding to multiple second cluster centers are extracted and used as trajectory model features, saved to a non-volatile database for use in subsequent steps.
[0131] Non-volatile databases refer to databases whose stored data will not be lost when the current is cut off.
[0132] By using a non-volatile database to store trajectory model features, it can be ensured that the target trajectory model features can be directly obtained from the database when performing anomaly detection on trajectory data in the future.
[0133] In some embodiments, determining the number of target clusters for the first cluster based on the trajectory similarity of the first cluster includes:
[0134] Based on the trajectory similarity of the first cluster, the first cluster is clustered according to the Canopy algorithm to determine the number of target clusters.
[0135] In this embodiment, a first similarity threshold and a second similarity threshold can be preset, wherein the first similarity threshold is greater than the second similarity threshold, and a first sample historical trajectory data is randomly selected from the elastic distributed dataset corresponding to the first cluster.
[0136] One approach is to use a random sampling algorithm to randomly select a historical trajectory data point from the elastic distributed dataset as the first sample historical trajectory data.
[0137] In practice, according to the Canopy algorithm, the trajectory similarity between the first sample historical trajectory data and other historical trajectory data in the elastic distributed dataset is calculated. A Canopy class buffer is created for historical trajectory data with a trajectory similarity less than the first similarity threshold, and historical trajectory data with a trajectory similarity greater than the second similarity threshold is deleted from the Canopy class buffer. The number of Canopy classes obtained by traversing all historical trajectory data in the elastic distributed dataset is the number of target clusters.
[0138] By using the Canopy algorithm, the number of target clusters is obtained. Based on the number of target clusters, the number of initial cluster centers required for subsequent clustering is determined, thus solving the problem that the number of target clusters cannot be determined in subsequent clustering processes and improving the clustering efficiency of subsequent clustering processes.
[0139] In some embodiments, after obtaining the trajectory to be detected in step 110 and before determining the similarity of the target trajectory based on the features of the trajectory to be detected and the target trajectory model in step 120, the abnormal trajectory detection method further includes:
[0140] By indexing the trajectory model feature database, target trajectory model features similar to the trajectory to be detected are found. The trajectory model feature database includes multiple trajectory model features.
[0141] The trajectory model feature database can be a non-volatile database.
[0142] The trajectory model features are the historical trajectory data corresponding to the second cluster center. The trajectory model feature database includes historical trajectory data corresponding to multiple second cluster centers.
[0143] The target trajectory model features are trajectory model features in the trajectory model feature database that are similar to the trajectory to be detected (e.g., trajectory model features with a similarity greater than a certain threshold).
[0144] In this embodiment, the index can be the time information corresponding to the model trajectory features. For example, the index corresponding to the trajectory model features can be the user's work hours, and the trajectory model features can be the historical trajectory data corresponding to the user's work hours.
[0145] When the trajectory to be detected is also the trajectory data of the user during their work hours, the corresponding trajectory model features can be found in the trajectory model feature database through indexing, and the trajectory model features similar to the trajectory to be detected can be used as the target trajectory model features.
[0146] The following is a specific embodiment to describe the abnormal trajectory detection method provided in this application.
[0147] likeFigure 4 As shown, the input historical trajectory dataset is cleaned, transformed, and decompressed to obtain a trajectory similarity matrix, which includes the pairwise trajectory similarity between historical trajectory data in the historical trajectory dataset.
[0148] Based on the average latitude and longitude of each point in the historical trajectory data, the historical trajectory data is classified according to geohash encoding, thus obtaining the first cluster of different geographical blocks.
[0149] like Figure 5 As shown, based on geohash encoding, the historical trajectory dataset is first clustered to obtain independent regional divisions based on the encoding, and then the first clusters of different geographical blocks are obtained, such as the Hefei trajectory cluster, the Shanghai trajectory cluster, and the Beijing trajectory cluster.
[0150] like Figure 4 As shown, based on multiple different first clusters, the distributed computing engine distributes the data to multiple kernel tasks, uses multiple kernels to parallelize the clustering algorithm, uses the Canopy algorithm to determine the number of target clusters, uses K-means++ to determine the initial set of cluster centers, and then uses the K-means clustering algorithm to determine multiple trajectory model features.
[0151] Based on the index, the trajectory model features with the same index as the trajectory to be detected are found in the trajectory model feature database. The trajectory model features similar to the trajectory to be detected are used as the target trajectory model features. Then, the state information of the trajectory to be detected is determined based on the target trajectory similarity between the trajectory to be detected and the target trajectory features.
[0152] The abnormal trajectory detection method provided in this application can be executed by an abnormal trajectory detection device. This application uses an abnormal trajectory detection device executing the abnormal trajectory detection method as an example to illustrate the abnormal trajectory detection device provided in this application.
[0153] This application also provides an abnormal trajectory detection device.
[0154] like Figure 7 As shown, the abnormal trajectory detection device includes:
[0155] The acquisition module 710 is used to acquire the trajectory to be detected;
[0156] The first processing module 720 is used to determine the similarity between the target trajectory and the model features of the trajectory to be detected and the target trajectory.
[0157] The second processing module 730 is used to determine the trajectory status information of the trajectory to be detected based on the similarity of the target trajectory. The trajectory status information is either an abnormal trajectory status or a normal trajectory status.
[0158] The target trajectory model features are obtained through the following steps:
[0159] Based on the trajectory location information of the historical trajectory dataset, the historical trajectory dataset is subjected to the first clustering to obtain at least one first cluster.
[0160] Based on the trajectory similarity of the first cluster, the first cluster is further clustered into a second cluster to obtain at least one second cluster.
[0161] Based on the second cluster center corresponding to the second cluster, the target trajectory model features are obtained.
[0162] According to the abnormal trajectory detection device provided in the embodiments of this application, historical trajectory data is divided into first clusters of different regions through first clustering, and then second clustering is performed based on trajectory similarity to obtain target trajectory model features corresponding to normal trajectory data in the same region. Target trajectory similarity is calculated to determine whether the trajectory to be detected is abnormal, which can improve the accuracy of abnormal trajectory detection.
[0163] In some embodiments, a processing module is used to process historical trajectory data in a first cluster through a distributed computing engine to obtain an elastic distributed dataset corresponding to the first cluster.
[0164] The first processing module 720 is used to perform a second clustering on the elastic distributed dataset corresponding to the first cluster based on the trajectory similarity of the first cluster, so as to obtain at least one second cluster.
[0165] In some embodiments, the first processing module 720 is further configured to determine the number of target clusters of the first cluster based on the trajectory similarity of the first cluster;
[0166] Based on the number of target clusters and the trajectory similarity of the first cluster, the historical trajectory data in the first cluster are clustered to obtain the second cluster.
[0167] In some embodiments, the first processing module 720 is further configured to obtain an initial cluster center trajectory set based on the target number of clusters, the initial cluster center trajectory set including the target number of initial cluster centers;
[0168] Based on the trajectory similarity between the initial cluster center trajectory set and the first cluster, K-means clustering is performed on the historical trajectory data in the first cluster.
[0169] In some embodiments, the first processing module 720 is further configured to perform clustering processing on the first cluster based on the trajectory similarity of the first cluster according to the Canopy algorithm, and determine the number of target clusters.
[0170] In some embodiments, the first processing module 720 is further configured to find target trajectory model features similar to the trajectory to be detected by indexing in the trajectory model feature database, the trajectory model feature database including multiple trajectory model features.
[0171] The abnormal trajectory detection device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.
[0172] The abnormal trajectory detection device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.
[0173] The abnormal trajectory detection device provided in this application embodiment can achieve... Figures 1 to 6 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.
[0174] In some embodiments, such as Figure 8 As shown, this application embodiment also provides an electronic device 800, including a processor 801, a memory 802, and a computer program stored in the memory 802 and executable on the processor 801. When the program is executed by the processor 801, it implements the various processes of the above-described abnormal trajectory detection method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0175] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0176] This application also provides a non-transitory computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described abnormal trajectory detection method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0177] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0178] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described abnormal trajectory detection method.
[0179] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0180] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0181] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0182] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
[0183] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0184] Although embodiments of this application have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the claims and their equivalents.
Claims
1. An abnormal trajectory detection method, characterized in that, include: Obtain the trajectory to be detected; Based on the features of the target trajectory and the model of the trajectory to be detected, the similarity of the target trajectory is determined; Based on the target trajectory similarity, the trajectory status information of the trajectory to be detected is determined, and the trajectory status information is either an abnormal trajectory status or a normal trajectory status. The target trajectory model features are obtained through the following steps: Based on the trajectory location information of the historical trajectory dataset, the historical trajectory dataset is subjected to a first clustering to obtain at least one first cluster. Based on the trajectory similarity of the first cluster, the first cluster is further clustered to obtain at least one second cluster; Based on the second cluster center corresponding to the second cluster, the features of the target trajectory model are obtained; After performing a first clustering on the historical trajectory dataset based on trajectory location information to obtain at least one first cluster, and before performing a second clustering on the first cluster based on trajectory similarity to obtain at least one second cluster, the method further includes: The historical trajectory data in the first cluster is processed by a distributed computing engine to obtain the elastic distributed dataset corresponding to the first cluster. The step of performing a second clustering on the first cluster based on the trajectory similarity of the first cluster to obtain at least one second cluster includes: Based on the trajectory similarity of the first cluster, the elastic distributed dataset corresponding to the first cluster is subjected to the second clustering to obtain at least one second cluster.
2. The abnormal trajectory detection method according to claim 1, characterized in that, The step of performing a second clustering on the first cluster based on the trajectory similarity of the first cluster to obtain at least one second cluster includes: Based on the trajectory similarity of the first cluster, the number of target clusters of the first cluster is determined; Based on the number of target clusters and the trajectory similarity of the first cluster, the historical trajectory data in the first cluster are clustered to obtain the second cluster.
3. The abnormal trajectory detection method according to claim 2, characterized in that, The second cluster is obtained by clustering historical trajectory data in the first cluster based on the number of target clusters and the trajectory similarity of the first cluster, including: Based on the target number of clusters, an initial cluster center trajectory set is obtained, which includes the target number of initial cluster centers; Based on the trajectory similarity between the initial cluster center trajectory set and the first cluster, K-means clustering is performed on the historical trajectory data in the first cluster to obtain the second cluster.
4. The abnormal trajectory detection method according to claim 2, characterized in that, Determining the target number of clusters for the first cluster based on the trajectory similarity of the first cluster includes: Based on the trajectory similarity of the first cluster, the first cluster is clustered according to the Canopy algorithm to determine the number of target clusters.
5. The abnormal trajectory detection method according to any one of claims 1-4, characterized in that, After acquiring the trajectory to be detected and before determining the similarity of the target trajectory based on the features of the trajectory to be detected and the target trajectory model, the method further includes: The target trajectory model features similar to the trajectory to be detected are found by indexing in the trajectory model feature database. The trajectory model feature database includes multiple trajectory model features.
6. An abnormal trajectory detection device, suitable for employing the abnormal trajectory detection method according to any one of claims 1-5, characterized in that, include: The acquisition module is used to acquire the trajectory to be detected; The first processing module is used to determine the similarity of the target trajectory based on the features of the target trajectory model and the trajectory to be detected; The second processing module is used to determine the trajectory status information of the trajectory to be detected based on the target trajectory similarity, wherein the trajectory status information is either an abnormal trajectory status or a normal trajectory status. The target trajectory model features are obtained through the following steps: Based on the trajectory location information of the historical trajectory dataset, the historical trajectory dataset is subjected to a first clustering to obtain at least one first cluster. Based on the trajectory similarity of the first cluster, the first cluster is further clustered to obtain at least one second cluster; Based on the second cluster center corresponding to the second cluster, the features of the target trajectory model are obtained.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the abnormal trajectory detection method as described in any one of claims 1-5.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the abnormal trajectory detection method as described in any one of claims 1-5.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the abnormal trajectory detection method as described in any one of claims 1-5.