A trajectory similarity analysis method based on spatiotemporal information extraction, electronic device and storage medium

Through the trajectory similarity analysis method based on spatiotemporal information extraction, the Transformer model and dimensionality reduction clustering technology are used to solve the problem of neglecting time factors in the existing technology, and a more accurate and efficient trajectory similarity evaluation is achieved.

CN118484481BActive Publication Date: 2025-05-02HARBIN INST OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410663869.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-27
Publication Date
2025-05-02
Estimated Expiration
2044-05-27

AI Technical Summary

Technical Problem

The existing trajectory similarity measurement methods mainly focus on spatial characteristics and ignore time factors, resulting in limited applicability to time-sensitive tasks.

Method used

A trajectory similarity analysis method based on spatiotemporal information extraction is proposed. By collecting and preprocessing spatiotemporal trajectory data, using the Transformer model with position code layer for processing, combined with T-SNE dimensionality reduction and K-Means clustering, the trajectory similarity analysis results based on spatiotemporal information extraction are obtained.

Benefits of technology

This method can more comprehensively and accurately evaluate trajectory similarity, improve the efficiency and accuracy of data analysis, and adapt to application needs with strong time correlation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118484481B_ABST
    Figure CN118484481B_ABST
Patent Text Reader

Abstract

A trajectory similarity analysis method, electronic device and storage medium based on spatiotemporal information extraction belong to the technical field of urban intelligent computing and data mining. In order to evaluate trajectory similarity more comprehensively and accurately, the present invention collects spatiotemporal trajectory data, performs data cleaning and encoding preprocessing, obtains the preprocessed spatiotemporal trajectory data and inputs it into a Transformer model with a position code layer for processing, and outputs the spatiotemporal trajectory characterization result; the spatiotemporal trajectory characterization result is subjected to T-SNE dimensionality reduction processing through cosine similarity, and then K-Means clustering is performed to obtain a visualization result of trajectory similarity analysis based on spatiotemporal information extraction. The present invention converts complex trajectory data into an easy-to-process vector form, and then uses a similarity analysis algorithm to efficiently compare and classify the trajectory data. This can not only improve the efficiency of data analysis, but also improve the accuracy and reliability of the analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of urban intelligent computing and data mining, and specifically relates to a trajectory similarity analysis method based on spatiotemporal information extraction, an electronic device and a storage medium. Background Art

[0002] Exploring trajectory similarity analysis is not only a basic research direction in trajectory research, but also a basic direction in trajectory data analysis. It plays an important role in application scenarios such as trajectory anomaly detection, trajectory clustering, trajectory anomaly identification, and path recommendation. It involves using a specific algorithm to calculate the similarity score between two or more spatiotemporal trajectories, which may lead to completely different results for different algorithms, so a lot of comparative analysis is required based on data characteristics. However, in many application scenarios, temporal features are indispensable because they can reveal the specific moment when the trajectory occurs, which is crucial to understanding the dynamics of the trajectory. Most of the current trajectory similarity measurement methods focus on spatial characteristics and ignore temporal factors, which limits their applicability to time-sensitive tasks. Summary of the invention

[0003] The problem to be solved by the present invention is to evaluate trajectory similarity more comprehensively and accurately, and propose a trajectory similarity analysis method based on spatiotemporal information extraction, an electronic device and a storage medium.

[0004] To achieve the above object, the present invention is implemented through the following technical solutions:

[0005] A trajectory similarity analysis method based on spatiotemporal information extraction includes the following steps:

[0006] S1. Collect spatiotemporal trajectory data, perform data cleaning and encoding preprocessing, and obtain preprocessed spatiotemporal trajectory data;

[0007] S2. The preprocessed spatiotemporal trajectory data obtained in step S1 is input into a Transformer model with a position code layer for processing, and the spatiotemporal trajectory representation result is output;

[0008] S3. The spatiotemporal trajectory representation result obtained in step S2 is subjected to T-SNE dimensionality reduction processing through cosine similarity, and then K-Means clustering is performed to obtain the visualization result of trajectory similarity analysis based on spatiotemporal information extraction.

[0009] Furthermore, the specific implementation method of step S1 includes the following steps:

[0010] S1.1 Data cleaning:

[0011] S1.1.1 Collect spatiotemporal trajectory data. First, check whether the user ID item and the acquisition time item have null values, delete the entire data with null values, and then filter the vacant positions of the longitude item or latitude item and fill the vacant positions with 0;

[0012] S1.1.2 delete duplicate data, ping-pong switching data, drift data and range error data to obtain the spatiotemporal trajectory data after data cleaning;

[0013] S1.2. Encode and preprocess the spatiotemporal trajectory data after data cleaning in step S1.1, convert the two-dimensional discrete spatial coordinates into a one-dimensional digital sequence, and obtain the continuous trajectory sequence data TR of each user after preprocessing, which is represented as {TR: tr = (pot_1, pot_2, ..., pot_n)}, where each trajectory point is represented as pot (lat, lon). Lat represents latitude, and lon represents longitude.

[0014] Furthermore, the specific implementation method of step S1.2 includes the following steps:

[0015] S1.2.1. First, traverse all spatiotemporal trajectory data, find out the ping-pong switching data with short switching time interval and frequent switching back and forth between two base station locations, and delete them;

[0016] S1.2.2. Process the drift data in the spatiotemporal trajectory data, set the urban drift speed threshold, filter the records of the same user, calculate the distance and time interval between the i-th record and the i+1-th record based on the longitude and latitude information, and obtain the speed v between the two points. If v is greater than the urban drift speed threshold, delete the i+1-th record and recalculate the speed between the i-th record and the i+2-th record, and so on, until the above process is applied to all spatiotemporal trajectory data;

[0017] S1.2.3. Filter the data with range errors, find out the data that does not fall within the reasonable range of the location where the data was acquired through the set longitude and latitude range, and delete it.

[0018] Furthermore, the specific implementation method of step S2 includes the following steps:

[0019] S2.1. Input the continuous trajectory sequence data obtained in step S1 into the Transformer model with a position code layer to obtain the position encoding of the trajectory. The position encoding is to map the position information to the embedding space of the model, and the expression is:

[0020]

[0021]

[0022] Among them, PE(pos,2i) is the sine value of the position encoding of position pos and dimension 2i; PE(pos,2i+1) is the cosine value of the position encoding of position pos and dimension 2i+1; pos is the index of the position, i is the index of the dimension, and d model is the embedding dimension of the model;

[0023] S2.2. The position encoding of the trajectory is fused with the self-attention layer tensor through the fusion function to obtain the output result of the self-attention layer. The fusion function is calculated based on the calculation method of the matrix dot product Attn(Q, K, V), and the expression is:

[0024]

[0025] Among them, Q is the query matrix, V is the value matrix, K is the key matrix, and d k is the dimension of the key vector;

[0026] S2.3. The output of the self-attention layer is input into the feedforward neural network layer for processing. The feedforward neural network layer consists of a simple fully connected layer and a ReLU activation function. After being processed by the encoder, the output is a tensor containing the high-level feature representation of the input trajectory data, expressed as:

[0027] Result (i) =ReLU(A i W+b);

[0028] Among them, A i is the data matrix of the i-th trajectory, W is the weight matrix, b is the bias vector, Result (i) is the output result of the i-th trajectory data after transformation and activation function;

[0029] S2.4. Use average pooling to reduce the dimension of the tensor containing the high-level feature representation of the input trajectory data and output the spatiotemporal trajectory representation result;

[0030] Set the dimension of the tensor TR containing the high-level feature representation of the input trajectory data to (N, C, H, W), where N is the batch size, C is the number of channels, H is the height, and W is the width;

[0031] Calculate the average value of each channel c within each pooling window;

[0032] Assume that the pooling window size is (K×K), then the global representation result TR global The calculation formula is as follows:

[0033]

[0034] in, is the global representation result corresponding to channel c, H' is the height after pooling; W' is the width after pooling, TR n,c,h,w is the element value of the nth sample, cth channel, hth row, wth column in the tensor TR;

[0035] By calculating the mean of all elements within each pooling window, the average value representing the regional features is obtained.

[0036] Furthermore, the specific implementation method of step S3 includes the following steps:

[0037] S3.1. The spatiotemporal trajectory representation results obtained in step S2 are processed by T-SNE dimensionality reduction through cosine similarity. For the i-th trajectory representation tr i and the jth trajectory representation tr j The similarity p (j|i) The expression is:

[0038]

[0039] Among them, ||tr i -tr j || represents the i-th trajectory tr i and the jth trajectory representation tr j The Euclidean distance between Characterize tr for the i-th trajectory i The variance of is set by the perplexity hyperparameter;

[0040] Then minimize the KL loss function C to optimize the T-SNE dimensionality reduction method, the expression is:

[0041]

[0042] Among them, P i To consider tr i The conditional probability distribution between other points in space. Similarly, Q i is the corresponding conditional probability distribution in low-dimensional space;

[0043] S3.2. After using T-SNE to reduce the dimension, the similarity result between trajectories is obtained by calculating the cosine similarity matrix between trajectories. i ,tr j ), the expression is:

[0044]

[0045] Finally, K-Means is used to perform clustering based on the cosine similarity matrix to divide all trajectory sequences into K clusters, so that each sample belongs to the cluster corresponding to the center of the cluster closest to it.

[0046] An electronic device includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of a trajectory similarity analysis method based on spatiotemporal information extraction when executing the computer program.

[0047] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for analyzing trajectory similarity based on spatiotemporal information extraction is implemented.

[0048] Beneficial effects of the present invention:

[0049] The trajectory similarity analysis method based on spatiotemporal information extraction described in the present invention converts complex trajectory data into a vector form that is easy to process, and then uses a similarity analysis algorithm to efficiently compare and classify the trajectory data. This can not only improve the efficiency of data analysis, but also improve the accuracy and reliability of the analysis.

[0050] The trajectory similarity analysis method based on spatiotemporal information extraction described in the present invention considers the trajectory similarity measurement of spatial and temporal information, evaluates the trajectory similarity more comprehensively and accurately, adapts to the application requirements with strong time correlation, and improves the efficiency of trajectory data analysis and processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 It is a flow chart of a trajectory similarity analysis method based on spatiotemporal information extraction according to the present invention;

[0052] Figure 2 It is the result curve of the index of the silhouette coefficient as the objective function of the clustering algorithm;

[0053] Figure 3 The Davies-Bouldin index is used as the index result curve of the objective function of the clustering algorithm;

[0054] Figure 4 The Calinski-Harabasz index is used as the indicator result curve of the objective function of the clustering algorithm;

[0055] Figure 5 This is the result image obtained by visualizing the dimensionality reduction of T-SNE;

[0056] Figure 6 This is the result image obtained by calculating cosine similarity after T-SNE dimensionality reduction. DETAILED DESCRIPTION

[0057] In order to make the purpose, technical solution and advantages of the present invention more clear, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention, that is, the specific embodiments described are only part of the embodiments of the present invention, rather than all of the specific embodiments. The components of the specific embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations, and the present invention can also have other embodiments.

[0058] Therefore, the following detailed description of the specific embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected specific embodiments of the present invention. Based on the specific embodiments of the present invention, all other specific embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.

[0059] In order to further understand the content, features and effects of the present invention, the following specific implementation methods are given as examples, and the attached Figure 1 -Attached Figure 6 The detailed instructions are as follows: Specific implementation method one:

[0061] A trajectory similarity analysis method based on spatiotemporal information extraction includes the following steps:

[0062] S1. Collect spatiotemporal trajectory data, perform data cleaning and encoding preprocessing, and obtain preprocessed spatiotemporal trajectory data;

[0063] Furthermore, the specific implementation method of step S1 includes the following steps:

[0064] S1.1 Data cleaning:

[0065] S1.1.1 Collect spatiotemporal trajectory data. First, check whether the user ID item and the acquisition time item have null values, delete the entire data with null values, and then filter the vacant positions of the longitude item or latitude item and fill the vacant positions with 0;

[0066] S1.1.2 delete duplicate data, ping-pong switching data, drift data and range error data to obtain the spatiotemporal trajectory data after data cleaning;

[0067] S1.2. Encode and preprocess the spatiotemporal trajectory data after data cleaning in step S1.1, convert the two-dimensional discrete spatial coordinates into a one-dimensional digital sequence, and obtain the continuous trajectory sequence data TR of each user after preprocessing, which is represented as {TR: tr = (pot_1, pot_2, ..., pot_n)}, where each trajectory point is represented as pot (lat, lon). Lat represents latitude, and lon represents longitude;

[0068] Furthermore, the specific implementation method of step S1.2 includes the following steps:

[0069] S1.2.1. First, traverse all spatiotemporal trajectory data, find out the ping-pong switching data with short switching time interval and frequent switching back and forth between two base station locations, and delete them;

[0070] S1.2.2. Process the drift data in the spatiotemporal trajectory data, set the urban drift speed threshold, filter the records of the same user, calculate the distance and time interval between the i-th record and the i+1-th record based on the longitude and latitude information, and obtain the speed v between the two points. If v is greater than the urban drift speed threshold, delete the i+1-th record and recalculate the speed between the i-th record and the i+2-th record, and so on, until the above process is applied to all spatiotemporal trajectory data;

[0071] S1.2.3. Filter the data with range errors, find out the data that does not belong to the reasonable range of the location where the data is obtained through the set longitude and latitude range, and delete it;

[0072] S2. Encode and preprocess the spatiotemporal trajectory data after data cleaning in step S1, convert the two-dimensional discrete spatial coordinates into a one-dimensional digital sequence, and obtain the continuous trajectory sequence data TR of each user after preprocessing, which is represented by {TR:tr=(pot_1, pot_2, ..., pot_n)}, where each trajectory point is represented by pot(lat, lon). Lat represents latitude, and lon represents longitude;

[0073] S2. The preprocessed spatiotemporal trajectory data obtained in step S1 is input into a Transformer model with a position code layer for processing, and the spatiotemporal trajectory representation result is output;

[0074] Furthermore, the specific implementation method of step S2 includes the following steps:

[0075] S2.1. Input the continuous trajectory sequence data obtained in step S1 into the Transformer model with a position code layer to obtain the position encoding of the trajectory. The position encoding is to map the position information to the embedding space of the model, and the expression is:

[0076]

[0077]

[0078] Among them, PE(pos,2i) is the sine value of the position encoding of position pos and dimension 2i; PE(pos,2i+1) is the cosine value of the position encoding of position pos and dimension 2i+1; pos is the index of the position, i is the index of the dimension, and d model is the embedding dimension of the model;

[0079] S2.2. The position encoding of the trajectory is fused with the self-attention layer tensor through the fusion function to obtain the output result of the self-attention layer. The fusion function is calculated based on the calculation method of the matrix dot product Attn(Q, K, V), and the expression is:

[0080]

[0081] Among them, Q is the query matrix, V is the value matrix, K is the key matrix, and d k is the dimension of the key vector;

[0082] Q represents the query vector after linear transformation of the input sequence, which is used to represent the similarity between the current time step or position and all other time steps or positions in the attention calculation;

[0083] K represents the key vector after linear transformation of the input sequence, which is used to calculate the similarity by dot product with the query vector;

[0084] V represents the value vector after linear transformation of the input sequence, which is used to perform weighted summation on the input sequence after calculating the attention weight to obtain the final output;

[0085] d k Factor used to scale the dot product to ensure numerical stability when computing dot product similarity. Typically, d k are the dimensions of the query and key vectors.

[0086] S2.3. The output of the self-attention layer is input into the feedforward neural network layer for processing. The feedforward neural network layer consists of a simple fully connected layer and a ReLU activation function. After being processed by the encoder, the output is a tensor containing the high-level feature representation of the input trajectory data, expressed as:

[0087] Result (i) =ReLU(A i W+b);

[0088] Among them, A iis the data matrix of the i-th trajectory, W is the weight matrix, b is the bias vector, Result (i) is the output result of the i-th trajectory data after transformation and activation function;

[0089] A i Contains all relevant features of the trajectory, specifically, A i is a two-dimensional matrix where the rows represent time steps and the columns represent the different features at that time step.

[0090] W is used to linearly transform the input features. Its dimension is usually (number of input features, number of output features), which is used to convert the input trajectory data into another representation or to reduce the dimension.

[0091] b is used to add an offset to the result after linear transformation. Its dimension should be consistent with the number of output features to facilitate element-by-element addition.

[0092] Result (i) is a tensor representing the high-level feature representation of the i-th track. The nonlinear expression capability of the model can be increased by using the ReLU activation function.

[0093] S2.4. Use average pooling to reduce the dimension of the tensor containing the high-level feature representation of the input trajectory data and output the spatiotemporal trajectory representation result;

[0094] Set the dimension of the tensor TR containing the high-level feature representation of the input trajectory data to (N, C, H, W), where N is the batch size, C is the number of channels, H is the height, and W is the width;

[0095] Calculate the average value of each channel c within each pooling window;

[0096] Assume that the pooling window size is (K×K), then the global representation result TR global The calculation formula is as follows:

[0097]

[0098] in, is the global representation result corresponding to channel c, H' is the height after pooling; W' is the width after pooling, TR n,c,h,w is the element value of the nth sample, cth channel, hth row, wth column in the tensor TR;

[0099] By calculating the mean of all elements within each pooling window, the average value representing the regional features is obtained;

[0100] S3. The spatiotemporal trajectory representation result obtained in step S2 is subjected to T-SNE dimensionality reduction processing through cosine similarity, and then K-Means clustering is performed to obtain a visualization result of trajectory similarity analysis based on spatiotemporal information extraction;

[0101] Furthermore, the specific implementation method of step S3 includes the following steps:

[0102] S3.1. The spatiotemporal trajectory representation results obtained in step S2 are processed by T-SNE dimensionality reduction through cosine similarity. For the i-th trajectory representation tr i and the jth trajectory representation tr j The similarity p (j|i) The expression is:

[0103]

[0104] Among them, ||tr i -tr j || represents the i-th trajectory tr i and the jth trajectory representation tr j The Euclidean distance between Characterize tr for the i-th trajectory i The variance of is set by the perplexity hyperparameter;

[0105] Then minimize the KL loss function C to optimize the T-SNE dimensionality reduction method, the expression is:

[0106]

[0107] Among them, P i To consider tr i The conditional probability distribution between other points in space. Similarly, Q i is the corresponding conditional probability distribution in low-dimensional space;

[0108] S3.2. After using T-SNE to reduce the dimension, the similarity result between trajectories is obtained by calculating the cosine similarity matrix between trajectories. i ,tr j ), the expression is:

[0109]

[0110] Finally, K-Means is used to perform clustering based on the cosine similarity matrix to divide all trajectory sequences into K clusters, so that each sample belongs to the cluster corresponding to the center of the cluster closest to it.

[0111] In the trajectory similarity analysis method based on spatiotemporal information extraction described in this embodiment, the trajectory sequence data is passed through the Transformer model to obtain trajectory representation, and then T-SNE dimension reduction and K-Means clustering are performed through cosine similarity to obtain clustering visualization results. The silhouette coefficient, Davies-Bouldin index and Calinski-Harabasz index are used as the objective function of the clustering algorithm, and the clustering effect is observed by observing the change effect of the above values, so as to obtain the best clustering result.

[0112] Because different K values ​​of the K-means algorithm have a great influence on the experimental results, this method selects the K value based on the evaluation indicators silhouette coefficient, Davies-Bouldin index and Calinski-Harabasz index, and sets the selection range K = [3,20]. The indicator results under different K values ​​are as follows Figure 2-Figure 4 As shown:

[0113] from Figure 2 -Attached Figure 4 As can be seen, under different numbers of clusters, the values ​​of different indicators are also different. The silhouette coefficient reaches the maximum value when K=20, the Davies-Bouldin index reaches the minimum value when K=18, and the Calinski-Harabasz index reaches the maximum value when K=18. In general, when K=18, the clustering effect is the best, because the silhouette coefficient value is also very high at this time. Therefore, this embodiment sets K=18 for K-Means clustering and visualizes the results.

[0114] Figure 5 This is the result of direct T-SNE dimensionality reduction visualization after the trajectory sequence data passes through the Transformer hidden layer. Figure 6 This is the result of calculating the cosine similarity after T-SNE dimensionality reduction. It can be seen that the trajectories with high similarity are divided into one cluster. Specific implementation method 2:

[0116] An electronic device includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of a trajectory similarity analysis method based on spatiotemporal information extraction when executing the computer program.

[0117] The computer device of the present invention may be a device including a processor and a memory, such as a single chip microcomputer including a central processing unit, etc. Moreover, the processor is used to implement the steps of the above-mentioned trajectory similarity analysis method based on spatiotemporal information extraction when executing the computer program stored in the memory.

[0118] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0119] The memory may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. Specific implementation method three:

[0121] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for analyzing trajectory similarity based on spatiotemporal information extraction is implemented.

[0122] The computer-readable storage medium of the present invention can be any form of storage medium that can be read by a processor of a computer device, including but not limited to a non-volatile memory, a volatile memory, a ferroelectric memory, etc. The computer-readable storage medium stores a computer program. When the processor of the computer device reads and executes the computer program stored in the memory, the steps of the above-mentioned trajectory similarity analysis method based on spatiotemporal information extraction can be implemented.

[0123] The computer program includes computer program code, which may be in source code form, object code form, executable file or some intermediate form, etc. The computer readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer readable media do not include electric carrier signals and telecommunication signals.

[0124] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0125] Although the present application has been described above with reference to specific embodiments, various modifications may be made thereto and parts thereof may be replaced with equivalents without departing from the scope of the present application. In particular, as long as there is no structural conflict, the various features in the specific embodiments disclosed in the present application may be used in combination with each other in any manner, and the fact that these combinations are not exhaustively described in this specification is only for the sake of omitting space and saving resources. Therefore, the present application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A trajectory similarity analysis method based on spatiotemporal information extraction, characterized in that: The steps include: S1. Collect spatiotemporal trajectory data, perform data cleaning and encoding preprocessing, and obtain preprocessed spatiotemporal trajectory data; The specific implementation method of step S1 includes the following steps: S1.1 Data cleaning: S1.1.1 Collect spatiotemporal trajectory data. First, check whether the user ID item and the acquisition time item have null values, delete the entire data with null values, and then filter the vacant positions of the longitude item or latitude item and fill the vacant positions with 0; S1.1.2 delete duplicate data, ping-pong switching data, drift data and range error data to obtain the spatiotemporal trajectory data after data cleaning; S1.

2. Encode and preprocess the spatiotemporal trajectory data after data cleaning in step S1.1, convert the two-dimensional discrete spatial coordinates into a one-dimensional digital sequence, and obtain the continuous trajectory sequence data TR of each user after preprocessing, which is represented by {TR:tr=(pot_1, pot_2, …, pot_n)}, where each trajectory point is represented by pot(lat, lon), lat represents latitude, and lon represents longitude; The specific implementation method of step S1.2 includes the following steps: S1.2.

1. First, traverse all spatiotemporal trajectory data, find out the ping-pong switching data with short switching time interval and frequent switching back and forth between two base station locations, and delete them; S1.2.

2. Process the drift data in the spatiotemporal trajectory data, set the urban drift speed threshold, filter the records of the same user, calculate the distance and time interval between the i-th record and the i+1-th record based on the longitude and latitude information, and obtain the speed v between the two points. If v is greater than the urban drift speed threshold, delete the i+1-th record and recalculate the speed between the i-th record and the i+2-th record, and so on, until the above process is applied to all spatiotemporal trajectory data; S1.2.

3. Filter the data with range errors, find out the data that does not belong to the reasonable range of the location where the data is obtained through the set longitude and latitude range, and delete it; S2. The preprocessed spatiotemporal trajectory data obtained in step S1 is input into a Transformer model with a position code layer for processing, and the spatiotemporal trajectory representation result is output; The specific implementation method of step S2 includes the following steps: S2.

1. Input the continuous trajectory sequence data obtained in step S1 into the Transformer model with a position code layer to obtain the position encoding of the trajectory. The position encoding is to map the position information to the embedding space of the model, and the expression is: Among them, PE(pos,2i) is the sine value of the position encoding of position pos and dimension 2i; PE(pos,2i+1) is the cosine value of the position encoding of position pos and dimension 2i+1; pos is the index of the position, i is the index of the dimension, and d model is the embedding dimension of the model; S2.

2. The position encoding of the trajectory is fused with the self-attention layer tensor through the fusion function to obtain the output result of the self-attention layer. The fusion function is calculated based on the calculation method of the matrix dot product Attn(Q, K, V), and the expression is: Among them, Q is the query matrix, V is the value matrix, K is the key matrix, and d k is the dimension of the key vector; S2.

3. The output of the self-attention layer is input into the feedforward neural network layer for processing. The feedforward neural network layer consists of a simple fully connected layer and a ReLU activation function. After being processed by the encoder, the output is a tensor containing the high-level feature representation of the input trajectory data, expressed as: Result (i) =ReLU(A i W+b); Among them, A i is the data matrix of the i-th trajectory, W is the weight matrix, b is the bias vector, Result (i) is the output result of the i-th trajectory data after transformation and activation function; S2.

4. Use average pooling to reduce the dimension of the tensor containing the high-level feature representation of the input trajectory data and output the spatiotemporal trajectory representation result; Set the dimension of the tensor TR containing the high-level feature representation of the input trajectory data to (N, C, H, W), where N is the batch size, C is the number of channels, H is the height, and W is the width; Calculate the average value of each channel c within each pooling window; Assume that the pooling window size is (K×K), then the global representation result TR global The calculation formula is as follows: in, is the global representation result corresponding to channel c, H' is the height after pooling; W' is the width after pooling, TR n,c,h,w is the element value of the nth sample, cth channel, hth row, wth column in the tensor TR; By calculating the mean of all elements within each pooling window, the average value representing the regional features is obtained; S3. The spatiotemporal trajectory representation result obtained in step S2 is subjected to T-SNE dimensionality reduction processing through cosine similarity, and then K-Means clustering is performed to obtain the visualization result of trajectory similarity analysis based on spatiotemporal information extraction.

2. The trajectory similarity analysis method based on spatiotemporal information extraction according to claim 1 is characterized in that: The specific implementation method of step S3 includes the following steps: S3.

1. The spatiotemporal trajectory representation results obtained in step S2 are processed by T-SNE dimensionality reduction through cosine similarity. For the i-th trajectory representation tr i and the jth trajectory representation tr j The similarity p (j|i) The expression is: Among them, ||tr i -tr j || represents the i-th trajectory tr i and the jth trajectory representation tr j The Euclidean distance between Characterize tr for the i-th trajectory i The variance of is set by the perplexity hyperparameter; Then minimize the KL loss function C to optimize the T-SNE dimensionality reduction method, the expression is: Among them, P i To consider tr i The conditional probability distribution between other points in space. Similarly, Q i is the corresponding conditional probability distribution in low-dimensional space; S3.

2. After using T-SNE to reduce the dimension, the similarity result between trajectories is obtained by calculating the cosine similarity matrix between trajectories. i ,tr j ), the expression is: Finally, K-Means is used to perform clustering based on the cosine similarity matrix to divide all trajectory sequences into K clusters, so that each sample belongs to the cluster corresponding to the center of the cluster closest to it.

3. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of a trajectory similarity analysis method based on spatiotemporal information extraction as described in any one of claims 1 to 2 when executing the computer program.

4. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the trajectory similarity analysis method based on spatiotemporal information extraction according to any one of claims 1 to 2 is implemented.

Citation Information

Patent Citations

  • Moving object spatio-temporal trajectory clustering method considering multi-dimensional semantics

    CN113408640A

  • Multi-modal spatial-temporal trajectory prediction method, device, equipment and medium

    CN117132002A

  • Spatial-temporal trajectory generation method and device, computer equipment and storage medium

    CN117332033A