Human spatio-temporal trajectory representation learning method based on feature fusion and related device
Through the quad-tree encoding and time-frequency domain fusion method, the problem that the existing trajectory representation learning method fails to effectively capture the spatial heterogeneity and global characteristics of the trajectory data is achieved, and a more accurate and robust trajectory representation is achieved.
Patent Information
- Application Number
- CN202510250736.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-20
AI Technical Summary
When processing trajectory data, existing trajectory representation learning methods fail to effectively capture the spatial heterogeneity and global characteristics of trajectory points, resulting in poor representation accuracy.
Quadtree encoding is used to process the trajectory data, generate discrete time domain signals, and convert them into frequency domain signals through discrete Fourier transform. Input time domain signals and frequency domain signals into a self-supervised learning architecture based on sequence encoding and decoding, perform time-frequency domain fusion, and generate more accurate trajectory representations.
Through quadtree encoding and time-frequency domain fusion, the spatial heterogeneity and global characteristics of trajectory data can be more effectively captured, improving the accuracy and robustness of trajectory representation.
Smart Images

Figure CN120180035A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of trajectory analysis, and relates to a method and related device for human spatio-temporal trajectory representation learning based on feature fusion. Background Art
[0002] Human spatio-temporal trajectory refers to the movement path and behavior information of an individual collected through smart phones, GNSS devices, etc. within a specific time and space. Trajectory representation learning is a process of converting the original trajectory data composed of longitude, latitude, and timestamp into a low-dimensional representation vector, and the generated vector can serve various downstream tasks, such as trajectory query, trajectory clustering, trajectory prediction, trajectory classification, trajectory anomaly detection, etc. The advantages of trajectory representation learning are as follows: First, to a certain extent, the vector representation can alleviate problems such as noise points generated during the signal acquisition and transmission of sensors, as well as missing original trajectory points, uneven sampling, and the existence of noise points. Second, compared with traditional similarity measurement methods, the vector representation can greatly improve the similarity calculation efficiency. In addition, since the trajectory is a kind of high-dimensional data containing space, time, etc., which is not conducive to machine learning processing, the trajectory data often cannot be directly used as the input data of the algorithm. However, trajectory representation learning can reduce the dimension of the high-dimensional trajectory data, which can not only standardize and simplify the trajectory data in form, but also extract valuable parts from the redundant original information, making the entire model more efficient. At the same time, trajectory representation is also the basis for constructing a general trajectory large model. Since it involves the input of the model, the quality of the trajectory representation will also directly affect the final effect of trajectory data mining. Therefore, it has very important practical significance to study the effective representation of trajectory data.
[0003] Current trajectory representation learning methods can be classified according to supervised learning and self-supervised learning. Supervised learning trains a model through labeled data (i.e., input-label pairs) to establish a mapping relationship from input data to target labels. However, such methods highly rely on manually labeled data, and the labeling process requires a large amount of time and human resources, especially in complex temporal scenarios such as trajectory data where the cost is particularly significant. In contrast, self-supervised learning can utilize auxiliary tasks to mine its own supervision information from large-scale unlabeled trajectory data, thus getting rid of the dependence on manual labeling. The representations generated through self-supervised learning can be applied to downstream tasks (such as trajectory prediction or classification) through fine-tuning based on a small batch of labeled data; in addition, the representations it generates can also be directly applied to unsupervised tasks (such as trajectory clustering), significantly reducing the training overhead of downstream tasks and improving computational efficiency. Trajectory representation learning methods based on self-supervised learning are further divided into two types of frameworks: those based on sequence encoding and decoding structures and those based on contrastive learning. Based on sequence encoding and decoding structures, such as Seq2Seq, it maps trajectories to a latent space based on a pre-training task of generating autoregression, and then reconstructs its original data from this latent space. In recent years, contrastive learning has been used to conduct research on trajectory representation learning. Contrastive learning trains a model by comparing similar and dissimilar sample pairs, so as to pull similar samples closer and push dissimilar samples farther away in a high-dimensional space. The core idea of contrastive learning is to use data augmentation to generate positive and negative sample pairs, and optimize the objective function (such as the InfoNCE loss function) to quantify the model's ability to distinguish sample similarity. Through contrastive learning, the model can learn more robust and effective representations, thus performing better in downstream tasks.
[0004] Although there are already various trajectory representation learning methods, the existing technologies still have significant defects and deficiencies, which are specifically reflected in the following aspects:
[0005] Uniform grid encoding: Existing methods usually map trajectory points to uniformly divided grid cells and encode them. This encoding method divides the entire area into non-overlapping grid cells of equal size. All grid cells form a spatial vocabulary, each grid cell is marked with a vocabulary ID, and two-dimensional trajectory points are converted into the grid cell IDs where they are located. However, this encoding method does not consider the spatial heterogeneity of the distribution of trajectory points. Spatial heterogeneity refers to the uneven distribution of trajectory data in space, where there is dense trajectory data in some areas.
[0006] RNN-based Trajectory Representation Learning Method: Existing methods usually use RNN to process input sequences, which can capture the temporal dependencies of trajectory points. However, due to the structural characteristics of RNN, especially in the case of long sequences, its ability to capture the global features of the entire trajectory is limited. Although Long Short Term Memory (LSTM) and Gate Recurrent Unit (GRU) can improve the poor performance of RNN in dealing with long-distance dependencies to a certain extent, they still cannot effectively integrate the global context of the entire trajectory, so the accuracy of the representation is poor. Summary of the Invention
[0007] The purpose of the present invention is to overcome the above-mentioned disadvantages of the prior art and provide a method and related device for human spatio-temporal trajectory representation learning based on feature fusion, which has a better effect on human spatio-temporal trajectory representation learning.
[0008] To achieve the above purpose, the present invention discloses a method for human spatio-temporal trajectory representation learning based on feature fusion, including:
[0009] Obtain the original trajectory sequence of the user;
[0010] Process the original trajectory sequence of the user through quadtree encoding to obtain a discretized trajectory time-domain signal;
[0011] Convert the discretized trajectory time-domain signal into a frequency-domain signal;
[0012] Input the discretized trajectory time-domain signal and the frequency-domain signal into a trained self-supervised learning architecture based on sequence encoding and decoding to obtain a time-frequency domain fused trajectory representation.
[0013] A further improvement of the method for human spatio-temporal trajectory representation learning based on feature fusion according to the present invention is as follows:
[0014] Further, the process of converting the discretized trajectory time-domain signal into a frequency-domain signal is as follows:
[0015] Perform discrete Fourier transform on the discretized trajectory time-domain signal to obtain the frequency-domain signal.
[0016] Further, the self-supervised learning architecture based on sequence encoding and decoding includes an RNN-based time-domain encoder, a Transformer-based frequency-domain encoder, a splicing layer, and a linear layer. Among them, the output end of the RNN-based time-domain encoder and the output end of the Transformer-based frequency-domain encoder are connected to the input end of the splicing layer, and the output end of the splicing layer is connected to the input end of the linear layer.
[0017] Further, the process of inputting the discretized trajectory time-domain signal and the frequency-domain signal into the trained self-supervised learning architecture based on sequence encoding and decoding to obtain the time-frequency domain fused trajectory representation is as follows:
[0018] Input the discretized trajectory time-domain signal into the time-domain encoder based on RNN to obtain the time-domain feature e TE (O); input the frequency-domain signal into the frequency-domain encoder based on Transformer to obtain the frequency-domain feature e FE (O) of the trajectory sequence O, and then through the splicing layer, splice the time-domain feature e TE (O) and the frequency-domain feature e FE (O) of the trajectory sequence O, and then reduce the dimension to d through the linear layer model , to obtain the time-frequency domain fused trajectory representation e(X i ).
[0019] Further, the time-frequency domain fused trajectory representation e(X i ) is:
[0020] e(O) = Linear(concat(e TE (O) · e FE (O))).
[0021] Further, the loss function of the self-supervised learning architecture based on sequence encoding and decoding during training is:
[0022]
[0023] where is the spatial proximity weight. The closer the region q is to the region where the decoder output trajectory point y i is located, the larger it is, ‖q - y i ‖2 is the spatial distance between the region center points, where θ > 0 represents the spatial distance scale coefficient.
[0024] The present invention discloses a human spatio-temporal trajectory representation learning system based on feature fusion, including:
[0025] An acquisition module for acquiring the original trajectory sequence of the user;
[0026] An encoding module for encoding the original trajectory sequence of the user through a quadtree to obtain a discretized trajectory time-domain signal;
[0027] A conversion module for converting the discretized trajectory time-domain signal into a frequency-domain signal;
[0028] A characterization module, configured to input the discretized trajectory time-domain signal and the frequency-domain signal into a trained self-supervised learning architecture based on sequence encoding and decoding, so as to obtain a time-frequency domain fused trajectory characterization.
[0029] A further improvement of the human spatio-temporal trajectory characterization learning system based on feature fusion according to the present invention lies in:
[0030] Further, the self-supervised learning architecture based on sequence encoding and decoding includes an RNN-based time-domain encoder, a Transformer-based frequency-domain encoder, a splicing layer, and a linear layer. Among them, the output end of the RNN-based time-domain encoder and the output end of the Transformer-based frequency-domain encoder are connected to the input end of the splicing layer, and the output end of the splicing layer is connected to the input end of the linear layer.
[0031] The present invention discloses a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the human spatio-temporal trajectory characterization learning method based on feature fusion are implemented.
[0032] The present invention discloses a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the human spatio-temporal trajectory characterization learning method based on feature fusion are implemented.
[0033] The present invention has the following beneficial effects:
[0034] When the human spatio-temporal trajectory characterization learning method based on feature fusion and related devices according to the present invention are specifically operated, the original trajectory sequence of the user is processed by quadtree encoding to obtain a discretized trajectory time-domain signal. Using the quadtree to divide the space can take into account the spatial distribution heterogeneity of trajectory points, and make the area with dense trajectory points have a higher spatial resolution without increasing the total amount of spatial units. In addition, the present invention inputs the discretized trajectory time-domain signal and the frequency-domain signal into a trained self-supervised learning architecture based on sequence encoding and decoding to obtain a time-frequency domain fused trajectory characterization, and complements each other through the fusion of time-domain features and frequency-domain features to improve the effect of trajectory characterization learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The specification drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0036] Figure 1 is a flowchart of the method of the present invention;
[0037] Figure 2 Schematic diagram of a self-supervised learning architecture based on sequence encoding and decoding;
[0038] Figure 3 System structure diagram of the present invention. Detailed implementation manners
[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0040] In the description of the present invention, it should be understood that the terms "include" and "comprise" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0041] It should also be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0042] It should be further understood that the term " / and" used in the specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present invention generally represents an "or" relationship between the contextually related objects.
[0043] It should be understood that although the terms first, second, third, etc. may be used in the embodiments of the present invention to describe preset ranges, etc., these preset ranges should not be limited to these terms. These terms are only used to distinguish the preset ranges from each other. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.
[0044] Depending on the context, as used herein, the word "if" can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detected (stated condition or event)" or "in response to detecting (stated condition or event)".
[0045] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. Generally, the components described and shown in the accompanying drawings herein can be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0046] Schematic diagrams of various structures according to the disclosed embodiments of the present invention are shown in the accompanying drawings. These figures are not drawn to scale, where certain details are enlarged for clarity of expression and certain details may be omitted. The shapes of the various regions and layers shown in the figures and their relative sizes and positional relationships are merely exemplary, and may actually deviate due to manufacturing tolerances or technical limitations, and those skilled in the art can additionally design regions / layers with different shapes, sizes and relative positions according to actual requirements.
[0047] Embodiment 1
[0048] Referring to Figure 1 , the method for human spatio-temporal trajectory characterization learning based on the fusion of time-domain and frequency-domain features according to the present invention includes the following steps:
[0049] 1) Obtain the original trajectory sequence of the user, and process the trajectory sequence of the user through quadtree encoding to obtain a discretized trajectory time-domain signal;
[0050] It should be noted that traditional position encoding methods, such as grid ID encoding, divide the entire area into non-overlapping grid cells of equal size. All grid cells form a spatial vocabulary, each grid cell is marked with a vocabulary ID, and two-dimensional trajectory points are converted into the ID of the grid cell where they are located. Due to its fixed and uniform division method, regardless of the spatial distribution density of the trajectory points, the size and shape of the grid remain the same. Therefore, it cannot capture the spatial heterogeneity of the trajectory points, and the trajectory information in the dense area of the trajectory points cannot be represented more finely, while the sparse area occupies too many grid resources;
[0051] Based on the above problems, the present application uses a quadtree to divide the geographical space, then encodes each divided area, and converts the original trajectory sequence into a discrete trajectory sequence, enabling the model to better learn the trajectory representation.
[0052] It should be noted that quadtree encoding is a spatial division method based on a tree structure. By recursively dividing a spatial area into smaller sub-areas to construct a tree structure, it can dynamically adjust the depth of the division according to the density of the trajectory data, avoiding excessive division in sparse areas, and improving the resolution of the dense area of the trajectory points without increasing the number of divided areas. Through the hierarchical structure of the tree, a certain point or area can be quickly located, especially in large-scale spatial data. Therefore, the quadtree can effectively improve the query efficiency of trajectory points.
[0053] The quadtree starts from the whole of a two-dimensional space and gradually divides it into four equal sub-areas. For each sub-area, if the number of data points contained in it reaches the division condition, then the sub-area is recursively divided into four smaller areas until the division condition is not met.
[0054] The quadtree spatial division and encoding process is as follows:
[0055] Initial division: The entire space is divided into four quadrants, that is, four sub-areas, according to the midpoint of the area. Each number represents a sub-area.
[0056] Recursive division: For each sub-area, if it reaches the division condition, the division condition is that the number of trajectory points exceeds the point threshold or the maximum division level is not reached, then it continues to be divided into four smaller areas. Each time it is divided, the sub-area is assigned a corresponding code.
[0057] Encoding trajectory points: Each trajectory point will fall into a certain leaf node of the quadtree according to its position in the space. The code of the leaf node is the spatial code of the point, and the length of the code depends on the specific division level where the trajectory point falls.
[0058] 2) Convert the discretized trajectory time-domain signal into a frequency-domain signal. Specifically, use the discrete Fourier transform to convert the discretized trajectory time-domain signal into a frequency-domain signal;
[0059] It should be noted that in natural language processing (NLP), the Seq2Seq model is often used to process discrete sequences of words. After the trajectory data is discretized, it is similar to regarding trajectory points as "words" and trajectory sequences as "sentences". The Seq2Seq model learns the transition relationships between trajectory points by encoding the trajectory sequence, just like capturing the dependencies between words in a language model. This method has been proven to be effective in time-series modeling. In addition, using the discretized trajectory sequence facilitates the construction of an efficient spatial index structure and makes it easier to perform tasks such as trajectory query and clustering.
[0060] Therefore, the present invention uses quadtree encoding to convert the trajectory sequence into a discretized time-domain signal, and converts the two-dimensional longitude and latitude of each trajectory stop point into a one-dimensional quadtree encoding ID.
[0061] To obtain the frequency-domain characteristics of the trajectory time-domain signal and thus capture the long-term trend of the trajectory, it is necessary to perform a frequency transformation on the trajectory time-domain signal. Common frequency transformations include the Discrete Fourier Transform (DFT), the Discrete Cosine Transform (DCT), and the Discrete Wavelet Transform (DWT). Among them, DCT is often used for image transformation, and DWT can perform time-domain analysis, but both of these methods only retain the real part components. In fact, discarding the imaginary part components may lead to information loss, thereby affecting the expression of the global pattern. Therefore, this application uses DFT for frequency transformation. DFT plays an important role in the field of digital signal processing. Given a sequence of length N, DFT converts it to:
[0062]
[0063] where j is the imaginary unit, N is the length of the input signal, x[n] is the input time-domain signal, the value of the nth sample point, and e -j(2π / N)kn is the Fourier basis function, and X[k] is the complex value of the kth frequency component, representing the amplitude and phase of the kth frequency in the spectrum, and X ∈ C k consists of the real part Re and the imaginary part Im, where X is:
[0064]
[0065] X = Re + jIm
[0066] The amplitude part and the phase part of X are respectively:
[0067]
[0068] The frequency signal can be obtained through discrete Fourier transform.
[0069] 3) Input the frequency-domain signal into the trained self-supervised learning architecture based on sequence encoding and decoding to obtain the time-frequency domain fusion trajectory representation.
[0070] The self-supervised learning architecture based on sequence encoding and decoding includes an RNN-based time-domain encoder, a Transformer-based frequency-domain encoder, a splicing layer, and a linear layer. Among them, the trajectory time-domain signal is input into the RNN-based time-domain encoder to obtain the time-domain feature e TE (O) of the trajectory sequence O; the frequency-domain signal is input into the Transformer-based frequency-domain encoder to obtain the frequency-domain feature e FE (O) of the trajectory sequence O, and then the time-domain feature e TE (O) and the frequency-domain feature e FE (O) of the trajectory sequence O are spliced through the splicing layer, and then the dimension is restored to d model through the linear layer to obtain the time-frequency domain fusion trajectory representation e(X i ) as:
[0071] e(O) = Linear(concat(e TE (O) · e FE (O)))
[0072] Reference Figure 2 , specifically, to improve the robustness of the trajectory representation learning model to trajectory noise, the pre-training task adopted is trajectory reconstruction. For this purpose, the present invention designs an upsampling strategy for processing the input of the encoder. Specifically, given an original trajectory sequence O = [o1, o2, o3,.., o n , where o n = (lon i , lat i ), a probability distribution P = [p0, p1, p2,.., p m is defined, where p i represents the probability of inserting i new trajectory points between every two adjacent trajectory points. For example, P = [0.4, 0.3, 0.2, 0.1] means that there is a 40% probability of not inserting new trajectory points, a 30% probability of inserting 1 new trajectory point, and so on. For the case where new trajectory points need to be inserted, the method of linear interpolation is used. Suppose at the trajectory point (lon i , lat i ) and the trajectory point (lon i+1 , lati+1 ) Insert k new trajectory points between them, then the coordinates (lon i,j , lat i,j ) of the j-th inserted trajectory point can be expressed as:
[0073]
[0074] Among them, j = 1, 2,..., k. In addition to upsampling, the present invention further performs Gaussian noise processing on the trajectory, and defines a set of perturbation rates D = [d0, d1, d2,.., d v , where d i represents the intensity of the perturbation. For each trajectory point (lon i , lat i ), after adding random noise, the perturbed trajectory point (lon′ i , lat′ i ) is obtained:
[0075] lon′ i = lon i + δ lon
[0076] lat′ i = lat i + δ lat
[0077] Among them, δ lon and δ lat are random variables that follow a normal distribution (i.e., Gaussian distribution) with a mean of 0 and a standard deviation of d i . The perturbation rate d i serves as the standard deviation of the Gaussian noise to control the intensity of the noise. The larger the standard deviation, the greater the amplitude of the added noise and the more obvious the perturbation. By adjusting the value of d i , multiple enhanced trajectory sequences are generated, providing rich inputs for the self-supervised pre-training task of the model, which helps the model better learn the global features of the trajectory.
[0078] The present invention adopts an RNN-based time-domain encoder. RNN is a deep learning model specifically designed for processing sequence data. Its internal loop structure allows information to be passed between different time steps of the sequence. Therefore, RNN can effectively capture the local dependencies and context information of the trajectory sequence. Since RNN may encounter problems of gradient vanishing or gradient explosion when processing long sequences, the present invention adopts a variant of RNN, GRU, to model the local features of the trajectory. Specifically, the trajectory sequence after upsampling and noise addition is encoded through a quadtree to obtain the time-domain signal of the trajectory, and is input into the GRU-based time-domain encoder. In the time-domain encoder, a deep recurrent neural network is adopted. For the time-domain signal T = [I1, I2, I3,.., In , at time step t, the hidden state of the l-th layer is represented as The time-domain feature e of the output of the encoder TE (O) is composed of the hidden states of all layers at the last time step, that is:
[0079]
[0080] The present invention adopts a Transformer-based frequency-domain encoder. To effectively capture the global features of trajectory data, the time-domain signal is first converted into a frequency-domain signal through the discrete Fourier transform. In the frequency domain, the global features of the trajectory sequence are closely related to each frequency component. Therefore, the present invention uses a Transformer encoder as the frequency-domain encoder, taking the frequency-domain signal as the input. The self-attention mechanism of the Transformer can effectively model the global dependencies in the sequence data, thereby extracting the frequency-domain features and comprehensively characterizing the global pattern of the trajectory. Specifically, in the input of the Transformer encoder, each frequency component of the trajectory is the sum of the frequency component embedding and the position embedding:
[0081] Embedding(f i ) = FrequencyEmbedding(f i ) + PositionEmbedding(f i )
[0082] Among them, the position embedding is used to identify the position of each frequency component. By assigning a unique sequential encoding to each frequency component, the model can identify and extract the low-frequency and high-frequency components of the trajectory. The model uses the attention mechanism to capture the contribution of each frequency component to the entire trajectory. The formula for attention calculation is:
[0083]
[0084] Among them, Q is the query matrix, K is the key matrix, V is the value matrix, d k = d v = d model / h = 256, which is the dimension of K.
[0085] The output of the frequency encoder is obtained by e FE (O) the position-wise feed-forward networks and layer normalization (Norm):
[0086] FFN(x) = max(0, xW1 + b1)W2 + b2
[0087] e FE(O) = LayerNorm(x + FFN(x))
[0088] Among them, x is the output of the previous layer's layer normalization, respectively represent the weight matrix and bias vector of the first-layer feed-forward neural network, respectively represent the weight matrix and bias vector of the second-layer feed-forward neural network, and the dimension of the internal layer is d ff = 512.
[0089] During the training process, a spatial loss function is used for model training. Different from the commonly used Negative Log Likelihood (NLL) loss function, the spatial loss function considers the spatial distance between trajectories when calculating the loss value. The spatial distance between the decoder output trajectory points and the target trajectory points determines the degree of loss penalty. The smaller the distance, the smaller the corresponding loss value of the penalty:
[0090]
[0091] Among them, is the spatial proximity weight. The closer the region q is to the region where the decoder output trajectory point y i is located, the larger, ‖q - y i ‖2 is the spatial distance between the center points of the regions, where θ > 0 is the spatial distance scale coefficient. A smaller θ will severely penalize regions with a larger spatial distance.
[0092] Verification experiment
[0093] To verify the performance of the present invention in the trajectory query task, four measure-based similarity metric methods (the first 4) and two deep learning models (the last 2) are selected as comparison methods:
[0094] LCSS compares the similarity between trajectories by finding the longest common subsequence of two trajectories. The larger the LCSS value, the greater the similarity between the trajectories.
[0095] EDR compares the similarity between trajectories by calculating the number of insert, delete, and replace operations required to transform one trajectory into another. The larger the EDR value, the smaller the similarity between the trajectories.
[0096] Hausdorff describes their similarity in geometric space by measuring the maximum and minimum distances between points on two trajectories and is sensitive to noise points.
[0097] Fréchet considers the time order of points on two trajectories and simulates the shortest leash length between a person leading a dog and the dog. It is a more intuitive trajectory similarity metric method.
[0098] SIMformer is a Transformer-based supervised learning model that obtains labels (true similarity values) through measure-based similarity metrics such as Hausdorff and Fréchet. During training, the similarity between the representations of pairwise trajectories and the true similarity values are used to calculate the loss function.
[0099] t2vec is a self-supervised learning model based on Seq2Seq. It uses uniform grid encoding to convert two-dimensional trajectory points into one-dimensional IDs and designs a spatial perception loss function to improve the accuracy of the representation.
[0100] Given a query trajectory, its similarity is calculated with all trajectories in the dataset to be queried. The ranking of the similarity between the query trajectory and the most similar trajectory among all pairwise trajectory similarity values is the similar trajectory query ranking. Specifically, randomly select n trajectories from the test set as the trajectories to be queried, denoted as J, and then select another m trajectories as the query trajectories, denoted as Q. For each trajectory in Q, two sub-trajectories T q and T q′ are created by taking odd and even points respectively, and two datasets D Q and D Q′ are constructed. The same operation is performed on the trajectories in J to obtain D J and D J′ . For each query trajectory T q , its similarity is calculated with all trajectories in the dataset D Q′ ∪D J′ , and then the ranking of the similarity between T q and T q′ is observed among all trajectory similarities. In an ideal situation, the similarity between T q and T q′ should rank first because they are generated from the same trajectory. Finally, the average ranking of the m trajectories in Q is used as the performance evaluation metric. In the experiment, the cosine function is used to measure the vector distance between the representations of two trajectories as the result of similarity calculation.
[0101] Referring to Table 1 (in Table 1, bold: best effect; underlined: second-best effect), the experimental results of the present invention show that the present invention can exhibit more superior performance compared to the comparative models.
[0102] Table 1
[0103]
[0104] Example 2
[0105] Refer to Figure 3, the human spatio-temporal trajectory representation learning system based on feature fusion according to the present invention includes:
[0106] An acquisition module for acquiring the original trajectory sequence of the user;
[0107] An encoding module for performing quadtree encoding on the original trajectory sequence of the user to obtain a discretized trajectory time-domain signal;
[0108] A conversion module for converting the discretized trajectory time-domain signal into a frequency-domain signal;
[0109] A representation module for inputting the discretized trajectory time-domain signal and the frequency-domain signal into a trained self-supervised learning architecture based on sequence encoding and decoding to obtain a time-frequency domain fused trajectory representation.
[0110] The process of converting the discretized trajectory time-domain signal into a frequency-domain signal is as follows:
[0111] Performing a discrete Fourier transform on the discretized trajectory time-domain signal to obtain the frequency-domain signal.
[0112] In this embodiment, the self-supervised learning architecture based on sequence encoding and decoding includes an RNN-based time-domain encoder, a Transformer-based frequency-domain encoder, a splicing layer, and a linear layer. Among them, the output end of the RNN-based time-domain encoder and the output end of the Transformer-based frequency-domain encoder are connected to the input end of the splicing layer, and the output end of the splicing layer is connected to the input end of the linear layer.
[0113] In this embodiment, the process of inputting the discretized trajectory time-domain signal and the frequency-domain signal into a trained self-supervised learning architecture based on sequence encoding and decoding to obtain a time-frequency domain fused trajectory representation is as follows:
[0114] Inputting the discretized trajectory time-domain signal into the RNN-based time-domain encoder to obtain the time-domain feature e TE (O) of the trajectory sequence O; inputting the frequency-domain signal into the Transformer-based frequency-domain encoder to obtain the frequency-domain feature e FE (O) of the trajectory sequence O, and then splicing the time-domain feature e TE (O) and the frequency-domain feature e FE (O) of the trajectory sequence O through the splicing layer, and then reducing the dimension to d through the linear layer model , to obtain the time-frequency domain fused trajectory representation e(X i ).
[0115] In this embodiment, the time-frequency domain fused trajectory representation e(X i ) is:
[0116] e(O) = Linear(concat(e TE (O)·e Fe (O)))。
[0117] In this embodiment, the loss function of the self-supervised learning architecture based on sequence encoding and decoding during training is:
[0118]
[0119] where is the spatial proximity weight. The closer the region q is to the region where the decoder output trajectory point y i is located, the larger it is, and ‖q - y i ‖2 is the spatial distance between the center points of the regions, where θ > 0 represents the spatial distance scale coefficient.
[0120] The division of modules in the embodiments of the present application is illustrative. It is only a logical function division. In actual implementation, there may be other division methods. In addition, in each embodiment of the present application, each functional module can be integrated in one processor, or can exist separately physically, or two or more modules can be integrated in one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules.
[0121] Embodiment III
[0122] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the human spatio-temporal trajectory representation learning method based on feature fusion. For example, it includes: obtaining the original trajectory sequence of the user; processing the original trajectory sequence of the user through quadtree encoding to obtain a discretized trajectory time-domain signal; converting the discretized trajectory time-domain signal into a frequency-domain signal, and inputting the discretized trajectory time-domain signal and the frequency-domain signal into the trained self-supervised learning architecture based on sequence encoding and decoding to obtain a time-frequency domain fused trajectory representation. Among them, the memory may include a memory, such as a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk memory, etc.; the processor, network interface, and memory are interconnected through an internal bus, and this internal bus can be an Industry Standard Architecture bus, a Peripheral Component Interconnect Standard bus, an Extended Industry Standard Architecture bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. The memory is used to store programs. Specifically, the program may include program code, and the program code includes computer operation instructions. The memory can include a memory and a non-volatile memory, and provides instructions and data to the processor.
[0123] Example 4
[0124] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the human spatio-temporal trajectory representation learning method based on feature fusion. For example, it includes: obtaining the original trajectory sequence of a user; processing the original trajectory sequence of the user through quadtree encoding to obtain a discretized trajectory time-domain signal; converting the discretized trajectory time-domain signal into a frequency-domain signal, and inputting the discretized trajectory time-domain signal and the frequency-domain signal into a trained self-supervised learning architecture based on sequence encoding and decoding to obtain a spatio-temporal frequency-domain fused trajectory representation. Specifically, the computer-readable storage medium includes but is not limited to, for example, volatile memory and / or non-volatile memory. The volatile memory may include random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include read-only memory (ROM), hard disk, flash memory, optical disc, magnetic disk, etc.
[0125] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0126] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more of the processes and / or blocks Figure 1 one or more of the blocks.
[0127] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one or more of the processes and / or blocks Figure 1The functions specified in one or more boxes.
[0128] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the steps of the functions specified in one or more processes and / or boxes Figure 1 One process or more processes and / or boxes Figure 1 The steps of the functions specified in one or more boxes.
[0129] Those skilled in the art will readily conceive of other embodiments of the present invention upon considering the specification and the disclosure of the invention. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include known common knowledge or conventional technical means in the technical field not disclosed by the present invention. The specification and examples are only illustrative, and the true scope and spirit of the present invention are pointed out by the following claims.
[0130] It should be understood that the present invention is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.
[0131] As described above, the above are only preferred embodiments of the present invention and do not impose any limitation on the present invention. Any simple modifications, changes, and equivalent structural changes made to the above embodiments according to the technical essence of the present invention still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A method for learning human spatiotemporal trajectory representation based on feature fusion, characterized in that: include: Get the user's original trajectory sequence; Processing the original trajectory sequence of the user through quadtree coding to obtain a discretized trajectory time domain signal; Converting the discretized trajectory time domain signal into a frequency domain signal; The discretized trajectory time domain signal and the frequency domain signal are input into a trained self-supervised learning architecture based on sequence encoding and decoding to obtain a trajectory representation of time-frequency domain fusion.
2. The method for learning human spatiotemporal trajectory representation based on feature fusion according to claim 1, characterized in that: The process of converting the discretized trajectory time domain signal into a frequency domain signal is as follows: Performing discrete Fourier transform on the discretized trajectory time domain signal to obtain the frequency domain signal.
3. The method for learning human spatiotemporal trajectory representation based on feature fusion according to claim 1, characterized in that: The self-supervised learning architecture based on sequence encoding and decoding includes an RNN-based time domain encoder, a Transformer-based frequency domain encoder, a splicing layer and a linear layer, wherein the output end of the RNN-based time domain encoder and the output end of the Transformer-based frequency domain encoder are connected to the input end of the splicing layer, and the output end of the splicing layer is connected to the input end of the linear layer.
4. The method for learning human spatiotemporal trajectory representation based on feature fusion according to claim 3 is characterized in that: The process of inputting the discretized trajectory time domain signal and the frequency domain signal into the trained self-supervised learning architecture based on sequence encoding and decoding to obtain the trajectory representation of time-frequency domain fusion is as follows: The discretized trajectory time domain signal is input into the RNN-based time domain encoder to obtain the time domain feature e of the trajectory sequence O TE (O); Input the frequency domain signal into the frequency domain encoder based on Transformer to obtain the frequency domain feature e of the trajectory sequence O FE (O), and then the temporal features e of the trajectory sequence O are transformed through the concatenation layer. TE (O) and the frequency domain features of trajectory sequence O FE (O) concatenates and then restores the dimension to d through the linear layer model , we get the trajectory representation e(X i ).
5. The method for learning human spatiotemporal trajectory representation based on feature fusion according to claim 4 is characterized in that: The trajectory representation of the time-frequency domain fusion is e(X i )for: e(O)=Linear(concat(e TE (O)·e FE (THE))).
6. The method for learning human spatiotemporal trajectory representation based on feature fusion according to claim 1, characterized in that: The loss function of the self-supervised learning architecture based on sequence encoding and decoding during training is: in, is the spatial proximity weight, the distance between region q and the decoder output trajectory point y i The closer the area is, The larger the i ‖2 is the spatial distance between the center points of the region, where θ>0 represents the spatial distance scale coefficient.
7. A human spatiotemporal trajectory representation learning system based on feature fusion, characterized in that: include: An acquisition module is used to obtain the user's original trajectory sequence; An encoding module, used for processing the original trajectory sequence of the user through quadtree encoding to obtain a discretized trajectory time domain signal; A conversion module, used for converting the discretized trajectory time domain signal into a frequency domain signal; The representation module is used to input the discretized trajectory time domain signal and the frequency domain signal into the trained self-supervised learning architecture based on sequence encoding and decoding to obtain a trajectory representation of time-frequency domain fusion.
8. The human spatiotemporal trajectory representation learning system based on feature fusion according to claim 7 is characterized in that: The self-supervised learning architecture based on sequence encoding and decoding includes an RNN-based time domain encoder, a Transformer-based frequency domain encoder, a splicing layer and a linear layer, wherein the output end of the RNN-based time domain encoder and the output end of the Transformer-based frequency domain encoder are connected to the input end of the splicing layer, and the output end of the splicing layer is connected to the input end of the linear layer.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the human spatiotemporal trajectory representation learning method based on feature fusion as described in any one of claims 1 to 6 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the human spatiotemporal trajectory representation learning method based on feature fusion as described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Multi-target flight path identification instrument based on dual-channel pre-training
CN117493948A
Physiological signal self-supervision representation learning method and system based on time-frequency reconstruction
CN119089378A
Expressway pedestrian and vehicle in-transit track matching method and device and medium
CN119402825A