Trajectory traffic mode classification method, device, electronic device and storage medium
By introducing a comparative learning framework and dynamic queue mechanism in traffic pattern classification, the shortcomings of the existing technology in processing high-dimensional complex trajectory data are solved, and more efficient and accurate traffic pattern classification is achieved.
Patent Information
- Application Number
- CN202510245695.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-03-04
AI Technical Summary
When processing high-dimensional and complex spatiotemporal trajectory data, existing traffic pattern classification methods are difficult to make full use of global information and timing dependencies, especially in small sample scenarios and complex classification tasks.
A new contrast learning framework is proposed, combining multi-layer convolutional network, Transformer encoder and gated recurrent unit (GRU), to dynamically expand the number and diversity of negative samples through dynamic queue mechanisms, improving the discriminant ability of feature representation and the classification performance of the model.
It significantly enhances the classification accuracy and robustness of the model, and can more effectively capture the complex spatio-temporal relationships of trajectory data, and is suitable for trajectory data analysis in complex scenarios.
Smart Images

Figure CN119740101B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent transportation technology, and more specifically, relates to a trajectory traffic mode classification method, device, electronic device and storage medium. Background Art
[0002] Trajectory data, as a sequence of spatial positions of moving objects that change over time, has wide application value in intelligent traffic management, logistics optimization, urban planning and other fields. In intelligent traffic management, trajectory data can be used to monitor traffic flow in real time, predict congested areas, and provide precise support for traffic scheduling and planning; in logistics optimization, by analyzing the movement paths of vehicles or goods, transportation routes can be optimized and distribution efficiency can be improved; in urban planning, trajectory data reveals the movement patterns of people and vehicles, providing an important basis for the rational layout of public facilities and the design of transportation networks. With the rapid development of technologies such as sensors, satellite positioning and social media, the scale and complexity of trajectory data continue to grow, and its potential value has become a research hotspot. Trajectory mining, as an important data analysis task, can extract traffic flow patterns and behavior patterns from massive trajectory data. For example, by analyzing the trajectories of urban vehicles, it is possible to identify major traffic corridors, grasp the distribution of traffic during peak hours, and provide decision support for the scientific formulation of traffic management strategies. This technology not only serves the transportation field, but is also widely used in scenarios such as environmental monitoring and behavior analysis, and is an important basis for data-driven decision-making.
[0003] As one of the core tasks of trajectory mining, traffic mode classification aims to identify and distinguish the modes of transportation used by mobile objects (such as cars, buses, bicycles, and walking). This task is of great significance in intelligent transportation systems, public transportation scheduling, and personalized navigation. For example, traffic mode classification can be used to monitor traffic status in real time, optimize signal control and road planning; classification results can support bus line design and scheduling strategy adjustment; accurate identification of traffic modes can also improve the accuracy of navigation path planning. Accurate traffic mode classification can not only optimize urban traffic planning and improve traffic operation efficiency, but also provide reliable data support for traffic flow modeling. In the construction of smart cities, traffic mode classification is one of the key technologies to promote the development of green and low-carbon cities.
[0004] Existing traffic mode classification methods mainly include traditional methods, machine learning-based methods, and deep learning-based methods. Traditional methods usually rely on statistical features such as speed, acceleration, and heading change rate to classify traffic modes by setting rules or thresholds. For example, low-speed trajectory points are judged as walking, and high-speed trajectories are judged as driving. However, such methods have difficulty in processing high-dimensional and complex spatiotemporal trajectory data, and cannot effectively capture implicit features, especially in scenarios with high similarity in traffic modes such as taxis and private cars. At the same time, they have poor adaptability to dynamic environments. Machine learning-based methods, such as support vector machines and random forests, improve classification performance by designing features such as path length and average speed. However, such methods rely on domain knowledge, have a long development cycle and lack flexibility, and have difficulty in utilizing the global information and temporal dependencies of trajectory data, limiting their applicability. In contrast, deep learning-based methods (such as convolutional neural networks and long short-term memory networks) have shown strong feature extraction capabilities in trajectory classification tasks. Convolutional neural networks extract local spatial features of trajectories through convolution kernels and capture geometric relationships between trajectory points, such as direction changes and local path patterns; long short-term memory networks model the time series dependencies of trajectories through memory mechanisms and are suitable for dynamically changing trajectory features. Some studies have also combined the two to capture both the spatial patterns of trajectories and model time-dependent characteristics, thereby achieving excellent performance in complex classification tasks. The end-to-end learning capability of deep learning can automatically extract high-dimensional feature representations of trajectories without the need for tedious manual feature design. It is particularly suitable for processing high-dimensional and complex trajectory data, while modeling the global characteristics of trajectories with the support of large-scale data.
[0005] Although existing traffic mode classification methods have made some progress in different scenarios, they still have shortcomings when processing high-dimensional and complex spatiotemporal trajectory data, including insufficient use of global information and temporal dependencies, and limited performance in small sample scenarios and complex classification tasks. In response to the above problems, contrastive learning has gradually become an effective method to solve feature representation and classification tasks in recent years. Contrastive Predictive Coding (CPC) is an emerging contrastive learning method that has achieved remarkable results in fields such as audio, image, and natural language processing. CPC efficiently learns potential representations under unsupervised conditions by predicting future segments of data, and can capture complex temporal and spatial features, providing a new technical path for the deep representation of trajectory data. At the same time, the dynamic queue mechanism proposed by Momentum Contrast (MoCo) significantly improves the sample utilization efficiency and discrimination ability of contrastive learning by introducing a fixed-length negative sample queue, dynamically updating the diversity of negative samples, and using the momentum encoder to generate stable feature representations.
[0006] Based on the advantages of contrastive learning, this paper proposes a new contrastive learning framework for efficiently learning trajectory feature representation and applying it to traffic mode classification tasks. The framework extracts high-dimensional feature representations of trajectories through deep neural networks, and effectively models the spatiotemporal dependencies of trajectories in combination with contrastive learning strategies. In order to further optimize the discriminative ability of feature representation, the framework innovatively introduces a dynamic queue mechanism to dynamically expand the number and diversity of negative samples, thereby improving the effect and efficiency of contrastive learning. The framework significantly enhances the classification accuracy and robustness of the model, and provides an efficient and reliable technical solution for trajectory data analysis in complex scenarios. Summary of the invention
[0007] In view of the shortcomings of existing traffic mode classification methods in processing high-dimensional complex trajectory data, the present invention provides a trajectory traffic mode classification method, which proposes a new contrastive learning framework for traffic mode classification tasks. The framework constructs a multi-layer convolutional network and a Transformer encoder through a deep neural network to extract high-dimensional feature representations of trajectory data, and combines the gated recurrent unit (GRU) to model the time series dependencies of the trajectory. In addition, the framework innovatively introduces a dynamic queue mechanism to dynamically expand the number and diversity of negative samples, significantly improving the discriminative ability of feature representation and the classification performance of the model.
[0008] The detailed technical scheme of the present invention is as follows:
[0009] A trajectory traffic mode classification method, the method comprising:
[0010] S1, performing preliminary processing operations on the original trajectory data to construct a trajectory data target set;
[0011] S2, inputting the trajectory data target set as training data into a multi-level trajectory feature encoder, and training the multi-level trajectory feature encoder by using a contrastive learning method in combination with an autoregressive model;
[0012] S3, using the trained multi-level trajectory feature encoder to extract high-dimensional feature representation of the trajectory data, and using it as input to train the MLP classifier so that it can learn the mapping relationship between trajectory features and traffic modes. After the training is completed, the high-dimensional feature representation of the trajectory data is input into the MLP classifier to classify the trajectory traffic mode;
[0013] Wherein, the S2 specifically includes:
[0014] S21, inputting the trajectory data target set into the multi-level trajectory feature encoder, using the shallow feature extraction module of the multi-level trajectory feature encoder to extract the local spatial features of the trajectory and convert them into low-dimensional potential feature representations of uniform dimensions, and using the deep feature extraction module of the multi-level trajectory feature encoder to obtain the global spatiotemporal dependency relationship between trajectory points based on the low-dimensional potential feature representation to generate a high-dimensional feature representation;
[0015] S22, inputting the high-dimensional feature representation into an autoregressive model, wherein the autoregressive model generates a contextual representation of each time step based on the dependency of the time series, constructing a positive sample based on the contextual representation of each time step and the true feature representation, and introducing a dynamic queue mechanism to construct a negative sample based on the contextual representation of each time step and the feature representation in the dynamic queue;
[0016] S23. Using the constructed positive samples and negative samples and based on the noise contrast estimation loss function, adopt a contrastive learning method to train the multi-level trajectory feature encoder.
[0017] According to a preferred embodiment of the present invention, S1 specifically includes: preprocessing the original trajectory data to construct a trajectory data set , the preprocessing operation includes: smoothing the noise points in the original trajectory data by using a dynamic window weighted average method, and correcting the geographical coordinates of the drift points in the original trajectory data by using a weighted regression method based on trajectory trends;
[0018] The method of smoothing the noise points in the original trajectory data by using a dynamic window weighted average method specifically includes:
[0019] By calculating the speed of adjacent trajectory points and comparing it with the preset maximum speed threshold Compare, exceed the maximum speed threshold The trajectory point is judged as a noise point; the calculation formula of its speed is as follows:
[0020] (1);
[0021] In formula (1): represents the velocity of the trajectory point; d( ) represents the trajectory point and The geographical distance between , Represents the trajectory points and timestamp;
[0022] Through the track point A dynamic window composed of several non-noise points before and after Weighted average, re-estimate its coordinate value; the smoothed coordinate calculation formula is:
[0023] (2);
[0024] In formula (2): , Represents the smoothed trajectory points The latitude and longitude coordinates of; Represents trajectory points A dynamic window containing several non-noise points before and after; , Represents dynamic windows The latitude and longitude coordinates of each valid track point in Represents a dynamic window The index of each valid trajectory point in , Represents a dynamic window The number of valid trajectory points within;
[0025] The weighted regression method based on trajectory trend corrects the geographic coordinates of the drift point in the original trajectory data, specifically including:
[0026] The geographical distance d( ) and dynamic threshold Compare and exceed the dynamic threshold The trajectory points of are judged as drift points, namely:
[0027] (3);
[0028] The corrected coordinates of the drift point are determined by fitting the trajectory trends of the trajectory points before and after the drift point. Assuming that the trajectory point is the drift point, and its normal trajectory points before and after are and , then:
[0029] (4);
[0030] (5);
[0031] In formula (4) and (5): Represents the corrected trajectory point The latitude coordinate of Represents the corrected trajectory point The longitude coordinates of Indicates the previous normal trajectory point The latitude coordinate of Indicates the previous normal trajectory point The longitude coordinates of Represents the next normal trajectory point The latitude coordinate of Represents the next normal trajectory point The longitude coordinates of , Both represent weights, which are set according to the time interval or geographical distance between the drift point and the normal trajectory points before and after; weight , The calculation formula is:
[0032] = = (6);
[0033] In formula (6): d( ) represents the trajectory point The next normal trajectory point The geographical distance between ) represents the trajectory point The previous normal trajectory point The geographical distance between ) represents the previous normal trajectory point The next normal trajectory point The geographical distance between
[0034] The trajectory data set Contains several tracks , each trajectory A series of trajectory points Composition, and trajectory points The core fields included are latitude and longitude coordinates and timestamp, namely:
[0035] (7);
[0036] (8);
[0037] (9);
[0038] In formulas (7)-(9): For all trajectories All satisfy the condition that the number of trajectory points is greater than or equal to the minimum threshold requirements; Represents trajectory points The latitude coordinate of Represents trajectory points The longitude coordinates of Represents trajectory points The timestamp of the
[0039] According to a preferred embodiment of the present invention, the S1 further comprises: The trajectory data in is subjected to standardization operation, data enhancement operation, shallow feature extraction operation and transposition operation in turn to obtain the trajectory data target set;
[0040] Among them, the trajectory data set The trajectory data in the data set is standardized, specifically: the trajectory length of the trajectory data is uniformly adjusted to 1024 points, that is, for trajectories with a length greater than 1024 points, they are reduced to 1024 points by random truncation, and for trajectories with a length less than 1024 points, they are supplemented based on a sliding window linear interpolation method in the data enhancement operation;
[0041] Based on the sliding window linear interpolation method, the standardized trajectory data is enhanced. Specifically, the number of target trajectory points is set. As the length of the sampling window, and calculate the number of current trajectory points ; According to the target trajectory points With the current track points The difference between ,Right now:
[0042] (10);
[0043] When the number of trajectory points to be inserted When interpolation is required, the sliding window method is used to traverse the trajectory points, with the window size set to 4 and the moving step size set to 3. For each group of four adjacent trajectory points , and , calculate the first trajectory point The change in orientation between the other three trajectory points and the average value , and its calculation formula is:
[0044] (11);
[0045] If the average value of the azimuth change Less than the preset threshold , it is considered that the trajectory characteristics of the area formed by the above four adjacent trajectory points change little and meet the interpolation conditions; when the interpolation conditions are met, at the trajectory point and Insert new track points between , whose coordinates are calculated by linear interpolation:
[0046]
[0047] The insertion of a new trajectory point increases the number of trajectory points by 1. Then, the sliding window continues to move forward, calculates the position change of the next set of adjacent trajectory points, and determines whether to continue interpolation. When a round of interpolation is completed, if the number of trajectory points has not reached the target number of trajectory points, , a new round of interpolation starts from the beginning, and the number of points inserted in each round is gradually reduced until the number of trajectory points reaches the target trajectory point number. ;
[0048] The shallow feature extraction operation is performed on the enhanced trajectory data. The extracted shallow features include speed , speed change ,position , Direction Change , angular velocity and curvature ,in:
[0049] = (13);
[0050] (14);
[0051] (15);
[0052] (16);
[0053] (17);
[0054] (18);
[0055] In formulas (13)-(18): Represents trajectory points The speed of movement; Represents trajectory points and The geographical distance between and Represents trajectory points and timestamp; Represents trajectory points and The speed change between Represents trajectory points azimuth; , Represents trajectory points The latitude and longitude coordinates of; Represents trajectory points directional changes; Represents trajectory points azimuth; Represents trajectory points Angular velocity; Represents trajectory points The curvature of
[0056] The transposition operation is specifically as follows: the original data with the shape of (trajectory_window, num_features) is transposed to form target data with the shape of (num_features, trajectory_window), where trajectory_window represents the number of trajectory points for each trajectory, and num_features represents the initial number of features; after the transposition is completed, the trajectory data is converted into a PyTorch tensor, and finally constructed as a trajectory data target set.
[0057] Preferably, in S21, the shallow feature extraction module of the multi-level trajectory feature encoder adopts a five-layer one-dimensional convolutional network structure, which is composed of a one-dimensional convolutional layer Conv1×1, a batch normalization layer BN and a ReLU activation function stacked in sequence, for gradually extracting multi-level features of trajectory data and generating a high-dimensional feature representation containing spatiotemporal information;
[0058] Assume that the original input of the shallow feature extraction module is , the convolution operation of each layer is expressed as:
[0059] (19);
[0060] In formula (19): Indicates the number of convolutional layers; Indicates Output features of the layer;
[0061] The overall effect of the shallow feature extraction module is expressed as:
[0062] (20);
[0063] In formula (20): represents a low-dimensional latent feature representation; Represents the trajectory data in the trajectory data target set; A feature encoder representing trajectory data;
[0064] The output of the shallow feature extraction module is a three-dimensional tensor with a shape of [batch_size, feature_dim, sequence_length], where batch_size=8 means that the current batch contains 8 trajectories, feature_dim=512 means the feature dimension of each time step, and sequence_length=32 means that the trajectory consists of 32 time steps after shallow feature extraction;
[0065] The deep feature extraction module of the multi-level trajectory feature encoder adopts a Transformer encoder, which is stacked by L layers, each layer including a multi-head self-attention module and a multi-layer perceptron module, for capturing the global dependency between trajectory subsequences; the three-dimensional tensor output by the shallow feature extraction module is flattened into a two-dimensional tensor, and the trajectory feature is converted from a two-dimensional sequence containing feature dimensions and time steps into a one-dimensional sequence; the flattened trajectory feature is divided into N equal-sized segments, each of which is mapped to a fixed dimension D through a linear projection layer to provide a unified high-dimensional input for the Transformer encoder; a one-dimensional position encoding is added to each segment , the final input sequence of the Transformer encoder is:
[0066] (twenty one);
[0067] In formula (21): represents the input sequence of the Transformer encoder, represents the trajectory feature subsequence after linear projection, represents the positional encoding, i.e., a learnable parameter matrix consistent with the length of the subsequence, which is used to embed the position information of the fragment;
[0068] The input sequence The multi-head self-attention module is used to calculate the correlation and attention weights between trajectory segments to model global dependencies. The calculation process is as follows:
[0069] (twenty two);
[0070] In formula (22): Represents the Transformer encoder The intermediate feature representation of the layer after multi-head attention calculation; Indicates -1 layer output; LN(·) is the layer normalization operation to ensure the stability of the input distribution;
[0071] The multi-layer perceptron module further extracts nonlinear features from the attention output and enhances the feature expression ability and network convergence through residual connections; the specific calculation is:
[0072] (twenty three);
[0073] In formula (23): Represents the Transformer encoder The final output feature representation of the layer, that is, the high-dimensional feature representation ;MLP It is a two-layer fully connected network that includes activation functions to enhance the nonlinear expression of features.
[0074] Preferably, in S22, the autoregressive model uses a gated recurrent unit GRU as a core component, and its input is a high-dimensional feature representation extracted by the multi-level trajectory feature encoder. ; The process of calculating the hidden state of the autoregressive model at each time step is as follows:
[0075] =GUR( )(twenty four);
[0076] In formula (24): Indicates the hidden state at the current moment, represents the dimension of the hidden layer; represents the hidden state of the previous time step; Represents the input features of the current time step, that is, the high-dimensional feature representation output by the multi-level trajectory feature encoder;
[0077] At each time step, the output hidden state of the autoregressive model is used as the context representation ,Right now:
[0078] (25);
[0079] Through the linear projection layer The context of each time step is represented as Mapped to feature space to get predicted feature representation ,Right now:
[0080] (26);
[0081] Based on the predicted feature representation and the true feature representation Construct positive samples;
[0082] Introducing dynamic queues , based on the predicted feature representation and dynamic queues Feature Representation Construct negative samples, where , u represents the queue The index of the samples stored in .
[0083] Preferably, in S23, the multi-level trajectory feature encoder is optimized and trained based on the noise contrast estimation loss function, and the optimization goal is to maximize the similarity of positive samples and minimize the similarity of negative samples. In the similarity measurement stage, the positive samples are Similarity and negative sample pairs The similarity is calculated by inner product:
[0084] (27);
[0085] (28);
[0086] In formula (27)-(28): Indicates the similarity of positive sample pairs; Represents the time step Trajectory data at represents the transposed vector of the predicted feature representation; Represents the time step The true feature representation of Represents the similarity of negative sample pairs; Represents the trajectory data input corresponding to the negative sample; Represents the negative sample feature representation;
[0087] The noise contrast estimation loss function is defined as:
[0088] (29);
[0089] In formula (29): represents the noise contrast estimation loss function, Represents the distribution of trajectory data Take the mathematical expectation.
[0090] Preferably, in S3, the MLP classifier includes an input layer, two hidden layers and an output layer, which uses a cross entropy loss function as a training target to measure the difference between the predicted value and the true label; assuming that the predicted output of the MLP classifier is , the true label is , Represents the number of categories, then the loss function is defined as:
[0091] (30);
[0092] In formula (30): represents the cross entropy loss; batch size Indicates the number of samples contained in a training batch; the number of categories Represents the total number of target categories in the classification task; the true category index The value range is , used to traverse all categories; indicator function 1=( =c) means when the sample The true category is The value is 1 when the value is 1, otherwise it is 0; prediction score (logits) Representation sample In category The output value on is not normalized by the softmax layer; the category index J is used for the summation calculation of the normalized denominator of the softmax layer, and its value range is , corresponding to the sample In category The prediction score on , used to calculate the normalized probability.
[0093] In another aspect of the present invention, a trajectory traffic mode classification device is provided, the device comprising: a data processing module, used to perform preliminary processing operations on raw trajectory data to construct a trajectory data target set; a model training module, used to input the trajectory data target set as training data into a multi-level trajectory feature encoder, and train the multi-level trajectory feature encoder using a contrastive learning method in combination with an autoregressive model; using the trained multi-level trajectory feature encoder to extract a high-dimensional feature representation of the trajectory data, and using this as input to train an MLP classifier so that it learns the mapping relationship between trajectory features and traffic modes. After the training is completed, the high-dimensional feature representation of the trajectory data is input into the MLP classifier to perform trajectory traffic mode classification.
[0094] In another aspect of the present invention, an electronic device is provided, comprising: at least one processor; and a memory, wherein the memory stores instructions, and when the instructions are executed by the at least one processor, the at least one processor executes the trajectory traffic mode classification method as described above.
[0095] In another aspect of the present invention, a computer-readable storage medium is provided, which stores executable instructions. When the instructions are executed, the machine executes the trajectory traffic mode classification method as described above.
[0096] Compared with the prior art, the present invention has the following beneficial effects:
[0097] (1) Innovative contrastive learning framework and dynamic queue mechanism. This paper proposes a new contrastive learning framework that combines multi-layer convolutional networks, Transformer encoders, and gated recurrent units (GRUs) to achieve efficient feature extraction and modeling of trajectory data. The framework innovatively introduces a dynamic queue mechanism to dynamically expand the number and diversity of negative samples, significantly improving the discriminative ability and optimization efficiency of feature representation. The framework can fully capture the complex spatiotemporal relationships of trajectory data, providing an efficient and accurate solution for feature learning and traffic mode classification of high-dimensional complex trajectory data.
[0098] (2) Multi-stage workflow design and efficient classification. The present invention designs a multi-stage workflow that includes three stages: data preprocessing and enhancement, multi-level feature encoding and model training, and traffic mode classification. In the data preprocessing and enhancement stage, the data quality and feature characterization capabilities are greatly improved by introducing data cleaning, sliding window linear interpolation methods, and shallow dynamic features (such as speed, orientation, curvature, etc.). On this basis, the high-dimensional feature representation of the trajectory is extracted using the trained encoder, and combined with the multi-layer perceptron (MLP) classifier, efficient and accurate traffic mode classification is achieved, which meets the application requirements in complex scenarios and provides reliable technical support for fields such as intelligent transportation, urban planning, and dynamic traffic management. BRIEF DESCRIPTION OF THE DRAWINGS
[0099] Figure 1 It is a framework diagram of the trajectory traffic mode classification method described in the present invention.
[0100] Figure 2 It is a schematic diagram of enhancing trajectory data based on the sliding window linear interpolation method in Example 1 of the present invention.
[0101] Figure 3 It is a structural diagram of the shallow feature extraction module in Example 1 of the present invention.
[0102] Figure 4 It is a structural diagram of the deep feature extraction module in Example 1 of the present invention.
[0103] Figure 5 It is a schematic diagram of constructing positive and negative samples based on the dynamic queue mechanism in Example 1 of the present invention.
[0104] Figure 6 It is a structural diagram of a multi-layer perceptron (MLP) classifier in Example 1 of the present invention. DETAILED DESCRIPTION
[0105] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0106] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.
[0107] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.
[0108] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.
[0109] In view of the shortcomings of existing traffic mode classification methods in processing high-dimensional complex trajectory data, this paper proposes a new contrastive learning framework for traffic mode classification tasks. The framework constructs a multi-layer convolutional network and Transformer encoder through a deep neural network to extract high-dimensional feature representations of trajectory data, and combines the gated recurrent unit (GRU) to model the time series dependencies of the trajectory. In addition, the framework innovatively introduces a dynamic queue mechanism to dynamically expand the number and diversity of negative samples, significantly improving the discriminative ability of feature representation and the classification performance of the model.
[0110] Example 1
[0111] This embodiment provides a trajectory traffic mode classification method. Figure 1 ,The overall process includes data preprocessing and enhancement stage, multi-level feature encoding and ,model training stage, and traffic mode classification stage.
[0112] In the data preprocessing and enhancement stage, the original trajectory data is first cleaned, including removing noise points, correcting drift points, and adjusting the order of data dimensions through transposition operations to meet the model input requirements. In order to increase data density and enhance feature continuity, a linear interpolation method based on a sliding window is used to achieve data enhancement, and the number of trajectory points is increased while keeping the key features of the trajectory unchanged. In addition, in the shallow feature enrichment stage, shallow dynamic features such as speed, azimuth, and curvature are introduced to further supplement the characterization capabilities of the original longitude and latitude and timestamps, providing more comprehensive feature support for subsequent modeling.
[0113] In the multi-level feature encoding and model training stage, the shallow spatial features of the trajectory are first extracted through a five-layer one-dimensional convolutional network (1D-CNN) and uniformly converted into a low-dimensional latent feature representation of a fixed dimension. Subsequently, a multi-layer Transformer encoder is used to model the global dependencies between trajectory points to further extract deep spatiotemporal features. At the same time, the gated recurrent unit (GRU) is combined as an autoregressive model to model the time series relationship of the trajectory and generate the contextual representation of the current moment. On this basis, the future time step prediction task is combined with the noise contrast estimation (InfoNCE) loss function for optimization to maximize the mutual information between the contextual representation and the predicted features. In addition, a dynamic queue mechanism is innovatively introduced to dynamically expand the number and diversity of negative sample sets, thereby effectively improving the discriminative ability and optimization efficiency of contrastive learning.
[0114] In the traffic mode classification stage, the trained encoder is used to extract high-dimensional feature representations of the trajectory, and these representations are input into the multi-layer perceptron (MLP) classifier for classification. Relying on high-quality feature representation, the classifier achieves efficient and accurate traffic mode classification, significantly improves the classification accuracy and robustness of the model, and meets the actual application needs in complex scenarios. Through this multi-stage workflow, the new contrastive learning framework proposed in this method can fully capture the complex spatiotemporal relationships in trajectory data, significantly improve the accuracy and adaptability of traffic mode classification, and provide an efficient and reliable technical solution for trajectory feature representation and classification.
[0115] Specifically, the trajectory traffic mode classification method includes:
[0116] S1. Perform preliminary processing operations on the original trajectory data to construct a trajectory data target set.
[0117] This method can be applied to common trajectory datasets, which usually contain multiple traffic modes, such as cars, motorcycles, walking, cycling, etc. The datasets cover multiple cities and different proportions of traffic modes to ensure the good performance of the model in various traffic mode classification tasks. The datasets are screened according to features such as the number of trajectory points and time intervals, and are ultimately used for model training, validation, and testing. In trajectory data processing, data quality has a crucial impact on model performance. The original trajectory data often contains noise points and drift points, which may be caused by the accuracy limitations of the acquisition equipment or environmental interference. Without preprocessing, these abnormal data will cause trajectory feature distortion and affect the model's accurate characterization of trajectory characteristics. Therefore, data cleaning is a necessary step before performing trajectory feature extraction and classification tasks. Preprocessing operations such as smoothing noise points and correcting drift points can effectively improve the quality and consistency of the data, providing a solid foundation for subsequent modeling.
[0118] In this embodiment, the original trajectory data is pre-processed, including pre-processing the original trajectory data to construct a trajectory data set. The preprocessing operation includes: smoothing the noise points in the original trajectory data by using a dynamic window weighted average method, and correcting the geographic coordinates of the drift points in the original trajectory data by using a weighted regression method based on trajectory trends. The specific operations are as follows:
[0119] 1) Smoothing noise points: In order to detect and process noise points in trajectory data, this method uses a dynamic window weighted average smoothing method based on the spatiotemporal characteristics of trajectory points. First, possible abnormal points are detected by analyzing the motion characteristics between adjacent trajectory points. Specifically, the speed of adjacent trajectory points is calculated and compared with the preset maximum speed threshold. Compare, exceed the maximum speed threshold The trajectory points of are determined to be noise points. The speed calculation formula is as follows:
[0120] (1);
[0121] In formula (1): represents the velocity of the trajectory point; d( ) represents the trajectory point and The geographical distance between , Represents the trajectory points and The timestamp of the
[0122] For the detected noise points, a weighted average method based on a dynamic window is used for smoothing. Assuming that the trajectory point is detected as a noise point, where For trajectory points The latitude coordinate of For trajectory points The longitude coordinates of the track point A dynamic window composed of several non-noise points before and after , and smooth its coordinates. The smoothed latitude and longitude coordinates are obtained by calculating the simple average of the non-noise points in the window, and the formula is as follows:
[0123] (2);
[0124] In formula (2), , Represents the smoothed trajectory points The latitude and longitude coordinates of; Represents trajectory points A dynamic window containing several non-noise points before and after; , Represents dynamic windows The latitude and longitude coordinates of each valid track point in the Represents a dynamic window The index of each valid trajectory point in , that is, Represents trajectory points The indexes of the valid points before and after. The coordinates of these points will be used to calculate the smoothing result. It is a dynamic window The number of valid trajectory points in the range is used to ensure that the correct number of points is considered during calculation.
[0125] Through the above-mentioned weighted average method, the influence of noise points on trajectory data can be effectively smoothed.
[0126] 2) Handling drift points: In the trajectory similarity measurement, drift points are usually abnormal trajectory points caused by GPS signal errors or poor reception conditions, which are manifested as significant deviations from the actual trajectory. In order to detect drift points, this method adopts a dynamic threshold strategy to identify abnormal points by analyzing the geographical distance between adjacent trajectory points. If the distance d( ) exceeds the dynamic threshold , then the trajectory point or is the drift point, that is:
[0127] (3).
[0128] The dynamic threshold It can be dynamically adjusted according to the data sampling frequency and the maximum possible speed between trajectory points to adapt to the characteristics of different trajectory data. For the detected drift points, this method uses a weighted regression method based on trajectory trends to correct their geographic coordinates instead of simple linear interpolation. Specifically, assuming that the trajectory point is the drift point, and its normal trajectory points before and after are and , then the corrected coordinates of the drift point are determined by fitting the trajectory trends of the previous and next trajectory points, that is:
[0129] (4);
[0130] (5);
[0131] In formula (4) and (5), Represents the corrected trajectory point The latitude coordinate of Represents the corrected trajectory point The longitude coordinates of Indicates the previous normal trajectory point The latitude coordinate of Indicates the previous normal trajectory point The longitude coordinates of Represents the next normal trajectory point The latitude coordinate of Represents the next normal trajectory point The longitude coordinates of , Both represent weights, which are set according to the time interval or geographical distance between the drift point and the normal track points before and after it. The closer the weight, the greater the proportion of the track point. , The calculation formula is:
[0132] = = (6);
[0133] In formula (6), d( ) represents the trajectory point The next normal trajectory point The geographical distance between ) represents the trajectory point The previous normal trajectory point The geographical distance between ) represents the previous normal trajectory point The next normal trajectory point This weighted regression method not only considers the spatial information of trajectory points, but also combines the time series relationship of trajectory points, so as to maintain the overall trend and coherence of the trajectory when correcting the drift points.
[0134] After data cleaning, the trajectory data set Contains several trajectory data that meet the quality standards. Each trajectory is recorded as , where each trajectory A series of trajectory points The number of trajectory points meets the minimum threshold Formally, the trajectory data set It is expressed as:
[0135] (7);
[0136] In formula (7), For all trajectories All satisfy the condition that the number of trajectory points is greater than or equal to the minimum threshold requirements.
[0137] Each track It can be further expressed as a sequence of trajectory points, namely:
[0138] (8).
[0139] Among them, the trajectory point The core fields included are latitude and longitude coordinates and timestamps, for example:
[0140] (9);
[0141] In formula (9), Represents trajectory points The latitude coordinate of Represents trajectory points The longitude coordinates of Represents trajectory points The timestamp of the
[0142] Cleaned trajectory data set It provides a standardized data foundation for subsequent feature extraction and modeling, while ensuring the consistency and accuracy of trajectory points.
[0143] In this embodiment, the initial processing operation of the original trajectory data also includes processing the trajectory data set The trajectory data in are subjected to standardization, data enhancement, shallow feature extraction and transposition operations in turn to obtain the trajectory data target set.
[0144] 3) For the trajectory data set The trajectory data in the data set is standardized, specifically: the trajectory length of the trajectory data is uniformly adjusted to 1024 points, that is, for trajectories with a length greater than 1024 points, they are reduced to 1024 points by random truncation, and for trajectories with a length less than 1024 points, they are supplemented based on the sliding window linear interpolation method in the data enhancement operation. That is, in the data preprocessing, the trajectory lengths of different data sets are first preliminarily screened, and the trajectory lengths are limited to a specific range (such as 200 to 2000 points) to eliminate abnormally short or long trajectories. On this basis, in order to meet the input requirements of the model (i.e., the shallow feature extraction module of the multi-level trajectory feature encoder in the subsequent steps), the trajectory length is further uniformly adjusted to 1024 points. Specifically, for trajectories with a length greater than 1024, they are reduced to 1024 points by random truncation; for trajectories with a length less than 1024, they are supplemented by a sliding window-based linear interpolation method in the subsequent data enhancement stage to ensure the uniformity of the trajectory length and the consistency of the model input. The fixed choice of trajectory length is not arbitrary, but is carefully considered in the experimental design. Fixing the trajectory length to 1024 points can fully preserve the temporal and spatial information of the trajectory, while effectively balancing the computational efficiency of the model and the adequacy of feature extraction. With this setting, the model can capture the global and local dynamic characteristics of the trajectory while avoiding the negative impact of excessive padding or feature redundancy on model performance, providing standardized input guarantee for subsequent feature extraction and modeling.
[0145] 4) Perform data enhancement operations on the standardized trajectory data based on the sliding window linear interpolation method.
[0146] In order to improve the model's perception of trajectory details, a sliding window-based linear interpolation method is proposed. By increasing the number of trajectory points, the temporal and spatial resolution of the trajectory is enhanced, thereby meeting the training requirements of the contrastive learning framework. In the multi-level trajectory feature encoder, one-dimensional convolution is used to extract trajectory features, but the convolution operation may cause the trajectory feature length to gradually shorten, especially when the number of trajectory points is small or the sampling is uneven, which will limit the effective learning of features by the entire model composed of the encoder and the autoregressive model. To solve this problem, the sliding window linear interpolation method increases the density of trajectory points while maintaining the overall structure and key features of the trajectory, avoiding information loss caused by feature reduction. Specific parameters are as follows: Figure 2 By introducing this method, the model not only enhances the ability to capture trajectory details, but also improves the understanding of global patterns, provides reliable support for feature extraction and representation learning, and effectively improves the training effect and classification performance of the model.
[0147] In order to achieve effective interpolation of the trajectory and meet the target number of points, first set the target trajectory points As the length of the sampling window, and calculate the number of current trajectory points According to the number of target trajectory points With the current track points The difference between ,Right now:
[0148] (10).
[0149] When the number of trajectory points to be inserted When , no interpolation is required. When interpolation is required, the sliding window method is used to traverse the trajectory points. The window size is set to 4, the moving step is set to 3, and for each group of four adjacent trajectory points , and , calculate the first trajectory point The change in orientation between the other three trajectory points and the average value . Azimuth change average The calculation formula is:
[0150] (11).
[0151] If the average value of the azimuth change Less than the preset threshold , it is considered that the trajectory characteristics of the area formed by the above four adjacent trajectory points change little and meet the interpolation conditions. When the interpolation conditions are met, at the trajectory point and Insert new track points between , whose coordinates are calculated by linear interpolation:
[0152]
[0153] The insertion of a new trajectory point increases the number of trajectory points by 1. Then, the sliding window continues to move forward, calculates the position change of the next set of adjacent trajectory points, and determines whether to continue interpolation. When a round of interpolation is completed, if the number of trajectory points has not reached the target number of trajectory points, , a new round of interpolation starts from the beginning, and the number of points inserted in each round is gradually reduced until the number of trajectory points reaches the target trajectory point number. Through this sliding window-based interpolation strategy, the key features of the trajectory are retained, while the density and uniformity of the trajectory points are significantly improved, providing better data support for subsequent feature extraction and training of the entire model.
[0154] 5) Perform shallow feature extraction operations on the enhanced trajectory data.
[0155] After data enhancement, in order to further improve the representation ability of trajectory features, a shallow feature enrichment stage is designed. In the trajectory similarity representation learning task, it is difficult to fully characterize the complex motion patterns and dynamic behaviors of the trajectory by relying solely on longitude, latitude and timestamp information. Although longitude and latitude data can provide basic location information, they have obvious limitations in describing the dynamic characteristics of trajectory movement, such as speed, direction and change. This deficiency may lead to insufficient feature expression in scenes with complex motion changes, thereby limiting the performance of the model. To solve this problem, a series of shallow dynamic features are introduced, including speed, direction and timestamp information. , speed change ,position , Direction Change , angular velocity and curvature ,in:
[0156] = (13);
[0157] (14);
[0158] (15);
[0159] (16);
[0160] (17);
[0161] (18);
[0162] In formulas (13)-(18): Represents trajectory points Speed of movement; d( ) represents the trajectory point and The geographical distance between and Represents trajectory points and timestamp; Represents trajectory points and The speed change between , Represents trajectory points and azimuth; , Represents trajectory points The latitude and longitude coordinates of; Represents trajectory points directional changes; Represents trajectory points Angular velocity; Represents trajectory points The curvature.
[0163] The above features describe the dynamic behavior of the trajectory from multiple dimensions. The speed and speed change reveal the speed of movement and the acceleration and deceleration trend, which helps to distinguish trajectories with similar spatial positions but different motion patterns. The orientation information further provides the directional characteristics of the object's movement, so that the trajectory similarity calculation can distinguish trajectories with similar spatial positions but different motion directions. Orientation change The angular velocity describes the frequency and magnitude of changes in direction, helping to distinguish sharp turns from smooth movements. It reflects the rate of change of direction per unit time and can capture rapid turns or direction adjustments in complex traffic scenarios. Although the orientation, orientation change, and angular velocity can better describe the direction change characteristics of the trajectory, curvature supplements the understanding of the curvature of the trajectory from a geometric perspective. Curvature not only reflects the change in direction, but also reveals the overall curvature trend of the trajectory, especially in scenarios with complex routes or frequent turns, it can more comprehensively describe the shape characteristics of the trajectory.
[0164] By extracting these shallow features, our method enhances the spatiotemporal understanding of trajectories from multiple perspectives, providing richer information for subsequent trajectory representation learning and similarity calculation.
[0165] 6) Data transposition and format optimization: After the above preprocessing, data enhancement and shallow feature extraction, the data is transposed to ensure that the dimensional order of the data matches the model input requirements. The original data shape is (trajectory_window, num_features), where trajectory_window represents the number of trajectory points for each trajectory and num_features represents the number of initial features. After transposition, the data shape is adjusted to (num_features, trajectory_window), arranging the data of each feature together so that the model can correctly parse each feature sequence. This transposition operation is particularly suitable for time series models because time series models need to rely on the temporal relationship between features for prediction. After the transposition is completed, the data is converted into a PyTorch tensor. The efficient computing power and GPU acceleration function of the PyTorch framework not only simplifies the subsequent training process, but also improves the training efficiency and performance of the model. This processing method ensures the accuracy and effectiveness of trajectory data in training, laying a solid foundation for the trajectory similarity representation learning of the model.
[0166] Based on the above operations, the processed trajectory data is finally constructed into a trajectory data target set.
[0167] S2. Input the trajectory data target set as training data into a multi-level trajectory feature encoder, and train the multi-level trajectory feature encoder using a contrastive learning method in combination with an autoregressive model.
[0168] It specifically includes: S21, inputting the trajectory data target set into the multi-level trajectory feature encoder, using the shallow feature extraction module of the multi-level trajectory feature encoder to extract the local spatial features of the trajectory and convert them into low-dimensional potential feature representations of unified dimensions, and using the deep feature extraction module of the multi-level trajectory feature encoder to obtain the global spatiotemporal dependency between trajectory points based on the low-dimensional potential feature representation to generate a high-dimensional feature representation; S22, inputting the high-dimensional feature representation into an autoregressive model, the autoregressive model generates a context representation of each time step based on the dependency of the time series, constructing positive samples based on the context representation of each time step and the true feature representation, and introducing a dynamic queue mechanism to construct negative samples based on the context representation of each time step and the feature representation in the dynamic queue; S23, using the constructed positive and negative samples, and based on the noise contrast estimation loss function, adopting a contrastive learning method to train the multi-level trajectory feature encoder.
[0169] That is, in this embodiment, the multi-level feature encoding and model training stage consists of the following core modules: a multi-level trajectory feature encoder, an autoregressive module for time-dependent modeling, a positive and negative sample construction module, and a contrastive learning optimization module. Among them, the multi-level trajectory feature encoder captures the local dynamic characteristics of the trajectory through a shallow feature extraction module, and uses a deep feature extraction module to model the global spatiotemporal dependency between trajectory points. The time-dependent modeling module uses a gated recurrent unit (GRU) to generate a contextual representation of the trajectory to capture time series dependencies. The positive and negative sample construction module combines a dynamic queue mechanism to generate positive sample pairs through context and real features, and uses a dynamic queue to expand the diversity of negative samples to improve the model's discriminative ability and robustness. The contrastive learning optimization module maximizes the similarity of positive sample pairs and minimizes the similarity of negative sample pairs based on the noise contrast estimation (InfoNCE) loss function, thereby optimizing the quality of trajectory feature representation. Through the collaborative work of these modules, the model achieves efficient feature extraction and training optimization of trajectory data.
[0170] Specifically, in step S21, the multi-level trajectory feature encoder consists of a shallow feature extraction module and a deep feature extraction module. The shallow feature extraction module uses a five-layer one-dimensional convolutional network (1D-CNN) to extract the local spatial features of the trajectory and fix it to a low-dimensional potential feature representation of uniform dimension to provide standardized input. The deep feature extraction module models the global dependencies of trajectory points through the Transformer encoder, uses the attention mechanism to capture the deep spatiotemporal characteristics of the trajectory, and further enhances the global representation capability of the features.
[0171] Furthermore, the shallow feature extraction module adopts a five-layer one-dimensional convolutional network (1D-CNN) structure to gradually extract the multi-level features of the trajectory data and generate a high-dimensional feature representation containing spatiotemporal information. This module is composed of a one-dimensional convolution layer Conv1×1, a batch normalization layer BN (BatchNorm1d) and a ReLU activation function stacked in sequence to convert the processed trajectory data into Mapping to low-dimensional latent feature representation .in, Represents the input feature dimension of the original trajectory data, which usually includes multiple feature dimensions such as longitude, latitude, and timestamp of the trajectory point. Through the multi-layer convolution structure, the encoder extracts features of different granularities at each layer, thereby capturing the multi-scale spatiotemporal characteristics of the trajectory. Assume that the original input of the shallow feature extraction module is , the convolution operation of each layer can be expressed as:
[0172] (19);
[0173] In formula (19), represents the number of convolutional layers, Indicates The output features of the layer.
[0174] Finally, the overall effect of the encoder can be expressed as:
[0175] (20);
[0176] In formula (20), represents the low-dimensional latent feature representation, represents the trajectory data in the trajectory data target set, A feature encoder representing trajectory data.
[0177] In the design of the convolutional network, the initial layer uses a larger convolution kernel (size is 10, stride is 2, padding is 4) to capture the global features of long time spans in the trajectory data. As the network goes deeper, the size of the convolution kernel gradually decreases. The convolution kernel size of the second layer is 8, stride is 2, and padding is 3, while the convolution kernels of the subsequent three layers are 4, stride is 2, and padding is 1, gradually focusing on the fine-grained change characteristics of the trajectory. Among them, the convolution kernel sizes of each layer are 10, 8, 4, 4, and 4, respectively, forming a top-down feature extraction path. Through this feature extraction strategy from global to local, the shallow feature extraction module can fully capture the multi-scale features of the trajectory data. Figure 3 , Figure 3Conv1×1 S-10 in the example indicates the use of one-dimensional convolution, where the convolution kernel size is 10 and S represents the size of the convolution kernel. In addition, the number of output channels of each convolution layer is kept at 512 to ensure sufficient expressiveness and flexibility when encoding complex features of trajectories. Batch normalization (BN) and ReLU activation functions are used after each convolution layer to accelerate the training process, improve convergence speed, and effectively alleviate the problems of gradient vanishing and gradient exploding, thereby enhancing the stability and generalization ability of the network.
[0178] To ensure the stability and efficiency of network training, the convolutional and linear layer weights of the shallow feature extraction module are initialized using the He method. He initialization dynamically adjusts the weight distribution according to the number of input features, which enhances the network's ability to capture trajectory features while significantly reducing the risk of gradient vanishing. The output of the shallow feature extraction module is a three-dimensional tensor with a shape of [batch_size, feature_dim, sequence_length]. Among them, batch_size=8 means that the current batch contains 8 trajectories, feature_dim=512 indicates the feature dimension of each time step, and sequence_length=32 indicates that the trajectory consists of 32 time steps after shallow feature extraction. The original trajectory with a length of 1024 is gradually downsampled to 32 through the convolution operation, which is achieved through the step size design of multi-layer one-dimensional convolution.
[0179] In order to further capture the global dependencies and deep spatiotemporal characteristics of trajectory data, the low-dimensional latent feature representation generated by the shallow feature extraction module needs to be reorganized and input into the Transformer encoder. The structure of the Transformer encoder is shown in the figure. Figure 4As shown in the figure. First, the three-dimensional tensor is flattened into a two-dimensional tensor with a shape of [8, 16384], where 16384=512×32. The flattening operation converts the trajectory features from a two-dimensional sequence (feature dimension and time step) to a one-dimensional sequence, which facilitates the subsequent partitioning operation. The flattened features are then divided into N=64 equal-sized segments. Each segment after the partition represents a subsequence with a feature dimension of M=256, that is, the original sequence is reorganized at a finer granularity. After the partition, the feature shape of a single trajectory becomes [64, 256], while the trajectory data of the entire batch is represented as [8, 64, 256]. This partitioning method can better preserve local information while providing structured input for capturing global dependencies. To improve the model's ability to express trajectory features, each segment is mapped to a fixed dimension D=1024 through a linear projection layer. After projection, the shape of the trajectory segment becomes [8, 64, 1024]. This fixed-dimensional design can provide a unified high-dimensional input for the Transformer encoder and enhance the representation ability of features. Subsequently, in order to preserve the temporal information of the trajectory, the deep feature extraction module adds a one-dimensional position encoding to the trajectory segment. Since the trajectory feature sequence after division and projection has lost the direct association between the original time steps, the position encoding The introduction of can explicitly provide position information for each segment, allowing the encoder to better model the spatiotemporal dependencies of the trajectory. The final input sequence is calculated by the following formula:
[0180] (twenty one);
[0181] In formula (21), represents the input sequence of the Transformer encoder, represents the trajectory feature subsequence after linear projection, represents the positional encoding, i.e., a learnable parameter matrix consistent with the length of the subsequence, which is used to embed the position information of the fragment.
[0182] After combining with position encoding, the shape of the trajectory feature sequence becomes [8, 64, 1024], which provides temporal context information for the subsequent processing of the Transformer encoder. The Transformer encoder is composed of L stacked layers, each of which contains a multi-head self-attention (MSA) module and a multi-layer perceptron (MLP) module, aiming to capture the global dependencies between trajectory subsequences. First, the input sequence calculates the correlation and attention weights between trajectory segments through the multi-head self-attention mechanism to model the global dependencies:
[0183] (twenty two);
[0184] In formula (22), Represents the Transformer encoder The intermediate feature representation of the layer after multi-head attention calculation, Indicates -1 layer output, LN(·) is the layer normalization operation to ensure the stability of the input distribution.
[0185] Subsequently, the multi-layer perceptron module further extracts nonlinear features from the attention output and enhances the feature expression ability and network convergence through residual connections. The specific calculation is:
[0186] (twenty three);
[0187] In formula (23), Represents the Transformer encoder The final output feature representation of the layer, that is, the high-dimensional feature representation ; Among them, MLP It is a two-layer fully connected network that includes activation functions (such as GELU) to enhance the nonlinear expression of features. After passing through the L-layer Transformer encoder, the shape of the trajectory features remains [8, 64, 1024]. This process effectively captures the long-distance dependencies between trajectory subsequences and extracts deeper trajectory dynamic information by alternately applying multi-head self-attention and perceptron modules.
[0188] Finally, to adapt to the subsequent time series modeling, the output sequence of the Transformer encoder is reshaped into the structure of the original trajectory. The specific operation is to splice the N=64 divided trajectory segments into a single sequence and reconstruct it into the shape [512, 32]. Subsequently, the final [8, 512, 32] is restored through the transposition operation, so that the spatial dimension and time series length of the trajectory data are consistent with the output of the shallow feature extraction module, while retaining the global spatiotemporal characteristics after deep feature extraction. After the trajectory data is extracted by the multi-level trajectory feature encoder to extract high-dimensional feature representation, the autoregressive model is responsible for further modeling the temporal dependency and dynamic change relationship of the trajectory. The three-dimensional tensor [8, 512, 32] output by the multi-level trajectory feature encoder represents the high-dimensional spatiotemporal features of each trajectory, where [512, 32] are the feature dimension and time series length, respectively. The autoregressive model takes the features extracted by the encoder as input, generates contextual representation based on the dependency of the time series, and provides global information support for subsequent tasks.
[0189] In step S22, the autoregressive model uses a gated recurrent unit (GRU) as a core component and inputs a high-dimensional feature representation extracted by a multi-level trajectory feature encoder. GRU effectively captures the long-range dependencies and dynamic change characteristics in the time series through its gating mechanism (update gate and reset gate), thereby modeling the temporal information of the trajectory. Specifically, the input dimension of GRU is set to be consistent with the encoder output feature dimension, that is, =512, the dimension of the hidden layer is set to , to enhance the modeling ability of trajectory dynamic characteristics. The process of GRU calculating the hidden state at each time step is as follows:
[0190] =GUR( )(twenty four);
[0191] In formula (24), Indicates the hidden state at the current moment, represents the hidden state at the previous time step, Represents the input features of the current time step, that is, the high-dimensional feature representation output by the multi-level trajectory feature encoder.
[0192] At each time step, the output hidden state of the GRU is used as the context representation ,Right now:
[0193] (25).
[0194] This contextual representation combines the historical information in the trajectory time series with the dynamic characteristics of the current moment, providing complete and high-quality temporal context information for subsequent prediction tasks (such as contrastive learning optimization or classification tasks). By integrating GRU as an autoregressive model module, the model can effectively capture the temporal dependencies of the trajectory, complementing the aforementioned multi-level trajectory feature encoder, further improving the dynamic feature modeling capabilities of trajectory data, and laying the foundation for improving the overall model performance.
[0195] Based on the above context representation , then construct positive and negative samples. In the autoregressive modeling of trajectory features, the context representation It is a dynamic representation learned from the input trajectory by the GRU module. In order to optimize the quality of feature representation, it is necessary to combine the contrastive learning strategy to maximize the similarity of positive samples by constructing positive and negative sample pairs, while minimizing the similarity of negative samples. Based on this goal, this method introduces a dynamic queue mechanism in the process of constructing positive and negative samples to enhance the diversity and coverage of negative samples, thereby further improving the effect of contrastive learning. Figure 5 As shown in the figure, the process of constructing positive and negative sample pairs after the dynamic queue mechanism is introduced. Generated by the GRU module based on the global dynamic information of the previous time step, the context representation The integration of trajectory dynamic information before the current time step is the key to predicting future features. The feature representation of the autoregressive model is achieved through a linear projection layer Representing the context Mapping back to feature space to generate predicted feature representation :
[0196] (26).
[0197] The construction of positive samples revolves around each time step Contextual representation of Expand. The positive sample consists of With the real feature representation The pairing structure of positive samples aims to capture the dependency between context and future features, providing supervision signals for the model to enhance its temporal dependency modeling capabilities. The construction of negative samples adopts a dynamic queue mechanism to expand the diversity and number of negative samples. The dynamic queue is a fixed-size storage structure ( ), which is used to save feature representations from previous batches. In each training iteration, the feature representations generated by the current batch are dynamically added to the queue, and the earliest feature representations are removed, implementing a "first in, first out" update strategy. For each time step , negative samples are represented by the predicted features With queue Feature representation in ( ) pairing. Among them, u represents the queue The sample index stored in , that is, the sample number selected from the historical data, satisfies , ensuring that the negative samples entered are consistent with the current prediction features Do not belong to the same positive sample pair. This mechanism ensures that negative samples come not only from the current batch, but also from historical batches, thereby increasing the coverage and diversity of negative samples. Through this positive and negative sample construction strategy, the model can capture real feature information through positive sample pairs at each time step, while distinguishing irrelevant features through negative sample pairs generated by dynamic queues, laying the foundation for efficient contrastive learning. This process ensures that the contextual representation can effectively connect to the real features of future time steps, laying the foundation for the construction of positive and negative sample pairs and the optimization of contrastive learning.
[0198] In the contrastive learning optimization stage, the core goal of the entire model is to improve the correlation between trajectory context representation and future feature representation by comparing the similarity of positive and negative sample pairs. By maximizing the similarity of positive sample pairs and minimizing the similarity of negative sample pairs, the entire model can effectively capture the temporal dependency and global dynamic characteristics of the trajectory. That is, in the similarity measurement stage, the positive sample pairs Similarity and negative sample pairs The similarity is calculated by inner product:
[0199] =exp (27);
[0200] =exp (28);
[0201] In formula (27)-(28): Indicates the similarity of positive sample pairs; Represents the time step Trajectory data at represents the transposed vector of the predicted feature representation; Represents the time step The true feature representation of Represents the similarity of negative sample pairs; Represents the trajectory data input corresponding to the negative sample; represents the negative sample feature representation.
[0202] The similarity metric of the positive sample pairs reflects the degree of match between the predicted features and the real future features. The model optimizes the ability of the context representation to capture the dynamics of future trajectories by maximizing the similarity of the positive sample pairs. The model uses a dynamic queue mechanism to generate negative sample feature representations. ,The dynamic queue significantly increases the number and diversity of negative samples by introducing ,data from historical batches, and overcomes the limitation that the similarity of negative samples in the same batch is too ,high.
[0203] In step S23, the multi-level trajectory feature encoder is optimized and trained based on the noise contrast estimation loss function. To achieve the contrast learning goal, the model uses the noise contrast estimation (InfoNCE) loss function as the optimization target. The updated loss function is defined as:
[0204] (29);
[0205] In formula (29), Represents the distribution of trajectory data Take the mathematical expectation, that is, calculate the mean value on the entire trajectory dataset to ensure that the loss function is applicable to all samples, the molecular part Represents the similarity score of the positive sample pair. The denominator is the sum of the positive sample pair score and the scores of all negative sample pairs in the dynamic queue.
[0206] By maximizing the scores of positive pairs and reducing the scores of negative pairs, the model can effectively learn contextual representations. and future features The correlation between the context representation and the future features is improved, thereby improving the quality of feature representation. At the same time, the dynamic queue mechanism expands the coverage of negative samples by introducing feature representations of historical batches, and ensures the diversity and consistency of negative samples through a first-in-first-out update strategy. This mechanism effectively improves the model's ability to distinguish irrelevant features, while enhancing the stability and robustness of the optimization process. Combined with the InfoNCE loss function, the model further optimizes the mutual information between context representation and future features, laying the foundation for high-quality representation of trajectory features.
[0207] In this embodiment, the training and verification of the model are specifically as follows: the model training is based on the PyTorch framework and is performed in the NVIDIAA100 GPU environment. During the training process, a batch size of 8 is used, the trajectory window size is set to 1024, the time step is 12, and negative sample pairs are generated through a dynamic queue mechanism. The optimization goal is to maximize the similarity of positive samples and minimize the similarity of negative samples. The optimizer uses dynamic Adam, the initial learning rate is 0.001, the first 5 epochs use a linear growth strategy, and then the learning rate is decayed every 5 epochs. The total number of training rounds is set to 50 rounds, and an early stopping mechanism is introduced to terminate the training in advance when the validation set loss does not decrease significantly within 5 consecutive epochs. The model automatically adjusts parameters during the training process to ensure the accuracy and stability of feature representation. In the verification stage, the model performance is evaluated through an independent verification set, mainly using accuracy and loss values as measurement indicators.
[0208] Based on the above, this method proposes a new contrastive learning framework, which constructs a multi-layer convolutional network and Transformer encoder through a deep neural network to extract high-dimensional feature representations of trajectory data, and combines the gated recurrent unit (GRU) to model the time series dependencies of the trajectory. In addition, the framework innovatively introduces a dynamic queue mechanism to dynamically expand the number and diversity of negative samples, significantly improving the discriminative ability of feature representation and the classification performance of the model.
[0209] S3. Use the trained multi-level trajectory feature encoder to extract the high-dimensional feature representation of the trajectory data, and use it as input to train the MLP classifier so that it can learn the mapping relationship between trajectory features and traffic modes. After the training is completed, the high-dimensional feature representation of the trajectory data is input into the MLP classifier to classify the trajectory traffic mode.
[0210] After the model training is completed, the high-dimensional feature representation extracted by the encoder is used for the traffic mode classification task. After being processed by the encoder, each trajectory is converted into a fixed-dimensional feature representation, which contains rich dynamic information and global spatiotemporal dependencies, providing a solid foundation for the classification task.
[0211] To complete the classification task, this method designs a multi-layer perceptron (MLP) classifier to map high-dimensional features to target categories. The classifier consists of an input layer, two hidden layers, and an output layer. The size of the input layer is consistent with the output dimension of the encoder and is used to receive high-dimensional feature representations. The first hidden layer contains 256 neurons and the second hidden layer contains 128 neurons. Both layers use the ReLU activation function to enhance the nonlinear modeling capability, and Dropout (the dropout rate is set to 0.5) is added after each layer to prevent overfitting. The output layer maps the output of the hidden layer to the number of target categories through a fully connected layer, and directly generates classification scores (logits) for subsequent loss calculations. Figure 6 shown.
[0212] The classifier is trained with the cross entropy loss function as the target, which is used to measure the difference between the predicted value and the true label. Assume that the predicted output of the classifier is , the true label is ,in is the number of categories, and the loss function is defined as:
[0213] (30);
[0214] In formula (30), represents the cross entropy loss, batch size Indicates the number of samples and categories contained in a training batch Represents the total number of target categories in the classification task; the actual category index The value range is , used to traverse all categories; indicator function 1=( =c) When the sample The true category is The value is 1 when the value is 1, otherwise it is 0; prediction score (logits) Representative samples In category The output value on , not normalized by softmax; category index Used for the summation calculation of the softmax normalized denominator, the value range is , corresponding to the sample Prediction score on category J , used to calculate the normalized probability.
[0215] The optimizer uses Adam, the learning rate is set to 0.001, and the weight decay value is set to 0.0001 to regularize the model. During the training process, the model optimizes the parameters through small batch stochastic gradient descent, and introduces an early stopping mechanism: when the loss in the verification process does not decrease for multiple consecutive rounds, the training is automatically terminated to prevent overfitting. The training process of the classifier is set to 50 training rounds. After the training is completed, the model performance is evaluated to measure its ability to classify traffic patterns. The evaluation indicators include classification accuracy and loss value. The entire training process is monitored by performance to ensure that the classifier has a high generalization ability when processing the complex spatiotemporal characteristics of trajectory data.
[0216] In summary, the present invention effectively improves the classification effect by first using a trained encoder to extract high-dimensional features of the trajectory and then training the MLP classifier. Compared with traditional methods such as support vector machine (SVM), recurrent neural network (RNN) and gated recurrent unit (GRU), this design has significant advantages in classification accuracy and efficiency. This method provides reliable technical support for the deep analysis of trajectory data and traffic mode classification tasks, reflecting its potential in complex scenarios.
[0217] Example 2
[0218] The present embodiment provides a trajectory traffic mode classification device, the device comprising: a data processing module, used to perform preliminary processing operations on raw trajectory data to construct a trajectory data target set; a model training module, used to input the trajectory data target set as training data into a multi-level trajectory feature encoder, and train the multi-level trajectory feature encoder using a contrastive learning method in combination with an autoregressive model; extracting a high-dimensional feature representation of the trajectory data using the trained multi-level trajectory feature encoder, and using the high-dimensional feature representation as input to train an MLP classifier so that the MLP classifier learns the mapping relationship between trajectory features and traffic modes; after the training is completed, inputting the high-dimensional feature representation of the trajectory data into the MLP classifier to perform trajectory traffic mode classification;
[0219] Among them, the model training module is specifically used to: input the trajectory data target set into the multi-level trajectory feature encoder, use the shallow feature extraction module of the multi-level trajectory feature encoder to extract the local spatial features of the trajectory and convert them into low-dimensional potential feature representations of unified dimensions, and use the deep feature extraction module of the multi-level trajectory feature encoder to obtain the global spatiotemporal dependency between trajectory points based on the low-dimensional potential feature representation to generate a high-dimensional feature representation; input the high-dimensional feature representation into the autoregressive model, the autoregressive model generates a context representation of each time step based on the dependency of the time series, constructs positive samples based on the context representation of each time step and the true feature representation, and introduces a dynamic queue mechanism to construct negative samples based on the context representation of each time step and the feature representation in the dynamic queue; use the constructed positive and negative samples, and based on the noise contrast estimation loss function, adopt a contrastive learning method to train the multi-level trajectory feature encoder.
[0220] Example 3
[0221] This embodiment provides an electronic device, including: at least one processor; and a memory, wherein the memory stores instructions, and when the instructions are executed by the at least one processor, the at least one processor executes the trajectory traffic mode classification method as described above.
[0222] Example 4
[0223] This embodiment also provides a computer-readable storage medium storing executable instructions, which, when executed, enable the machine to perform the trajectory traffic mode classification method as described above.
[0224] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the technical solution of the present invention, and are not intended to limit the specific implementation methods of the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the claims of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A trajectory traffic mode classification method, characterized in that: The method comprises: S1, performing preliminary processing operations on the original trajectory data to construct a trajectory data target set; S2, inputting the trajectory data target set as training data into a multi-level trajectory feature encoder, and training the multi-level trajectory feature encoder by contrastive learning method in combination with an autoregressive model; S3, using the trained multi-level trajectory feature encoder to extract high-dimensional feature representation of the trajectory data, and using it as input to train the MLP classifier so that it can learn the mapping relationship between trajectory features and traffic modes. After the training is completed, the high-dimensional feature representation of the trajectory data is input into the MLP classifier to classify the trajectory traffic mode; Wherein, the S2 specifically includes: S21, inputting the trajectory data target set into the multi-level trajectory feature encoder, using the shallow feature extraction module of the multi-level trajectory feature encoder to extract the local spatial features of the trajectory and convert them into low-dimensional potential feature representations of uniform dimensions, and using the deep feature extraction module of the multi-level trajectory feature encoder to obtain the global spatiotemporal dependency relationship between trajectory points based on the low-dimensional potential feature representation to generate a high-dimensional feature representation; S22, inputting the high-dimensional feature representation into an autoregressive model, wherein the autoregressive model generates a contextual representation of each time step based on the dependency of the time series, constructing a positive sample based on the contextual representation of each time step and the true feature representation, and introducing a dynamic queue mechanism to construct a negative sample based on the contextual representation of each time step and the feature representation in the dynamic queue; S23. Using the constructed positive samples and negative samples and based on the noise contrast estimation loss function, adopt a contrastive learning method to train the multi-level trajectory feature encoder.
2. The trajectory traffic mode classification method according to claim 1 is characterized in that: S1 specifically includes: preprocessing the original trajectory data to construct a trajectory data set , the preprocessing operation includes: smoothing the noise points in the original trajectory data by using a dynamic window weighted average method, and correcting the geographical coordinates of the drift points in the original trajectory data by using a weighted regression method based on trajectory trends; The method of smoothing the noise points in the original trajectory data by using a dynamic window weighted average method specifically includes: By calculating the speed of adjacent trajectory points and comparing it with the preset maximum speed threshold Compare, exceed the maximum speed threshold The trajectory point is judged as a noise point; the calculation formula of its speed is as follows: (1); In formula (1): represents the velocity of the trajectory point; d( ) represents the trajectory point and The geographical distance between , Represents the trajectory points and timestamp; Through the track point A dynamic window composed of several non-noise points before and after Weighted average, re-estimate its coordinate value; the smoothed coordinate calculation formula is: (2); In formula (2): , Represents the smoothed trajectory points The latitude and longitude coordinates of; Represents trajectory points A dynamic window containing several non-noise points before and after; , Represents dynamic windows The latitude and longitude coordinates of each valid track point in Represents a dynamic window The index of each valid trajectory point in , Represents a dynamic window The number of valid trajectory points within; The weighted regression method based on trajectory trend corrects the geographic coordinates of the drift point in the original trajectory data, specifically including: The geographical distance d( ) and dynamic threshold Compare and exceed the dynamic threshold The trajectory points of are judged as drift points, namely: (3); The corrected coordinates of the drift point are determined by fitting the trajectory trends of the trajectory points before and after the drift point. Assuming that the trajectory point is the drift point, and its normal trajectory points before and after are and , then: (4); (5); In formula (4) and (5): Represents the corrected trajectory point The latitude coordinate of Represents the corrected trajectory point The longitude coordinates of Indicates the previous normal trajectory point The latitude coordinate of Indicates the previous normal trajectory point The longitude coordinates of Represents the next normal trajectory point The latitude coordinate of Represents the next normal trajectory point The longitude coordinates of , Both represent weights, which are set according to the time interval or geographical distance between the drift point and the previous and next normal trajectory points; Weight , The calculation formula is: = = (6); In formula (6): d( ) represents the trajectory point The next normal trajectory point The geographical distance between ) represents the trajectory point The previous normal trajectory point The geographical distance between ) represents the previous normal trajectory point The next normal trajectory point The geographical distance between The trajectory data set Contains several tracks , each trajectory A series of trajectory points Composition, and trajectory points The core fields included are latitude and longitude coordinates and timestamp, namely: (7); (8); (9); In formulas (7)-(9): For all trajectories All satisfy the condition that the number of trajectory points is greater than or equal to the minimum threshold requirements; Represents trajectory points The latitude coordinate of Represents trajectory points The longitude coordinates of Represents trajectory points The timestamp of the 3. The trajectory traffic mode classification method according to claim 2 is characterized in that: The S1 specifically includes: The trajectory data in is subjected to standardization operation, data enhancement operation, shallow feature extraction operation and transposition operation in turn to obtain the trajectory data target set; Among them, the trajectory data set The trajectory data in is standardized, specifically: The trajectory length of the trajectory data is uniformly adjusted to 1024 points, that is, for trajectories with a length greater than 1024 points, they are reduced to 1024 points by random truncation, and for trajectories with a length less than 1024 points, they are supplemented based on a sliding window linear interpolation method in a data enhancement operation; The data enhancement operation is performed on the standardized trajectory data based on the sliding window linear interpolation method, specifically: Set the target track points As the length of the sampling window, and calculate the number of current trajectory points ; According to the target trajectory points With the current track points The difference between ,Right now: (10); When the number of trajectory points to be inserted No interpolation is required when . When interpolation is required, the sliding window method is used to traverse the trajectory points, with the window size set to 4 and the moving step size set to 3. For each group of four adjacent trajectory points , and , calculate the first trajectory point The change in orientation between the other three trajectory points and the average value , and its calculation formula is: (11); If the average value of the azimuth change Less than the preset threshold , it is considered that the trajectory characteristics of the area formed by the above four adjacent trajectory points change little and meet the interpolation conditions; When the interpolation conditions are met, at the trajectory point and Insert new track points between , whose coordinates are calculated by linear interpolation: The insertion of a new trajectory point increases the number of trajectory points by 1; Then, the sliding window continues to move forward, calculates the position change of the next set of adjacent trajectory points, and determines whether to continue interpolation; when a round of interpolation is completed, if the number of trajectory points has not reached the target number of trajectory points, , a new round of interpolation starts from the beginning, and the number of points inserted in each round is gradually reduced until the number of trajectory points reaches the target trajectory point number. ; The shallow feature extraction operation is performed on the enhanced trajectory data. The extracted shallow features include speed , speed change ,position , Direction Change , angular velocity and curvature ,in: speed The calculation formula is: = (13); In formula (13): Represents trajectory points The speed of movement; d( ) represents the trajectory point and The geographical distance between and Represents trajectory points and timestamp; Speed Change for: (14); In formula (14): Represents trajectory points and The speed change between position The calculation formula is: (15); In formula (15): Represents trajectory points azimuth; , Represents trajectory points The latitude and longitude coordinates of; Change of direction The calculation formula is: (16); In formula (16): Represents trajectory points directional changes; Represents trajectory points azimuth; Angular velocity The calculation formula is: (17); In formula (17): Represents trajectory points Angular velocity; Curvature The calculation formula is: (18); In formula (18): Represents trajectory points The curvature of The transposition operation is specifically as follows: Transpose the original data of shape (trajectory_window, num_features) to form target data of shape (num_features, trajectory_window), where trajectory_window represents the number of trajectory points for each trajectory and num_features represents the number of initial features; After the transposition is completed, the trajectory data is converted into a PyTorch tensor, which is finally constructed into a trajectory data target collection.
4. The trajectory traffic mode classification method according to claim 3 is characterized in that: In S21, the shallow feature extraction module of the multi-level trajectory feature encoder adopts a five-layer one-dimensional convolutional network structure, which is composed of a one-dimensional convolutional layer Conv1×1, a batch normalization layer BN and a ReLU activation function stacked in sequence, for gradually extracting multi-level features of trajectory data and generating a high-dimensional feature representation containing spatiotemporal information; Assume that the original input of the shallow feature extraction module is , the convolution operation of each layer is expressed as: (19); In formula (19): Indicates the number of convolutional layers; Indicates Output features of the layer; The overall effect of the shallow feature extraction module is expressed as: (20); In formula (20): represents a low-dimensional latent feature representation; Represents the trajectory data in the trajectory data target set; A feature encoder representing trajectory data; The output of the shallow feature extraction module is a three-dimensional tensor with a shape of [batch_size, feature_dim, sequence_length], where batch_size=8 means that the current batch contains 8 trajectories, feature_dim=512 means the feature dimension of each time step, and sequence_length=32 means that the trajectory consists of 32 time steps after shallow feature extraction; The deep feature extraction module of the multi-level trajectory feature encoder adopts a Transformer encoder, which is stacked by L layers, each layer including a multi-head self-attention module and a multi-layer perceptron module, for capturing the global dependency between trajectory subsequences; First, the low-dimensional latent feature representation generated by the shallow feature extraction module is Reorganize, specifically: Flattening the three-dimensional tensor output by the shallow feature extraction module into a two-dimensional tensor, converting the trajectory features from a two-dimensional sequence containing feature dimensions and time steps into a one-dimensional sequence; The flattened trajectory features are divided into N segments of equal size, each of which is mapped to a fixed dimension D through a linear projection layer to provide a unified high-dimensional input for the Transformer encoder; Add a one-dimensional positional encoding to each segment , the final input sequence of the Transformer encoder is: (21); In formula (21): represents the input sequence of the Transformer encoder, represents the trajectory feature subsequence after linear projection, represents the positional encoding, i.e., a learnable parameter matrix consistent with the length of the subsequence, which is used to embed the position information of the fragment; The input sequence The correlation and attention weights between trajectory segments are calculated through a multi-head self-attention module to model global dependencies; The calculation process is as follows: (22); In formula (22): Represents the Transformer encoder The intermediate feature representation of the layer after multi-head attention calculation; Indicates -1 layer output; LN(·) is the layer normalization operation to ensure the stability of the input distribution; The multi-layer perceptron module further extracts nonlinear features from the attention output and enhances the feature expression ability and network convergence through residual connections; the specific calculation is: (23); In formula (23): Represents the Transformer encoder The final output feature representation of the layer, that is, the high-dimensional feature representation ;MLP It is a two-layer fully connected network that includes activation functions to enhance the nonlinear expression of features.
5. The trajectory traffic mode classification method according to claim 4 is characterized in that: In S22, the autoregressive model uses a gated recurrent unit GRU as a core component, and its input is a high-dimensional latent space feature representation extracted by the multi-level trajectory feature encoder. ; The process of calculating the hidden state of the autoregressive model at each time step is as follows: =GUR( )(24); In formula (24): Indicates the hidden state at the current moment, represents the dimension of the hidden layer; represents the hidden state of the previous time step; Represents the input features of the current time step, that is, the high-dimensional feature representation output by the multi-level trajectory feature encoder; At each time step, the output hidden state of the autoregressive model is used as the context representation ,Right now: (25); Through the linear projection layer The context of each time step is represented as Mapped to feature space to get predicted feature representation ,Right now: (26); Based on the predicted feature representation and the true feature representation Construct positive samples; Introducing dynamic queues , based on the predicted feature representation and dynamic queues Feature Representation Construct negative samples, where , u represents the queue The index of the samples stored in .
6. The trajectory traffic mode classification method according to claim 5, characterized in that: In S23, the multi-level trajectory feature encoder is optimized and trained based on the noise contrast estimation loss function, and the optimization goal is to maximize the similarity of positive samples and minimize the similarity of negative samples, wherein: In the similarity measurement stage, the similarity measurement refers to the model predicting feature representation by calculating and the true feature representation The similarity between them is used to measure the matching degree between the predicted results and the actual results; Positive sample pair The similarity of is calculated by inner product as: =exp (27); In formula (27): Indicates the similarity of positive sample pairs; Represents the time step Trajectory data at represents the transposed vector of the predicted feature representation; Represents the time step The true feature representation of Negative sample pairs The similarity of is calculated by inner product as: =exp (28); In formula (28): Represents the similarity of negative sample pairs; Represents the trajectory data input corresponding to the negative sample; Represents the negative sample feature representation; The noise contrast estimation loss function is defined as: (29); In formula (29): represents the noise contrast estimation loss function, Represents the distribution of trajectory data Take the mathematical expectation.
7. The trajectory traffic mode classification method according to claim 1, characterized in that: In S3, the MLP classifier includes an input layer, two hidden layers and an output layer, which uses the cross entropy loss function as a training target to measure the difference between the predicted value and the true label; assuming that the predicted output of the MLP classifier is , the true label is , Represents the number of categories, then the loss function is defined as: (30); In formula (30): represents the cross entropy loss; batch size Indicates the number of samples contained in a training batch; the number of categories Represents the total number of target categories in the classification task; True Category Index The value range is , used to traverse all categories; indicator function 1=( =c) means when the sample The true category is The value is 1 when the value is 1, otherwise it is 0; prediction score (logits) Representation sample In category The output value on is not normalized by the softmax layer; The category index J is used for the summation of the normalized denominator of the softmax layer, and its value range is , corresponding to the sample Prediction score on category J , used to calculate the normalized probability.
8. A trajectory traffic mode classification device, characterized in that: The device comprises: The data processing module is used to perform preliminary processing operations on the original trajectory data to construct a trajectory data target set; A model training module, used to input the trajectory data target set as training data into a multi-level trajectory feature encoder, and train the multi-level trajectory feature encoder by using a contrastive learning method in combination with an autoregressive model; A trajectory traffic mode classification module is used to extract high-dimensional feature representation of trajectory data using the trained multi-level trajectory feature encoder, train an MLP classifier, and input the high-dimensional feature representation into the trained MLP classifier to perform trajectory traffic mode classification; Wherein, the model training module is specifically used for: Inputting the trajectory data target set into the multi-level trajectory feature encoder, using the shallow feature extraction module of the multi-level trajectory feature encoder to extract the local spatial features of the trajectory and convert them into low-dimensional latent feature representations of uniform dimensions, and using the deep feature extraction module of the multi-level trajectory feature encoder to obtain the global spatiotemporal dependency relationship between trajectory points based on the low-dimensional latent feature representation to generate a high-dimensional feature representation; Inputting the high-dimensional feature representation into an autoregressive model, the autoregressive model generates a contextual representation of each time step based on the dependency of the time series, constructing a positive sample based on the contextual representation of each time step and the true feature representation, and introducing a dynamic queue mechanism to construct a negative sample based on the contextual representation of each time step and the feature representation in the dynamic queue; The constructed positive samples and negative samples are used, and based on the noise contrast estimation loss function, a contrastive learning method is adopted to train the multi-level trajectory feature encoder.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and A memory storing instructions, which, when executed by the at least one processor, enable the at least one processor to perform the trajectory traffic mode classification method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The machine stores executable instructions, which, when executed, enable the machine to perform the trajectory traffic mode classification method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Strip mine truck scheduling method based on road network-track combined comparative learning
CN116629531A
Intelligent automobile multi-mode non-autoregressive trajectory prediction method and device and medium
CN119293749A