Dual-branch transformer destination prediction method based on multi-modal data fusion

Through the dual-branch transformer method of multimodal data fusion, the trajectory sequence and image features are extracted and fused, and the problem of insufficient extraction and fusion of multimodal feature in the prior art is solved, which significantly improves the accuracy and reliability of destination prediction.

CN120030432APending Publication Date: 2025-05-23HUNAN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411862805.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing destination prediction methods lack technical means to simultaneously extract and effectively fuse multimodal features of trajectory sequences and trajectory images, resulting in limitations, insufficient accuracy and reliability when dealing with long-distance dependence problems.

Method used

The dual-branch transformer method using multimodal data fusion is used to extract the trajectory sequence and image features through quad-tree embedding and trajectory imageization methods, and the feature fusion is performed using the cross attention mechanism, and finally the destination prediction is performed through the Softmax classifier.

Benefits of technology

It significantly improves the accuracy and reliability of destination prediction, can dig deeper into trajectory feature information, and has a stronger ability to deal with long-distance dependence problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030432A_ABST
    Figure CN120030432A_ABST
Patent Text Reader

Abstract

The invention discloses a dual-branch transformer destination prediction method based on multi-modal data fusion. The method comprises the following steps of S1, data preprocessing, S2, candidate destination construction, S3, dual-branch extraction of multi-modal features, S4, feature fusion and S5, destination prediction. According to the method, based on two kinds of data of a vehicle track sequence and a track image, track sequence features and image features are extracted through a transformer and a visual transformer, namely, a Transform model and a Vision Transform model; secondly, proposing a middle-term fusion method based on a cross attention mechanism, obtaining fusion features, obtaining probability distribution of each destination according to the fusion features, and ranking the probability distribution; then, Top-K destinations are constructed as a destination prediction result, according to the multi-mode double-branch destination prediction method, sequence feature information and image feature information in trajectory data are extracted, feature fusion is carried out through a fusion module, and therefore the destinations are predicted more accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent transportation system and artificial intelligence technology, and relates to a variety of vehicle trajectory data processing and deep learning methods, especially transformers, i.e. Transformer models, vision transformers, i.e. Vision Transformer, quadtree embedding methods, trajectory visualization methods and cross-attention methods. Background Art

[0002] With the popularization and development of mobile devices, communication technology, and the Internet, massive amounts of mobile GPS trajectory data are generated every day. These data contain the behavioral patterns of mobile objects, including: personal behavioral habits, group trends, etc. Mining these trajectory data for destination prediction can provide strong support for autonomous driving, intelligent transportation, vehicle scheduling, and urban planning. Current destination prediction methods are mainly divided into two categories: (1) converting trajectory data into time series data and then using time series analysis models to extract its time features to predict the destination; (2) converting trajectory data into trajectory images and then using computer vision technology to extract its time features to predict the destination. However, there is currently a lack of research on combining the features extracted by these two methods for destination prediction.

[0003] The simultaneous extraction of multimodal features from trajectory sequences and trajectory images will face complex long-distance dependency problems. Existing temporal feature extraction models, including RNN and image feature extraction models, including CNN, have certain limitations in dealing with long-distance dependency problems. The Transformer model calculates the weight of each position in the input data through the self-attention mechanism, allowing the model to focus on the key information in the input data, thereby effectively handling long-distance dependencies. In addition, the self-attention mechanism can be calculated in parallel, which improves the training efficiency of the model.

[0004] Existing destination prediction methods only consider the temporal information of the trajectory sequence and the spatial features of the trajectory image separately when extracting feature information from the trajectory, and lack methods to extract the features of these two modalities at the same time. In addition, there is a lack of technical means to effectively fuse the feature information of these two modalities. Cross-attention is a strategy for multimodal feature fusion, which can establish an attention mechanism between features of different modalities to capture the interaction between modalities. Specifically, the cross-attention mechanism imposes the attention weight of one modality on the features of another modality, allowing the model to focus on the key information between different modalities. In this way, the cross-attention mechanism can effectively fuse multimodal features and improve the accuracy and reliability of destination prediction. Summary of the invention

[0005] The present invention aims to provide a method for extracting two modal information, trajectory sequence information and image spatial feature information, from trajectory data, and performing multimodal fusion to predict the destination. The method mainly includes two aspects: (1) extracting two modal information, trajectory sequence information and image spatial feature information; and (2) performing multimodal fusion to predict the destination.

[0006] To achieve the above object, the present invention provides the following technical solution: a dual-branch transformer destination prediction method based on multimodal data fusion, the method comprising the following steps:

[0007] S1. Data preprocessing: collect real trajectory data, perform data cleaning and standardization to ensure data quality and consistency;

[0008] S2, construction of candidate destinations: extract all the alighting locations from the trajectory data, perform cluster analysis on these locations, group similar alighting locations, and form a representative candidate destination set CPD;

[0009] S3. Dual-branch extraction of multimodal features:

[0010] 3.1. In the extraction of sequence feature branches of trajectories, the trajectory data is first converted into a quadtree embedding sequence using the quadtree embedding method;

[0011] 3.2. In the extraction of image feature branches of trajectories, the trajectory data is first converted into trajectory images using the trajectory imaging method;

[0012] S4, feature fusion: construct a cross-attention based trajectory multimodal feature fuser to fuse the previously extracted sequence features and image features to capture the mutual relationship and complementary information between the two modalities;

[0013] S5. Predict destination: Through a Softmax classifier, the probability of the candidate destination is output based on the fusion of multimodal feature information, and the candidate destinations ranked top k in probability are returned as the predicted destinations.

[0014] Preferably, in step S3, a transformer model, ie, an encoder part of a transformer, is used to extract sequence information of a quadtree embedding sequence.

[0015] Preferably, in step S3, a visual transformer model, namely, Vision Transformer (VIT), is used to extract image feature information from the trajectory image.

[0016] Compared with the prior art, the present invention has the following beneficial effects:

[0017] 1. Based on two types of data, vehicle trajectory sequence and trajectory image, the present invention extracts trajectory sequence features and image features through two models, namely, Transformer and Vision Transformer, respectively; secondly, a mid-term fusion method based on the cross-attention mechanism is proposed to obtain fusion features, and then the probability distribution of each destination is obtained according to the fusion features and ranked; then, the Top-K destinations are constructed as destination prediction results;

[0018] 2. The multimodal dual-branch destination prediction method extracts sequence feature information and image feature information from trajectory data and fuses features through a fusion module to more accurately predict the destination. Compared with traditional methods, this model can more deeply explore the feature information of the trajectory and significantly improve the accuracy and reliability of the prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a flow chart of the destination prediction model of the present invention;

[0020] Figure 2 This is a diagram of the cross-attention feature fusion module of the present invention. DETAILED DESCRIPTION

[0021] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0022] A dual-branch transformer destination prediction method based on multimodal data fusion, the method comprising the following steps:

[0023] S1. Data preprocessing:

[0024] 1.1. Collect historical trajectory data of vehicles from the urban traffic monitoring system. The data should include the travel coordinates and timestamp of each vehicle;

[0025] 1.2. Clean the collected data to eliminate outliers, including wrong coordinates and unreasonable timestamps;

[0026] 1.3. Standardize the data so that different features are on the same scale to facilitate model training;

[0027] S2. Construction of candidate destinations:

[0028] 2.1 Traverse the trajectory data and extract the end point coordinates from the itinerary of each trajectory to form a set of end point positions;

[0029] 2.2 Use the Mean-Shift clustering algorithm to cluster similar destinations in the destination set to form a candidate destination set CPD. The formula of Mean-Shift clustering is as follows:

[0030]

[0031] Among them, x is the center point; x i is a point within the bandwidth; n is the number of trajectory destination points within the bandwidth, G(x) is a unit kernel function, W(x i ) represents the weight of each sample;

[0032] S3. Dual-branch extraction of multimodal features:

[0033] 3.1. Trajectory Sequence Feature Branch Extraction

[0034] 3.1.1 First, the entire urban area is divided into multiple levels of grids. On the basis of the hierarchical grids, a quadtree is constructed. Each node of the quadtree represents a grid area, and each leaf node represents the smallest grid unit. Each grid area is recursively divided into four sub-areas to construct a complete quadtree. For each point on each trajectory, the leaf node of the quadtree is determined and the unique identifier of the node is recorded. Then, one-hot encoding technology is used to convert each node identifier into a high-dimensional vector. The one-hot encoding of each trajectory point is concatenated in chronological order of the trajectory to form a quadtree embedding sequence of a trajectory. In this way, each trajectory is converted into a sequence composed of one-hot encodings.

[0035] 3.1.2 A sequence feature extractor based on a transformer model encoder is proposed. The extractor first divides the quadtree embedding sequence into the input parts of the transformer model and performs position encoding on each input part. Then, the trajectory sequence features in the quadtree embedding sequence are extracted through the transformer model encoder. The encoder uses a multi-head attention mechanism, a feedforward neural network, residual connections, and layer normalization to extract high-level trajectory sequence feature information S from the quadtree embedding sequence.

[0036] 3.2. Trajectory Image Feature Branch Extraction

[0037] 3.2.1 First, the entire urban area is divided into hierarchical grids, and the grid size and number are consistent with step 3.1.1; then each track point in each trajectory is mapped to the corresponding grid, the grid value of the track point without mapping is 0, and the grid value of the track point mapping is set according to the grayscale spatiotemporal coding method. The formula of grayscale spatiotemporal coding is as follows:

[0038]

[0039] Among them, T is the trajectory data, l k is the kth trajectory point in T, l k →I (i,j) This is represented by k In the grid area corresponding to the subscript position i,j of the grid matrix I; after all the trajectory points are mapped to P, they are converted into trajectory images;

[0040] 3.2.2 An image feature extractor based on the visual transformer model, namely VisionTransformer, VIT, is proposed. The extractor first divides the trajectory image into image blocks of fixed size, and then uses a multi-layer transformer encoder to process the input image blocks. Each encoder layer contains a multi-head self-attention mechanism and a feedforward neural network; the multi-head self-attention mechanism can capture the long-distance dependencies between image blocks, and the feedforward neural network is used to further extract features. After being processed by the multi-layer encoder, the high-level image feature information I of the trajectory is finally output;

[0041] S4. Feature fusion:

[0042] 4.1. A multimodal fusion module based on a cross-attention network is proposed. Its basic working mechanism is to impose the attention weight of one modality on the features of another modality to capture the interaction between the two modalities. In step 3, the trajectory sequence feature information S has been obtained, which is a matrix of shape M, DS, where M is the number of sequence features and DS is the dimension of each feature; and the image feature information I is a matrix of shape N, DI, where N is the number of sequence features and DI is the dimension of each feature; then the attention weight A of the image feature on the sequence feature is calculated respectively. I→S And the attention weight A of sequence features on image features S→I , the formula is as follows:

[0043]

[0044] Where W Q and W K is the initialization weight matrix, used to initialize the query,key of the input vector; k is the intermediate feature dimension; then the attention weight is applied to calculate the weighted sequence features and weighted image features, the formula is as follows:

[0045] S'=A I→S IW V (5)

[0046] I′=A S→I SWV (6)

[0047] Where W Q An initialization weight matrix is ​​used to initialize the value of the input vector; after obtaining the weighted sequence feature S' and the weighted image feature I', they are combined to obtain the fusion feature F;

[0048] F = Concat(S′+I′) (7)

[0049] 4.2. Use a multi-layer perceptron, i.e., MLP network, to deeply fuse the fused feature information F to further enhance the feature expression capability;

[0050] S5. Predicted destination:

[0051] 5.1 Through a Softmax classifier, the probability of the candidate destination is output based on the fusion of multimodal feature information;

[0052] 5.2 Return the top k candidate destinations ranked by probability as the predicted destinations;

[0053] In specific operations, the model structure and parameters are customized and optimized according to actual application requirements and data characteristics.

[0054] In summary, based on two kinds of data, vehicle trajectory sequence and trajectory image, the present invention extracts trajectory sequence features and image features through two models, namely, Transformer and Vision Transformer, respectively; secondly, a mid-term fusion method based on the cross-attention mechanism is proposed to obtain fusion features, and then the probability distribution of each destination is obtained according to the fusion features and ranked; then, the Top-K destinations are constructed as destination prediction results.

[0055] This multimodal dual-branch destination prediction method extracts sequence feature information and image feature information from trajectory data and performs feature fusion through a fusion module to more accurately predict the destination. Compared with traditional methods, this model can more deeply explore the feature information of the trajectory and significantly improve the accuracy and reliability of the prediction.

[0056] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.

[0057] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A dual-branch transformer destination prediction method based on multimodal data fusion, characterized by: The method comprises the following steps: S1. Data preprocessing: collect real trajectory data, perform data cleaning and standardization to ensure data quality and consistency; S2, construction of candidate destinations: extract all the alighting locations from the trajectory data, perform cluster analysis on these locations, group similar alighting locations, and form a representative candidate destination set CPD; S3. Dual-branch extraction of multimodal features: 3.

1. In the extraction of sequence feature branches of trajectories, the trajectory data is first converted into a quadtree embedding sequence using the quadtree embedding method; 3.

2. In the extraction of image feature branches of trajectories, the trajectory data is first converted into trajectory images using the trajectory imaging method; S4. Feature fusion: Construct a cross-attention based trajectory multimodal feature fuser to fuse the previously extracted sequence features and image features to capture the mutual relationship and complementary information between the two modalities; S5. Predict destination: Through a Softmax classifier, the probability of the candidate destination is output based on the fusion of multimodal feature information, and the candidate destinations ranked top k in probability are returned as the predicted destinations.

2. The dual-branch transformer destination prediction method based on multimodal data fusion according to claim 1 is characterized in that: In step S3, the transformer model, i.e., the encoder part of the Transformer, is used to extract the sequence information of the quadtree embedding sequence.

3. The dual-branch transformer destination prediction method based on multimodal data fusion according to claim 1 is characterized in that: In step S3, a vision transformer model, namely, Vision Transformer (VIT), is used to extract image feature information from the trajectory image.