Pre-training large model traffic flow prediction method based on double-activation domain bridging and space-time self-attention

Through dual-activated domain bridge and space-time self-attention mechanism, the large language model is fine-tuned, which solves the problems of high computational complexity and domain gap in traffic flow prediction, and achieves more efficient and accurate traffic flow prediction, adapts to different computing resources and scenarios, and enhances the model's adaptability and prediction capabilities.

CN120260282APending Publication Date: 2025-07-04NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510501570.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art has problems in traffic flow prediction with high computational complexity, insufficient data adaptability, insufficient temporal and spatial characteristics fusion, and a domain gap between large language models and traffic flow data, resulting in insufficient prediction accuracy and low efficiency.

Method used

The dual-activated domain bridge module and space-time self-attention mechanism are designed, and the large language model is fine-tuned through multiple embedding, dual-activated domain bridge, LoRA strategy and partial attention freezing methods to enhance its adaptability and modeling ability to traffic flow data.

Benefits of technology

It improves the accuracy and robustness of traffic flow prediction, simplifies the model deployment process, enhances the ability to capture multi-grained spatiotemporal features, adapts to different computing resources and scenarios, and has good scalability and versatility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260282A_ABST
    Figure CN120260282A_ABST
Patent Text Reader

Abstract

The invention discloses a pre-training large model traffic flow prediction method based on double-activation domain bridging and space-time self-attention, and the method comprises the steps: designing a space-time feature multi-embedding module, carrying out the fusion of time, space and feature embedding, and constructing the multi-granularity representation of traffic flow data; designing a double-activation field bridging module, and aligning traffic flow data with the pre-training model space representation through a double-path activation mechanism; designing a dynamic gating mechanism, and adjusting the weight of the double activation branches through a softmax function; a LoRA strategy is combined with a partial attention freezing method to carry out fine tuning on the large language model; a pre-training multi-granularity space-time Transform module is designed, a space-time self-attention mechanism is used, and the modeling capability of the model for multi-granularity space-time features in traffic flow data is enhanced; and designing an output regression layer, and outputting a final prediction result by using a DADB module in combination with the convolutional layer. According to the method, the technical gap problem of the pre-training model and the traffic flow data can be effectively solved, and scientific decision support is provided for optimization of a smart city intelligent traffic system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of traffic flow prediction, and particularly relates to a traffic flow prediction method for a pre-trained large model based on dual-activated domain bridging and spatio-temporal self-attention. Background Art

[0002] Traffic flow prediction plays a crucial role in urban planning. Traditional urban planning often relies on historical data and macro trend analysis, lacking accurate judgment on the future traffic development trend. In the current era of information explosion, with the help of traffic flow analysis technology, urban managers can plan road construction more scientifically, optimize the layout of traffic infrastructure, reasonably plan the location and scale of facilities such as parking lots, charging stations, and bus hubs, and improve the overall operation efficiency of the city. In addition, in the event of emergencies such as bad weather, accurate prediction can also provide support for emergency management, helping relevant departments quickly formulate the optimal response plan and reducing the risk of traffic paralysis.

[0003] However, the traffic flow prediction task also faces many challenges. Time series data often has complex spatio-temporal dependencies. The data in the same sequence is affected by different historical trends and is also interfered by multi-dimensional external factors such as weather and holidays, resulting in insufficient prediction accuracy of the model.

[0004] Traditional traffic flow prediction methods are mainly based on statistical models, including autoregressive moving average models, regression models, and probability-based methods. These methods usually rely on the statistical characteristics of historical traffic flow data and describe the traffic flow change trend by establishing mathematical formulas to achieve the prediction of future traffic flow. Among them, the autoregressive moving average model describes the short-term and long-term trends of time series through autoregression and moving average. The Kalman filter method in the probability-based method uses a recursive update mechanism to dynamically estimate traffic flow. However, these methods based on mathematical statistics often rely on the assumption of data stationarity and are only effective when the traffic flow data has a high degree of stability, with low ability to handle dynamic time. In reality, traffic flow data is highly non-stationary dynamic data, showing a complex dynamic mixture of periodicity, trend, and suddenness, making it difficult for traditional models to accurately fit the long-term and short-term changes of traffic flow data. In addition, traditional statistical methods are also difficult to model complex non-linear relationships. In reality, traffic flow data is affected by various non-linear factors such as weather and seasons, and traditional statistical methods cannot accurately describe these non-linear interference factors. At the same time, traditional statistical methods have poor robustness to data missing and noise in the data, and the prediction performance will be greatly reduced in the face of emergencies.

[0005] With the development of computer technology, deep learning technology has gradually matured, and researchers have begun to attempt to use neural networks to process the complex dynamic spatio-temporal dependencies in traffic flow data. Among them, methods based on Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), Graph Neural Network (GNN), etc. have been widely used. The CNN technology uses the local receptive field and weight sharing mechanism, which has advantages in extracting spatial features and is suitable for capturing local correlations in the road network structure. However, CNN mainly relies on convolutional kernels of fixed size, making it difficult to model long-distance spatio-temporal dependencies and having a weak ability to capture the temporal dynamic changes of sequence data. The RNN can capture short-term and long-term dependencies in time series through the recursive calculation of hidden states. However, the standard RNN is prone to the problems of gradient vanishing or gradient explosion when processing long sequence data, resulting in the model being difficult to learn long-distance time-dependent information. Variants of RNN, such as Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU), alleviate this problem through the gating mechanism. However, they have a high computational complexity, are easily restricted by the sequence length during training, and are difficult to efficiently process large-scale traffic data. In addition, the RNN structure is difficult to parallelize the calculation and may face challenges of limited computing resources in actual deployment. The GNN has a powerful modeling ability for non-Euclidean space data and has been widely applied in the field of traffic flow prediction. The GNN method models the adjacency matrix of spatial nodes in the traffic network as a graph structure, enabling each traffic station to be used as a graph node in the graph structure, while the propagation relationship of traffic flow is modeled through the edges in the graph structure. This method effectively describes the complex topological relationships in the traffic network and captures the dynamic changes in traffic flow data to a certain extent. However, the GNN overly relies on the static adjacency matrix in traffic flow to model the graph structure, and the constructed graph is a relatively static graph structure. The relationships between spatial nodes in the traffic flow network are highly dynamic, such as spatio-temporal heterogeneity, that is, two spatially adjacent spatial nodes do not necessarily have similar time trends, while two nodes that have no relationship in terms of location may also have similar time trends. The GNN-based model is difficult to accurately capture the dynamically evolving spatio-temporal relationships. At the same time, the GNN method has a high computational complexity, especially in large-scale urban traffic prediction tasks, with large storage and computational costs.

[0006] In view of the limitations of traditional deep learning methods in traffic flow prediction, Transformer-based methods have been introduced into the field of traffic flow prediction to enhance the ability to model long-distance spatio-temporal dependencies. Transformer adopts the attention mechanism, which can globally model the spatio-temporal relationship of input data, overcoming the deficiencies of CNN in capturing long-distance dependencies and the gradient vanishing problem of RNN in processing long sequence data. In addition, since Transformer can process the entire input sequence in parallel during calculation, it has higher computational efficiency than RNN, making it one of the mainstream methods in traffic flow prediction tasks in recent years. To further improve the performance of Transformer in traffic flow prediction, CNN has begun to be combined with Transformer to construct graph Transformer, enabling the model to combine the topological structure modeling ability of GNN and the long-term dependence capture ability of Transformer, thereby improving the prediction accuracy in the spatial and temporal dimensions of traffic flow prediction tasks respectively. This type of method realizes the explicit modeling of road networks by incorporating graph convolution operations into the Transformer structure, improving the prediction ability in the spatial dimension. However, this method has a high computational cost on large-scale traffic networks, and at the same time, due to the quadratic complexity of the attention calculation of Transformer with the growth of the input scale, there are still scalability problems in processing ultra-large-scale traffic data. In addition, spatio-temporal enhanced Transformer reduces the time complexity by introducing local attention mechanisms and hierarchical modeling methods. However, such methods often require fine-tuning of hyperparameters and may need to be adjusted specifically in different scenarios, resulting in limited model generalization ability. Generally speaking, although Transformer and its variants have demonstrated powerful spatio-temporal modeling capabilities in traffic flow prediction, they still face challenges such as high computational complexity, insufficient data adaptability, and insufficient spatio-temporal feature fusion.

[0007] In recent years, the technology of large language models has developed rapidly and has become a hot topic in artificial intelligence technology. However, despite the powerful modeling capabilities of large language models, they still face challenges in traffic flow prediction tasks. Pre-trained large language models are mainly pre-trained based on natural language corpora, and there are significant domain differences between their representation spaces and traffic flow sequence data. At the same time, large language models process discrete symbol sequences during training, with clear context relationships and grammatical structures, while traffic flow data is a continuous value sequence with obvious dynamics. This creates a domain gap between the processing capabilities of large language models and traffic flow data. This domain gap results in the inability to directly and effectively transfer pre-trained knowledge to traffic prediction tasks, limiting the application effect of large language models in traffic flow prediction. Therefore, in practical applications, how to efficiently adapt large language models to enhance their performance in traffic flow prediction tasks remains a direction worthy of in-depth research. Summary of the Invention

[0008] The purpose of the present invention is to address the problems existing in the above-mentioned prior art, and provide a pre-trained large model traffic flow prediction method based on dual-activation domain bridging and spatio-temporal self-attention. An innovative dual-activation domain bridging module is designed and introduced into the traffic flow prediction task based on large language models, solving the problem of the domain gap of large language models for traffic flow data. A multi-granularity spatio-temporal attention mechanism is also designed to enhance the large language model's ability to capture multi-granularity spatio-temporal features and improve the large language model's modeling ability for spatio-temporal data. Combining the LoRA strategy with the partial attention freezing method to fine-tune the large language model further improves the adaptability of the large language model to traffic flow data. Finally, an innovative solution more suitable for traffic flow prediction tasks based on large language models is obtained.

[0009] The technical solution to achieve the purpose of the present invention is: a pre-trained large model traffic flow prediction method based on dual-activation domain bridging and spatio-temporal self-attention, the method comprising:

[0010] Step 1, design multiple embeddings of spatio-temporal features to construct a multi-granularity representation of traffic flow data;

[0011] Step 2, design a dual-activation domain bridging and dynamic gating mechanism;

[0012] Step 3, enhance the embedded features through the dual-activation domain bridging designed in Step 2 to align the traffic flow data with the representation space of the pre-trained large prediction model;

[0013] Step 4, introduce the LoRA strategy and the partial attention freezing method, and add a low-rank matrix to the attention layer of the pre-trained large prediction model to fine-tune the large language model;

[0014] Step 5: Design a pre-trained multi-granularity spatio-temporal Transformer, combined with the spatio-temporal self-attention mechanism, to enhance the modeling ability of multi-granularity spatio-temporal dependency relationships;

[0015] Step 6: Use dual-activation domain bridging to enhance feature representation and combine it with the convolutional layer to output the final traffic flow prediction result.

[0016] Furthermore, the design of spatio-temporal feature multiple embeddings in Step 1 to construct a multi-granularity representation of traffic flow data specifically includes:

[0017] Step 1.1: Determine the input data and target data for the traffic flow prediction task;

[0018] For the given past T time steps, the traffic flow input data is represented as where T represents the time step length, N represents the number of spatial nodes, F represents the feature dimension of the data, and x t represents the traffic flow data at time step t;

[0019] For the future P time steps, the traffic flow target data is represented as where P represents the time step length;

[0020] Step 1.2: Generate a time embedding representation to perform an embedding representation of the time features of the traffic flow input data and convert the time information into a learnable time representation;

[0021] Step 1.3: Generate a spatial embedding representation to characterize the spatial features of the traffic flow input data;

[0022] Step 1.4: Generate a feature embedding representation to map the traffic flow input data to a high-dimensional representation to adapt to the model input;

[0023] Step 1.5: Concatenate the time embedding, spatial embedding, and feature embedding to construct a complete spatio-temporal feature representation, which is used as the input for the subsequent large prediction model to achieve the traffic flow prediction task.

[0024] Furthermore, Step 2 specifically includes:

[0025] Step 2.1: Design a dual-activation domain bridging module;

[0026] First, perform feature extraction on the input feature X of the traffic flow data T using a dual-branch structure;

[0027] (1) ReLU branch processing: First, perform channel conversion using a two-dimensional 1×1 convolution and normalize the data using a batch normalization layer. Then, perform a non-linear transformation on the output after batch normalization using the ReLU activation function, and calculate the square of the output to enhance the non-linear expression ability of the features, expressed as:

[0028] Branch ReLU =(ReLU(BN(Conv2d(X T )))) 2

[0029] where Conv2d represents a two-dimensional 1×1 convolution, BN represents batch normalization processing, and Branch ReLU represents the result of ReLU branch processing;

[0030] (2) GELU branch processing: First, perform channel conversion using a two-dimensional 1×1 convolution and perform normalization using batch normalization. Then, perform a non-linear transformation using the GELU activation function, and calculate the square of the output to adapt to a more complex data distribution and enhance the adaptability to abnormal patterns, expressed as:

[0031] Branch GELU =(GELU(BN(Conv2d(X T )))) 2

[0032] In the formula, Branch GELU represents the result of GELU branch processing;

[0033] Next, set the number of channels in the intermediate layer to twice the number of final output channels, and add a Dropout layer after the calculations of the ReLU branch and the GELU branch;

[0034] Step 2.2, design a dynamic gating fusion mechanism to make full use of the advantages of the dual-branch structure;

[0035] Transform the input feature X T using a two-dimensional 1×1 convolution and normalize it using the Softmax activation function to ensure that the sum of the weights of the two branches is 1, obtaining the initial weight matrix W B :

[0036] W B =Softmax(Conv2d(X T ))

[0037] Split the initial weight matrix W B into w r and w gTwo weight matrices are multiplied by the outputs of the ReLU branch and the GELU branch respectively, and then weighted summation is performed to obtain the output data output:

[0038] output = w r ·Branch ReLU + w g ·Branch GELU

[0039] During the operation of the dual-activation domain bridging module, the dynamic gating mechanism dynamically adjusts the importance of the two branches according to the nature of the input features: when the data exhibits obvious linear features, the weight w r of the ReLU branch will be automatically increased; when complex non-linear features and abnormal patterns appear, the influence w g of the GELU branch will be automatically increased;

[0040] Step 2.3, design different variants of the dual-activation domain bridging module, including:

[0041] (1) The basic version of the dual-activation domain bridging module, which adopts a two-layer 1×1 convolution structure. The module is shown as the following formula:

[0042] Branch ReLU = (ReLU(BN(2Conv2d(X T )))) 2

[0043] Branch GELU = (GELU(BN(2Conv2d(X T )))) 2

[0044] Among them, 2Conv2d represents a two-layer 1×1 convolution structure;

[0045] (2) The deep version of the dual-activation domain bridging module, which adopts a four-layer 1×1 convolution structure. The module is shown as the following formula:

[0046] Branch ReLU = (ReLU(BN(4Conv2d(X T )))) 2

[0047] Branch GELU = (GELU(BN(4Conv2d(X T )))) 2

[0048] Among them, 4Conv2d represents a four-layer 1×1 convolution structure;

[0049] (3) Linear version dual-activation domain bridging module, which replaces the two-layer convolutional structure with a single-layer linear mapping and uses layer normalization instead of batch normalization. The module is shown in the following formula:

[0050] Branch ReLU =(ReLU((LN(X T )))) 2

[0051] Branch GELU =(GELU((LN(X T )))) 2

[0052] Among them, LN represents the layer normalization operation.

[0053] Furthermore, the specific method of enhancing the embedded features by the dual-activation domain bridging designed in step 2 in step 3 is: using the basic version dual-activation domain bridging module and the linear version dual-activation domain bridging module to enhance the embedded features.

[0054] Furthermore, GPT-2 is used as the pre-trained large language model in step 3. This model is based on the deep Transformer structure and consists of L layers. Each layer contains a multi-head self-attention mechanism and a feed-forward neural network;

[0055] In the specific implementation process, first, the input data is normalized. Subsequently, the input data enters the multi-layer Transformer structure. In the calculation process of the multi-head self-attention in each layer, each attention head independently calculates the attention weights of the input sequence, and then performs weighted summation. Finally, the results of the multi-heads are combined through a linear mapping to form the final attention representation; the input data P of the pre-trained large language model ToPLM is used as the hidden state H of the first layer l , and the calculation process of each layer is as follows:

[0056]

[0057] Among them, H h-1 represents the hidden state of the (h - 1)-th layer, represents the hidden state of the h-th layer, LN is layer normalization, MHA is multi-head self-attention, FFN is the feed-forward neural network, and the hidden state of the last layer is used as the output P of the pre-trained large language model FromPLM .

[0058] Furthermore, step 4 specifically includes:

[0059] Step 4.1, fine-tuning the pre-trained large language model by combining the LoRA strategy with the partial attention freezing method, specifically including:

[0060] Introduce low-rank matrix adjustment parameters in the self-attention layer of the Transformer to adapt it to traffic flow prediction tasks; among them, the LoRA strategy uses two small trainable matrices A and B as parameter replacements for the full-parameter matrix:

[0061] ΔW = A × B

[0062] W new = W orig + ΔW

[0063] where ΔW is the calculation result of LoRA parameters, and W orig is the original pre-trained weight of the pre-trained large language model, and W new is the weight of the pre-trained large language model after fine-tuning by the LoRA strategy;

[0064] For the first U - L layers of the pre-trained large language model, only the LoRA parameters and layer normalization parameters are updated, and the remaining parameters are frozen; for the subsequent U layers, the LoRA parameters, layer normalization parameters, and attention mechanism parameters are allowed to be updated, but the feed-forward network layer is frozen;

[0065] Step 4.2, perform reverse transformation on the output P FromPLM of the pre-trained large language model through the linear dual activation domain bridging module, namely the DADBL module, to project the features of the pre-trained large language model back to the original data space; then, use residual connection to add the output features of the pre-trained large language model to the original input feature X FromPLM to retain the original information and obtain the output data X FromPLM :

[0066] X FromPLM = P FromPLM + DADBL(P FromPLM )

[0067] Furthermore, step 5 specifically includes:

[0068] Step 5.1, combine the output data X FromPLM of the large language model in step 4 with the feature-enhanced representation X Feature obtained in step 3, and then perform feature projection through the basic dual activation domain bridging module, namely the DADB module, to calculate the query Q, key K, and value V matrices:

[0069] Q, K, V = DADB(X FromPLM + X Feature )

[0070] Step 5.2, introduce a learnable spatio-temporal encoding Emb TSA through random initialization to enhance the spatio-temporal relationship modeling ability:

[0071]

[0072] This spatio-temporal encoding vector is used to represent the relative position relationship between the time step and the spatial nodes. During the model training process, this spatio-temporal encoding is automatically updated through gradients and learned; where T represents the time step length, N represents the number of spatial nodes, F represents the feature dimension of the data, and B represents the data batch;

[0073] Step 5.3, calculate the dot product of the query Q and the key K, and normalize it through Softmax to obtain the attention weight α:

[0074]

[0075] where d k is the dimension of the key K, and α ij represents the attention weight between time step i and time step j, Q i represents the query Q of time step i, and K j represents the key K of time step j;

[0076] Step 5.4, apply the attention weight α ij to the value V, and combine it with the spatio-temporal encoding Emb TSA for operation to obtain the output data X after spatio-temporal self-attention processing output :

[0077] X output = TSA(Emb TSA , V, α)

[0078] where TSA represents the spatio-temporal self-attention mechanism.

[0079] Furthermore, step 6 specifically includes:

[0080] Step 6.1, perform enhancement processing on the spatio-temporal self-attention output data X output obtained in step 5 through the deep version of the dual activation domain bridging module, and use its dual-branch structure to extract linear mode and non-linear mode information respectively:

[0081] X enhanced = DADB2(X output )

[0082] where DADB2 represents the deep version of the dual activation domain bridging module, and X enhanced represents the enhanced feature;

[0083] Step 6.2, use two-dimensional convolution to reduce the dimension of the enhanced feature X enhanced and map it to the final prediction space to obtain the prediction result of traffic flow

[0084]

[0085] Compared with the prior art, the significant advantages of the present invention are as follows:

[0086] (1) The present invention improves the problems such as the domain gap between the expression space of the large language model and the traffic flow data in the traffic flow prediction task based on the large language model, and provides a more efficient and accurate method.

[0087] (2) The present invention adopts multiple spatio-temporal embeddings to capture periodicity and heterogeneity, and realizes a comprehensive characterization of the periodicity and regional heterogeneity of traffic flow at different time scales and spatial scales.

[0088] (3) The present invention innovatively designs a dual-activation domain bridging module to solve the problem of the domain gap between the expression space of the large language model and the traffic flow data.

[0089] (4) The present invention designs a variety of dual-activation domain bridging modules, including three dual-activation modules: the basic version (DADB), the deep version (DADB2), and the linear version (DADBL), which facilitates flexible deployment on high-performance servers or resource-constrained devices to adapt to different computing requirements.

[0090] (5) The present invention effectively alleviates the distribution difference between the corpus of the pre-trained large language model and the traffic data through the domain bridging of DADB, and avoids the reverse interference of the pre-trained knowledge during application.

[0091] (6) The present invention designs a multi-granularity spatio-temporal attention mechanism to enhance the large language model's ability to capture multi-granularity spatio-temporal features.

[0092] (7) The present invention combines the Lora strategy with the partial attention freezing method to fine-tune the large language model, further improving the adaptability of the large language model to traffic flow data.

[0093] (8) The present invention adopts two-stage regression prediction to improve accuracy and robustness. First, DADB is used to enhance features in the output layer, and then convolution is used to generate predictions, taking into account both the linear trend and non-linear fluctuations of the traffic flow, especially showing more robustness in complex patterns.

[0094] (9) The present invention completes the feature alignment, pre-training fine-tuning, and regression prediction in series within the same framework without intermediate manual intervention, simplifying the model online deployment and maintenance process.

[0095] (10) The present invention has good scalability and generality. The modular structure can easily replace the underlying pre-trained large language model or be extended to other spatio-temporal sequence tasks, such as weather prediction or power grid load prediction.

[0096] The present invention will be further described in detail below with reference to the accompanying drawings. Description of the Drawings

[0097] Figure 1 It is a flowchart of a traffic flow prediction method for a pre-trained large model based on dual activation domain bridging and spatio-temporal self-attention of the present invention.

[0098] Figure 2 It is a flowchart of using a pre-trained large model with dual activation domain bridging and spatio-temporal self-attention for traffic flow prediction of the present invention.

[0099] Figure 3 It is a framework diagram of a traffic flow prediction for a pre-trained large model based on dual activation domain bridging and spatio-temporal self-attention of the present invention. Detailed Embodiment

[0100] In order to make the objectives, technical solutions and advantages of the present application more clear, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0101] It should be noted that if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, such descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first", "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between various embodiments may be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.

[0102] In one embodiment, in combination with Figure 1 , a traffic flow prediction method for a pre-trained large model based on dual activation domain bridging and spatio-temporal self-attention is provided, and the method includes:

[0103] Step 1, design a spatio-temporal feature multiple embedding to construct a multi-granularity representation of traffic flow data;

[0104] Step 2, design a dual activation domain bridging and dynamic gating mechanism;

[0105] Step 3, enhance the embedded features through the dual activation domain bridging designed in Step 2 to align the traffic flow data with the representation space of the pre-trained large prediction model;

[0106] Step 4, introduce the LoRA strategy and partial attention freezing method, add a low-rank matrix to the attention layer of the pre-trained large prediction model to fine-tune the large language model;

[0107] Step 5, design a pre-trained multi-granularity spatio-temporal Transformer, combine the spatio-temporal self-attention mechanism to enhance the modeling ability of multi-granularity spatio-temporal dependence relationships;

[0108] Step 6, use dual activation domain bridging to enhance feature representation and combine with the convolutional layer to output the final traffic flow prediction result.

[0109] Furthermore, in one embodiment, the design of spatio-temporal feature multiple embedding in Step 1 to construct a multi-granularity representation of traffic flow data specifically includes:

[0110] Step 1.1, determine the input data and target data of the traffic flow prediction task;

[0111] For a given past T time steps, the traffic flow input data is represented as where T represents the time step length, N represents the number of spatial nodes, F represents the feature dimension of the data, and x t represents the traffic flow data at time step t;

[0112] For the future P time steps, the traffic flow target data is represented as where P represents the time step length;

[0113] Step 1.2, generate a time embedding representation, perform an embedding representation on the time features of the traffic flow input data, and convert the time information into a learnable time representation;

[0114] Step 1.3, generate a spatial embedding representation to characterize the spatial features of the traffic flow input data;

[0115] Step 1.4, generate a feature embedding representation to map the traffic flow input data to a high-dimensional representation to adapt to the model input;

[0116] Step 1.5, concatenate the time embedding, spatial embedding, and feature embedding to construct a complete spatio-temporal feature representation, which is used as the input to the subsequent large prediction model to achieve the traffic flow prediction task.

[0117] Here, preferably, in some embodiments, the time feature embedding in Step 1.2 adopts a dual time feature encoding method, encodes the date information and week information respectively through two independent embedding matrices to generate a time embedding matrix; through an index mapping method, maps the time stamp T to the corresponding time embedding matrix to generate a learnable date embedding matrix W Dayand the week embedding matrix W Week to obtain the time embedding representation Emb Temporal :

[0118]

[0119] where B represents the data batch;

[0120] Step 1.3 specifically includes:

[0121] Assign an independent embedding vector to each traffic node, and generate a learnable node embedding matrix W S to represent the spatial information of different monitoring stations;

[0122] Then, perform a linear transformation through the learnable parameter w s and the bias term b s for mapping and calculate the spatial embedding representation Emb Spatial :

[0123]

[0124] Step 1.4 specifically includes:

[0125] Set the learnable parameter w f and the bias term b f Perform a linear transformation on the input feature X to map it to a unified embedding space, and the calculation method is as follows:

[0126]

[0127] where Emb Feature is the feature embedding representation;

[0128] Step 1.5 specifically includes:

[0129] Through the Concat operation, splice the time, space, and feature embedding matrices along the feature dimension to generate the final spatio-temporal feature embedding fusion representation Emb Fusion :

[0130] Emb Fusion = Concat(Emb Feature , Emb Spatial , Emb Temporal ).

[0131] Furthermore, in one of the embodiments, Step 2 specifically includes:

[0132] Step 2.1, design a Dual Activation Domain Bridge (DADB);

[0133] First, perform feature extraction on the input features X of traffic flow data T using a dual-branch structure;

[0134] (1) ReLU branch processing: First, perform channel conversion using a 2D 1×1 convolution and normalize the data using a batch normalization layer. Then, perform a non-linear transformation on the output after batch normalization through the ReLU activation function, and calculate the square of the output to enhance the non-linear expression ability of the features, expressed as:

[0135] Branch ReLU = (ReLU(BN(Conv2d(X T )))) 2

[0136] where Conv2d represents a 2D 1×1 convolution, BN represents batch normalization processing, and Branch ReLU represents the processing result of the ReLU branch;

[0137] (2) GELU branch processing: First, perform channel conversion using a 2D 1×1 convolution and perform normalization using batch normalization. Then, perform a non-linear transformation using the GELU activation function, and calculate the square of the output to adapt to a more complex data distribution and enhance the adaptability to abnormal patterns, expressed as:

[0138] Branch GELU = (GELU(BN(Conv2d(X T )))) 2

[0139] In the formula, Branch GELU represents the processing result of the GELU branch;

[0140] Next, set the number of channels in the intermediate layer to twice the number of final output channels, and add a Dropout layer after the calculations of the ReLU branch and the GELU branch to avoid overfitting, improve the generalization performance of the model, and enhance the expression ability of the module;

[0141] Step 2.2, design a dynamic gating fusion mechanism to make full use of the advantages of the dual-branch structure;

[0142] Perform a transformation on the input feature X through a 2D 1×1 convolution, and use the Softmax activation function for normalization to ensure that the sum of the weights of the two branches is 1, obtaining the initial weight matrix W T : B :

[0143] W B = Softmax(Conv2d(X T ))

[0144] Split the initial weight matrix \(W\) B into \(w\) r and \(w\) g Two weight matrices, which are multiplied by the outputs of the ReLU branch and the GELU branch respectively, and then weighted and summed to obtain the output data output:

[0145] output = \(w\) r ·Branch ReLU + \(w\) g ·Branch GELU

[0146] During the operation of the dual-activation domain bridging module, the dynamic gating mechanism will dynamically adjust the importance of the two branches according to the nature of the input features: when the data exhibits obvious linear features, the weight \(w\) of the ReLU branch will be automatically increased r ; when complex non-linear features and abnormal patterns appear, the influence of the GELU branch \(w\) will be automatically increased g to enhance the adaptability of the dual-activation domain bridging module to complex features.

[0147] Step 2.3, design different variants of the dual-activation domain bridging module, including:

[0148] (1) Basic dual-activation domain bridging module, using a two-layer 1×1 convolutional structure, which provides effective feature extraction ability while ensuring computational efficiency. Suitable for scenarios with low requirements for computing resources, such as edge computing devices or real-time prediction tasks. The module is shown as the following formula:

[0149] Branch ReLU = (ReLU(BN(2Conv2d(X T )))) 2

[0150] Branch GELU = (GELU(BN(2Conv2d(X T )))) 2

[0151] Among them, 2Conv2d represents a two-layer 1×1 convolutional structure;

[0152] (2) Deep dual-activation domain bridging module, using a four-layer 1×1 convolutional structure to build deeper feature extraction levels. By gradually expanding and contracting the number of channels, the ability to model complex spatio-temporal features is improved. Suitable for scenarios with sufficient computing resources and high requirements for accuracy, such as offline analysis and large-scale traffic flow prediction tasks. The module is shown as the following formula:

[0153] BranchReLU = (ReLU(BN(4Conv2d(X T )))) 2

[0154] Branch GELU = (GELU(BN(4Conv2d(X T )))) 2

[0155] where 4Conv2d represents a four-layer 1×1 convolutional structure;

[0156] (3) Linear dual-activation domain bridging module, which replaces the two-layer convolutional structure with a single-layer linear mapping to reduce the computational cost. Layer normalization is used instead of batch normalization to improve the stability of the model during mini-batch training. It is applicable to feature projection tasks, such as scenarios of dimensionality reduction, fusion, and feature transformation of traffic data. The module is shown in the following formula:

[0157] Branch ReLU = (ReLU((LN(X T )))) 2

[0158] Branch GELU = (GELU((LN(X T )))) 2

[0159] where LN represents the layer normalization operation.

[0160] Furthermore, in one embodiment, combining Figure 2 and Figure 3 , the enhanced embedding features by the dual-activation domain bridging designed in step 2 in step 3 are specifically: using the basic dual-activation domain bridging module and the linear dual-activation domain bridging module to enhance the embedding features.

[0161] Preferably, in some embodiments, step 3 specifically includes:

[0162] Step 3.1, enhancing the spatio-temporal feature embedding fusion representation Emb Temporal generated in step 1 through the basic dual-activation domain bridging module to extract deeper feature information. The dual-activation domain bridging module extracts different patterns of data through the dual-activation mechanism and uses the dynamic gating mechanism to adjust the contribution ratio of the two activation functions to enhance the feature expression ability, obtaining an intermediate representation H Feature of enhanced features:

[0163] H Feature = DADB(Emb Fusion )

[0164] Among them, DADB represents the basic dual-activation domain bridging module;

[0165] Step 3.2: Through residual connection, maintain the original feature information of the embedded data and prevent gradient vanishing, improve the training stability, and obtain the feature-enhanced representation X Feature :

[0166] X Feature = Emb Fusion + H Feature

[0167] Step 3.3: Through the linear dual-activation domain bridging module, perform dimension alignment on the feature-enhanced representation X Feature to make it meet the input requirements of the pre-trained large language model, and obtain the input data P of the pre-trained large language model ToPLM :

[0168]

[0169] Among them, DADBL represents the linear dual-activation domain bridging module, F d is the feature dimension of the PLM, N represents the number of spatial nodes, T represents the time step length, and B represents the data batch.

[0170] Preferably, in some embodiments, GPT-2 is used as the pre-trained large language model in Step 3. This model is based on the deep Transformer structure and consists of L layers, each layer containing a multi-head self-attention mechanism and a feed-forward neural network;

[0171] In the specific implementation process, first, the input data is normalized to ensure the numerical stability of the input data. Subsequently, the input data enters the multi-layer Transformer structure. In the calculation process of the multi-head self-attention of each layer, each attention head independently calculates the attention weights of the input sequence, and then performs weighted summation. Finally, the results of the multi-heads are merged through a linear mapping to form the final attention representation; the input data P of the pre-trained large language model ToPLM is used as the hidden state H l of the first layer, and the calculation process of each layer is as follows:

[0172]

[0173] Among them, H h-1 represents the hidden state of the (h - 1)-th layer, represents the hidden state of the h-th layer, LN is layer normalization, MHA is multi-head self-attention, FFN is the feed-forward neural network, and the hidden state of the last layer is used as the output P FromPLM .

[0174] Further, in one of the embodiments, in combination with Figure 2 and Figure 3 , step 4 specifically includes:

[0175] Step 4.1, fine-tuning the pre-trained large language model through the LoRA strategy combined with the partial attention freezing method, specifically including:

[0176] Introduce low-rank matrix adjustment parameters in the self-attention layer of Transformer to adapt it to the traffic flow prediction task; among them, the LoRA strategy uses two small trainable matrices A and B as parameter replacements for the full parameter matrix, significantly reducing the number of parameters while maintaining the computational efficiency of the model:

[0177] ΔW = A × B

[0178] W new = W orig + ΔW

[0179] where ΔW is the calculation result of LoRA parameters, W orig is the pre-trained original weight of the pre-trained large language model, and W new is the weight of the pre-trained large language model after fine-tuning by the LoRA strategy;

[0180] Only the LoRA parameters and layer normalization parameters of the first U-L layers of the pre-trained large language model are updated, and the remaining parameters are frozen to retain the pre-trained knowledge of the PLM; the last U layers allow the update of LoRA parameters, layer normalization parameters, and attention mechanism parameters, but the feed-forward network layer is frozen to avoid excessive deviation from the original pre-trained distribution;

[0181] Step 4.2, perform reverse transformation on the output P FromPLM of the pre-trained large language model through the linear version of the dual activation domain bridging module, namely the DADBL module, to project the features of the pre-trained large language model back to the original data space; then use residual connection to add the output features of the pre-trained large language model to the original input feature X FromPLM to retain the original information and improve the generalization ability of the model, and obtain the output data X FromPLM :

[0182] X FromPLM = P FromPLM + DADBL(P FromPLM ).

[0183] Further, in one of the embodiments, step 5 specifically includes:

[0184] Step 5.1, combine the output data X FromPLM of the large language model in step 4 with the feature enhancement representation X obtained through step 3Feature Combine them, and then perform feature projection through the basic dual-activation domain bridging module, i.e., the DADB module, to calculate the query Q, key K, and value V matrices:

[0185] Q, K, V = DADB(X FromPLM + X Feature )

[0186] Step 5.2, introduce a learnable spatio-temporal encoding Emb TSA through random initialization to enhance the spatio-temporal relationship modeling ability:

[0187]

[0188] This spatio-temporal encoding vector is used to represent the relative position relationship between time steps and spatial nodes. During the model training process, this spatio-temporal encoding is automatically updated through gradients and learned to enhance the model's perception ability of long-term spatio-temporal dependence relationships; where T represents the time step length, N represents the number of spatial nodes, F represents the feature dimension of the data, and B represents the data batch;

[0189] Step 5.3, calculate the dot product of the query Q and the key K, and perform Softmax normalization to obtain the attention weight α:

[0190]

[0191] where d k is the dimension of the key K, and α ij represents the attention weight between time step i and time step j, Q i represents the query Q of time step i, and K j represents the key K of time step j;

[0192] Step 5.4, apply the attention weight α ij to the value V, and perform an operation in combination with the spatio-temporal encoding Emb TSA to obtain the output data X output after spatio-temporal self-attention processing:

[0193] X output = TSA(Emb TSA , V, α)

[0194] where TSA represents the spatio-temporal self-attention mechanism.

[0195] Through spatio-temporal self-attention calculation, the model can effectively model long-term dependence relationships across time steps and learn the interaction patterns of different spatial positions.

[0196] Furthermore, in one of the embodiments, step 6 specifically includes:

[0197] Step 6.1, enhance the spatio-temporal self-attention output data X obtained in Step 5 through the Deep Version Dual Activation Domain Bridging Module, and use its dual-branch structure to extract linear pattern and non-linear pattern information respectively, so as to improve the expression ability of features: output X = DADB2(X

[0198] X enhanced =DADB2(X output )

[0199] where DADB2 represents the Deep Version Dual Activation Domain Bridging Module, and X enhanced represents the enhanced feature;

[0200] Step 6.2, perform dimensionality reduction on the enhanced feature X enhanced by using two-dimensional convolution and map it to the final prediction space to obtain the prediction result of traffic flow

[0201]

[0202] The final output of the model is the predicted value of traffic flow for the next P time steps, which can be used for intelligent traffic scheduling and optimization.

[0203] In summary, the present invention designs a spatio-temporal feature multiple embedding module to fuse time, space, and feature embedding to construct a multi-granularity representation of traffic flow data; designs a dual activation domain bridging module to align traffic flow data with the spatial representation of the pre-trained model through a dual-path activation mechanism; designs a dynamic gating mechanism to adjust the weights of the dual activation branches through the softmax function; uses the LoRA strategy combined with the partial attention freezing method to fine-tune the large language model; designs a pre-trained multi-scale spatio-temporal Transformer module (Pretrained Multi-scale Spatio-Temporal Transformer, PMSTT), uses the spatio-temporal self-attention mechanism to enhance the model's ability to model multi-granularity spatio-temporal features in traffic flow data; designs an output regression layer, and uses the DADB module combined with the convolutional layer to output the final prediction result. The present invention effectively overcomes the technical gap problem between the pre-trained model and traffic flow data, and provides scientific decision-making support for the optimization of the intelligent transportation system in smart cities.

[0204] In one embodiment, a traffic flow prediction system based on a pre-trained large model with dual activation domain bridging and spatio-temporal self-attention is provided, and the system includes:

[0205] The first module is used to design spatio-temporal feature multiple embedding to construct a multi-granularity representation of traffic flow data;

[0206] The second module is used to design the dual activation domain bridging and the dynamic gating mechanism.

[0207] The third module is used to enhance the embedded features through the designed dual-activation domain bridge, aligning the traffic flow data with the representation space of the pre-trained large prediction model;

[0208] The fourth module is used to introduce the LoRA strategy and the partial attention freezing method, adding a low-rank matrix to the attention layer of the pre-trained model to fine-tune the large language model;

[0209] The fifth module is used to design a pre-trained multi-granularity spatio-temporal Transformer, combining the spatio-temporal self-attention mechanism to enhance the modeling ability of multi-granularity spatio-temporal dependence relationships;

[0210] The sixth module is used to enhance the feature expression through the dual-activation domain bridge and combine the convolutional layer to output the final prediction result.

[0211] For the specific limitations of the pre-trained large model traffic flow prediction system based on the dual-activation domain bridge and spatio-temporal self-attention, reference can be made to the limitations of the pre-trained large model traffic flow prediction method based on the dual-activation domain bridge and spatio-temporal self-attention in the above text, which will not be elaborated here. Each module in the above pre-trained large model traffic flow prediction system based on the dual-activation domain bridge and spatio-temporal self-attention can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0212] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the pre-trained large model traffic flow prediction method based on the dual-activation domain bridge and spatio-temporal self-attention.

[0213] For the specific limitations of each step, reference can be made to the limitations of the pre-trained large model traffic flow prediction method based on the dual-activation domain bridge and spatio-temporal self-attention in the above text, which will not be elaborated here.

[0214] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it implements the pre-trained large model traffic flow prediction method based on the dual-activation domain bridge and spatio-temporal self-attention.

[0215] For the specific limitations of each step, reference can be made to the limitations of the pre-trained large model traffic flow prediction method based on the dual-activation domain bridge and spatio-temporal self-attention in the above text, which will not be elaborated here.

[0216] In summary, the present invention designs a dual-activation domain bridging module to achieve efficient alignment and fusion of traffic sequence data with the representation ability of large language models, effectively overcoming the domain gap. The dual-activation domain bridging module realizes the efficient alignment and fusion of features through a parallel dual-branch structure and a dynamic gating mechanism. In its dual-branch parallel architecture, the ReLU branch captures linear relationships and simple non-linear patterns in the data by virtue of its sparsity and gradient propagation characteristics; the GELU branch can handle complex non-linear features and abnormal patterns by introducing a probability distribution method. In addition, a multi-granularity spatio-temporal attention mechanism based on a pre-training strategy is designed, which enhances the ability of large language models to capture multi-granularity spatio-temporal features through learnable spatio-temporal encoding and differential feature projection strategies.

[0217] In the traffic flow prediction task based on large language models, designing a dual-activation domain bridging module and a multi-granularity spatio-temporal attention mechanism is an innovative approach. Specifically, after multiple embeddings and embedding fusion of traffic flow data, the traffic flow data is aligned with the processing capabilities of the large language model through the basic dual-activation domain bridging module and the linear dual-activation domain bridging module. The large language model is fine-tuned using the LoRA strategy and the attention freezing strategy. The output of the large language model is fed into the multi-granularity spatio-temporal attention mechanism to enhance the ability to capture multi-granularity spatio-temporal features. Finally, the prediction of future traffic flow data is obtained through an output regression layer with the deep dual-activation domain bridging module as the core. This method has the following advantages: 1. Diversified feature extraction ability. Through the ReLU and GELU dual-branch parallel architecture, the model can capture both linear relationships and simple non-linear patterns in the data, as well as handle complex non-linear features and abnormal patterns, effectively making up for the limitations of a single activation function, thereby enhancing the comprehensive understanding of traffic flow data. 2. Adaptive feature fusion. Using a dynamic gating mechanism and calculating the weights of the two branches with Softmax, the ReLU branch dominates when the data presents linear features, while the GELU branch plays a greater role when the data contains complex patterns. This adaptive feature enables the model to make optimal adjustments according to the characteristics of the input data, improving the ability of feature alignment and fusion. 3. Enhanced model expressiveness and generalization ability. The number of channels is doubled in the intermediate layer to expand the feature space, and the dropout mechanism is combined to prevent overfitting. This design not only improves the feature expression ability of the model but also enhances its adaptability to different traffic flow patterns, making the prediction results more stable and accurate. 4. Flexibility to adapt to different computing requirements. Three variants, namely the basic dual-activation domain bridging module, the deep dual-activation domain bridging module, and the linear dual-activation domain bridging module, are designed to be applicable to different computing resources and application scenarios. 5. Efficient pre-trained knowledge transfer. The LoRA strategy is used for efficient fine-tuning. While maintaining the pre-trained knowledge, the attention layer parameters are adjusted through low-rank matrices, enabling the large language model to further adapt to traffic flow data. 6. Enhanced multi-granularity spatio-temporal feature modeling ability. Through the spatio-temporal self-attention mechanism, combined with learnable spatio-temporal encoding and compact projection strategies, the model can capture the relative position relationships across time steps and spatial nodes, deeply mining the long-term and short-term dependencies in traffic flow data.

[0218] In summary, the traffic flow prediction method of the pre-trained large model based on dual-activated domain bridging and spatio-temporal self-attention proposed by the present invention not only solves the domain gap problem of large language models when facing traffic flow data through dual-activated domain bridging, enhancing the spatial dimension modeling ability of large language models, but also further enhances the ability of large language models to capture multi-granularity spatio-temporal features through the multi-granularity spatio-temporal attention mechanism, providing new ideas and methods for the development of research in the field of traffic flow prediction.

[0219] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A traffic flow prediction method for a pre-trained large model based on dual-activated domain bridging and spatio-temporal self-attention, characterized in that, The method includes: Step 1, design multiple embeddings of spatio-temporal features to construct multi-granularity representations of traffic flow data; Step 2, design a dual-activation domain bridging and dynamic gating mechanism; Step 3, enhance the embedded features through the dual-activation domain bridging designed in Step 2, and align the traffic flow data with the representation space of the pre-trained large prediction model; Step 4, introduce the LoRA strategy and partial attention freezing method, and add a low-rank matrix to the attention layer of the pre-trained large prediction model to fine-tune the large language model; Step 5, design a pre-trained multi-granularity spatio-temporal Transformer, combine spatio-temporal self-attention mechanism to enhance the modeling ability of multi-granularity spatio-temporal dependence relationships; Step 6, enhance feature expression through dual-activation domain bridging, and combine convolutional layers to output the final traffic flow prediction result.

2. The traffic flow prediction method of the pre-trained large model based on dual-activated domain bridging and spatio-temporal self-attention according to claim 1, wherein The design of multiple embeddings of spatio-temporal features in Step 1 to construct multi-granularity representations of traffic flow data specifically includes: Step 1.1, determine the input data and target data of the traffic flow prediction task; For a given past T time steps, the traffic flow input data is represented as where T represents the time step length, N represents the number of spatial nodes, F represents the feature dimension of the data, and x t represents the traffic flow data at time step t; For the next P time steps, the traffic flow target data is expressed as where P represents the time step length; Step 1.2, generate time embedding representations, perform embedding representations on the time features of traffic flow input data, and convert time information into learnable time representations; Step 1.3, generate spatial embedding representations to characterize the spatial features of traffic flow input data; Step 1.4, generate feature embedding representations to map traffic flow input data to high-dimensional representations to adapt to model input; Step 1.5, concatenate time embeddings, spatial embeddings, and feature embeddings to construct a complete spatio-temporal feature representation, which is used as the input for the subsequent large prediction model to achieve the traffic flow prediction task.

3. The traffic flow prediction method of the pre-trained large model based on dual-activated domain bridging and spatio-temporal self-attention according to claim 2, wherein In step 1.2, the time feature embedding adopts a dual-time feature encoding method. The date information and week information are respectively encoded through two independent embedding matrices to generate a time embedding matrix. Through the index mapping method, the timestamp T is mapped to the corresponding time embedding matrix to generate a learnable date embedding matrix W Day and week embedding matrix W Week , obtaining the time embedding representation Emb Temporal : In the formula, B represents the data batch; Step 1.3 specifically includes: Assign an independent embedding vector to each traffic node and generate a learnable node embedding matrix W through random initialization S , to represent the spatial information of different monitoring stations; Then, a linear transformation is adopted to perform mapping through learnable parameters w s and bias term b s to calculate the spatial embedding representation Emb of the node Spatial : Step 1.4 specifically includes: Set the learnable parameter w f and the bias term b f Perform a linear transformation on the input feature X to map it to a unified embedding space, and the calculation method is as follows: where Emb Feature is the feature embedding representation; Step 1.5 specifically includes: Through the Concat operation, the time, space, and feature embedding matrices are concatenated along the feature dimension to generate the final spatio-temporal feature embedding fusion representation Emb Fusion : Emb Fusion = Concat(Emb Feature , Emb Spatial , Emb Temporal ).

4. The traffic flow prediction method of the pre-trained large model based on dual-activated domain bridging and spatio-temporal self-attention according to claim 1, wherein Step 2 specifically includes: Step 2.1, design a dual-activation domain bridging module; First, the input features X of the traffic flow data T are subjected to feature extraction using a dual-branch structure; (1) ReLU branch processing: First, use a two-dimensional 1×1 convolution for channel conversion and use a batch normalization layer to normalize the data. Then, perform a non-linear transformation on the output after batch normalization through the ReLU activation function, and calculate the square of the output, expressed as: Branch ReLU = (ReLU(BN(Conv2d(X T )))) 2 Among them, Conv2d represents a two-dimensional 1×1 convolution, BN represents batch normalization processing, and Branch ReLU represents the processing result of the ReLU branch; (2) GELU branch processing: First, use a two-dimensional 1×1 convolution for channel conversion and use batch normalization for normalization. Then, use the GELU activation function for non-linear transformation, and calculate the square of the output, expressed as: Branch GELU = (GELU(BN(Conv2d(X T )))) 2 In the formula, Branch GELU represents the processing result of the GELU branch; Next, set the number of channels in the intermediate layer to twice the number of final output channels, and add a Dropout layer after the calculations of the ReLU branch and GELU branch; Step 2.2, design a dynamic gating fusion mechanism to make full use of the advantages of the dual-branch structure; Transform the input feature X through a two-dimensional 1×1 convolution T and normalize it using the Softmax activation function to ensure that the sum of the weights of the two branches is 1, obtaining the initial weight matrix W B : W B = Softmax(Conv2d(X T )) Split the initial weight matrix W B into w r and w g Two weight matrices, which are multiplied by the outputs of the ReLU branch and the GELU branch respectively, and then weighted and summed to obtain the output data output: output = w r ·Branch ReLU +w g ·Branch GELU During the operation of the bridging module in the dual-activation domain, the dynamic gating mechanism dynamically adjusts the importance of the two branches according to the nature of the input features: when the data exhibits obvious linear features, the weight w of the ReLU branch will be automatically increased r ; when complex non-linear features and abnormal patterns appear, the influence w of the GELU branch will be automatically increased g ; Step 2.3, design different variants of the dual-activation domain bridging module, including: (1) Basic version of the dual-activation domain bridging module, using a two-layer 1×1 convolution structure, and the module is shown as the following formula: Branch ReLU = (ReLU(BN(2Conv2d(X T )))) 2 Branch GELU =(GELU(BN(2Conv2d(X T )))) 2 Among them, 2Conv2d represents a two-layer 1×1 convolution structure; (2) Deep version of the dual-activation domain bridging module, using a four-layer 1×1 convolution structure, and the module is shown as the following formula: Branch ReLU = (ReLU(BN(4Conv2d(X T )))) 2 Branch GELU =(GELU(BN(4Conv2d(X T )))) 2 Among them, 4Conv2d represents a four-layer 1×1 convolution structure; (3) Linear version dual-activation domain bridging module, which replaces the two-layer convolutional structure with a single-layer linear mapping and uses layer normalization instead of batch normalization. The module is shown in the following formula: Branch ReLU = (ReLU((LN(X T )))) 2 Branch GELU = (GELU((LN(X T )))) 2 Among them, LN represents the layer normalization operation.

5. The traffic flow prediction method of the pre-trained large model based on dual-activated domain bridging and spatio-temporal self-attention according to claim 4, wherein In step 3, the enhanced embedding features by the dual-activation domain bridging designed in step 2 are specifically: using the basic version dual-activation domain bridging module and the linear version dual-activation domain bridging module to enhance the embedding features.

6. The traffic flow prediction method of the pre-trained large model based on dual-activated domain bridging and spatio-temporal self-attention according to claim 5, characterized in that Step 3 specifically includes: Step 3.1, enhance the spatio-temporal feature embedding fusion representation Emb generated in Step 1 through the basic version of the dual-activation domain bridging module Temporal to obtain an intermediate representation H with enhanced features Feature : H Feature = DADB(Emb Fusion ) Among them, DADB represents the basic version dual-activation domain bridging module; Step 3.2, through residual connection to maintain the original feature information of the embedded data and prevent gradient disappearance, obtaining the feature enhanced representation X Feature : X Feature = Emb Fusion + H Feature Step 3.3, perform dimensional alignment on the feature enhanced representation X Feature through the linear version dual activation domain bridging module to make it meet the input requirements of the pre-trained large language model, and obtain the input data P ToPLM : Among them, DADBL represents the linear version of the dual-activation domain bridging module, F d is the feature dimension of the PLM, N represents the number of spatial nodes, T represents the time step length, and B represents the data batch.

7. The traffic flow prediction method of the pre-trained large model based on dual-activated domain bridging and spatio-temporal self-attention according to claim 1, characterized in that In step 3, GPT-2 is used as the pre-trained large language model. This model is based on the deep Transformer structure and consists of L levels. Each level contains a multi-head self-attention mechanism and a feed-forward neural network; In the specific implementation process, the input data is first normalized. Subsequently, the input data enters a multi-layer Transformer structure. During the calculation of multi-head self-attention in each layer, each attention head independently calculates the attention weights of the input sequence, and then performs weighted summation. Finally, the results of multiple heads are combined through a linear mapping to form the final attention representation; the input data P of the pre-trained large language model ToPLM is used as the hidden state H of the first layer l , and the calculation process of each layer is as follows: Among them, H h-1 represents the hidden state of the (h-1)-th layer, represents the hidden state of the h-th layer, LN is layer normalization, MHA is multi-head self-attention, FFN is a feed-forward neural network, and the hidden state of the last layer is used as the output P of the pre-trained large language model FromPLM .

8. The traffic flow prediction method for the pre-trained large model based on dual-activated domain bridging and spatio-temporal self-attention according to claim 4 or 7, characterized in that, Step 4 specifically includes: Step 4.1, fine-tuning the pre-trained large language model by combining the LoRA strategy with the partial attention freezing method, specifically including: Introducing low-rank matrix adjustment parameters in the self-attention layer of the Transformer to make it adapt to the traffic flow prediction task; among them, the LoRA strategy uses two small trainable matrices A and B as parameters to replace the full parameter matrix: ΔW = A × B W new = W orig + ΔW Among them, ΔW is the calculation result of LoRA parameters, and W orig is the original pre-trained weight of the pre-trained large language model, and W new is the weight of the pre-trained large language model after fine-tuning with the LoRA strategy; For the first U-L layers of the pre-trained large language model, only the LoRA parameters and the layer normalization parameters are updated, and the remaining parameters are kept frozen; for the subsequent U layers, the LoRA parameters, the layer normalization parameters, and the attention mechanism parameters are allowed to be updated, but the feed-forward network layer is frozen; Step 4.2, perform inverse transformation on the output P of the pre-trained large language model through the linear double-activation domain bridging module, namely the DADBL module, to project the pre-trained large language model features back to the original data space; then adopt residual connection to add the output features of the pre-trained large language model and the original input feature X FromPLM to retain the original information and obtain the output data X FromPLM : FromPLM ​ X FromPLM = P FromPLM + DADBL(P FromPLM )。 9. The traffic flow prediction method of the pre-trained large model based on dual-activated domain bridging and spatio-temporal self-attention according to claim 1 or 4, characterized in that Step 5 specifically includes: Step 5.1: Take the output data X of the large language model in Step 4 FromPLM and combine it with the feature-enhanced representation X obtained in Step 3 Feature After that, perform feature projection through the basic double-activation domain bridging module, i.e., the DADB module, to calculate the query Q, key K, and value V matrices: Q, K, V = DADB(X FromPLM + X Feature ) Step 5.2, introduce a learnable spatio-temporal encoding Emb through random initialization TSA to enhance the spatio-temporal relationship modeling ability: This spatio-temporal encoding vector is used to represent the relative position relationship between the time step and the spatial node. During the model training process, this spatio-temporal encoding is automatically updated through the gradient and learned; among them, T represents the time step length, N represents the number of spatial nodes, F represents the feature dimension of the data, and B represents the data batch; Step 5.3, calculating the dot product of the query Q and the key K, and normalizing it through Softmax to obtain the attention weight α: where d k is the dimension of key K, and α ij represents the attention weight between time step i and time step j, Q i represents the query Q at time step i, and K j represents the key K at time step j; Step 5.4, apply the attention weight α ij to the value V and combine it with the spatio-temporal encoding Emb TSA for operation to obtain the output data X processed by spatio-temporal self-attention output : X output = TSA(Emb TSA , V, α) Among them, TSA represents the spatio-temporal self-attention mechanism.

10. The traffic flow prediction method of the pre-trained large model based on dual-activated domain bridging and spatio-temporal self-attention according to claim 1 or 4, characterized in that, Step 6 specifically includes: Step 6.1, enhance the spatio-temporal self-attention output data X obtained in Step 5 through the deep version double-activation domain bridging module output and use its double-branch structure to extract linear pattern and non-linear pattern information respectively: X enhanced = DADB 2(X output ) Among them, DADB2 represents the deep version of the dual activation domain bridging module, and X enhanced represents the enhanced feature; Step 6.2, perform two-dimensional convolution on the enhanced feature X enhanced to reduce the dimension and map it to the final prediction space, obtaining the prediction result of traffic flow

Citation Information

Cited By

  • Solar radiation forecasting method based on satellite data and multi-spatio-temporal scale super-division model

    CN120635598A

  • Forecasting method of solar radiation based on satellite data and multi-time and space scale super-resolution model

    CN120635598B

  • Traffic flow prediction method, system and equipment based on hierarchical semantic prompt large language model

    CN121260014A

  • Traffic prediction method based on distillation big language model

    CN121305878A