Large-scale highway network traffic flow diagram fine-tuning large-model prediction method
By building a heterogeneous graph of the expressway network, combining the graph convolution network and the temporal convolution network, the spatial and temporal characteristics are captured, and a large language model is used for fine-tuning, the problem of temporal dependence in a large-scale expressway network is solved, and more efficient traffic flow prediction is achieved.
Patent Information
- Application Number
- CN202510454563.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-04-11
AI Technical Summary
Existing traffic flow prediction methods cannot effectively capture space-time dependencies in large-scale highway networks, rely on large-scale historical data, and cannot respond to rapid changes in traffic flow in real time, especially in complex environments.
A heterogeneous graph of the highway network is constructed, and a graph convolution network GCN and a time convolution network TCN are used to capture spatial and temporal features, combined with a large language model LLM for fine-tuning, and traffic flow prediction is performed through adaptive preference gates and part of the attention layer freezing strategy.
It improves the accuracy and efficiency of traffic flow prediction, enhances the adaptability and generalization capabilities of the model, and can adapt to different time and space data prediction tasks.
Smart Images

Figure CN120412262A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of traffic flow prediction, and particularly to a prediction method for fine-tuning a large model of traffic flow maps of a large-scale highway network. Background Art
[0002] Highway traffic flow prediction mainly relies on traditional traffic flow detection devices and prediction models based on historical data. Traditional traffic flow monitoring devices include geomagnetic sensors, radar speed detectors, laser speed measurement devices, cameras, and floating car data, etc. Geomagnetic sensors calculate vehicle flow and speed by sensing the interference of vehicle metal on the magnetic field; radar speed detectors and laser speed measurement devices measure the speed of vehicles by emitting radio frequency signals or laser beams, and calculate the flow by the driving time and distance of the vehicles; cameras mainly collect traffic flow data in real time through image monitoring, and extract information such as traffic density and vehicle speed with the help of image processing technology; floating car data (Floating Car Data, FCD) collects information about traffic flow by real-time monitoring of the GPS positions of vehicles, and then conducts flow prediction.
[0003] Although these traditional traffic flow monitoring devices and methods can provide traffic flow information to a certain extent, they still face many problems. First of all, these methods often rely on the layout of physical sensors. In special sections such as highway tunnel groups and tunnel entrances and exits, traditional devices face the problem of scarce data. Secondly, most traditional traffic flow prediction methods are based on statistical analysis of historical traffic data, using methods such as time series analysis and regression models. These methods usually rely on relatively simple statistical models, ignoring the complex spatio-temporal dynamics in traffic flow and unable to fully consider the influence of complex environmental factors. In addition, existing traffic flow prediction methods mainly rely on macroscopic traffic flow parameters (such as flow, speed, occupancy, etc.) for modeling, but these macroscopic parameters have great limitations in practical applications.
[0004] In actual prediction applications, most existing methods predict based on local data without analyzing large-scale data. In this way, the model will be greatly restricted and unable to effectively capture the spatio-temporal dependence relationships between various highway sections in the road network. At the same time, these methods rely on statistical models and traditional machine learning algorithms and are unable to handle complex traffic flow change patterns over a long time span. Especially in sections where the highway changes frequently, traditional models are difficult to adapt to the rapidly changing traffic flow patterns. Therefore, although existing methods perform well in some small-scale and static environments, they cannot respond in real time to the rapid changes in traffic flow when dealing with traffic flow prediction tasks in large-scale and complex environments, which limits their application.
[0005] Therefore, although traditional traffic flow prediction methods are still effective for small-scale road network data, they have significant deficiencies in the application of large-scale road networks. In a complex large-scale environment, how to effectively handle spatio-temporal dependencies, how to obtain complete and real-time traffic flow data, and how to improve the generalization ability of the model are the core issues that current methods need to solve, which also strongly demands the proposal and development of new traffic flow prediction methods, especially advanced models that can combine large-scale traffic data and have strong adaptability and real-time performance.
[0006] In recent years, with the rapid development of deep learning technology, more and more research has attempted to apply deep learning to the field of traffic flow prediction. These methods mainly model and predict time series data by constructing complex deep neural network models, such as recurrent neural networks (RNNs), long short-term memory networks (LSTMs), graph neural networks (GNNs), etc. In particular, models such as LSTM and GNN can effectively capture complex dependencies in time series data and have strong modeling and generalization abilities, but these methods still have certain limitations. Although models such as LSTM can handle long-term dependencies in time series data, their performance is not ideal when dealing with spatially dependent data. Although graph neural networks (GNNs) can handle spatial structure data to a certain extent, they usually have limited extraction of spatial information from data and cannot effectively combine complex environmental factors for modeling. Most existing deep learning models rely on large-scale historical data for training, and their performance will also decline in the case of scarce or unavailable data. Summary of the Invention
[0007] The purpose of the present invention is to provide a large-scale highway network traffic flow graph fine-tuning large model prediction method to solve the problem that existing deep learning models need to rely on large-scale historical data for training to maintain performance.
[0008] To solve the above technical problems, the present invention provides a large-scale highway network traffic flow graph fine-tuning large model prediction method, including the steps of:
[0009] S1: Construct a heterogeneous graph of highway network traffic flow, where the nodes in the heterogeneous graph are toll station nodes and intersection nodes, and the edges in the heterogeneous graph are the traffic flows between the nodes;
[0010] S2: Use the graph convolutional network GCN to aggregate the neighbor information of the nodes, capture the spatial characteristics of the highway network traffic flow, and calculate the spatial weights of the edges between the nodes;
[0011] S3: Add time information to the heterogeneous graph, use the Temporal Convolutional Network (TCN) combined with Chebyshev filtering and dilated convolution to capture the temporal features of the highway network traffic flow, and calculate the temporal weights of the edges between nodes;
[0012] S4: Concatenate the spatial features and temporal features to generate a unified sentence vector;
[0013] S5: Input the unified sentence vector into the Large Language Model (LLM), and perform traffic flow prediction through instruction fine-tuning and partial attention layer freezing strategy.
[0014] Further, the heterogeneous graph is represented as:
[0015] G = (V, E)
[0016] where V is the set of nodes, including the toll station nodes U and intersection nodes I; E is the set of edges, including the traffic flow between toll station nodes and intersection nodes, the traffic flow between toll station nodes, and the traffic flow between intersection nodes.
[0017] Further, the spatial features output by each graph convolutional layer of the Graph Convolutional Network (GCN) are represented as:
[0018]
[0019] where: σ is the activation function, is the set of neighbor nodes of node u, Wuv is the spatial weight of the edge between node u and node v; h (l) is the input of the l-th layer of graph convolution.
[0020] Further, the temporal features output by each convolutional layer of the Temporal Convolutional Network (TCN) are represented as:
[0021]
[0022] where: is the temporal weight of the edge between node u and node v, is the dilated convolution representation vector, is the traffic flow prediction score at the |T t |-d l t-th time step, yt is the Chebyshev filtering representation vector;
[0023] The dilated convolution representation vector is represented as:
[0024]
[0025] where: x t is the traffic flow feature at the t-th time step; d lis the dilation rate of the l-th layer; is the kernel function; Tt is the vectorized representation of the traffic flow on the highway network corresponding to the time in the time series of t; is the |Tt|-d-th l input of t time steps, indicating the time steps skipped by the convolution kernel in the dilated convolution;
[0026] The time convolutional network TCN uses the first-kind Chebyshev filter for tensor filtering, and the transfer function of the first-kind Chebyshev filter is expressed as:
[0027]
[0028] where: ε is the ripple coefficient of the filter, used to control the gain fluctuation in the passband; is the n-th order Chebyshev polynomial, ω c is the cut-off frequency of the filter, s is the input parameter of the Chebyshev filter;
[0029] By passing parameters, H(s) is the result of the graph convolution calculation H(G i,j );
[0030] For the signal sequence of the ρ-layer time convolution input on C i channels the weight of the convolution layer is w i the output after Chebyshev filtering is:
[0031]
[0032] G i,j is the graph convolution;
[0033] The traffic flow prediction score is expressed as:
[0034]
[0035] where: α is the decay factor of the traffic flow in the time dimension; is the time difference between the time step t and |Tt|-d l and t; η is the parameter for controlling the time decay; χ is the time weight threshold.
[0036] Furthermore, the unified sentence vector uf is expressed as:
[0037] uf = [gf; tf]
[0038] where, gf is the vector representation of the spatial feature h (l+1) ; the tf is the vector representation of the time feature ; [;] represents the vector concatenation operation.
[0039] Furthermore, the method further includes embedding the tokens through pointwise convolution; for the spatial vectors, an adaptive vector embedding is adopted
[0040]
[0041] where σ represents the activation function, is a learnable parameter, Par S is the spatial vector embedding;
[0042] For the temporal vectors, a linear layer is used to encode the input data into the day of the month and the day of the week, and the two are independently embedded, and then the two embeddings are added together to obtain the temporal embedding representation
[0043] Embedding of the day of the month is:
[0044]
[0045] Embedding of the day of the week
[0046]
[0047] where, and are the position information generated by performing absolute position encoding on each provincial traffic data at the resolutions of "month" and "week"; and are the learnable temporal embeddings for the day of the week and the day of the month.
[0048] Furthermore, the method further includes using an adaptive preference gate bidirectional gated recurrent unit to process the input uf sequence u1, u2, …, u t , …, u[[ID=4|8]] T from front to back and from back to front respectively, and the forward sequence output is:
[0049]
[0050] The reverse sequence output is:
[0051]
[0052] By concatenating the hidden states of the forward sequence output and the reverse sequence output, the hidden state is:
[0053]
[0054] Then, the degree of attention paid to different features is dynamically adjusted, and an additional gating mechanism is introduced. At the tth moment, the update process of the preference gate is:
[0055] g t =σ(W g ·u t +U g ·h t-1 )
[0056] Among them, g t is the current moment, W g and U g is the learnable weight, x t The input at the current moment.
[0057] Furthermore, the large language model LLM adopts an encoder and decoder architecture;
[0058] The encoder converts the input sequence X into a contextual representation by fine-tuning the instruction:
[0059] Henc=Encoder(X,θ enc )
[0060] Among them, H enc is the encoded context hidden representation, θ enc is the encoder parameter;
[0061] X=Concat(P,H)
[0062] Where P is a matrix composed of time vectors, expressed as, P = {W1|W2|…|W t |…|W T}; H is the historical traffic data matrix composed of uf sequence according to column vectors H = {u1|u2|…|u t |…|u T};
[0063] The decoder outputs the prediction result based on the encoder:
[0064] Y=Decoder(Henc,Y <t θ dec )
[0065] Among them, Y is the output sequence of the decoder, Y <t is the output of the decoder before time step t, θ dec Decoder parameters.
[0066] Further, during the training of the large language model LLM, the multi-head attention layer and the feed-forward network of the large language model LLM are frozen during training, so that the first F layers of the large language model LLM are frozen, and the parameters of the last U multi-head attention layers are unfrozen.
[0067] Further, the loss function of the large language model LLM is:
[0068]
[0069]
[0070] where the regularization term λ∑κ∈θκ 2 is L2 regularization, θ is the learnable parameter of the model, λ is the regularization hyperparameter, y is the true value of the training data, is the smoothed true value of the v-th prediction target, is the predicted probability of the v-th target, σ is the label smoothing factor, |V| is the number of the graph point set; ω is the coefficient that adjusts the overall loss function and "penalizes" the negative sampling pairs;
[0071]
[0072] where: cos(·,·) is the distance function, and the calculation formulas for the u-th and i-th prediction targets are:
[0073]
[0074] is the indicator function, and the calculation formula for the smoothed true values of the u-th and i-th prediction targets is:
[0075]
[0076] where N r is the number of randomly sampled pairs in is the point representation of the heterogeneous graph in the last hidden layer, as the indicator function, when then it means and have different labels; when then it indicates that and have the same label.
[0077] The beneficial effects of the present invention are as follows: By combining the spatio-temporal graph structure and the deep learning model, the accuracy and efficiency of prediction are improved, and the problem that the traditional technology cannot well capture the spatio-temporal dependence relationship in traffic flow data is solved; At the same time, based on this, a traffic flow prediction method of a unified sentence vector large language model improved by an adaptive preference gate is designed, so that the adaptability and generalization ability of the model on various traffic flow data sets are significantly improved, and it has good scalability and universality, and can adapt to different spatio-temporal data prediction tasks. Description of the Drawings
[0078] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The same reference numerals are used to represent the same or similar parts in these drawings. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0079] Figure 1 is a flowchart of an embodiment of the present invention;
[0080] Figure 2 Principle diagram of the adaptive preference gate;
[0081] Figure 3 Principle diagrams of the reset gate, update gate and hidden state. Detailed Embodiments
[0082] As Figure 1 shown, the large model prediction method for fine-tuning the traffic flow map of a large-scale highway network includes the steps of:
[0083] S1: Construct a heterogeneous graph of the traffic flow of the highway network. The nodes in the heterogeneous graph are toll station nodes and intersection nodes, and the edges in the heterogeneous graph are the traffic volumes between each node; The heterogeneous graph is represented as:
[0084] G=(V,E)
[0085] where V is the set of nodes, including the toll station nodes U and intersection nodes I; E is the set of edges, including the traffic volume between the toll station node and the intersection node, the traffic volume between the toll station nodes, and the traffic volume between the intersection nodes.
[0086] The heterogeneous graph edge adjacency matrix is used to represent the edge relationship in the graph. There are N u toll station nodes and N i intersection nodes in the graph, and the following adjacency matrix is defined:
[0087] Intersection-intersection adjacency matrix
[0088] Toll station-intersection adjacency matrix where aui = 1 indicates that there is traffic flow between toll station u and intersection i.
[0089] Toll station - toll station adjacency matrix
[0090] S2: Use the graph convolutional network GCN to aggregate the neighbor information of nodes, capture the spatial characteristics of the traffic flow in the highway network, and calculate the spatial weights of the edges between nodes.
[0091] Based on the edge relationships of these heterogeneous graphs, this step uses the graph convolutional network (GCN) to aggregate neighbor information and update the representation of nodes. The update of the node representation can be achieved through the graph convolutional layer, and the update rule of the node representation (i.e., spatial characteristics) is as follows:
[0092]
[0093] where: σ is the activation function, is the set of neighbor nodes of node u, W uv is the spatial weight of the edge between node u and node v. Different types of edges have different weights, and these weights are obtained through spatial graph learning; h (l) is the input of the l - th layer of graph convolution.
[0094] S3: Add time information to the heterogeneous graph, and use the time convolutional network TCN combined with Chebyshev filtering and dilated convolution to capture the time characteristics of the traffic flow in the highway network, and calculate the time weights of the edges between nodes.
[0095] In existing traffic flow prediction models, the data used is usually dynamically small - scale, that is, all interaction behaviors between toll stations only change within a small range and still have the same importance in the long - term range. However, the road network traffic flow is dynamically changing, and the time factor plays a crucial role in traffic flow prediction. In the long - term range, the traffic flow will change significantly over time, so the time sensitivity of historical road network traffic flow needs to be considered.
[0096] The present invention adds time information (i.e., time stamps) to the graph structure to enhance the prediction accuracy. The goal is to predict the traffic flow at the next time step and for a period of time in the future given the traffic flow data of the past M time steps.
[0097] Input data
[0098] X = {x1, x2, …, x t …, x M}
[0099] where each in the X sequence is a one - dimensional input, C i is the input size of the time feature map, x tis the traffic flow feature at the t-th time step, including traffic flow data of different vehicle types.
[0100] Output data
[0101] Y = {y1, y2, …, y t …y M}
[0102] In the multi-dimensional tensor C o is the output size of the time feature map, where each y t is the predicted traffic flow corresponding to the time step.
[0103] Next, use TCN (Temporal Convolutional Network) to capture the temporal dependencies in the traffic flow data through convolutional operations. Apply the graph convolution G α to the multi-dimensional tensor input. For each of the M time steps t, perform ρ layers of temporal convolutions, which can be computed in parallel
[0104] For nodes u and v, the calculation method of the graph convolution G u,v is as follows:
[0105]
[0106] where: The softmax function is an activation function. For elements z i and z j in the tensor Z, the softmax function formula is:
[0107]
[0108] X = {x1, x2, …, x t …, x M} is the input data, is the heterogeneous graph adjacency matrix at the corresponding moment, and σ is the activation function (using ReLU as above).
[0109] is the edge weight of the l-th layer, is the edge weight of the (l + 1)-th layer:
[0110]
[0111] TCN will perform convolutional operations on this multi-dimensional tensor. This model uses dilated convolutions to capture temporal dependencies. The size of each convolutional kernel is k, and the dilation rate is d. First, use Chebyshev filters with good spectral characteristics to perform tensor filtering in cases with strict frequency response requirements.
[0112] The Chebyshev filter controls the gain characteristics of the filter by appropriately selecting the Chebyshev coefficients and minimizes the fluctuations in the frequency response. Chebyshev polynomials are a class of orthogonal polynomials commonly used for approximating functions and constructing filters.
[0113] Generally, the recurrence formula for the Chebyshev polynomials of the first kind T n (x) is:
[0114] T0(x) = 1
[0115] T1(x) = x
[0116] T n (x) = 2xT n-1 (x) - T n-2 (x), n ≥ 2
[0117] In filter design, Chebyshev polynomials are used to construct the frequency response, especially to control the "ripple" phenomenon in graph convolutional filtering, i.e., the gain fluctuations in the signal passband. For TCN, the Chebyshev Type I filter is used, and its transfer function is expressed as:
[0118]
[0119] where: ε is the ripple coefficient of the filter, which controls the gain fluctuations in the passband, is the nth-order Chebyshev polynomial, ω c is the cut-off frequency of the filter, and s is the input parameter of the Chebyshev filter.
[0120] Applied to this model, through passing parameters, H(s) is the result of the graph convolutional calculation H(G i,j ) as described below.
[0121] For a signal sequence input to the ρ-layer temporal convolution on C i channels, the weights of the convolutional layer are w i (learnable parameters), and the output after Chebyshev filtering is:
[0122]
[0123] G i,j corresponds to the generated graph convolution G u,v described above, x i and y i are the compositional parameters of xt and yt.
[0124] To capture dependencies over long time horizons, the TCN stacks multiple convolutional layers, each with a different dilation rate. Let the dilation rate of the l-th layer be d l , and the kernel function be According to the vectorized representation of the highway traffic flow corresponding to time in the time series of N as T0, T1, …, TN, then the operator can be expressed as:
[0125]
[0126] is the result after T passes through the F operator
[0127] Perform the dilated convolution operation on each element Tt:
[0128]
[0129] is the input at the (|Tt| - d)-th l t time steps, indicating the time steps skipped by the convolutional kernel in the dilated convolution. After performing the dilated convolution in this way, the corresponding dilated convolution for each time step predicts the traffic flow is calculated, that is:
[0130]
[0131] Then the time information will be combined with the representation of the point-edge interaction, is the time weight, is the dilated convolution representation vector of the h-th layer, is the Chebyshev filter representation vector of the h-th layer. As shown in the formula:
[0132]
[0133] The behaviors of different time series have different impacts on the prediction. Therefore, time weighting is used to weight the traffic flow. The closer the time of the traffic flow is to the current, the more it can represent the true value of the prediction for a period of time later. This is achieved by introducing a decay factor, as shown in the formula:
[0134]
[0135] represents the traffic flow prediction score, which is used to measure the correlation between specific time steps Tm and Tn. Intuitively, it represents the degree of influence of the traffic flow state at a certain historical time step m on the future time step n. Since the influence of the distance in time on the prediction is different, a time decay factor α is used to adjust this weight.
[0136] where α is the attenuation factor, representing the time distance of traffic from the current time. The attenuation method used in this method is exponential attenuation, that is:
[0137]
[0138] α m,n represents the attenuation factor of traffic flow in the time dimension.
[0139] δ mn represents the time difference between time steps m and n (i.e., the absolute value of the time interval);
[0140] η represents the parameter controlling time attenuation, similar to the role of standard deviation. The larger its value, the slower the time attenuation, and the historical data still has influence in a longer time. When its value is small, it means that the earlier historical data has less influence, and the model pays more attention to the recent data.
[0141] χ represents the time weight threshold, which sets a minimum weight threshold for filtering irrelevant time steps. Only when is greater than or equal to χ, it is considered that there is sufficient influence between time steps i and j, otherwise it is directly set to 0. This is equivalent to introducing a pruning mechanism to reduce the influence of irrelevant historical data on the current prediction and improve the calculation efficiency. The updated is as follows:
[0142]
[0143] In this way, the time weight of traffic flow can be combined with the spatial weight, so that this method can accurately make predictions that change over time.
[0144] S4: Concatenate the spatial features and time features to generate a spatial vector; concatenate the spatial weight and time weight to generate a time vector; then concatenate the spatial vector and time vector to construct a unified sentence vector; The unified sentence vector, Unified Sentence Vector (USV) is the framework of this model for fusing spatio-temporal graph embedding vectors. After being reset and updated by an adaptive preference gate bidirectional Bi-GRU (bidirectional gated recurrent unit), and then performing Seq-to-Seq instruction fine-tuning, a sentence vector with prediction information is generated.
[0145] The unified sentence vector uf is expressed as:
[0146] uf = [gf; tf]
[0147] where gf is the spatial vector, that is, after being processed by a multi-layer spatial heterogeneous graph, the spatial weight Wuv and the spatial feature representation h of the last layer are obtained (l+1), through the adaptive recurrent gate, the spatial vector gf in the USV splicing is obtained; the tf is the time vector, that is, through multi-layer time convolution processing, the time weights are obtained and the time feature representation Through the adaptive recurrent gate, the time vector tf in the USV splicing is obtained; [;] represents the concatenation operation of vectors.
[0148] The bidirectional adaptive recurrent gate Bi-GRU is a recurrent neural network (RNN) structure commonly used in natural language processing, which can capture the forward and backward dependencies in the sequence. Compared with the traditional GRU (gated recurrent unit), the adaptive Bi-GRU in this method enhances the memory ability of the model by running GRU units in two directions (forward and backward) respectively.
[0149] The adaptive preference gate consists of three gating mechanisms: the reset gate, the update gate, and the candidate hidden state. The update formula for the hidden state h t is the formula:
[0150]
[0151] where: z t is the update gate, which controls the combination of the previous hidden state and the current candidate hidden state. r t is the reset gate, which controls how to reset the previous hidden state. h t is the candidate hidden state, which is calculated through a non-linear activation function (tanh). w is the learnable parameter of Bi-GRU and is obtained during training.
[0152] In order to unify the time-space vectors and generate a format that can be used by the deep learning model, it is necessary to embed the tokens through pointwise convolution. In this model, the principle of token embedding is: converting the input data x i into the embedding
[0153] Emb i = Cov Point (x i , θ i )
[0154] where, Emb i represents the embedded token. Cov Point represents the pointwise convolution operation using a 1×1 convolution kernel. x iis the input data (i.e., the input vector for pointwise convolution, which is the spatial and temporal representation in the above text), C i is the input hidden dimension (consistent with the input size C of the temporal feature map in the previous text i in size), θ i represents the learnable parameter for each point in pointwise convolution.
[0155] For spatial correlation, the model uses adaptive vector embedding
[0156]
[0157] where σ represents the activation function, is the learnable parameter, Par S is the spatial vector embedding, that is, the embedded representation of the spatial vector generated corresponding to the spatial weight in the above text (which is an embedded representation, with the order changed).
[0158] To retain the time information in the unified sentence vector, a linear layer is used to encode the input data into the number of days in a month and the number of days in a week, and the two are independently embedded. Absolute position encoding is performed on each provincial traffic data at the resolutions of "month" and "week", and the generated position information is and the embedding of the number of days in a month and the embedding of the number of days in a week are calculated as follows:
[0159]
[0160] where and are the learnable time embeddings for the number of days in a week and the number of days in a month, that is, the embedded representations of the time vectors generated corresponding to the time weights in the above text. By adding these two embeddings, the time embedding representation
[0161] S5: Input the unified sentence vector into the large language model LLM, and perform traffic flow prediction through instruction fine-tuning and the partial attention layer freezing strategy.
[0162] Fine-tuning the large language model LLM requires dividing the input into two parts:
[0163] Spatio-temporal encoding part: The spatio-temporal weight matrix P (i.e., composed of the weights generated by the time-space graph in the above text), P = {W1|W2|…|W t |…|W T}.
[0164] Context part: The representation uf of the USV is a sequence u1, u2, …, u t , …, u T , and this sequence is formed into a historical traffic flow data matrix H = {u1|u2|…|u t |…|u T} by column vectors.
[0165] The final USV input sequence can be expressed as the formula:
[0166] X = Concat(P, H)
[0167] Concat(·) represents the matrix concatenation operation.
[0168] In a pre-trained large language model (LLM), to make it suitable for traffic prediction tasks, the model defines the time step at each position of the traffic flow data in the above method as a token, so that the spatio-temporal embedding layer converts these tokens into a spatio-temporal representation aligned with the LLM.
[0169] The Instruction Fine-Tuning of this model uses the LLM to guide the model to perform specific traffic flow prediction tasks, which can enhance the generalization ability and task understanding ability of the large model and optimize the effect of traffic flow prediction.
[0170] The LLM adopts an encoder-decoder architecture. The goal in the instruction fine-tuning stage is to adjust the model parameters to optimize its understanding ability of spatio-temporal encoding and context. The encoder converts the input sequence X into a context representation, as shown in the formula:
[0171] H enc = Encoder(X, θ enc )
[0172] where H enc is the encoded context hidden representation, and θ enc is the encoder parameter.
[0173] The decoder generates prediction results based on the encoder output, as shown in the formula:
[0174] Y = Decoder(H enc , Y <t ; θ dec )
[0175] where Y is the output sequence generated by the decoder, Y <t is the output of the decoder before time step t, and θ dec is the decoder parameter.
[0176] This method introduces FusionConv to project traffic features into the dimensions required by the LLM. FusionConv integrates spatial and temporal embeddings to represent USV uniformly:
[0177] K F = FConv(Y; θ f )
[0178] where K F represents the result of performing FusionConv on the unified sentence vector, and the resulting representation is used as the input to MHA and LN. θ f represents the learnable parameters of FConv.
[0179] This model is specifically designed with parameter freezing to improve the accuracy of traffic prediction by partially freezing a large language model (LLM). The multi-head attention layer (MHA) and the feed-forward network (FFN) are frozen during training because these layers contain the most important parts learned by the LLM. The model keeps the first F layers frozen and unfreezes the parameters of the last U layers of the multi-head attention layer because these attention layers effectively handle the spatio-temporal dependencies in the data. Therefore, the parameter-frozen LLM can adapt to the traffic flow prediction task while maintaining the knowledge acquired during pre-training.
[0180] Generally, layer normalization (LN) is placed at the input of each sub-module. LN is more suitable for vector inputs with unified representations compared to pre-activation in residual networks. This model also adds an additional layer normalization after the last multi-head attention layer.
[0181]
[0182] μ and σ represent the mean and standard deviation respectively, ⊙ represents element-wise multiplication, and γ and β are learnable scaling and translation parameters. is the output after the first LN, and then is fed into the MHA multi-head attention:
[0183]
[0184] Each head i is computed by the attention mechanism:
[0185]
[0186] W are the parameters of the attention mechanism and are learned during training.
[0187] Attention(·) is calculated as follows:
[0188]
[0189] For Through the second LN, we get
[0190]
[0191] Next Enter the feed-forward network FFN:
[0192]
[0193] In the first F layers of the frozen parameter LLM, the multi-head attention and the feed-forward layer are frozen. First, calculate through MHA and then add the fused USV:
[0194]
[0195] Then, through normalization and FFN, and then add the one that has passed through the frozen MHA Get the output K to the next layer i+1 :
[0196]
[0197] where i ∈ {1, …, F - 1}, and at the same time K i+1 = [K F + PosEmb], and PosEmb represents the learnable position encoding. represents the intermediate representation of the i-th layer after applying the frozen multi-head attention (MHA) and the first unfrozen layer normalization (LN). K i represents the final representation after applying the unfrozen LN and the frozen feed-forward network (FFN).
[0198] Similarly, unfreeze the MHA in the last U layers of the LLM, so that the model can capture the spatio-temporal dependence of the traffic flow data:
[0199]
[0200] where represents the intermediate representation of the F + U - 1 layer after applying the unfrozen MHA and the second frozen LN, and K F+U represents the final output after applying the unfrozen LN and the frozen FFN.
[0201] 2.3 Learning and Optimization Methods
[0202] This model designs a new loss function as shown in the equation:
[0203]
[0204] where L soft such as Equation (1-22):
[0205]
[0206] The regularization term λ∑ k∈θκ 2 is L2 regularization, which penalizes the model parameters to prevent overfitting. This can control the model complexity, limit the magnitude of the parameters, effectively avoid overfitting on the training set, while enhancing the generalization ability of the model, making it more robust on the test set and improving its prediction ability for unseen data. The Gaussian distribution function is selected as the soft label optimization objective, as shown in the equation:
[0207]
[0208] where θ are the learnable parameters of the model, λ is the regularization hyperparameter, y is the true value of the training data, is the smoothed true value of the v-th prediction target, is the predicted probability of the v-th target, σ is the label smoothing factor, and |V| is the number of graph point sets. ω is the coefficient that adjusts the overall loss function and "penalizes" the negative sampling pairs.
[0209] Label smoothing avoids overfitting caused by the model relying too much on "hard labels" or "absolute labels". This model smooths the true label y using the Gaussian distribution factor σ. In previous prediction tasks, the labels were hard, while Gaussian label smoothing transforms these hard labels into a distribution, so that the model no longer completely depends on the exact labels, but on the soft label y S . Through such smoothing, the prediction results of the model are more "fuzzy" and more robust, reducing the dependence on hard labels. Especially when dealing with noisy data, the model can better generalize when facing incomplete or sparse data.
[0210] σ controls the degree of smoothing. A smaller σ value indicates a lower degree of label smoothing, and the actual label is closer to the hard label; a larger σ value means the label is more "relaxed", and the model accepts a wider range of labels. This makes the model not overly dependent on extreme labels during training and avoids overfitting. To reduce the overfitting of the model to noise, since the data is usually sparse, traffic flow data may not be completely accurate or may contain noise. Label smoothing can reduce the over-dependence of the model on noisy labels by introducing a certain degree of uncertainty for each label, thereby enhancing the generalization ability of the model.
[0211] L sam The loss function is as shown in the equation:
[0212]
[0213] Where: cos(·,·) is a distance function, and the calculation formulas for the u-th and i-th prediction targets are as follows:
[0214]
[0215] Cosine similarity is insensitive to the length of vectors and only focuses on directions. Therefore, it is suitable for vectorized representation in traffic flow prediction, that is, the similarity of embedded vectors in a high-dimensional space. is an indicator function, and the calculation formulas for the true values of the smoothed u-th and i-th prediction targets are as follows:
[0216]
[0217] In the L sam loss function, where N r is the number of random samples in for is the node representation of the heterogeneous graph in the last hidden layer, As an indicator function, when it indicates that and have different labels, and when it indicates that and have the same label
[0218] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not restrictive. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A prediction method for fine-tuning large models of traffic flow maps of large-scale highway networks, characterized in that, Including the steps: S1: Construct a heterogeneous graph of highway network traffic flow. The nodes in the heterogeneous graph are toll station nodes and intersection nodes, and the edges in the heterogeneous graph are the traffic flows between each node. S2: Use the graph convolutional network GCN to aggregate the neighbor information of nodes, capture the spatial characteristics of highway network traffic flow, and calculate the spatial weights of the edges between nodes. S3: Add time information to the heterogeneous graph, and use the temporal convolutional network TCN combined with Chebyshev filtering and dilated convolution to capture the temporal characteristics of highway network traffic flow, and calculate the temporal weights of the edges between nodes. S4: Concatenate the spatial features and temporal features to generate a unified sentence vector. S5: Input the unified sentence vector into the large language model LLM, and perform traffic flow prediction through instruction fine-tuning and partial attention layer freezing strategy.
2. The large-scale highway network traffic flow map fine-tuning large model prediction method according to claim 1, wherein The heterogeneous graph is represented as: G=(V,E) where V is the set of nodes, including the toll station nodes U and intersection nodes I; E is the set of edges, including the traffic flows between toll station nodes and intersection nodes, the traffic flows between toll station nodes, and the traffic flows between intersection nodes.
3. The large-scale highway network traffic flow map fine-tuning large model prediction method according to claim 2, characterized in that The spatial features output by each graph convolutional layer of the graph convolutional network GCN are represented as: where: σ is the activation function, is the set of neighbor nodes of node u, W uv is the spatial weight of the edge between node u and node v; h (l) is the input of the l-th layer graph convolution.
4. The large-scale highway network traffic flow map fine-tuning large model prediction method according to claim 3, wherein, The temporal features output by each convolutional layer of the temporal convolutional network TCN are represented as: Wherein: is the time weight of the edge between node u and node v, is the dilated convolution representation vector, is the traffic flow prediction score at the |T t |-d l t time steps, y t is the Chebyshev filter representation vector; The dilated convolution representation vector is expressed as: Where: x t is the traffic flow feature at the t-th time step; d l is the dilation rate of the l-th layer; is the kernel function; T t is the vectorized representation of the time corresponding to the traffic flow of the highway network in the time series of t; is the input at the |T t |-d l t time steps, indicating the time steps skipped by the convolutional kernel in the dilated convolution; The temporal convolutional network TCN uses the first-kind Chebyshev filter for tensor filtering, and the transfer function of the first-kind Chebyshev filter is represented as: where ε is the ripple coefficient of the filter, which is used to control the gain fluctuation in the passband; is the nth-order Chebyshev polynomial, ω c is the cut-off frequency of the filter, and s is the input parameter of the Chebyshev filter; By passing parameters, H(s) is the result of graph convolution calculation H(G i,j ); For the signal sequence of the ρ-layer temporal convolution input on C i channels The weight of the convolutional layer is w i The output after Chebyshev filtering is: G i,j is for graph convolution; The traffic flow prediction score is expressed as: Where: α is the decay factor of traffic flow in the time dimension; is the time difference between time step t and |T t |-d l t; η is the parameter for controlling time decay; χ is the time weight threshold.
5. The large-scale highway network traffic flow map fine-tuning large model prediction method according to claim 4, characterized in that The unified sentence vector u f is expressed as: u f = [g f ; t f where g f is the vector representation of the spatial feature h (l+1) ; the t f is the vector representation of the temporal feature ; [;] represents the concatenation operation of vectors.
6. The large-scale highway network traffic flow map fine-tuning large model prediction method according to claim 5, characterized in that, The method further includes embedding the tokens through pointwise convolution; for the spatial vectors, an adaptive vector embedding is adopted : Among them, σ represents the activation function, is a learnable parameter, Par S is the spatial vector embedding; For the time vector, a linear layer is used to encode the input data into the day of the month and the day of the week. These two are independently embedded, and then the two embeddings are added together to obtain the time embedding representation Embedding the number of days in January is as follows: Embedding of the number of days in a week Among them, and are the position information generated by performing absolute position encoding on each provincial traffic data at the resolutions of "month" and "week"; and are learnable time embeddings for the number of days in a week and the number of days in a month.
7. The large-scale highway network traffic flow map fine-tuning large model prediction method according to claim 6, characterized in that, The method further includes using an adaptive preference gate bidirectional gated recurrent unit to process the input u from front to back and from back to front respectively f sequence u1, u2, …, u t , …, u T and the forward sequence output is: The reverse sequence output is: The hidden state is obtained by concatenating the hidden states of the forward sequence output and the reverse sequence output It is as follows: Then dynamically adjust the attention degree to different features, introduce an additional gating mechanism. At the t-th moment, the update process of the preference gate is the formula: g t = σ(W g · u t + U g · h t-1 ) where g t is the current time, W g and U g are learnable weights, and x t is the input at the current time.
8. The large-scale expressway network traffic flow map fine-tuning large model prediction method according to claim 7, characterized in that, The large language model LLM adopts an encoder-decoder architecture; The encoder converts the input sequence X into a context representation through instruction fine-tuning: H enc = Encoder(X, θ enc ) Among them, H enc is the encoded context hidden representation, and θ enc is the encoder parameter; X = Concat(P,H) Among them, P is a matrix composed of time vectors, expressed as P = {W1|W2|…|W t |…|W T}; H is a historical traffic data matrix formed by arranging the u f sequence in column vectors, H = {u1|u2|…|u t |…|u T}; The decoder predicts the result according to the output of the encoder: Y = Decoder(H enc , Y <t ; θ dec ) where Y is the output sequence of the decoder, and Y <t is the output of the decoder before time step t, and θ dec are the decoder parameters.
9. The large-scale highway network traffic flow map fine-tuning large model prediction method according to claim 8, wherein, During the training of the large language model LLM, the multi-head attention layer and the feed-forward network of the large language model LLM are frozen during training, so that the first F layers of the large language model LLM are frozen, and the parameters of the last U layers of the multi-head attention layer are unfrozen.
10. The large-scale highway network traffic flow map fine-tuning large model prediction method according to claim 9, characterized in that, The loss function of the large language model LLM is: Among them, the regularization term λ∑ κ∈θ κ 2 is L2 regularization, θ is the learnable parameter of the model, λ is the regularization hyperparameter, y is the true value of the training data, is the smoothed true value of the v-th prediction target, is the predicted probability of the v-th target, σ is the label smoothing factor, |V| is the number of graph point sets; ω is the coefficient that adjusts the overall loss function and "penalizes" the negative sampling pairs; where: cos(·,·) is the distance function, and the calculation formulas for the u-th and i-th prediction targets are: For the indicator function, the calculation formula for the true values of the smoothed $u$-th and $i$-th prediction targets is as follows: Among them, N r is the number of random sampling pairs in is the point representation of the heterogeneous graph in the last hidden layer, as an indicator function, when then it means and have different labels; when then it indicates and have the same label.
Citation Information
Patent Citations
Multi-source traffic data processing method and device and electronic equipment
CN117931977A
Urban traffic flow prediction method and system fusing spatial-temporal characteristics and large language model
CN119649599A
Short-term traffic flow prediction method based on causal gated-low-pass graph convolutional network
US20240029556A1
Cited By
Large-scale road network traffic control method based on deep reinforcement learning large model
CN120954238A
Large-scale road network traffic control method based on deep reinforcement learning large model
CN120954238B