Traffic flow prediction method and system based on space-time large language model, and medium
By using the large language model of space-time in traffic flow prediction, combining the spatiotemporal encoder and alignment module of recurrent neural network and graph neural network, the hallucination problem of large language model in traffic flow prediction is solved, and the accuracy and interpretability of prediction are improved.
Patent Information
- Application Number
- CN202510008362.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-03
AI Technical Summary
The existing traffic flow prediction methods based on large language models have hallucinations problems, especially in the case of sparse data and low-quality data, resulting in low prediction accuracy.
The traffic flow prediction method based on the large language model of space-time is adopted, and the spatiotemporal encoder is constructed through recurrent neural networks and graph neural networks, and the predicted traffic flow data is encoded to obtain the first spatiotemporal features, and the feature alignment is performed through the preset alignment module to obtain the second spatiotemporal features. Then, an action chain instruction is constructed, and the traffic flow prediction results are generated based on the second space-time characteristics and action chain instructions through a large language model.
By designing spatiotemporal encoder and alignment modules, we help large language models understand and explore complex spatiotemporal patterns, alleviate illusion problems, realize accurate spatiotemporal prediction under frozen large language models, and improve the accuracy of traffic flow prediction.
Smart Images

Figure CN120071606A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a traffic flow prediction method, system, and medium based on a spatio-temporal large language model. Background Art
[0002] With the rapid development of the Internet of Vehicles, the penetration rate of autonomous driving is getting higher and higher, showing great potential for achieving high-level intelligent transportation. The realization of high-level intelligent transportation depends on the accurate perception of the dynamic nature of the large-scale environment. Fully understanding the dynamic characteristics of the traffic system that change continuously over time and space, predicting and analyzing the spatio-temporal patterns and trends of the traffic system is of great significance to urban intelligence. Specifically, accurate traffic flow prediction can effectively improve the driving efficiency of vehicles and reduce congestion to achieve high-quality intelligent transportation.
[0003] With the development of deep learning, traffic flow prediction has evolved from using convolutional neural networks for region-grid-based traffic flow prediction to graph neural network (GNN-based) traffic flow prediction that considers the non-Euclidean distance of the road network, indicating an increasing emphasis on capturing network structure changes in traffic flow prediction tasks. In particular, the spatial graph structure of traffic road networks shows significant changes at different time periods, and the property differences of the same traffic area at different times affect the surrounding traffic areas and even the whole, which has been proven in many studies. At the same time, the large language model (LLM) with powerful learning ability and generalization has gradually become a research hotspot, and adding GNN-based to the large language model significantly improves the accuracy of spatio-temporal prediction. However, integrating GNN-based into the large language model and applying it to traffic flow prediction still poses challenges.
[0004] While the large language model has achieved amazing results in various industries, there is inevitably the problem of hallucinations. Problems such as data sparsity and data quality cause the general large language model to be prone to the problem of "talking nonsense seriously" in different downstream tasks. Specifically, reasoning using the large language model directly for spatio-temporal prediction will produce hallucinations, and partial fine-tuning of the large language model in the absence of high-quality data will also produce hallucinations, but full fine-tuning will incur unaffordable computational costs. Summary of the Invention
[0005] To solve the above technical problems, the purpose of the present invention is to provide a traffic flow prediction method, system, and medium based on a spatio-temporal large language model with high accuracy.
[0006] To achieve the above purpose, one aspect of the embodiments of this application proposes a traffic flow prediction method based on a spatio-temporal large language model, including the following steps:
[0007] Obtain the traffic flow data to be predicted;
[0008] Construct a spatio-temporal encoder based on a recurrent neural network and a graph neural network, and encode the traffic flow data to be predicted through the spatio-temporal encoder to obtain first spatio-temporal features;
[0009] Perform feature alignment on the first spatio-temporal features through a preset alignment module to obtain second spatio-temporal features, so as to align the dimensions of the spatio-temporal encoder and the large language model;
[0010] Construct an action chain instruction, and obtain a traffic flow prediction result through the large language model according to the second spatio-temporal features and the action chain instruction.
[0011] In some embodiments, the encoding of the traffic flow data to be predicted through the spatio-temporal encoder to obtain first spatio-temporal features specifically includes:
[0012] Input the traffic flow data to be predicted into the spatio-temporal encoder;
[0013] Perform spatial encoding on the traffic flow data to be predicted through the graph neural network to obtain spatial features;
[0014] Perform temporal encoding on the traffic flow data to be predicted through the recurrent neural network to obtain temporal features;
[0015] Perform non-linear connection on the spatial features and the temporal features to obtain the first spatio-temporal features.
[0016] In some embodiments, the encoding of the traffic flow data to be predicted through the recurrent neural network to obtain temporal features specifically includes:
[0017] Perform feature extraction on the traffic flow data to be predicted through the gated dilated convolutional layer in the recurrent neural network to obtain multi-scale temporal features;
[0018] Perform correlation merging on the multi-scale temporal features through the cross-scale correlation layer in the recurrent neural network to obtain the temporal features.
[0019] In some embodiments, the performing feature alignment on the first spatio-temporal features through a preset alignment module to obtain second spatio-temporal features specifically includes:
[0020] Input the first spatio-temporal features into the alignment module, and then project the first spatio-temporal features to obtain spatio-temporal tokens;
[0021] Perform mapping analysis on the spatio-temporal tokens to obtain a hidden representation;
[0022] Map and concatenate the spatio-temporal marker and the hidden representation to obtain the second spatio-temporal feature.
[0023] In some embodiments, the constructing action chain instruction specifically includes:
[0024] Construct a logical guidance template and a step-by-step decomposition guidance template;
[0025] Obtain regional information, time information, and weather data from the traffic flow data to be predicted, and determine marker characters;
[0026] Construct the action chain instruction according to the logical guidance template, the step-by-step decomposition guidance template, the regional information, the time information, the weather data, and the marker characters.
[0027] In some embodiments, the multi-scale time feature is obtained by the following formula:
[0028]
[0029] Where, represents the multi-scale time feature, represents the convolutional kernel and and represent the one-dimensional convolutional kernel and and represent the bias and k represents the number of one-dimensional convolutional kernels and k = [1, 2, 3, 4], l represents the number of layers, D f represents the number of the first channels, D tr represents the number of the second channels, represents the spatial feature, [·] represents the splicing operation, ⊙ represents the Hadamard product, σ 3 and σ 4 represent the activation function.
[0030] In some embodiments, the spatio-temporal marker and the hidden representation are mapped and concatenated by the following formula to obtain the second spatio-temporal feature:
[0031]
[0032] Where, represents the second spatio-temporal feature, Ξ pr represents the spatio-temporal marker, Γ hi represents the hidden representation and C represents the initial number of channels, N represents the number of spatial nodes, D L represents the number of the third channels, σ 5 represents the activation function, W 1 、W2 and W 3 represents a mapping operation and D b represents the number of the fourth channel, β represents the output step size, and [·,·] represents a concatenation operation.
[0033] To achieve the above object, another aspect of the embodiments of the present application provides a traffic flow prediction system based on a spatio-temporal large language model, including:
[0034] A data acquisition module, configured to acquire traffic flow data to be predicted;
[0035] A spatio-temporal feature encoding module, configured to construct a spatio-temporal encoder based on a recurrent neural network and a graph neural network, and encode the traffic flow data to be predicted through the spatio-temporal encoder to obtain first spatio-temporal features;
[0036] A feature alignment module, configured to perform feature alignment on the first spatio-temporal features through a preset alignment module to obtain second spatio-temporal features, so as to align the dimensions of the spatio-temporal encoder and the large language model;
[0037] An instruction construction and result prediction module, configured to obtain a traffic flow prediction result through the large language model according to the second spatio-temporal features and the action chain instructions.
[0038] To achieve the above object, another aspect of the embodiments of the present application provides an electronic device, where the electronic device includes a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing connection communication between the processor and the memory, and when the program is executed by the processor, it implements the traffic flow prediction method based on the spatio-temporal large language model as described above.
[0039] To achieve the above object, another aspect of the embodiments of the present application provides a storage medium, where the storage medium is a computer-readable storage medium for computer-readable storage, and the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the traffic flow prediction method based on the spatio-temporal large language model as described above.
[0040] The beneficial effects of the present invention are as follows: The traffic flow prediction method and system based on the spatio-temporal large language model of the present invention first obtain the traffic flow data to be predicted, then construct a spatio-temporal encoder based on a recurrent neural network and a graph neural network, encode the traffic flow data to be predicted through the spatio-temporal encoder to obtain the first spatio-temporal features, and then perform feature alignment on the first spatio-temporal features through a preset alignment module to obtain the second spatio-temporal features. Finally, an action chain instruction is constructed, and the traffic flow prediction result is obtained through the large language model according to the second spatio-temporal features and the action chain instruction. The present invention designs a spatio-temporal encoder to help the large language model fully understand and mine complex spatio-temporal patterns in different scenarios, and then introduces an action chain instruction and an alignment module to alleviate the hallucination of the large language model, enabling accurate spatio-temporal prediction under the frozen large language model and improving the accuracy of traffic flow prediction based on the large language model. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following introduces the drawings required to be used in the embodiments of the present invention. It should be understood that the drawings introduced below are only for conveniently and clearly expressing some embodiments of the technical solutions in the present invention, and those skilled in the art can also obtain other drawings based on these drawings without creative efforts.
[0042] Figure 1 It is a flowchart of the steps of a traffic flow prediction method based on a spatio-temporal large language model provided by an embodiment of the present invention;
[0043] Figure 2 It is a schematic flowchart of a traffic flow prediction method based on a spatio-temporal large language model provided by an embodiment of the present invention;
[0044] Figure 3 It is a schematic structural diagram of a traffic flow prediction system based on a spatio-temporal large language model provided by an embodiment of the present invention;
[0045] Figure 4 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] In order to make the objectives, technical solutions and advantages of the present application clearer, the following further details the present application in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. When the following description involves the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application detailed in the appended claims.
[0047] It can be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if", "when" as used herein may be interpreted as "when...", "while...", or "in response to determining".
[0048] The terms "at least one", "multiple", "each", "any one", etc. used in this application, at least one includes one, two or more than two, multiple includes two or more than two, each refers to each of the corresponding multiple, and any one refers to any one of the multiple.
[0049] With the rapid development of the vehicle-to-everything (V2X) network, the penetration rate of autonomous driving is getting higher and higher, showing great potential for achieving high-level intelligent transportation. The realization of high-level intelligent transportation depends on the accurate perception of the dynamic nature of the large-scale environment. Fully understanding the continuously changing dynamic characteristics of the traffic system across time and space, predicting and analyzing the spatio-temporal patterns and trends of the traffic system is of great significance to urban intelligence. Specifically, accurate traffic flow prediction can effectively improve the driving efficiency of vehicles and reduce congestion to achieve high-quality intelligent transportation.
[0050] With the development of deep learning, from traffic flow prediction based on convolutional neural networks for regional grids to traffic flow prediction using graph neural network (GNN-based) considering the non-Euclidean distance of the road network, it shows the increasing attention of the traffic flow prediction task to capturing the changes in the network structure. In particular, the spatial graph structure of the traffic road network shows significant changes at different time periods, and the property differences of the same traffic area at different times affect the surrounding traffic areas and even the whole, which has been proved in many studies. At the same time, the large language model (LLM) with powerful learning ability and generalization has gradually become a research hotspot. Incorporating GNN-based into the large language model significantly improves the accuracy of spatio-temporal prediction. However, there are still challenges in integrating GNN-based into the large language model and applying it to traffic flow prediction.
[0051] While large language models have achieved amazing results in various industries, the problem of hallucinations inevitably exists. Problems such as data sparsity and data quality lead to the problem that general large language models are prone to "talking nonsense seriously" in different downstream tasks. Specifically, reasoning using large language models directly for spatio-temporal prediction will produce hallucinations, and partial fine-tuning of large language models in the absence of high-quality data will also produce hallucinations, but full fine-tuning will cause unaffordable computational costs.
[0052] To this end, the embodiments of the present invention propose a traffic flow prediction method based on a spatio-temporal large language model. First, traffic flow data to be predicted is obtained. Then, a spatio-temporal encoder is constructed based on a recurrent neural network and a graph neural network. The traffic flow data to be predicted is encoded by the spatio-temporal encoder to obtain first spatio-temporal features. Furthermore, a preset alignment module is used to perform feature alignment on the first spatio-temporal features to obtain second spatio-temporal features. Finally, an action chain instruction is constructed, and the large language model is used to obtain a traffic flow prediction result according to the second spatio-temporal features and the action chain instruction. The present invention designs a spatio-temporal encoder to help the large language model fully understand and mine complex spatio-temporal patterns in different scenarios, and then introduces an action chain instruction and an alignment module to alleviate the hallucinations of the large language model, enabling accurate spatio-temporal prediction under the condition of freezing the large language model and improving the accuracy of traffic flow prediction based on the large language model.
[0053] Refer to Figure 1 , Figure 1 FIG. is a flowchart of the steps of a traffic flow prediction method based on a spatio-temporal large language model provided by an embodiment of the present invention. The embodiments of the present invention propose a traffic flow prediction method based on a spatio-temporal large language model, and the method includes steps S101 to S104:
[0054] S101. Obtain traffic flow data to be predicted;
[0055] S102. Construct a spatio-temporal encoder based on a recurrent neural network and a graph neural network, and encode the traffic flow data to be predicted by the spatio-temporal encoder to obtain first spatio-temporal features;
[0056] Further as an optional implementation manner, the step of encoding the traffic flow data to be predicted by the spatio-temporal encoder to obtain first spatio-temporal features can be specifically divided into the following steps S1021 to S1024:
[0057] S1021. Input the traffic flow data to be predicted into the spatio-temporal encoder;
[0058] S1022. Perform spatial encoding on the traffic flow data to be predicted by the graph neural network to obtain spatial features;
[0059] In some alternative embodiments, although large language models exhibit remarkable proficiency in language processing, they still face challenges in understanding the evolving patterns inherent in spatio-temporal data. To overcome this limitation, embodiments of the present invention first utilize a graph neural network (GNN) to extract spatial features G, which can be specifically formulated as the following equation:
[0060]
[0061] where X represents traffic flow, W k represents learnable parameters, represents an adaptive adjacency matrix, and σ represents an activation function. The extracted spatial features include feature information such as the relationships between traffic flows at different intersections.
[0062] S1023. Perform temporal encoding on the traffic flow data to be predicted through a recurrent neural network to obtain temporal features;
[0063] Specifically, after extracting the spatial feature representation, the spatio-temporal encoder helps the large language model capture and understand temporal features through temporal encoding. Embodiments of the present invention adopt an integrated multi-scale and multi-core temporal convolutional network to further extract temporal features of the traffic flow X, so as to effectively capture the complex temporal dependencies between different temporal resolutions and enhance the large language model's ability to understand the redundant temporal dynamics in spatio-temporal data.
[0064] Further as an alternative implementation manner, step S1023 can be specifically divided into the following steps S10231 and S10232:
[0065] S10231. Extract features from the traffic flow data to be predicted through a gated dilated convolutional layer in a recurrent neural network to obtain multi-scale temporal features;
[0066] Further as an alternative implementation manner, the multi-scale temporal features are obtained through the following equation:
[0067]
[0068] where, represents the multi-scale temporal features, represents the convolutional kernel and and represent one-dimensional convolutional kernels and and represent the biases and k represents the number of one-dimensional convolutional kernels and k = [1, 2, 3, 4], l represents the number of layers, D f represents the first channel number, D tr represents the second channel number, represents the spatial feature, [·] represents the concatenation operation, ⊙ represents the Hadamard product, σ 3 and σ 4 Represents the activation function.
[0069] Specifically, the temporal encoding consists of two key components: a gated dilated convolutional layer and a cross-scale association layer. The gated dilated convolutional layer is formalized as follows:
[0070]
[0071] Apply four one-dimensional convolution kernels and as well as and Extract temporal features at multiple scales. In the feature extraction process, [·] converts the output of four one-dimensional convolution kernels from Splice to Where N represents the number of spatial nodes, the activation function and the Hadamard product ⊙ are used to control the information retention rate. Residual connections between different layers and convolution kernels It is used to effectively capture the temporal dependencies of various scales, thereby enriching the temporal representation of the spatial feature G and avoiding the gradient disappearance (G tr The resulting multi-scale temporal features can reflect the temporal evolution perceived at different scales.
[0072] S10232. Correlation merging of multi-scale time features is performed through a cross-scale correlation layer in a recurrent neural network to obtain a time feature.
[0073] Specifically, by introducing a cross-scale correlation layer, multi-scale temporal features are merged The correlation between different scales in , thereby retaining the feature information of each scale, can be formalized as follows:
[0074]
[0075] in, represents the convolution kernel and b s Indicates bias and D s and D tr Each indicates a different number of channels.
[0076] S1024: nonlinearly connect the spatial feature and the temporal feature to obtain a first spatiotemporal feature.
[0077] Specifically, after the l-layer temporal encoding, the expressions of the above gated dilated convolutional layer and the cross-scale association layer are combined through nonlinear connections to obtain the final first spatiotemporal feature
[0078] It should be noted that through the above modeling of space and time, the embodiments of the present invention can, according to different traffic flow prediction tasks or traffic flow prediction scenarios, use an adaptive explicit graph structure and spatial information enhancement, and then fully analyze the time dependencies at different time scales, providing strong support for the large language model to learn different spatio-temporal evolution patterns.
[0079] S103. Align the features of the first spatio-temporal features through a preset alignment module to obtain second spatio-temporal features, so as to align the dimensions of the spatio-temporal encoder and the large language model;
[0080] It should be noted that effectively aligning spatio-temporal information with text can play a key role in enabling the large language model to understand spatio-temporal evolution patterns. The embodiments of the present invention fuse different evolution patterns through an alignment module, integrate the context features of text and spatio-temporal data, and use lightweight token alignment to extract more expressive and interpretable multi-level semantic features.
[0081] Further as an optional implementation manner, the step of aligning the features of the first spatio-temporal features through a preset alignment module to obtain second spatio-temporal features can be specifically divided into the following steps S1031 to S1033:
[0082] S1031. Input the first spatio-temporal features into the alignment module, and then project the first spatio-temporal features to obtain spatio-temporal tokens;
[0083] Specifically, the embodiments of the present invention adopt an indirect prediction strategy, that is, by generating prediction tokens that contribute to the prediction process, and then analyzing the hidden representation of the tokens through mapping to change the result distribution to generate more accurate prediction values, avoiding the misperception caused by directly predicting future spatio-temporal values. First, project the first spatio-temporal feature S r as follows:
[0084] Ξ pr = S r W p + b p
[0085] where Ξ pr represents the spatio-temporal token and W p represents the mapping operation and b p represents the bias and
[0086] S1032. Perform mapping analysis on the spatio-temporal tokens to obtain the hidden representation;
[0087] S1033. Perform mapping and concatenation on the spatio-temporal tokens and the hidden representation to obtain the second spatio-temporal features.
[0088] As a further optional implementation, the spatio-temporal markers and hidden representations are mapped and concatenated through the following formula to obtain the second spatio-temporal feature:
[0089]
[0090] Among them, represents the second spatio-temporal feature, Ξ pr represents the spatio-temporal marker, Γ hi represents the hidden representation and C represents the initial number of channels, N represents the number of spatial nodes, D L represents the third number of channels, σ 5 represents the activation function, W 1 、W 2 and W 3 represent the mapping operation and D b represents the fourth number of channels, β represents the output step size, and [·,·] represents the concatenation operation.
[0091] Specifically, there are two challenges in the instruction fine-tuning large language model used in the embodiments of the present invention. First, its spatio-temporal prediction usually depends on numerical data, and its structure and pattern focus on semantic and syntactic relationships, which are different from the natural language that the large language model is good at processing; second, the large language model usually uses multi-classification loss to predict vocabulary, thus generating a probability distribution of potential results, rather than the continuous value distribution required for the prediction task. Therefore, the embodiments of the present invention adopt an indirect prediction strategy, and use this method to utilize the rich spatio-temporal context features hidden in the spatio-temporal markers to capture dynamic spatio-temporal dependencies, thereby helping the large language model to achieve more accurate spatio-temporal prediction.
[0092] S104. Construct an action chain instruction, and obtain a traffic flow prediction result through the large language model according to the second spatio-temporal feature and the action chain instruction.
[0093] It should be noted that since both time and space information contain rich and valuable semantic details, it helps the large language model to understand the spatio-temporal patterns in a specific context. For example: There are also significant differences in the traffic patterns between different times and different regions in the traffic system. The introduction text of time and space information can help the large language model better analyze these differences, but problems such as data sparsity will affect the expression of information, thereby generating hallucinations and misleading the large language model's understanding of spatio-temporal evolution. Therefore, how to accurately represent space and time information through text has a great impact on the performance of the large language model.
[0094] As a further optional implementation, the step of constructing the action chain instruction can be specifically divided into the following steps S1041 to S1043:
[0095] S1041. Construct a logical guidance template and a step-by-step decomposition guidance template;
[0096] Specifically, the embodiments of the present invention design a brand-new instruction text named CoA. First, a logical guidance template "Let us predict step by step" is constructed to help the large language model gradually decompose the prediction problem. Then, a detailed step-by-step decomposition guidance template ("Separately. Then,") is provided to the large language model to calibrate the thinking process of the model. Finally, the prediction steps are decomposed to construct a complete prediction action chain for the large language model. By constructing the logical guidance template and the step-by-step decomposition guidance template, the embodiments of the present invention can improve the interpretability of the large language model while alleviating the hallucination of the large language model.
[0097] S1042. Obtain regional information, time information, and weather data from the traffic flow data to be predicted, and determine the marking characters;
[0098] S1043. Construct an action chain instruction according to the logical guidance template, the step-by-step decomposition guidance template, the regional information, the time information, the weather data, and the marking characters.
[0099] Specifically, multi-scale time information markers (such as the spatio-temporal marker Ξ pr ) and regional information introductions are integrated into the instruction text of the large language model. The time information includes the specific moments corresponding to different time steps and the day of the week. The regional information includes cities, specific administrative regions, and point-of-interest (POI) data near different nodes, etc. In addition, weather data is added for assistance. By integrating these different information, the embodiments of the present invention construct an action chain instruction to help CoA encapsulate these understandings in a complex spatio-temporal context and enhance the ability of the large language model to recognize and assimilate spatio-temporal evolution paradigms in different regions and time ranges.
[0100] Furthermore, the embodiments of the present invention use special markers such as <Ca_start>, <Ca_HIS>,..., <Ca_HIS>, and <Ca_end> in the instruction to help the large language model accurately identify different parts and reduce the ambiguity in input and output, thereby improving the understanding ability of the model. Among them, <Ca_start> and <Ca_end> respectively represent the start and end identifiers of the spatio-temporal marker, and <Ca_HIS> represents the placeholder of the spatio-temporal marker. Based on the above alignment, the large language model can effectively distinguish different spatio-temporal evolution patterns, thereby improving its spatio-temporal prediction ability in the traffic flow prediction scenario.
[0101] In summary, the processing flow of the traffic flow prediction method based on the spatio-temporal large language model in the embodiments of the present invention is as Figure 2As shown, first, the traffic flow data to be predicted is received, and then a spatio-temporal encoder including a recurrent neural network and a graph neural network is used to encode the temporal and spatial information to enrich the expression of spatio-temporal features and learn the relative order relationships in the time series and spatial nodes. After encoding the input traffic flow data to be predicted, action chain instructions are used to force the large language model to decompose the traffic flow prediction process. Secondly, an alignment module is added between the spatio-temporal encoder and the large language model to extract more expressive and interpretable multi-layer semantic features and ensure the consistency among the spatio-temporal encoder, the action chain instructions, and the large language model. Finally, the large language model generates the traffic flow prediction result.
[0102] The above describes the traffic flow prediction method based on the spatio-temporal large language model of the embodiments of the present invention. It can be recognized that compared with the traffic flow prediction methods in the prior art, the embodiments of the present invention have the following advantages:
[0103] First, by encoding the temporal and spatial information in the traffic flow data to be predicted through a spatio-temporal encoder, the expression of spatio-temporal features can be enriched more, and while taking into account generalization, the prediction performance of the large language model is improved;
[0104] Second, an alignment module is added to the spatio-temporal encoder and the large language model, and an indirect prediction strategy is adopted. By generating spatio-temporal markers that contribute to the prediction process, and then analyzing the hidden representations of the spatio-temporal markers through mapping to change the result distribution to generate more accurate prediction values, avoiding the misperception caused by directly predicting future spatio-temporal values. Through this method, the rich spatio-temporal context features hidden in the spatio-temporal markers are utilized to capture the dynamic spatio-temporal dependencies, thereby improving the spatio-temporal prediction accuracy of the traffic flow of the large language model;
[0105] Third, a logical guidance template and a step-by-step decomposition guidance template are added to the action chain instructions, and regional information, time information, and weather data are integrated, and context information is distinguished by marker characters, which can enable the large language model to reason and predict according to the action chain, and output the paradigm chain according to the specified traffic flow prediction scenario and the provided information according to the decomposed actions.
[0106] Referring to Figure 3 , the embodiments of the present invention also provide a traffic flow prediction system based on a spatio-temporal large language model, including:
[0107] A data acquisition module for acquiring traffic flow data to be predicted;
[0108] A spatio-temporal feature encoding module for constructing a spatio-temporal encoder based on a recurrent neural network and a graph neural network, and encoding the traffic flow data to be predicted through the spatio-temporal encoder to obtain the first spatio-temporal feature;
[0109] A feature alignment module, configured to perform feature alignment on the first spatio-temporal feature through a preset alignment module to obtain a second spatio-temporal feature, so as to align the dimensions of the spatio-temporal encoder and the large language model;
[0110] An instruction construction and result prediction module, configured to obtain a traffic flow prediction result through the large language model according to the second spatio-temporal feature and the action chain instruction.
[0111] The content in the above embodiments of the traffic flow prediction method based on the spatio-temporal large language model is applicable to the embodiments of the traffic flow prediction system based on the spatio-temporal large language model. The functions specifically implemented by the embodiments of the traffic flow prediction system based on the spatio-temporal large language model are the same as those of the above embodiments of the traffic flow prediction method based on the spatio-temporal large language model, and the beneficial effects achieved are also the same as those of the above embodiments of the traffic flow prediction method based on the spatio-temporal large language model.
[0112] An embodiment of the present invention also provides an electronic device, which includes: a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing connection communication between the processor and the memory. When the program is executed by the processor, it implements the above traffic flow prediction method based on the spatio-temporal large language model. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0113] As Figure 4 shown is a schematic hardware structure diagram of the electronic device provided by an embodiment of the present invention. Referring to Figure 4 , an embodiment of the present invention provides an electronic device, including:
[0114] A processor 1001, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of the present invention;
[0115] A memory 1002, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1002 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1002 and are called by the processor 1001 to execute the traffic flow prediction method based on the spatio-temporal large language model of the embodiments of the present invention;
[0116] An input / output interface 1003 for implementing information input and output;
[0117] A communication interface 1004 for implementing communication interaction between this device and other devices, which can achieve communication through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0118] A bus 1005 for transmitting information between various components of the device (such as a processor 1001, a memory 1002, an input / output interface 1003, and a communication interface 1004);
[0119] Among them, the processor 1001, the memory 1002, the input / output interface 1003, and the communication interface 1004 achieve communication connections with each other inside the device through the bus 1005.
[0120] An embodiment of the present invention also provides a storage medium, which is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the above-mentioned traffic flow prediction method based on the spatio-temporal large language model.
[0121] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0122] An embodiment of the present invention also discloses a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 1 the method shown.
[0123] In some alternative embodiments, the functions / operations recited in the block diagrams may not occur in the order presented in the operational illustrations. For example, depending on the functions / operations involved, two blocks shown in succession may actually be executed substantially simultaneously or the blocks may sometimes be executed in the reverse order. Additionally, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logic flows presented herein. Alternative embodiments are contemplated in which the order of various operations is altered and in which sub-operations described as part of a larger operation are performed independently.
[0124] Furthermore, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the above-described functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for an understanding of the present invention. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skill of an engineer. Accordingly, those of ordinary skill in the art can implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the particular concepts disclosed are illustrative only and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0125] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above methods in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as a USB flash drive, a portable hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.
[0126] The logic and / or steps represented in the flowchart or otherwise described herein can, for example, be considered a definitional sequence of executable instructions for implementing logical functions, which can be embodied specifically in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. As used in this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0127] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.
[0128] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.
[0129] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.
Claims
1. A traffic flow prediction method based on a spatiotemporal large language model, characterized in that: The following steps are involved: Obtain traffic flow data to be predicted; Constructing a spatiotemporal encoder based on a recurrent neural network and a graph neural network, encoding the traffic flow data to be predicted by the spatiotemporal encoder to obtain a first spatiotemporal feature; Performing feature alignment on the first spatiotemporal features through a preset alignment module to obtain second spatiotemporal features, so as to align the spatiotemporal encoder with the large language model dimension; An action chain instruction is constructed, and a traffic flow prediction result is obtained according to the second spatiotemporal feature and the action chain instruction through the large language model.
2. The traffic flow prediction method based on spatiotemporal large language model according to claim 1 is characterized in that: The encoding of the traffic flow data to be predicted by the spatiotemporal encoder to obtain the first spatiotemporal feature specifically includes: Inputting the traffic flow data to be predicted into the spatiotemporal encoder; Performing spatial encoding on the traffic flow data to be predicted by using the graph neural network to obtain spatial features; Time encoding the traffic flow data to be predicted is performed by the recurrent neural network to obtain a time feature; The spatial feature and the temporal feature are nonlinearly connected to obtain the first spatiotemporal feature.
3. The traffic flow prediction method based on spatiotemporal large language model according to claim 2 is characterized in that: The time encoding of the traffic flow data to be predicted by the recurrent neural network to obtain time features specifically includes: Extracting features of the traffic flow data to be predicted through a gated dilated convolutional layer in the recurrent neural network to obtain multi-scale time features; The multi-scale time features are correlated and merged through a cross-scale association layer in the recurrent neural network to obtain the time features.
4. The traffic flow prediction method based on spatiotemporal large language model according to claim 1 is characterized in that: The step of aligning the first spatiotemporal features by a preset alignment module to obtain a second spatiotemporal feature specifically includes: Inputting the first spatiotemporal feature into the alignment module, and then projecting the first spatiotemporal feature to obtain a spatiotemporal label; Performing mapping analysis on the spatiotemporal markers to obtain hidden representations; The spatiotemporal mark and the hidden representation are mapped and concatenated to obtain the second spatiotemporal feature.
5. The traffic flow prediction method based on spatiotemporal large language model according to claim 1 is characterized in that: The instructions for constructing an action chain specifically include: Build logical boot templates and step-by-step decomposition boot templates; Acquire area information, time information and weather data from the traffic flow data to be predicted, and determine marking characters; The action chain instruction is constructed according to the logic guidance template, the step-by-step decomposition guidance template, the area information, the time information, the weather data and the marking character.
6. The traffic flow prediction method based on spatiotemporal large language model according to claim 3 is characterized in that: The multi-scale time feature is obtained by the following formula: in, represents the multi-scale time feature, represents the convolution kernel and and represents a one-dimensional convolution kernel and and Indicates bias and k represents the number of one-dimensional convolution kernels and k = [1, 2, 3, 4], l represents the number of layers, D f Indicates the number of the first channel, D tr Indicates the number of the second channel, represents the spatial feature, [·] represents the concatenation operation, ⊙ represents the Hadamard product, σ3 and σ4 represent activation functions.
7. The traffic flow prediction method based on spatiotemporal large language model according to claim 4 is characterized in that: The second spatiotemporal feature is obtained by mapping and concatenating the spatiotemporal tag and the hidden representation through the following formula: in, represents the second spatiotemporal feature, pr represents the space-time label, Γ hi represents the hidden representation and C represents the number of initial channels, N represents the number of spatial nodes, and D L represents the number of the third channel, σ5 represents the activation function, W1, W2 and W3 represent the mapping operation and D b represents the number of the fourth channel, β represents the output step size, and [·,·] represents the cascade operation.
8. A traffic flow prediction system based on a spatiotemporal language model, characterized in that: include. A data acquisition module, used to acquire traffic flow data to be predicted; A spatiotemporal feature encoding module, used for constructing a spatiotemporal encoder based on a recurrent neural network and a graph neural network, encoding the traffic flow data to be predicted through the spatiotemporal encoder to obtain a first spatiotemporal feature; A feature alignment module, configured to perform feature alignment on the first spatiotemporal features through a preset alignment module to obtain second spatiotemporal features, so as to align the spatiotemporal encoder and the large language model dimension; The instruction construction and result prediction module is used to obtain the traffic flow prediction result according to the second spatiotemporal feature and the action chain instruction through the large language model.
9. An electronic device, characterized in that: The electronic device includes a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for realizing connection and communication between the processor and the memory. When the program is executed by the processor, the steps of the traffic flow prediction method based on the spatiotemporal large language model as described in any one of claims 1 to 7 are realized.
10. A storage medium, the storage medium being a computer-readable storage medium, used for computer-readable storage, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the traffic flow prediction method based on the spatiotemporal large language model as described in any one of claims 1 to 7.
Citation Information
Patent Citations
An online video behavior detection system and method based on space-time context analysis
CN109409307A
User trajectory tracking prediction method and device, electronic equipment and storage medium
CN115952816A
Traffic flow prediction method and system based on trend space-time diagram convolution, and medium
CN116895157A
Model reasoning method and device based on graph data structure, equipment and medium
CN117077791A
Multi-source traffic data processing method and device and electronic equipment
CN117931977A
Cited By
Data processing method, device and equipment for acquiring airport delay information
CN120338205A