Highway flow prediction method based on multi-modal fusion and sequence decomposition
Through the methods of multimodal fusion and sequence decomposition, the problems of low prediction accuracy and computational efficiency of traditional neural network models in highway traffic flow prediction are solved, and more efficient and accurate traffic flow prediction is achieved.
Patent Information
- Application Number
- CN202510844343.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-09-26
Smart Images

Figure CN120708420A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of traffic flow prediction, and in particular to a highway flow prediction method based on multimodal fusion and sequence decomposition. Background Art
[0002] In the field of intelligent transportation systems and traffic flow prediction, existing systems are widely built on cloud computing systems and use centralized neural network models such as long short-term memory networks (LSTMs), gated recurrent units (GRUs), deep distributed neural networks (DDNNs), graph convolutional networks (GCNs), and hierarchical graph convolutional networks (HGCNs). Although systems implemented using this technical approach have achieved the goal of traffic flow prediction to a certain extent, they generally suffer from the following drawbacks:
[0003] Prediction accuracy is insufficient: Traditional neural network models, such as LSTM, GRU, and DDNN, are suitable for processing datasets with single features. When faced with highway traffic flow data, due to its inherent multi-temporal dependency characteristics and the influence of external events such as weather and accidents, these models have difficulty extracting clear and unified traffic flow characteristics, resulting in large errors in prediction results. While GCN and HGCN network models can extract spatial feature information from complex traffic flow data, they seriously miss information when processing temporal features, especially long-term dependency features.
[0004] Low computational efficiency: Traditional neural network models, such as LSTM and GRU, require step-by-step gradient calculations during backpropagation, resulting in video memory usage proportional to the sequence length. This resource consumption is particularly significant in long-time series tasks involving large-scale traffic data. Furthermore, the recursive structure of LSTM cannot be computed in parallel due to its strict temporal dependencies. Training speed is limited by sequence length, making it difficult to efficiently process long time series inputs. The number of parameters in these models (such as the LSTM gate units) increases exponentially with the number of network layers, further exacerbating memory consumption. Stacking multiple layers significantly increases video memory requirements, limiting their practicality in resource-constrained scenarios. Summary of the Invention
[0005] In view of the above existing problems, the present invention is proposed.
[0006] Therefore, the present invention provides a highway traffic flow prediction method based on multimodal fusion and sequence decomposition to solve the problem in the existing technology that it is difficult to extract obvious and unified traffic flow characteristics due to the inherent time domain multiple dependence characteristics of traditional neural network models and the characteristics affected by external events such as weather and accidents, resulting in large errors in the prediction results.
[0007] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0008] In a first aspect, the present invention provides a highway traffic flow prediction method based on multimodal fusion and sequence decomposition, which comprises:
[0009] Obtaining historical monitoring data of highway traffic flow, meteorological condition data, and sampling time-scale data, extracting short-term and long-term dependence features of traffic flow in the monitoring data, and constructing multimodal feature input samples based on the meteorological data and time-scale data;
[0010] Based on the multimodal feature input sample, a dual-channel neural network prediction model is called to extract trend term and residual term features respectively, and a traffic flow prediction result is calculated through a time-frequency conversion and time series decomposition module;
[0011] The prediction results of the trend term and the residual term features are combined, and the model parameters are adjusted using a feedback mechanism based on multimodal coding parameter optimization to generate a final traffic flow prediction model;
[0012] Based on the trained prediction model and combined with highway traffic flow monitoring data, it supports accurate prediction of regional traffic flow and inter-regional traffic flow at specified times and periods in the future.
[0013] As a preferred solution of the highway traffic flow prediction method based on multimodal fusion and sequence decomposition of the present invention, wherein: the highway traffic flow historical monitoring data is the highway traffic flow historical monitoring data, which is a triple, and a traffic flow sampling matrix is constructed according to the number of nodes in the sampling area;
[0014] The meteorological data is extracted based on the distribution of flow sampling points, and the original meteorological condition data of the sampling points are then discretized based on the sampling time scale to obtain a meteorological condition vector.
[0015] As a preferred solution of the highway traffic flow prediction method based on multimodal fusion and sequence decomposition of the present invention, wherein: the dual-channel neural network prediction model includes a sample construction layer, a feature extraction layer and an output conversion layer;
[0016] The sample construction layer includes a time series decomposition module and a multimodal fusion sample construction module;
[0017] The feature extraction layer includes two feature extraction channels, which are used to process the trend term and the residual term extracted from the sample respectively;
[0018] The output conversion layer combines and converts the output results of the two processing channels and outputs the final prediction result.
[0019] As a preferred embodiment of the highway traffic flow prediction method based on multimodal fusion and sequence decomposition of the present invention, the time series decomposition module extracts a trend term dominated by time characteristics and a residual term containing random event response characteristics from the traffic flow sampling data sequence through low-pass filtering time series decomposition, including the following steps:
[0020] For sampling points within the traffic monitoring area, the sampling data obtained within a sampling period containing L sampling times constitutes a traffic flow sampling sequence;
[0021] The sliding window average filtering algorithm is used for the elements in the sequence, and the head and tail elements of the sequence are filled with themselves to obtain the trend item sequence T v Sum remainder sequence S v ;
[0022] According to the obtained trend item sequence T v and the remainder sequence S v , forming a trend item matrix and a remainder matrix, and merging the trend item matrix and the remainder matrix to form an original traffic sample matrix.
[0023] As a preferred solution of the highway traffic flow prediction method based on multimodal fusion and sequence decomposition of the present invention, wherein: the feature extraction layer performs feature extraction through the FEDformer channel;
[0024] The FEDformer channel converts the trend item data in the time domain into the frequency domain for processing through Fourier transform, selects the frequency components through random sampling with optional training, and performs attention calculation through the frequency enhanced attention link;
[0025] The filtered frequency domain signal is restored to the time series through inverse Fourier transform to obtain the model output.
[0026] As a preferred solution of the highway traffic prediction method based on multimodal fusion and sequence decomposition of the present invention, wherein: the algorithm of the FEDformer channel includes the core algorithm of the frequency enhancement module and the core algorithm of the frequency enhancement attention module;
[0027] The FEDformer channel is a multi-stage series structure, and the T_recon output of the last stage constitutes the output T of the FEDformer channel. out ;
[0028] The FEDformer channel contains not only the trend term part of the forecast output T out , also includes the remainder set {S1,S2,…,S n}, where n is related to the number of series.
[0029] As a preferred solution of the highway traffic prediction method based on multimodal fusion and sequence decomposition described in the present invention, wherein: the residual part of the multimodal feature input sample uses a gated recurrent neural network model for feature extraction; the input item will aggregate the residual remainder output by the FEDformer channel through a splicing algorithm.
[0030] As a preferred embodiment of the highway traffic flow prediction method based on multimodal fusion and sequence decomposition of the present invention, the traffic flow in the latter part of the sample time series is predicted by the trained traffic flow prediction model and compared with the real sample set in the test set to calculate various evaluation indicators;
[0031] If the indicator exceeds the error limit, the parameters in the multimodal feature encoding module and the FEDformer channel are adjusted and the training process of temporal decomposition and feature extraction of the multimodal feature input samples is repeated until the evaluation indicator meets the expected target.
[0032] In a second aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the highway traffic prediction method based on multimodal fusion and sequence decomposition as described in the first aspect of the present invention is implemented.
[0033] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the highway traffic prediction method based on multimodal fusion and sequence decomposition as described in the first aspect of the present invention.
[0034] The beneficial effects of the present invention are as follows: by performing time series decomposition on the input data, selecting a suitable wavelet basis to perform time-frequency conversion on the time series, and performing weight calculation in the frequency domain, the authenticity of the prediction results is improved; at the same time, the data is smoothed by using time series decomposition and a suitable length truncation is performed in the frequency domain, which reduces the complexity of the calculation, improves the processing speed of the prediction model, and improves the timeliness; the model can fully capture the various modal correlations of traffic flow, such as time scale and weather scale, and effectively improve the prediction accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0036] Figure 1This is a flowchart of the highway traffic flow prediction method based on multimodal fusion and sequence decomposition.
[0037] Figure 2 Schematic diagram of the dual-channel neural network prediction model architecture and data flow.
[0038] Figure 3 This is a flowchart of the internal architecture and data flow of the feature extraction module. DETAILED DESCRIPTION
[0039] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0040] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0041] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.
[0042] Reference Figures 1 to 3 , is an embodiment of the present invention, which provides a highway traffic flow prediction method based on multimodal fusion and sequence decomposition, comprising the following steps:
[0043] S1: Original sampling data arrangement:
[0044] Highway traffic flow monitoring data and meteorological data for highway-covered areas are collected from different sampling channels. Meteorological data changes slowly, and its sampling period is significantly longer than that of traffic flow monitoring data. Therefore, it is necessary to discretize the traffic flow sampling data based on its time scale to construct comprehensive input data that integrates temporal, spatial, and event characteristics.
[0045] S1.1: Road traffic flow sampling data collation
[0046] Road traffic flow sampling data is obtained based on highway vehicle traffic statistics and represents the total number of vehicles passing a sampling point within a sampling time slice. Flow data can be represented as a triple (x, tau, v), where x is the total number of vehicles passing the sampling point; tau is the data sampling time; and v is the sampling point number. When conducting flow monitoring and flow forecasting, the flow component is of interest, so it is abbreviated as xv, tau.
[0047] Assume that a sample period contains L sampling time slices, then the historical traffic flow data collected by a sampling point numbered v in a sample period constitutes a sampling sequence X v =(x v,1 ,x v,2 ,…,x v,l …,x v,L ); Assuming that there are N nodes deployed in the area, the traffic flow sampling matrix X can be constructed, and the structure is shown as follows:
[0048]
[0049] In the subscripts, v and l represent the sampling point number and the time number in the vector of the sampled data respectively. Since the sampling time slices are of equal length, the specific sampling time can be calculated by the number l:
[0050] tau l =tau1+l×sampling time slice length
[0051] The traffic flow sampling matrix X represents the statistical data of all N traffic detection nodes within L consecutive sampling time slices. The rows of X are the data of N sampling points in the lth sampling time slice, and the columns of X are the traffic flow monitoring data of the vth sampling point in the sampling period L.
[0052] S1.2 Sampling time stamp data arrangement
[0053] Extract the sampling time scale from the road traffic flow sampling data and construct the original time scale sample matrix for the traffic flow monitoring matrix X, which is recorded as TAU raw :
[0054]
[0055] S1.3 Weather Conditions Data Collection
[0056] Meteorological condition sampling data is obtained from China Weather Network. The original sampling data includes weather conditions, temperature, wind direction and speed, etc. The flow prediction method described in this patent focuses on weather conditions, sampling locations and sampling times. Based on the distribution of flow sampling points, the original meteorological condition data of the sampling points are extracted, and then discretized according to the sampling time scale to obtain the meteorological condition vector W v =(w v,1 ,w v,2 ,…,w v,L ). Construct the original meteorological sample matrix of all sampling points in the traffic flow monitoring area, denoted as W raw :
[0057]
[0058] S2: Constructing a dual-channel neural network prediction model
[0059] Traffic flow changes on highways have both short-term dependency characteristics of 24 hours and long-term dependency characteristics with weekly, monthly, and annual cycles. In order to efficiently extract both long-term and short-term dependency characteristics of traffic monitoring data while taking into account the impact of meteorological conditions, a dual-channel neural network prediction model is constructed to perform time series decomposition and time-frequency conversion on input samples, and perform feature extraction in the frequency domain and time domain respectively. The model architecture and internal data flow are as follows: Figure 2 As shown in , it consists of a sample construction layer, a feature extraction layer, and an output conversion layer. The sample construction layer includes a time series decomposition module and a multimodal fusion sample construction module; the feature extraction layer includes two feature extraction channels, which are used to process the trend term and the residual term extracted from the sample respectively; the output conversion layer merges and converts the output results of the two processing channels and outputs the final prediction result. The module structure and internal data flow are shown in Figure 2 , the internal structure and data flow of the timing decomposition module can be found in Figure 3 .
[0060] S2.1 Time series decomposition of traffic flow monitoring data
[0061] The time series decomposition module is used to process traffic monitoring data. Through low-pass filtering and time series decomposition, it extracts trend terms dominated by time characteristics and residual terms containing random event response characteristics from the traffic flow sampling data sequence. These are used to construct the sample set of the dual-channel neural network model. The algorithm is described as follows:
[0062] For the sampling point v in the traffic monitoring area, the sampling data obtained in a sampling period containing L sampling times constitutes the traffic flow sampling sequence X v =(x v,1 ,x v,2 ,…,x v,l …,x v,L ), the elements in the sequence are filtered using a sliding window average filter algorithm with k=2, and the head and tail elements of the sequence are filled with themselves to obtain the trend item sequence T v Sum remainder sequence S v , as shown in formula (4,5), where the operator AvgPool represents the sliding average convolution and k is the length of the convolution kernel:
[0063] T v =(t v,l ),t l =AvgPool(x l ,k=2),l=(1,2,…,L)
[0064] S v =(s v,l ),s l =xl -t l ,l=(1,2,…,L)
[0065] According to the above formula, the trend item sequence and the remainder item sequence of the sampling sequence of all N sampling points in the area within this period are calculated to form the trend item matrix T and the remainder item matrix S; based on T and S, the original traffic sample matrix X is constructed. raw , as shown below:
[0066] X raw =CAT[T,S], v=(1,2,…,N)
[0067] in,
[0068] S2.2 Constructing feature input samples for multimodal feature fusion
[0069] Original traffic sample matrix X raw , original meteorological sample matrix W raw , original time-scale sample matrix TAU raw After being processed by the multimodal feature encoding module, it is compounded to construct the traffic monitoring feature input sample X of multimodal feature fusion in , used as the input of the dual-channel neural network model. The multimodal feature encoding module contains three encoder channels, which are used for X raw 、W raw and TAU raw Encoding conversion. Figure 2 The parameterized operation of the encoding conversion is shown as follows, X in It is constructed from the output of the code converter as shown below:
[0070] X embed =GLU(Θ×X raw )
[0071] W embed =OneHot(W raw )×W w +b w
[0072] TAU embed =Linear(TAU raw )×W tau +b tau
[0073] X in =X enbed +W embed +TAU embed =CAT[T in ,S in ]
[0074] Where Θ is the convolution kernel of the graph neural network. raw Mapped into a high-dimensional vector OneHot (W raw ) represents a one-hot encoding encoder. W raw The meteorological condition samples in cannot be directly used for neural network processing and need to be encoded. Since the weather state value range in the weather state expression value range belongs to the set (sunny, cloudy, light rain, moderate rain, heavy rain, light snow, moderate snow, other), the one-hot encoding method is used to convert it into a numerical type. w is the trainable weight matrix, b w is the bias term;
[0075] The Linear operator encodes the sampling time in segments to form a four-dimensional vector of year, month, day, and hour, and then transforms it into a high-dimensional vector through linear mapping. tau is the trainable weight matrix, b tau is the bias term;
[0076] X in In order to complete the composite traffic monitoring feature input sample matrix, traffic flow data, time data and weather data are integrated, which can be divided into two parts: trend item input and residual item input;
[0077] S2.3 Time-Frequency Conversion and Feature Extraction
[0078] The trend item of the flow monitoring feature input sample matrix is extracted through the FEDformer channel.
[0079] The FEDformer channel uses Fourier transform to convert time-domain trend data into the frequency domain for processing. Frequency components are then selected through optional trained random sampling, and attention calculations are performed via the frequency-enhanced attention phase, reducing the distraction of redundant information and improving the targeted nature of feature extraction. Compared to traditional Transformer algorithms, this reduces the algorithmic complexity of signal feature extraction to linear complexity and improves the model's ability to capture long-term dependencies. The filtered frequency-domain signal is then restored to a time series through an inverse Fourier transform to obtain the model output. This process can be cascaded by adding a time series decomposition algorithm feedback module to enhance time-domain feature extraction capabilities. The residual matrix generated by the time series decomposition in the feedback phase is merged into the residual aggregation and feature extraction channel.
[0080] FEDformer frequency domain processing flow diagram Figure 2 The algorithm in FEDformer is expressed as follows:
[0081]
[0082] Among them, T in is the flow monitoring feature input sample matrix X in The trend term of the image is transformed by the FFT operator and transferred to the frequency domain for feature extraction; k is a trainable algorithm parameter. The random sampling and attention mechanism model are used to remove redundant information interference and extract T in The time-dependent characteristics of W iFFT is a trainable weight. In the multi-stage cascade structure of the FEDformer channel, the T_recon output of the last stage constitutes the output T of the FEDformer channel. out .
[0083] S2.4 Residual aggregation and feature extraction
[0084] The residual part of the flow monitoring feature input sample matrix is extracted using the gated recurrent neural network (GRU) model. In this process, the input items are aggregated through the splicing algorithm to aggregate the residual items output by the FEDformer channel, such as Figure 3 The algorithm is shown in the following formula:
[0085] S′ in =CAT[S in ,S1,S2,…,S n ]
[0086] S out =GRU(S′ in )
[0087] Among them, S in is the flow monitoring feature input sample matrix X in The remainder of S1, S2, ..., S n It is the residual term output by the FEDformer channel during the iteration process; the input S' of the gated recurrent neural network unit is obtained by merging in ;S out It is the output of the GRU unit and is used to construct the final output of the dual-channel neural network prediction model;
[0088] S2.5 Merge Output
[0089] The output conversion layer combines the time-frequency conversion with the output of the feature extraction channel and the residual processing channel to form the final output of the dual-channel neural network prediction model:
[0090] X out =T out +S out
[0091] S3: Prediction model training
[0092] S3.1 Data Allocation:
[0093] In the training dataset used in the data loading module of this article, the first 70% of the sample time series is used as the training set for model training; the middle 20% of the sample time series is used as the validation set for model verification; and the last 10% of the sample time series is used as the test set for model testing. The window size is set to 12, meaning that one hour of data is used for training at a time. 32 batches are used for each training session.
[0094] S3.2 Regional traffic prediction model test:
[0095] The test results are evaluated using mean square error (MSE), mean absolute square error (MAE), mean absolute percentage error (MAPE), and coefficient of determination (R-Square), as shown in the following formula:
[0096]
[0097] in, is the prediction result of the regional prediction model; y i is the real traffic data in the test set; n is the number of samples in the test set.
[0098] Using the trained regional prediction model, traffic flow in the latter part of the sample time series is predicted and compared with the actual sample set in the test set. Evaluation metrics are calculated as described above. If a metric exceeds the error limit, the parameters in the multimodal feature encoding module and the FEDformer channel are adjusted and the training process of S2 is repeated until the evaluation metrics meet the expected targets.
[0099] When measured by mean square error (MSE), mean absolute percentage error (MAPE), and coefficient of determination (R-Square), the MSE and MAPE of regional prediction results have been significantly reduced. In particular, in the actual traffic flow data test of multiple highways in the Guanzhong area of Shaanxi Province, the results showed that the model has low latency and excellent prediction effect. The training time and prediction time were improved by 5% to 10% and 10% respectively compared with the current state-of-the-art methods, and the prediction effect surpassed all existing methods.
[0100] S4: Highway Traffic Flow Prediction
[0101] Using the trained prediction model, we can predict regional and inter-regional traffic flows for a specified sampling period after time T. For example, to predict traffic flow characteristics during period T+k, the prediction model output includes the predicted traffic flow values for each node at time T+k.
[0102] This embodiment also provides a computer device, which is suitable for the highway traffic prediction method based on multimodal fusion and sequence decomposition, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement the highway traffic prediction method based on multimodal fusion and sequence decomposition proposed in the above embodiment.
[0103] The computer device may be a terminal, comprising a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner may be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a button, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse.
[0104] This embodiment also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the highway traffic flow prediction method based on multimodal fusion and sequence decomposition proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, disk or optical disk.
[0105] In summary, the present invention improves the authenticity of the prediction results by decomposing the input data into time series, selecting a suitable wavelet basis to perform time-frequency conversion on the time series, and performing weight calculation in the frequency domain. At the same time, the data is smoothed by using time series decomposition and a suitable length truncation is performed in the frequency domain, which reduces the complexity of the calculation, improves the processing speed of the prediction model, and improves the timeliness. The model can fully capture the various modal correlations of traffic flow, such as time scale and weather scale, and effectively improve the prediction accuracy of the model.
[0106] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A highway traffic flow prediction method based on multimodal fusion and sequence decomposition, characterized by: include, Obtaining historical monitoring data of highway traffic flow, meteorological condition data, and sampling time-scale data, extracting short-term and long-term dependence features of traffic flow in the monitoring data, and constructing multimodal feature input samples based on the meteorological data and time-scale data; Based on the multimodal feature input sample, a dual-channel neural network prediction model is called to extract trend term and residual term features respectively, and a traffic flow prediction result is calculated through a time-frequency conversion and time series decomposition module; The prediction results of the trend term and the residual term features are combined, and the model parameters are adjusted using a feedback mechanism based on multimodal coding parameter optimization to generate a final traffic flow prediction model; Based on the trained prediction model and combined with highway traffic flow monitoring data, it supports accurate prediction of regional traffic flow and inter-regional traffic flow at specified times and periods in the future.
2. The highway traffic flow prediction method based on multimodal fusion and sequence decomposition according to claim 1 is characterized by: The highway traffic flow historical monitoring data is the highway traffic flow historical monitoring data, which is a triple, and a traffic flow sampling matrix is constructed based on the number of nodes in the sampling area; The meteorological data is extracted based on the distribution of flow sampling points, and the original meteorological condition data of the sampling points are then discretized based on the sampling time scale to obtain a meteorological condition vector.
3. The highway traffic flow prediction method based on multimodal fusion and sequence decomposition according to claim 2 is characterized by: The dual-channel neural network prediction model includes a sample construction layer, a feature extraction layer and an output conversion layer; The sample construction layer includes a time series decomposition module and a multimodal fusion sample construction module; The feature extraction layer includes two feature extraction channels, which are used to process the trend term and the residual term extracted from the sample respectively; The output conversion layer combines and converts the output results of the two processing channels and outputs the final prediction result.
4. The highway traffic flow prediction method based on multimodal fusion and sequence decomposition according to claim 3 is characterized by: The time series decomposition module extracts trend items mainly based on time characteristics and residual items containing random event response characteristics from the traffic flow sampling data sequence through low-pass filtering time series decomposition, including the following steps: For sampling points within the traffic monitoring area, the sampling data obtained within a sampling period containing L sampling times constitutes a traffic flow sampling sequence; The sliding window average filtering algorithm is used for the elements in the sequence, and the head and tail elements of the sequence are filled with themselves to obtain the trend item sequence T v Sum remainder sequence S v ; According to the obtained trend item sequence T v and the remainder sequence S v , forming a trend item matrix and a remainder matrix, and merging the trend item matrix and the remainder matrix to form an original traffic sample matrix.
5. The highway traffic flow prediction method based on multimodal fusion and sequence decomposition according to claim 4 is characterized by: The feature extraction layer performs feature extraction through the FEDformer channel; The FEDformer channel converts the trend item data in the time domain into the frequency domain for processing through Fourier transform, selects the frequency components through random sampling with optional training, and performs attention calculation through the frequency enhanced attention link; The filtered frequency domain signal is restored to the time series through inverse Fourier transform to obtain the model output.
6. The highway traffic flow prediction method based on multimodal fusion and sequence decomposition according to claim 5 is characterized by: The algorithm of the FEDformer channel includes the core algorithm of the frequency enhancement module and the core algorithm of the frequency enhancement attention module; The FEDformer channel is a multi-stage series structure, and the T_recon output of the last stage constitutes the output T of the FEDformer channel. out ; The FEDformer channel contains not only the trend term part of the forecast output T out , also includes the remainder set {S1,S2,…,S n }, where n is related to the number of series.
7. The highway traffic flow prediction method based on multimodal fusion and sequence decomposition according to claim 6 is characterized by: The residual part of the multimodal feature input sample is extracted using a gated recurrent neural network model; the input item is aggregated by the residual remainder of the FEDformer channel output through a splicing algorithm.
8. The highway traffic flow prediction method based on multimodal fusion and sequence decomposition according to claim 7 is characterized in that: The trained traffic flow prediction model is used to predict the traffic flow in the latter part of the sample time series, and compared with the real sampling sample set in the test set to calculate various evaluation indicators; If the indicator exceeds the error limit, the parameters in the multimodal feature encoding module and the FEDformer channel are adjusted and the training process of temporal decomposition and feature extraction of the multimodal feature input samples is repeated until the evaluation indicator meets the expected target.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the highway traffic prediction method based on multimodal fusion and sequence decomposition according to any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the highway traffic flow prediction method based on multimodal fusion and sequence decomposition according to any one of claims 1 to 8 are implemented.
Citation Information
Cited By
Abnormal traffic flow prediction method and system based on multi-scale spatial-temporal feature fusion
CN121583097A