Bus travel time prediction method based on multi-component decomposition and attention fusion

CN121581277BActive Publication Date: 2026-08-28OCEAN UNIV OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511643747.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-11
Publication Date
2026-08-28
Estimated Expiration
2045-11-11

AI Technical Summary

Technical Problem

[0003]近年来深度学习模型取得显著突破:通过时空卷积网络捕捉路网拓扑关联,但其未对行程时间的成分特性(等待时间与行驶时间)进行显式解耦,难以适配两者在时间动态上的本质差异;引入时空自注意力机制增强特征交互,却未融合公交动态调度信息(如发车频率的时段性变化),导致对不同运营场景(如工作日与节假日)的适应性有限

Benefits of technology

本申请中,通过日期选择实现历史数据与实时运营场景的动态匹配,提供适配性输入;设计双成分编码结构分别提取乘车时间与等待时间的特征,精准捕捉二者差异化的时空动态;引入注意力融合机制动态加权多成分特征,增强对异质信息的融合能力;并创新性地提出成分内部与外部对齐损失,通过多尺度匹配策略确保成分预测与全局结果的一致性;在提升预测精度的同时,显著增强对复杂城市公交运营场景的适应性与鲁棒性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121581277B_ABST
    Figure CN121581277B_ABST
Patent Text Reader

Abstract

The application discloses a bus travel time prediction method based on multi-component decomposition and attention fusion, comprising the following steps: S1, date selection; S2, component coding; S3, feature fusion: the travel attention is used for weighted fusion of the riding features and the waiting features, a self-attention mechanism is used for modeling long-range dependence between different road sections and stations, and global fusion features are output; the corresponding time coding is calculated by using the historical average time, and dot product operation is performed on the global fusion features to generate comprehensive features finally used for prediction; and S4, loss alignment and prediction. The application respectively models the riding time and the waiting time, fuses offline and online departure frequency data to enhance the waiting time prediction, constructs a multi-component loss alignment mechanism, ensures the consistency of component-level and overall travel-level prediction, and finally improves the prediction accuracy and scene adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of intelligent transportation systems technology, and specifically relates to a method for predicting bus travel time based on multi-component decomposition and attention fusion. Background Technology

[0002] High-precision public transport travel time prediction is a core foundation for improving the operational efficiency and service reliability of public transport networks. Currently, mainstream research relies heavily on multi-source data such as traffic signals, passenger flow, and weather data. However, in small and medium-sized cities or resource-constrained scenarios, multi-source data often suffers from high acquisition costs, low quality, or even difficulty in obtaining such data. Meanwhile, research relying solely on trajectory data obtained from devices like vehicle-mounted GPS is limited, and these studies fail to fully explore the spatiotemporal patterns hidden within the trajectory data, such as congestion transmission and departure frequency regularity. This leaves the prediction accuracy based on single-source data needing improvement.

[0003] In recent years, deep learning models have made significant breakthroughs: they capture road network topology relationships through spatiotemporal convolutional networks, but they do not explicitly decouple the component characteristics of travel time (waiting time and travel time), making it difficult to adapt to the essential differences between the two in terms of temporal dynamics; they introduce spatiotemporal self-attention mechanisms to enhance feature interaction, but do not integrate dynamic public transport scheduling information (such as the time-dependent changes in departure frequency), resulting in limited adaptability to different operating scenarios (such as weekdays and holidays). However, while current mainstream solutions continue to focus on extending network depth and optimizing feature combinations, they have not yet broken through the core limitations of component characteristic decoupling and dynamic scheduling integration, and have failed to specifically address the mechanism differences between waiting time and travel time and the adaptation of scheduling information to prediction models. Summary of the Invention

[0004] To address or at least alleviate one or more of the above problems, a bus travel time prediction method based on multi-component decomposition and attention fusion is proposed. This method models travel time and waiting time separately; it integrates offline and online departure frequency data to enhance waiting time prediction; and it constructs a multi-component loss alignment mechanism to ensure consistency between component-level and overall travel-level predictions, ultimately improving prediction accuracy and scenario adaptability.

[0005] To achieve the above objectives, according to the first aspect of this application, a method for predicting bus travel time based on multi-component decomposition and attention fusion is provided, comprising the following steps: S1. Date Selection: The AGNES algorithm is used to cluster historical departure data to generate a set of departure frequency patterns for different operating modes; the real-time departure data of the current date is matched with the clustering patterns, the most similar historical date data is selected, and an adaptive input sequence is constructed. S2, Ingredient Code: The spatiotemporal convolutional layer extracts road segment travel features from the input sequence and outputs travel features that include spatial neighborhood association and temporal continuity; the masked convolutional layer extracts waiting features of the origin station and transfer station from the input sequence and outputs waiting features, wherein the mask matrix is ​​predefined by the station type; S3, Feature Fusion: The travel features and waiting features are weighted and fused using travel attention. The long-range dependencies between different road segments and stations are modeled using a self-attention mechanism, and a global fused feature is output. The corresponding time code is calculated using historical average time and then a dot product operation is performed with the global fused feature to generate the final comprehensive feature for prediction. S4. Loss Alignment and Prediction: The system calculates ride loss and waiting loss separately, monitors the accuracy of the tiered prediction, coordinates the consistency between the tiered prediction and the total travel time prediction, and finally outputs the total travel time prediction result.

[0006] To achieve the above objectives, according to a second aspect of this application, a bus travel time prediction system based on multi-component decomposition and attention fusion is provided, the bus travel time prediction system comprising: Date selection module: The AGNES algorithm is used to cluster historical departure data to generate a set of departure frequency patterns for different operating modes; the real-time departure data of the current date is matched with the clustering patterns to select the most similar historical date data and construct an adaptive input sequence; Component encoding module: Extracts road segment travel features from the input sequence through a spatiotemporal convolutional layer, and outputs travel features that include spatial neighborhood association and temporal continuity; extracts waiting features of origin stations and transfer stations from the input sequence through a masked convolutional layer, and outputs waiting features, wherein the mask matrix is ​​predefined by the station type; Feature fusion module: The travel features and waiting features are weighted and fused using travel attention. The long-range dependencies between different road segments and stations are modeled using a self-attention mechanism, and the global fused features are output. The corresponding time code is calculated using historical average time and the global fused features are then multiplied to generate the final comprehensive features for prediction. Loss Alignment and Prediction Module: Calculates ride loss and waiting loss separately, supervises the accuracy of component-level prediction, coordinates the consistency between component-level prediction and total travel time prediction, and finally outputs the total travel time prediction result.

[0007] To achieve the above objectives, according to a third aspect of this application, a computer-readable storage medium is provided, storing a computer program that, when executed by a processor, is used to implement the bus travel time prediction method based on multi-component decomposition and attention fusion as described above.

[0008] By adopting the above technical solution, this application has the following beneficial effects compared with the prior art: In this application, a date selection method is used to dynamically match historical data with real-time operational scenarios, providing adaptable input. A dual-component coding structure is designed to extract features of travel time and waiting time respectively, accurately capturing the different spatiotemporal dynamics of the two. An attention fusion mechanism is introduced to dynamically weight multi-component features, enhancing the ability to fuse heterogeneous information. Furthermore, an innovative internal and external alignment loss is proposed, and a multi-scale matching strategy is used to ensure the consistency between component predictions and global results. While improving prediction accuracy, this method significantly enhances the adaptability and robustness to complex urban public transport operation scenarios.

[0009] The specific embodiments of this application will be described in further detail below with reference to the accompanying drawings. Attached Figure Description

[0010] The accompanying drawings, which form part of this application, are used to provide a further understanding of the application. The illustrative embodiments and descriptions of the application are used to explain the application, but do not constitute an undue limitation of the application. Obviously, the drawings described below are merely some embodiments, and those skilled in the art can obtain other drawings based on these drawings without creative effort.

[0011] In the attached diagram: Figure 1 This is a schematic diagram illustrating the steps of the bus travel time prediction method based on multi-component decomposition and attention fusion in this specific embodiment. Figure 2 This is a logical diagram illustrating the bus travel time prediction method based on multi-component decomposition and attention fusion in this specific embodiment. Figure 3 This is a flowchart illustrating the bus travel time prediction method based on multi-component decomposition and attention fusion in this specific embodiment. Figure 4 This is a schematic diagram of the bus travel time prediction system based on multi-component decomposition and attention fusion in this specific embodiment. Detailed Implementation

[0012] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments will be clearly and completely described below with reference to the accompanying drawings. The following embodiments are used to illustrate this application, but are not intended to limit the scope of this application.

[0013] Please refer to Figures 1-3 This application provides a method for predicting bus travel time based on multi-component decomposition and attention fusion, including the following steps: S1. Date Selection: The AGNES algorithm is used to cluster historical departure data to generate a set of departure frequency patterns for different operating modes; the real-time departure data of the current date is matched with the clustering patterns, the most similar historical date data is selected, and an adaptive input sequence is constructed. S2, Ingredient Code: The spatiotemporal convolutional layer extracts road segment travel features from the input sequence and outputs travel features that include spatial neighborhood association and temporal continuity; the masked convolutional layer extracts waiting features of the origin station and transfer station from the input sequence and outputs waiting features, wherein the mask matrix is ​​predefined by the station type; S3, Feature Fusion: The travel features and waiting features are weighted and fused using travel attention. The long-range dependencies between different road segments and stations are modeled using a self-attention mechanism, and a global fused feature is output. The corresponding time code is calculated using historical average time and then a dot product operation is performed with the global fused feature to generate the final comprehensive feature for prediction. S4. Loss Alignment and Prediction: The system calculates ride loss and waiting loss separately, monitors the accuracy of the tiered prediction, coordinates the consistency between the tiered prediction and the total travel time prediction, and finally outputs the total travel time prediction result.

[0014] It should be noted that the execution entity of the bus travel time prediction method based on multi-component decomposition and attention fusion in this embodiment is a bus travel time prediction system based on multi-component decomposition and attention fusion. This system can be an electronic device, a component in an electronic device, an integrated circuit, or a chip. The electronic device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, etc., and non-mobile electronic devices can be servers and personal computers, etc., which are not specifically limited in this application. The following description uses a server as the execution entity to illustrate the bus travel time prediction method based on multi-component decomposition and attention fusion in this embodiment.

[0015] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0016] In some embodiments, the step of clustering historical departure data using the AGNES algorithm to generate a set of departure frequency patterns for different operating modes includes: Construct a departure frequency matrix; The AGNES algorithm is used to perform hierarchical clustering with Euclidean distance as a measure of date similarity, and the optimal number of clusters is adaptively determined based on the silhouette coefficient to obtain the cluster set. Calculate the centroid vector for each cluster to characterize the average departure frequency distribution under this operating mode.

[0017] Specifically, historical departure frequency data is collected by collecting all daily departure records for two weeks (14 days) prior to the prediction date. This data is then divided into fixed time slots. First, the date set D = { …, The entire day's departure records are used to construct the departure frequency matrix F ∈ . The data consists of 42 columns corresponding to 30-minute intervals from 4:00 to 24:00. This ensures data coverage of weekdays, non-working days, and special operating modes, addressing the issue of traditional methods neglecting fluctuations in operating modes.

[0018] Specifically, offline clustering employs hierarchical clustering to discover potential operational patterns. First, the AGENS algorithm (agglomerated hierarchical clustering) is used to calculate date similarity, with Euclidean distance as the metric: dist( ) = This algorithm can adaptively determine the optimal number of clusters K based on the silhouette coefficients of different input data, thus obtaining the cluster set { This method is suitable for situations where multiple departure modes coexist, overcomes the sensitivity of K-means to initial centers, and improves the robustness of mode partitioning.

[0019] Specifically, the cluster center feature calculation is based on the different categories defined by the AGENS algorithm. The centroid vector of each category reflects the average departure frequency distribution under that mode, and can reflect the typical characteristics of each mode category. Therefore, for each cluster... Calculate the centroid vector: = (k = 1,2,…K);

[0020] Specifically, it dynamically matches real-time operational status and inputs the predicted date. Observed period [4:00-] Departure frequency ∈ (t is the current time slot sequence), and extract the corresponding segments of the centroid vectors for each category: = [1:t].

[0021] In some embodiments, the step of matching real-time departure data for the current date with the clustering pattern, filtering out the most similar historical date data, and constructing an adaptive input sequence includes: Obtain the departure frequency vector for the predicted date during the observed time period; Extract the corresponding segments of the centroid vectors of each cluster; Calculate the Euclidean distance between the observation vector and each truncated centroid, select the cluster corresponding to the minimum distance, and use all historical dates contained in the cluster as the historical input set for waiting time prediction; Based on whether the current date is a workday, the set of workdays or non-workdays from the past two weeks is used as the historical input for travel time.

[0022] Specifically, similarity quantification calculations are performed to select the optimal historical reference set and calculate the distance between the observed vector and the truncated centroid: = (k = 1,2…K); Select the cluster corresponding to the minimum distance = The included dates serve as the input set for the waiting time prediction module, improving the spatial relevance of the input data to the current scene.

[0023] Specifically, the adaptation sequence is constructed to generate the model input data structure: (Selection) All historical dates are used as historical inputs for the waiting time prediction module; it is determined whether the current date is a working day. If it is a working day, the working days of the past two weeks are used as historical inputs for the travel time module; otherwise, the set of non-working days of the past two weeks are used as historical inputs for the travel time module.

[0024] In some embodiments, the extraction of waiting features of the origin station and transfer station from the input sequence through a masked convolutional layer, and the output of the waiting features, includes: Construct a binary mask matrix , in, = ; By performing masked convolution operations, feature extraction is performed based on site characteristics to obtain waiting features. , represented as: = ReLU(Conv2d( , ); in, , These are the waiting time input vector and the mask input, respectively. This represents a convolutional kernel that is spatially adjacent to three segments and temporally continuous for six steps, through the mask matrix. Hard-coded site prior knowledge guides the model to focus on feature extraction for origin and transfer stations.

[0025] Specifically, the spatial mask matrix is ​​constructed, site-aware feature filters are designed, and a binary mask matrix is ​​defined. ∈ ; in, = ; The model is guided to focus on key areas by hard-coded prior knowledge.

[0026] Specifically, the driving time module performs convolution operations to extract features from the driving time vector, performing the following convolution: = ReLU(Conv2d( ), ); in, It is the driving time input vector. This indicates the kernel size, with 3 spatially adjacent segments × 6 consecutive temporal steps, and 32 output channels.

[0027] Specifically, the waiting time module performs a masked convolution operation to extract features based on site characteristics, and performs the following masked convolution: = ReLU(Conv2d( , ); in, These are the waiting time input vector and the mask input, respectively. This indicates the kernel size, with 3 spatially adjacent segments × 6 consecutive temporal steps, and 32 output channels.

[0028] Specifically, the road segment-level time prediction will be obtained from... and Inputting the entry travel time module and the waiting time module respectively yields a road segment-level travel time feature of size 40×1. With a segment-level waiting time feature of size 40×1 .

[0029] In some embodiments, the weighted fusion of the travel features and waiting features through travel attention includes: Introduce a trainable parameter α, whose initial value is randomly initialized in the range [0,1], and constrain the waiting time weight β=1-α to ensure that α+β=1; By linearly weighting the feature space using travel time weight α and waiting time weight β, the travel features and waiting features are initially fused, as shown below: = α + β ; in, Indicates the characteristics after fusion. For travel time characteristics, This refers to the waiting time characteristic; During model training, the parameter α is dynamically optimized using the backpropagation algorithm to adaptively learn and balance the contributions of driving time and waiting time to the total travel time.

[0030] Specifically, traditional bus travel time prediction methods have two typical limitations: first, they fail to decouple waiting time prediction from travel time prediction, ignoring the essential differences in their generation mechanisms; second, although decoupling is achieved, the waiting time is simply added to the travel time during final integration, ignoring the differences in the contribution of each component. To address this issue, this embodiment proposes a fusion method based on a parameterized weighting mechanism. Specifically, in the initial stage, the travel time weight α is randomly initialized within the range [0,1], and the waiting time weight β is constrained to satisfy β=1-α, thereby ensuring the coupling relationship of α+β=1; during model training, α and β are dynamically optimized and adjusted using the backpropagation algorithm to adaptively capture the differences in the contribution of travel time and waiting time.

[0031] Specifically, the travel time weight α and waiting time weight β obtained in the previous step are used to perform linear weighting of the feature space: = α + β ; in, For travel time characteristics, To address the waiting time feature, dynamic weights α and β are used to achieve feature decoupling and rebalancing in complex traffic scenarios.

[0032] Specifically, a multi-head attention mechanism is used to model spatiotemporal dependencies. Traditional public transport travel time prediction methods suffer from significant technical bottlenecks in spatiotemporal dependency modeling, particularly in their inability to capture long-distance station associations. This is attributed to the inherent receptive field limitations of convolutional neural networks and the temporal fragmentation defects of recurrent neural networks. The generated feature vectors are then used... As the input tensor for the attention mechanism, h = 6 attention heads are configured, and the score of each attention head is calculated as follows: using 3 sets of learnable projection matrices. , , ,Will Decompose into queries ( ),key( ),value( Subtensor, calculation and The matrix product of (key transpose) yields the cross-segment-time step association score; this is then scaled and normalized, and divided by a scaling factor. Then, the attention weights are normalized using the Softmax function to obtain the score for the head. The calculation formula is as follows: = Softmax ( ) ; To avoid excessively large scores that could cause the softmax gradient to vanish, d = 6 is set. Then, the scores from the six attention heads are concatenated to obtain the run-length order feature vector. : =Concat( , ,… ).

[0033] In some embodiments, the step of performing a dot product operation between the historical average time-based ... Based on historical data, calculate each road segment Average driving time And the average waiting time of road sections with waiting time. Average waiting time for road sections with no waiting time Set to 0; Calculate the statistical travel time for each road segment. = + and based on The start time of the journey for each segment is calculated cumulatively. ; Construct a spatiotemporal mask matrix This corresponds to 40 road segments and 36 time intervals divided into 30-minute intervals starting at 6:00 AM; the spatiotemporal mask matrix... Start time of the journey in each section The value at the corresponding time interval position is set to 1, and the rest are 0; The global fusion features Element-wise multiplication with the spatiotemporal mask matrix NM yields the final feature enhanced with contextual information. .

[0034] Specifically, spatiotemporal masking enhancement is based on run-length level feature vectors. The process involves performing spatiotemporal masking enhancement to fuse key temporal context information. This allows for adaptive amplification or suppression of the contributions of waiting and traveling components in critical scenarios such as originating stations, transfer stations, and peak / special periods. Specifically, the technical process includes the following steps: First, historical data is divided into weekday and non-weekday categories. For the travel time of 40 road segments, the average travel time value for each road segment is calculated separately for each category as the statistical travel time value for that road segment. For the 40 road segments with waiting times, 100 start times were randomly selected throughout the day, and the average waiting time for each of these 100 start times was taken as the statistical waiting time value for that road segment. For other road segments with no waiting time, this value is set to 0. The statistical travel time is then added to the corresponding element of the statistical waiting time at the preceding bus stop to obtain the result. By obtaining the time consumed for each route segment, the start time of the journey for each segment can be determined. The calculation method is as follows: ; Then, the day is divided into 36 different intervals, starting from 6:00 AM and lasting for half an hour, to construct an initial timeframe. ∈ The new mask matrix is ​​initialized to all zeros, and then combined with the calculated... The value of the element in the middle will The value of the interval containing the departure time of the corresponding road segment is assigned as 1, and the result is the assigned value. The result will be and Performing element-wise multiplication yields the masked run-length feature vector. .

[0035] Specifically, the run-length feature vector after masking The input enters the trip prediction module, which outputs the trip-level time feature with a size of 40×1. And the predicted value y, and then compare the predicted value with the actual value. Comparative calculation of trip prediction loss , represented as: = |y - |

[0036] In some embodiments, the separate calculation of fare loss and waiting loss, the supervision of component prediction accuracy, and the coordination of consistency between component prediction and total travel time prediction, ultimately outputting the total travel time prediction result, include: Internal alignment: Randomly select n road segments and their segment-level travel time features F r Calculate its relationship with the corresponding true value Loss as driving loss Simultaneously, the segment-level waiting time characteristics F of m randomly selected road segments are analyzed. w Calculate its relationship with the corresponding true value Loss as waiting loss ; driving losses and waiting loss Each component is backpropagated separately to optimize feature extraction independently; External alignment: Four non-overlapping consecutive segments are randomly selected from the total trip; for each segment, the sum of the segment-level travel time feature values ​​within that segment is accumulated. The sum of the characteristic values ​​of waiting time at the road segment level and the sum of time characteristic values ​​at the travel level. ; Calculate the alignment loss for each segment , represented as: = | - ( ) |; The total inter-component alignment loss is obtained by summing the losses of the four sub-segments. , is represented as; = ; in, The number of consecutive subarrays; Global optimization: Reduce trip prediction losses Damage during driving With waiting loss and overall component alignment loss We obtain the total global loss by weighted summation. , represented as: = + ( + ) + ; in, and This is a hyperparameter.

[0037] Specifically, to accurately capture the temporal dynamics and error characteristics of both travel time and waiting time, independent loss functions are used for supervised optimization of each component, employing... The loss (mean absolute error) is used to ensure consistency in the optimization objective while adapting to the different operational attributes of the two. For a travel time F of size 40×1 road segments... r And the waiting time F at the segment level of 40×1 w Randomly select n road segments, calculate the predicted travel time for each of these n road segments, and then compare the predicted travel time with the actual travel time. loss As the internal loss of the travel time prediction module; the predicted waiting time of m road segments is randomly selected, and the difference between the predicted waiting time and the actual waiting time of these m road segments is calculated. loss As the internal loss of the waiting time prediction module; and The process is backpropagated to the waiting time encoding module and the driving time encoding module respectively, so as to achieve independent optimization of the two component feature extraction processes and ensure that each captures its own differentiated time dynamic patterns.

[0038] Specifically, to ensure consistency between component-level prediction and global trip-level prediction at different granularities and to avoid the accumulation of local errors leading to overall prediction deviation, this step designs an inter-component alignment loss module based on multi-scale random segment accumulation consistency constraints. Specifically, firstly, over a trip with a total length of 40 road segments, four non-overlapping consecutive segments (their lengths can be dynamically changed) are randomly selected. For each segment, the segment-level travel time feature vector is... Road segment level waiting time feature vector and travel-level time feature vectors The predicted values ​​within this sub-segment are summed to obtain the sum of the prediction times for each component within that sub-segment, denoted as follows: , and Then, the global travel prediction value within this segment is calculated. Sum of component predictions ( The absolute difference between the segments is used as the alignment loss for that segment: = | - ( ) |; Iterate through all four random segments, sum their alignment losses, and finally obtain the total inter-component alignment loss. : =

[0039] This loss function explicitly forces the model to maintain consistency between component-level and trip-level predictions during the learning process by directly constraining the sum of component predictions to be equal to the overall prediction on any local segment. This effectively suppresses error accumulation and improves the accuracy and robustness of the final total trip time prediction result.

[0040] Specifically, global loss collaborative optimization involves aligning the loss between components. With trip-level losses Intra-component loss ( and By merging these, we obtain the total global loss: = + ( + ) + ; in, and The loss weight hyperparameters, determined through validation set tuning, are used to balance the optimization priorities of different loss terms. Backpropagation is performed to the entire prediction system (including component encoding, feature fusion module, and alignment module) to achieve collaborative optimization of component-level and global-level predictions, and finally outputs a total travel time prediction result that balances local accuracy and global consistency.

[0041] Specifically, the training model gradually converges, and the network model with the highest experimental accuracy is saved.

[0042] Based on the same inventive concept, please refer to Figure 4 This application also provides a bus travel time prediction system based on multi-component decomposition and attention fusion, the bus travel time prediction system comprising: Date selection module: The AGNES algorithm is used to cluster historical departure data to generate a set of departure frequency patterns for different operating modes; the real-time departure data of the current date is matched with the clustering patterns to select the most similar historical date data and construct an adaptive input sequence; Component encoding module: Extracts road segment travel features from the input sequence through a spatiotemporal convolutional layer, and outputs travel features that include spatial neighborhood association and temporal continuity; extracts waiting features of origin stations and transfer stations from the input sequence through a masked convolutional layer, and outputs waiting features, wherein the mask matrix is ​​predefined by the station type; Feature fusion module: The travel features and waiting features are weighted and fused using travel attention. The long-range dependencies between different road segments and stations are modeled using a self-attention mechanism, and the global fused features are output. The corresponding time code is calculated using historical average time and the global fused features are then multiplied to generate the final comprehensive features for prediction. Loss Alignment and Prediction Module: Calculates ride loss and waiting loss separately, supervises the accuracy of component-level prediction, coordinates the consistency between component-level prediction and total travel time prediction, and finally outputs the total travel time prediction result.

[0043] In some embodiments, the date selection module is used to filter historical date data similar to the current scenario's operating mode, constructing an adaptive input sequence as input to the component encoding module; the component encoding module includes a travel time encoding unit and a waiting time encoding unit, which respectively extract travel time features and waiting time features in the trip, and the input sequence output by the date selection module is sent to the two encoding units respectively; the feature fusion module includes a travel attention submodule and a feature splicing layer, which is used to fuse the output features of the travel time encoding unit and the waiting time encoding unit to capture the spatiotemporal dependencies of the global trip; the loss alignment module includes an internal alignment submodule and an external alignment submodule, the internal alignment submodule is used to optimize the component-level prediction error of travel time and waiting time, and the external alignment submodule is used to coordinate the consistency between component-level prediction and total trip time prediction, and finally output the prediction result.

[0044] In some embodiments, the date selection module includes an offline clustering unit and an online matching unit: the offline clustering unit uses the AGNES algorithm to cluster the departure data (data1, data2...data14) of the past two weeks to generate a set of departure frequency patterns for different operating modes; the online matching unit matches the real-time departure data of the current date with the clustering patterns, and filters out the most similar historical date data to ensure the scenario adaptability of the input sequence.

[0045] In some embodiments, in the component encoding module, the travel time encoding unit uses a spatiotemporal convolutional layer to extract road segment travel features. The input is a historical travel time series, and the output is a travel feature H that includes spatial neighborhood correlation and temporal continuity. r The waiting time encoding unit combines the time features of the starting station (Start Stop) and the transfer station (Transfer Stop), and outputs the waiting feature H through a masked convolutional layer (activating only valid waiting station features). w The mask matrix is ​​predefined by the site type.

[0046] In some embodiments, the travel attention submodule of the feature fusion module... and We perform weighted fusion, model the long-range dependencies between different road segments and stations using a self-attention mechanism, and output global fusion features. ; Calculate the corresponding time code using historical average time and Perform a dot product operation to generate the final composite features used for prediction.

[0047] In some embodiments, within the loss alignment and prediction module, the internal alignment submodule calculates the riding loss and waiting loss, respectively supervising... and The component-level prediction accuracy is improved; the external alignment submodule ensures the consistency between the sum of component features and global features through multi-scale projection, and finally predicts the total travel time by outputting the total travel time.

[0048] Based on the same inventive concept, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the bus travel time prediction method based on multi-component decomposition and attention fusion as described above.

[0049] The program product of this application for implementing the above method may employ a portable compact disk read-only memory and include program code, and may run on a terminal device, such as a personal computer. However, the program product of this application is not limited thereto. In this application, the readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.

[0050] It should be noted that a computer-readable storage medium may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable storage medium may also be any readable medium other than a readable storage medium, capable of transmitting, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0051] The above are merely preferred embodiments of this application and are not intended to limit this application in any way. Although this application has disclosed preferred embodiments as described above, it is not intended to limit this application. Any person skilled in the art can make some modifications or alterations to the above-mentioned technical content to create equivalent embodiments without departing from the scope of the technical solution of this application. The implementation schemes in the above embodiments can also be further combined or replaced. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of this application without departing from the content of the technical solution of this application shall still fall within the scope of this application.

Claims

1. A bus travel time prediction method based on multi-component decomposition and attention fusion, characterized in that, Includes the following steps: S1. Date Selection: The AGNES algorithm is used to cluster historical departure data to generate a set of departure frequency patterns for different operating modes; the real-time departure data of the current date is matched with the departure frequency patterns, the most similar historical date data is selected, and an adaptive input sequence is constructed. S2, Ingredient Code: The system extracts road segment travel features from the input sequence using a spatiotemporal convolutional layer, outputting travel features that include spatial neighborhood correlation and temporal continuity; it also extracts waiting features for origin stations and transfer stations from the input sequence using a masked convolutional layer, outputting waiting features, where the mask matrix is ​​predefined by the station type; including: Construct a binary mask matrix , in, = ; By performing masked convolution operations, feature extraction is performed based on site characteristics to obtain waiting features. , represented as: ; Among them, X w , These are the waiting time input vector and the mask input, respectively. This represents a convolutional kernel that is spatially adjacent to three segments and temporally continuous for six steps, through the mask matrix. Hard-coded site prior knowledge guides the model to focus on feature extraction from origin and transfer stations; S3, Feature Fusion: The travel features and waiting features are weighted and fused using travel attention, including: Introduce a trainable parameter Trainable parameters The initial values ​​are randomly initialized in the range [0,1], and the waiting time weight is constrained. In order to protect ; Weighted by driving time Weighted by waiting time Linear weighting is performed in the feature space to initially fuse the travel features and waiting features, as shown below: ; in, Indicates the characteristics after fusion. For travel time characteristics, This refers to the waiting time characteristic; During model training, the parameters are processed using the backpropagation algorithm. Dynamic optimization is performed to adaptively learn and balance the contributions of driving time and waiting time to the total trip time; The self-attention mechanism is used to model the long-range dependencies between different road segments and stations, and global fusion features are output, including: The fused features As input, three sets of learnable projection matrices are used respectively. , , Generate query, key, and value tensors; The calculation method for each attention point is as follows: ; Among them, the definition =6; The outputs of the six attention heads are concatenated to obtain a global fusion feature that models long-range spatiotemporal dependence. , represented as: ; in, , ,… These represent the outputs of attention heads 1 through 6 respectively; The corresponding time code is calculated using historical average time, and a dot product operation is performed with the global fusion feature to generate the final comprehensive feature used for prediction; including: Based on historical data, calculate each road segment Average driving time And the average waiting time of road sections with waiting time. Average waiting time for road sections with no waiting time Set to 0; Calculate the statistical travel time for each road segment. and based on The start time of the journey for each segment is calculated cumulatively. ; Construct a spatiotemporal mask matrix This corresponds to 40 road segments and 36 time intervals divided into 30-minute intervals starting at 6:00 AM; the spatiotemporal mask matrix... Start time of the journey in each section The value at the corresponding time interval position is set to 1, and the rest are 0; The global fusion features Element-wise multiplication with the spatiotemporal mask matrix NM yields the final feature enhanced with contextual information. ; S4. Loss Alignment and Prediction: The system calculates ride loss and waiting loss separately, monitors the accuracy of the tiered prediction, coordinates the consistency between the tiered prediction and the total travel time prediction, and finally outputs the total travel time prediction result.

2. The method according to claim 1, characterized in that, The AGNES algorithm is used to cluster historical departure data to generate a set of departure frequency patterns for different operating modes, including: Construct a departure frequency matrix; The AGNES algorithm is used to perform hierarchical clustering with Euclidean distance as a measure of date similarity, and the optimal number of clusters is adaptively determined based on the silhouette coefficient to obtain the cluster set. Calculate the centroid vector for each cluster to characterize the average departure frequency distribution under this operating mode.

3. The method according to claim 1, characterized in that, The step of matching the real-time departure data of the current date with the departure frequency pattern, filtering out the most similar historical date data, and constructing an adaptive input sequence includes: Obtain the departure frequency vector for the predicted date during the observed time period; Extract the corresponding segments of the centroid vectors of each cluster; Calculate the Euclidean distance between the observation vector and each truncated centroid, select the cluster corresponding to the minimum distance, and use all historical dates contained in the cluster as the historical input set for waiting time prediction; Based on whether the current date is a workday, the set of workdays or non-workdays from the past two weeks is used as the historical input for travel time.

4. The method according to claim 1, characterized in that, The process involves separately calculating travel loss and waiting loss, supervising the accuracy of tiered predictions, coordinating the consistency between tiered predictions and total travel time predictions, and finally outputting the total travel time prediction result, including: Internal alignment: Randomly select n road segments and their segment-level travel time features F r Calculate its relationship with the corresponding true value Loss as a passenger loss Simultaneously, the segment-level waiting time characteristics F of m randomly selected road segments are analyzed. w Calculate its relationship with the corresponding true value Loss as waiting loss ; driving losses and waiting loss Each component is backpropagated separately to optimize feature extraction independently; External alignment: Randomly select four non-overlapping consecutive segments from the total journey; for each segment, calculate the sum of the travel times for that segment. Sum of segment-level waiting times and the sum of the travel times at the sub-segment level. ; Calculate the alignment loss for each segment , represented as: ; The total inter-component alignment loss is obtained by summing the losses of the four sub-segments. , is represented as; ; in, The number of consecutive subarrays; Global optimization: Reduce trip prediction losses Damage during driving With waiting loss and overall component alignment loss We obtain the total global loss by weighted summation. , represented as: ; in, and This is a hyperparameter.

5. A bus travel time prediction system based on multi-component decomposition and attention fusion, characterized in that, The bus travel time prediction system includes: Date selection module: The AGNES algorithm is used to cluster historical departure data to generate a set of departure frequency patterns for different operating modes; the real-time departure data of the current date is matched with the departure frequency patterns, the most similar historical date data is selected, and an adaptive input sequence is constructed. Component encoding module: Extracts road segment travel features from the input sequence through a spatiotemporal convolutional layer, outputting travel features that include spatial neighborhood correlation and temporal continuity; extracts waiting features of origin stations and transfer stations from the input sequence through a masked convolutional layer, outputting waiting features, wherein the mask matrix is ​​predefined by the station type; including: Construct a binary mask matrix , in, = ; By performing masked convolution operations, feature extraction is performed based on site characteristics to obtain waiting features. , is represented as: ; Among them, X w , These are the waiting time input vector and the mask input, respectively. This represents a convolutional kernel that is spatially adjacent to three segments and temporally continuous for six steps, through the mask matrix. Hard-coded site prior knowledge guides the model to focus on feature extraction from origin and transfer stations; Feature fusion module: This module performs weighted fusion of the travel features and waiting features using travel attention, including: Introduce a trainable parameter Trainable parameters The initial values ​​are randomly initialized in the range [0,1], and the waiting time weight is constrained. In order to protect ; Weighted by driving time Weighted by waiting time Linear weighting is performed in the feature space to initially fuse the travel features and waiting features, as shown below: ; in, Indicates the characteristics after fusion. For travel time characteristics, This refers to the waiting time characteristic; During model training, the parameters are processed using the backpropagation algorithm. Dynamic optimization is performed to adaptively learn and balance the contributions of driving time and waiting time to the total trip time; The self-attention mechanism is used to model the long-range dependencies between different road segments and stations, and global fusion features are output, including: The fused features As input, three sets of learnable projection matrices are used respectively. , , Generate query, key, and value tensors; The calculation method for each attention point is as follows: ; Among them, the definition =6; The outputs of the six attention heads are concatenated to obtain a global fusion feature that models long-range spatiotemporal dependence. , is represented as: ; in, , ,… These represent the outputs of attention heads 1 through 6 respectively; The corresponding time code is calculated using historical average time, and a dot product operation is performed with the global fusion feature to generate the final comprehensive feature used for prediction; including: Based on historical data, calculate each road segment Average driving time And the average waiting time of road sections with waiting times. Average waiting time for road sections with no waiting time Set to 0; Calculate the statistical travel time for each road segment. and based on The start time of the journey for each segment is calculated cumulatively. ; Construct a spatiotemporal mask matrix This corresponds to 40 road segments and 36 time intervals divided into 30-minute intervals starting at 6:00 AM; the spatiotemporal mask matrix... Start time of the journey on the middle section The value at the corresponding time interval position is set to 1, and the rest are 0; The global fusion features Element-wise multiplication with the spatiotemporal mask matrix NM yields the final feature enhanced with contextual information. ; Loss Alignment and Prediction Module: Calculates ride loss and waiting loss separately, supervises the accuracy of component-level prediction, coordinates the consistency between component-level prediction and total travel time prediction, and finally outputs the total travel time prediction result.

6. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it is used to implement the bus travel time prediction method based on multi-component decomposition and attention fusion as described in any one of claims 1-4.