Cross-layer low-rank fusion space-time traffic flow prediction modeling method

By adopting a cross-layer low-rank fusion spatiotemporal traffic flow prediction modeling method, the problems of information fragmentation and feature redundancy between layers are solved, the stability and interpretability of traffic flow prediction are realized, and the long-term prediction performance is improved.

CN121564962APending Publication Date: 2026-02-24NANJING AUDIT UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511737764.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-25
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing traffic flow prediction methods suffer from fragmented information between layers, feature redundancy, unstable decoding processes, and insufficient model interpretability when dealing with non-Euclidean graph-structured traffic networks. As a result, they struggle to achieve both long-term prediction performance and stability in complex scenarios.

Method used

A cross-layer low-rank fusion spatiotemporal traffic flow prediction modeling method is adopted. Through standardization, missing data completion, temporal embedding, spatial embedding and temporal convolution enhancement, a two-stage attention modeling structure of nodes and agents is generated. Adjacent layer output pairing, feature concatenation and rank configuration are performed. Low-rank bottleneck compression, gating weight registration and gating residual superposition are performed to form a cross-layer low-rank bottleneck and gating residual fusion structure. Temporal dimension aggregation and routing weight initialization are performed to generate a terminal hybrid expert steady-state decoding structure.

Benefits of technology

It improves the consistency and stability of cross-layer information fusion, is applicable to multi-horizon prediction, and enhances the accuracy and interpretability of traffic flow prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121564962A_ABST
    Figure CN121564962A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of intelligent traffic and deep learning modeling, and particularly relates to a cross-layer low-rank fusion space-time traffic flow prediction modeling method. The method comprises the following steps: firstly, carrying out standardization, complementation and coding processing on historical traffic data, time labels and road topology, and constructing a space-time embedding representation; then, through node aggregation and agent number adaptive estimation, a double-stage attention modeling structure is established so as to reduce the calculation complexity; efficient fusion of cross-layer information is realized through output pairing and feature splicing of adjacent layers in combination with low-rank bottleneck compression and gating residual fusion; and a robust prediction result is generated by adopting a hybrid expert decoding structure and combining time dimension aggregation, routing weight initialization and multi-target constraint. According to the method, the problems of high-dimensional redundancy, calculation bottleneck and distribution drift in large-scale traffic flow prediction are effectively solved, and the prediction precision and the model efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent transportation and deep learning modeling technology, and in particular to a cross-layer low-rank fusion spatiotemporal traffic flow prediction modeling method. Background Technology

[0002] With the development of Intelligent Transportation Systems (ITS), urban road sensors have generated massive amounts of traffic flow data. How to accurately predict future traffic conditions from multi-source spatiotemporal data is a key issue in achieving intelligent traffic scheduling, alleviating congestion, and optimizing travel decisions.

[0003] Traditional traffic flow prediction methods mainly include statistical model-based methods (such as ARIMA and VAR) and deep learning methods based on neural networks. Statistical models can capture short-term linear features of time series, but they struggle to handle the complex nonlinearities and spatial dependencies in traffic networks. Subsequently, recurrent neural networks (RNNs, LSTMs) and convolutional neural networks (CNNs) were introduced to model temporal dependencies and local spatial features. However, these methods are limited in their performance when dealing with traffic networks with non-Euclidean graph structures, failing to fully capture the dynamic relationships between different road nodes.

[0004] In recent years, Graph Neural Networks (GNNs) and Transformer architectures have been widely used for spatiotemporal modeling. Graph Convolutional Networks (GCNs) can characterize the topological dependencies of traffic networks, while spatiotemporal Transformers possess powerful temporal modeling and long-dependency learning capabilities. Existing research methods such as DCRNN, STGCN, and GraphWaveNet have improved prediction performance to some extent, but still have the following shortcomings: 1. Information fragmentation between layers: Although the multi-layer encoder structure can gradually extract high-order features, there is a lack of effective information fusion between layers, which makes it difficult to coordinate deep representation and shallow semantics.

[0005] 2. High-dimensional feature redundancy: Spatiotemporal features may have dimensional redundancy and repeated calculations at different levels, resulting in excessively large model parameters and decreased generalization performance.

[0006] 3. Unstable decoding process: Traditional Transformer decoders often use global feedforward networks (FFN), which are prone to accumulating errors in multi-step prediction, leading to a significant decrease in long-term prediction accuracy.

[0007] 4. Insufficient model interpretability and stability: When there is heterogeneity in multi-scale traffic flow features, the feature fusion and attention allocation mechanisms of existing models cannot guarantee stable output.

[0008] Existing technologies have proposed structures such as the Time Lag Aware Spatiotemporal Transformer (TLAST) to explicitly capture the temporal differences between historical and current moments. These models achieve hierarchical modeling of temporal and spatial dependencies through time lag embedding and spatial proxy attention mechanisms. However, these methods primarily focus on spatiotemporal feature interactions within a single layer, with limited handling of cross-layer feature fusion and information compression. Furthermore, they lack dedicated mechanisms for prediction stability and expert routing optimization at the decoding end, making it difficult to balance long-term prediction performance and model interpretability in complex scenarios.

[0009] Furthermore, this invention belongs to the field of intelligent transportation and deep learning modeling technology. Existing methods for modeling historical multi-node traffic sequences, time labels, and road topology-driven approaches typically revolve around standardization, missing data completion, label encoding, temporal and spatial embedding construction, and temporal convolutional enhancement processing, combined with node aggregation and attention initialization. These methods suffer from limitations such as a lack of unified registration for adjacent layer output pairing, a lack of adaptive linkage in rank configuration registration and low-rank bottleneck compression and reconstruction, and a lack of traceable constraints in gating weight registration and gating residual superposition. Existing methods often directly perform node-to-proxy attention calculation and aggregation, proxy-to-node backpropagation, feedforward transformation, and normalization processing after the standardization embedding and temporal convolutional enhancement structure. In scenarios involving adjacent layer output pairing and cross-layer low-rank fusion, insufficient coordination exists between routing weight initialization and temperature parameter registration oscillations, horizon-weighted robust loss construction, and sparsity and utilization constraint injection operations, making it difficult to achieve stable implementation of the end-point hybrid expert steady-state decoding structure. For the joint processing of time tags and adjacent layer output pairing and rank configuration registration, existing technologies generally have shortcomings in the consistency of cross-stage processing such as alignment index connection, feature splicing adjudication, gating weight registration and residual ratio injection, as well as normalization and balancing. It is difficult to form a continuous process from alignment index to feature splicing to rank configuration registration to gating residual superposition to routing weight initialization to expert selection to readout shaping to record registration in application scenarios of adjacent layer output pairing and time dimension aggregation. This results in insufficient stability and consistency of the end-stage hybrid expert steady-state decoding structure. Summary of the Invention

[0010] To address the aforementioned technical problems, this invention provides a cross-layer low-rank fusion spatiotemporal traffic flow prediction modeling method, comprising: The system acquires historical multi-node traffic sequences, time labels, and road topology, performs standardization, missing data completion, and label encoding, and performs temporal embedding, spatial embedding construction, and temporal convolutional enhancement processing to generate standardized embedding and temporal convolutional enhancement structures. Based on the standardized embedding and temporal convolutional enhancement structure, node aggregation, agent number adaptive estimation and attention initialization are performed. Node-to-agent attention calculation and aggregation, agent-to-node backpropagation and feedforward transformation, and normalization are performed to generate a two-stage attention modeling structure for nodes and agents. The two-stage attention modeling structure of nodes and agents is obtained. Adjacent layer output pairing, feature splicing and rank configuration registration are performed. Low-rank bottleneck compression, reconstruction and gating weight registration are performed. Gating residual superposition, normalization and balancing are performed to generate a cross-layer low-rank bottleneck and gating residual fusion structure. Based on the cross-layer low-rank bottleneck and gated residual fusion structure, time-dimensional aggregation and routing weight initialization are performed, temperature parameter registration, expert selection and feedforward calculation are performed, residual ratio injection, readout shaping and horizon-weighted robust loss construction, sparsity and utilization constraint injection are performed, and a terminal hybrid expert steady-state decoding structure is generated.

[0011] Furthermore, historical multi-node traffic sequences, time tags, and road topology include: The historical multi-node traffic sequence is a multi-channel time series composed of speed, flow, occupancy and related observations continuously reported by road monitoring terminals, including node identifiers, sampling times and observation types; the time tags include intraday location, intraweek location, holiday attributes, weather conditions and event trigger attributes, forming tag fields and tag vocabulary; the road topology describes the connectivity, direction, hierarchy and adjacency of road segments, and includes road segment length, number of lanes, speed limit, signal control phase and historical coupling degree.

[0012] Furthermore, the process of standardization, missing test completion, and label encoding also includes: Standardization processing includes extracting the central trend and dispersion amplitude based on historical distribution statistics to generate a standardized parameter set and performing dimensionless mapping and scale unification; Missing test completion includes prioritizing backfilling based on the same node and the same tier, and generating an aligned index by interpolating adjacent nodes in the topology.

[0013] Furthermore, the temporal convolution enhancement process also includes: The temporal convolution enhancement process includes constructing a one-dimensional convolution stack using a temporal convolutional network, performing slice-by-slice filtering for each node and each time slice, and switching to a shared channel and registering a shared node flag when a node is detected to be invalid for an extended period.

[0014] Furthermore, the process of performing node aggregation, adaptive estimation of the number of agents, and attention initialization also includes: The node aggregation includes generating an aggregation weight field based on topological adjacency weights, quality scores, and shared node tags, and enabling a buffered connection strategy at the endpoints of the sliding window. The adaptive estimation of the number of agents includes generating an agent size score based on the structural diversity index, the time fluctuation index and the quality score, and using a smooth interpolation strategy to generate a soft gating ratio. The attention initialization includes reading multi-scale temporal channels from the temporal convolutional enhancement structure to construct a temporal prior surface, reading adjacency weights from spatial embedding to construct a topological prior surface, and assigning temporal bias and topological bias weight ratios to each head of the multi-head attention.

[0015] Furthermore, the process of performing node-to-agent attention calculation and aggregation, and agent-to-node backhaul and feedforward transformation also includes: The node-to-agent attention calculation and aggregation includes aligning node-level queries with agent-level key values ​​and sending them into a multi-head channel, and implementing gate control suppression for abnormally high correlation. The agent-to-node backhaul includes restoration based on the node-to-agent mapping registration and agent ratio, taking into account the adjacent relationship weight and time prior surface. When the agent coverage is insufficient, a temporary neighborhood expansion is triggered. The feedforward transformation includes performing nonlinear mapping from the feedforward network and reading the initial value of the residual superposition ratio from the quality score and updating it smoothly within the window.

[0016] Furthermore, the process of pairing adjacent layer outputs, concatenating features, and registering rank configuration also includes: The adjacent layer output pairing includes checking the node dimension and time dimension piece by piece based on the alignment index. When there is a gap, the nearest agent is retrieved from the node to agent mapping registration and the gap is filled based on the adjacent relationship weight. The feature splicing adopts the splicing order of node field first and proxy field second and performs channel homogenization. When the window score is low, it enters the weight reduction path. The rank configuration registration includes generating rank scores based on channel stability and mutual information characterization, and registering compressible dimensions and candidate rank sources.

[0017] Furthermore, the process of performing low-rank bottleneck compression, reconstruction, and gating weight registration also includes: The low-rank bottleneck compression and reconstruction includes channel decomposition using an orthogonal decomposition paradigm, principal component truncation following the suppression flag in the rank configuration registration, and reconstruction weight stacks based on proxy source mapping and adjacency relation weights established in the node dimension and time dimension, respectively. The gating weight registration includes generating initial gating values ​​based on bottleneck representation, channel statistics of reconstruction representation, shared constraints, weight reduction factors in quality logs, and residual ratio registration, and then reducing or relaxing them according to the suppression flag in the rank configuration snapshot.

[0018] Furthermore, the process of gating residual superposition, normalization, and balancing also includes: The gated residual superposition includes weighting the bottleneck representation, reconstruction representation and node domain feedforward representation according to the gate weight and residual ratio, first performing short window mixing in the time dimension and then performing neighborhood diffusion in the node dimension. The normalization includes migrating parameters from the previous stable window and updating them lightly, and the balancing process includes balancing the energy level and channel proportion of the fused channel and the original splicing buffer.

[0019] Furthermore, the process of temperature parameter registration, expert selection, and feedforward calculation also includes: The temperature parameter registration includes setting the temperature scale, with the initial temperature value taken from the balance statistics record and adjusted according to the window score; The expert selection process includes selecting active experts from the candidate expert set based on the route weight initialization table and route temperature registration, and generating a sparse mask and an expert percentage list. The feedforward computation includes performing nonlinear mapping using a hybrid expert structure.

[0020] The key innovations of this invention include: (1) The linkage between adjacent layer output pairing and feature splicing, rank configuration registration and low-rank bottleneck compression and reconstruction, in the same link, from adjacent layer output pairing to feature splicing, and then rank configuration registration drives low-rank bottleneck compression and reconstruction to generate cross-layer low-rank fusion sequence.

[0021] (2) The controlled fusion mechanism of gating weight registration and gating residual superposition and normalization and balancing is to use gating weight registration to constrain the gating residual superposition, and to form a cross-layer low-rank bottleneck and gating residual fusion structure after normalization and balancing.

[0022] (3) The end joint process of time dimension aggregation and routing weight initialization and temperature parameter registration, expert selection and feedforward calculation and residual ratio injection and readout shaping and horizon weighted robust loss construction and sparsity and utilization constraint injection operations generates the end hybrid expert steady-state decoding structure.

[0023] The following are its main beneficial effects: (1) For the channel obtained by pairing the outputs of adjacent layers, feature splicing and rank configuration registration jointly drive low-rank bottleneck compression and reconstruction. The cross-layer low-rank fusion sequence can be directly called in the subsequent gating weight registration and gating residual superposition, which improves the consistency of cross-layer alignment and compression reconstruction and is suitable for the scenario of pairing the outputs of adjacent layers.

[0024] (2) The gated residual superposition is performed under the gated weight registration constraint, and the cross-layer low-rank bottleneck and gated residual fusion structure is obtained after normalization and balancing. This structure is referenced as a routing facet in time dimension aggregation, routing weight initialization and temperature parameter registration to form a stable channel ratio and energy level balancing, which is suitable for controlled reading after cross-layer low-rank fusion.

[0025] (3) Through time dimension aggregation, routing weight initialization and temperature parameter registration, it enters expert selection and feedforward calculation, then through residual ratio injection and readout shaping, combined with horizon weighted robust loss construction and sparse and utilization constraint injection operations, the output end hybrid expert steady-state decoding structure provides an end-to-end decoding link for continuous window operation of historical multi-node traffic sequences, time labels and road topology, which is suitable for multi-horizon prediction scenarios. Attached Figure Description

[0026] Figure 1 The system architecture of a cross-layer low-rank fusion spatiotemporal traffic flow prediction modeling method provided in this application embodiment; Figure 2 A flowchart illustrating a cross-layer low-rank fusion spatiotemporal traffic flow prediction modeling method provided in an embodiment of this application; Figure 3 A software architecture diagram provided for an embodiment of this application; Figure 4 A preferred internal structure diagram provided for an embodiment of this application; Figure 5 This is a preferred internal structure diagram of the end-to-end hybrid expert steady-state decoding module provided in the embodiments of this application. Detailed Implementation

[0027] like Figure 1 As shown, Figure 1 This application provides a system architecture for a cross-layer low-rank fusion spatiotemporal traffic flow prediction modeling method. The architecture begins with an input data preprocessing module (101) to receive raw historical traffic data; it is then connected to a spatiotemporal embedding module (102) responsible for constructing a unified spatiotemporal representation. The output of the spatiotemporal embedding module (102) is fed into an encoder module (103) and a cross-layer low-rank fusion module (104), respectively. The encoder module (103) includes a proxy attention mechanism for efficient spatial modeling, and the cross-layer low-rank fusion module (104) is used to achieve efficient compression and fusion of information between layers. The outputs of the above modules are collectively input into a steady-state MoE decoding module (105) for robust decoding through multi-expert collaboration. Finally, a prediction output module (106) generates the future traffic flow prediction results. This block diagram clearly illustrates the complete pipeline of this method from data preprocessing, feature extraction, fusion encoding to steady-state decoding.

[0028] 1. Input data preprocessing module (101) normalizes the original traffic sensor sequence, fills in missing values, segments it into sliding windows, and generates a time index.

[0029] The output is a standardized input sequence and time stamp.

[0030] 2. The spatiotemporal embedding module (102) encodes time tags into time embeddings and uses the road topology map to generate node embeddings, which are then spliced ​​together to form input features.

[0031] 3. The encoder module (103) contains a multi-layer Transformer structure, each layer consisting of temporal convolutional augmentation (TCN), adaptive spatial proxy attention, and a feedforward layer.

[0032] The encoder outputs interlayer feature sequences F1, F2, ..., FL.

[0033] 4. The cross-layer low-rank fusion module (104) concatenates the outputs of adjacent layers, achieving cross-layer feature compression and reconstruction through a linear low-rank bottleneck structure, and outputs fused features using gated residuals and normalization operations. This module introduces inter-layer feature interaction and parameter compression based on TLAST, significantly improving the consistency of network expression and computational efficiency.

[0034] 5. Decoder module (105) aggregates the fused features in the time dimension and inputs them into the Stable MoE feedforward network, dynamically selects experts according to weights and outputs the prediction result Y.

[0035] The decoding end uses only a single-layer expert structure to ensure the stability of multi-step prediction.

[0036] 6. Prediction output module (106) outputs the predicted values ​​for multiple future time steps (such as traffic flow / speed for the next 12 steps) and calculates the loss function for backpropagation training.

[0037] In a preferred embodiment, referring to Figure 1 This is a flowchart illustrating a cross-layer low-rank fusion spatiotemporal traffic flow prediction modeling method provided in an embodiment of the present invention. The flowchart may include at least steps S100-S400: S100: Obtain historical multi-node traffic sequences, time labels and road topology, perform standardization, missing data completion and label encoding processing, perform temporal embedding, spatial embedding construction and temporal convolution enhancement processing, and generate standardized embedding and temporal convolution enhancement structures. S200, based on standardized embedding and temporal convolution enhancement structure, performs node aggregation, adaptive estimation of agent number and attention initialization processing, performs node-to-agent attention calculation and aggregation, agent-to-node backpropagation and feedforward transformation, normalization processing, and generates a two-stage attention modeling structure for nodes and agents. S300: Obtain the two-stage attention modeling structure of nodes and agents, perform adjacent layer output pairing, feature splicing and rank configuration registration processing, perform low-rank bottleneck compression, reconstruction and gate weight registration processing, perform gate residual superposition and normalization and balancing processing, and generate a cross-layer low-rank bottleneck and gate residual fusion structure. S400, based on the cross-layer low-rank bottleneck and gated residual fusion structure, performs time-dimensional aggregation and routing weight initialization processing, performs temperature parameter registration, expert selection and feedforward calculation processing, performs residual ratio injection, readout shaping and horizon-weighted robust loss construction, sparsity and utilization constraint injection processing, and generates end-end hybrid expert steady-state decoding structure.

[0038] S100: Obtain historical multi-node traffic sequences, time labels and road topology, perform standardization, missing data completion and label encoding processing, perform temporal embedding, spatial embedding construction and temporal convolution enhancement processing, and generate standardized embedding and temporal convolution enhancement structures. Based on the historical multi-node traffic sequence, the time tags, and the road topology as inputs for this step, clock alignment and node identifier mapping are first completed at the access end, constructing a unified timeline and index set for the node and time layers. The historical multi-node traffic sequence is a multi-channel time series composed of speed, flow, occupancy, and related observations continuously reported by road monitoring terminals, including node identifiers, sampling times, and observation types. The time tags include intraday location, intraweek location, holiday attributes, weather conditions, and event trigger attributes, forming tag fields and tag dictionaries. The road topology describes road segment connectivity, direction, hierarchy, and adjacency relationships, and includes metadata such as road segment length, number of lanes, speed limits, signal control phases, and historical coupling. Specifically, the access end segments the raw data into fixed windows. If clock jumps, duplicate times, or cross-day boundaries are detected, an abnormal window is registered and a rollback strategy is triggered. The rollback content includes the window index, node set, anomaly type, and repair status, recorded in the quality log and event log for subsequent checks.

[0039] Specifically, standardization processing performs dimensionless mapping and scale unification for similar observations at the same node within a sliding window. Before processing, the central trend and dispersion magnitude are extracted from historical distribution statistics to generate a standardized parameter set. Then, the original observations are mapped to a unified scale space, and a standardized sequence is output. When an observation is identified as an outlier, it is replaced by robust estimates from neighboring time slices and similar nodes, and the replacement source, replacement ratio, and trigger threshold are written to the quality log. Missing data completion runs across hole moments and breakpoint periods. The process involves generating missing data markers, calling backfilling of the same node and level first, and selecting candidates from adjacent nodes in the topology according to their adjacent weights and completing the interpolation when backfilling fails. The completion process and the original markers are retained together to form an alignment index. If the proportion of missing data in a window exceeds the upper limit, the window is marked with insufficient observations and enters the deweighting channel, and a quality score is generated simultaneously. Furthermore, after standardization and completion, a standardized sequence, missing test markers, alignment index, and quality score are generated and written to the metadata buffer. The alignment index will provide a sampling and aggregation reference position in the subsequent node aggregation and agent number adaptive estimation in S200, and the quality score is used for weight cold start during attention initialization.

[0040] Furthermore, the label encoding constructs a trainable index mapping for the time labels, encoding intraday and intraweekly positions into dense vectors. Holiday attributes, weather ratings, and event trigger attributes are registered in the label vocabulary and transcribed into a dense format. Segment start and end tags are introduced at cross-day and cross-week boundaries to prevent cross-boundary mismatches. When multiple label conflicts exist, they are rearranged according to priority rules, and the reasons for conflict resolution are recorded. These records are persisted along with window metadata. Subsequently, a time embedding is constructed, consisting of a periodic channel and an event channel: the periodic channel generates multi-scale phase characterizations for intraday and intraweekly cycles, and the event channel generates short-term memory characterizations for abnormal weather and activity states. The two channels are concatenated in the time dimension and then nonlinearly projected to obtain the time embedding vector. Understandably, the time embedding and label encoding output are strictly aligned with the alignment index, forming a time semantic alignment sequence. This sequence will be used as a time scale source registration item in the rank configuration registration of S300, and will also provide time priors in the attention initialization phase of S200.

[0041] Specifically, spatial embedding constructs node vector representations and adjacency weights around the road topology. First, the topology direction is normalized, and self-loop edges are registered. Repair edges are created for broken or suspended nodes and written to the topology log. Then, multi-source similarity is obtained by integrating road segment length, intersection type, signal control phase, and historical traffic coupling. Adjacency weights are generated using hierarchical mapping, and the weights are sparsified to retain the core adjacency set. When a multi-layer road structure exists, inter-layer edges are registered with inter-layer types to ensure that both intra-layer and inter-layer channels are accessible downstream. Further, spatial embedding is normalized and bias-removed to alleviate weight stacking at dense nodes. When topology updates cause structural breaks, temporary subgraph construction and cold-start node relocation are triggered, and the relevant processes and time consumption are recorded in the topology log and maintenance log. Spatial and temporal embeddings are then concatenated hourly with the normalized sequence to generate candidate input tensors. These candidate input tensors completely follow alignment indices in both the time and node dimensions, incurring no additional alignment overhead for subsequent calculations, and are entered into the fusion stack as a precursor to the normalized embedding.

[0042] After completing the above characterization, temporal convolutional enhancement processing is performed. The framework adopts a temporal convolutional network (TCN), constructing a one-dimensional stack of convolutions along the time direction. It covers short-range and mid-range dynamics through dilated receptive fields and causal padding. Specifically, the processing chain includes several parallel sub-branches, each selecting different receptive field spans and strides. The outputs of each sub-branch are merged in the time dimension and integrated by a gating unit and a normalization unit. A learnable temporal alignment offset registration unit is placed before convolution to mitigate phase errors at window boundaries. A residual mapping unit is placed after convolution to reduce accumulated errors caused by distribution drift. When a high proportion of interpolation traces exist in the input window, a stronger smoothing kernel is automatically switched, and relevant thresholds and selection records are entered into the quality log. Understandably, the temporal convolutional augmentation process performs slice-by-slice filtering for each node at runtime. The filtered temporal features and the original representation are then fed into the fusion stack together, preventing short-range details from being overwritten. When a node remains in an invalid state for an extended period, the sub-branch switches to a shared channel and registers a shared node tag. This tag will be referenced in the attention initialization and node-to-proxy attention calculation in S200. Through the above processing, a temporal convolutional augmentation structure is generated for the encoder side. This structure retains multi-scale channels in the time dimension and shared tags and quality weights in the node dimension.

[0043] Furthermore, standardization, missing test completion, label encoding, temporal embedding, and spatial embedding construction run sequentially in a unified pipeline, with anomaly handling and rollback strategies implemented throughout. When the quality score falls below a threshold, the corresponding window enters a weighted branch, and the weighting weight, along with the score, is written to the metadata buffer and made visible to the downstream attention initialization phase. When a structural change occurs in the topology log, spatial embedding is immediately reconstructed, triggering temporal embedding realignment. The reconstruction process and realignment time are recorded in the operations log. When label and topology conflicts occur simultaneously in the same window, topology conflicts are handled first, and label reordering is frozen. The frozen state is carried downstream along with the window metadata to avoid repeated handling across different stages. The training and inference phases share this pipeline. Data arrival events trigger passive execution, while topology updates and holiday template changes trigger active execution. Trigger records are stored in the event log and propagated to subsequent steps.

[0044] The output of this step is a standardized embedding and a temporal convolutional augmentation structure. The standardized embedding includes a spatiotemporal representation constructed through standardization, missing test completion, label encoding, temporal embedding, and spatial embedding, while also carrying alignment indices, quality scores, and label conflict resolution records. The temporal convolutional augmentation structure is a multi-scale temporal feature set obtained after processing the temporal dimension by a temporal convolutional network, carrying shared node labels and residual mapping registration. Both serve as inputs to S200, where the standardized embedding is used for node aggregation and adaptive estimation of the number of agents, and the temporal convolutional augmentation structure is used for attention initialization and node-to-agent attention calculation and aggregation. Simultaneously, the temporal semantic alignment sequence in the standardized embedding is used in the rank configuration registration in S300 to register the temporal scale source of low-rank bottlenecks across layers. The temporal convolutional augmentation structure is invoked in the temporal dimension aggregation and route weight initialization stage of S400 to generate initial route temperature and residual ratio references, and provides multi-scale temporal priors in the feedforward calculation and readout shaping stages, achieving stable connection with subsequent main steps.

[0045] In summary, the technical effects of this step are as follows: through unified input representation construction and temporal convolution enhancement, the original observations, time labels, and road topology are semantically aligned on the same time axis and node axis to form a robust temporal structure, thereby obtaining high-quality input that is adapted to downstream attention modeling, low-rank fusion, and steady-state decoding.

[0046] S200, based on standardized embedding and temporal convolution enhancement structure, performs node aggregation, adaptive estimation of agent number and attention initialization processing, performs node-to-agent attention calculation and aggregation, agent-to-node backpropagation and feedforward transformation, normalization processing, and generates a two-stage attention modeling structure for nodes and agents. The standardized embedding and the temporal convolutional enhancement structure are used as inputs to this step. The standardized embedding carries the alignment index, quality score, and label conflict handling records, while the temporal convolutional enhancement structure carries multi-scale temporal channels, shared node labels, and residual mapping registration. Specifically, the alignment index is first read to complete the unified alignment of the node dimension and the temporal dimension. If a window out-of-bounds error or missing node mapping is found, an anomaly is registered in the event log and a rollback rule is triggered. The rollback rule prioritizes migrating the alignment parameters from the previous stable window, and the migration process and difference statistics are written to the quality log. The node aggregation stage then begins. Semantically, node aggregation refers to the learnable weight synthesis of multiple node representations of the same local road subgraph. The weights come from topological adjacency weights, quality scores, and shared node labels. These three are normalized and mapped on the same scale to generate an aggregated weight field. The aggregation operation runs on a sliding window along the time axis, with a buffering and connection strategy enabled at the window endpoints. This strategy suppresses boundary spikes by fading in and out at the endpoints. When long-term sparse observations of a subgraph are detected, node aggregation switches to a reduced-weight channel and attaches a reduced-weight flag to the topological log. This reduced-weight flag directly affects the initial facet of the query weights in the subsequent attention initialization stage.

[0047] Furthermore, an adaptive estimation of the number of agents is performed. Semantically, this estimation dynamically determines the size of the agent set based on the complexity and temporal fluctuation of the current window and subgraph. This process first extracts the structural diversity index and temporal fluctuation index from the local representation after node aggregation, and then constructs an agent size score by combining the quality score. When the score is within the threshold range, a smooth interpolation strategy is enabled to generate a soft-gated proportion between two integer sizes, which is then transcribed into a discrete agent set and its proportion during agent sampling. If the score is below the lowest threshold range, the system runs with the minimum agent size and registers a safety net flag, which is used to limit the burst of downstream attention. If the score exceeds the highest threshold range, the system runs with the maximum agent size and registers an expansion flag, which is used to indicate that a higher rank space is reserved in the feature stitching stage of S300. Subsequently, agent initialization is established. Agent initialization extracts representative spatiotemporal segments from the node aggregation output and completes candidate selection according to the joint criteria of geographical coverage and temporal diversity. Candidates are downweighted in segments with lower quality scores. The selection results, along with the agent size, are written to the metadata buffer to form three records: agent set, agent proportion, and agent source mapping. These three records can be directly referenced in subsequent sub-steps of this step.

[0048] After completing the proxy set, attention initialization begins. Semantically, attention initialization refers to setting a learnable initial distribution for query weights, key weights, and value weights, and injecting temporal and topological priors into the initial facets. Specifically, the multi-scale temporal channels of the temporal convolutional enhancement structure are read to construct temporal prior facets; spatial adjacency weights are read to construct topological prior facets; the two types of priors are concatenated at the channel dimension and then nonlinearly projected to obtain the initial attention facets. The attention mechanism employs multi-head attention (MHA). After the initial establishment of the multi-head channels, a weight ratio of temporal and topological biases is assigned to each head. When a quality score is found to be in a de-weighted channel, the temporal bias of that head decreases and is written to the bias log. If the node sharing flag is true, the head corresponding to that node enables shared channel constraints, and the constraint coefficients are registered in the shared table. After initialization, the node-to-agent attention calculation and aggregation process begins. During runtime, node-level queries and agent-level key-value pairs are aligned along the timeline and sent to a multi-head channel. Within each channel, correlation is calculated independently by head, and aggregation is performed outside the channel. To suppress occasional spikes, gate control suppression is implemented for abnormally high correlations occurring in a single head. The trigger threshold and suppression ratio for gate control suppression are recorded in the quality log. When the agent set is in an extended flag state, the number of channels is temporarily increased. Channel increases are automatically recycled after the current window ends, and the recycling time is recorded in the operation and maintenance log. The output of the node-to-agent attention calculation and aggregation is an agent aggregation representation and a node-to-agent mapping registration. The agent aggregation representation is used for subsequent agent-to-node backhaul, and the node-to-agent mapping registration is stored in the metadata buffer for use in S300's adjacent layer output pairing and feature concatenation, serving as a source index for cross-layer pairing.

[0049] Subsequently, proxy-to-node backhaul is performed. Semantically, proxy-to-node backhaul refers to restoring the proxy aggregation representation to the node domain based on the node-to-proxy mapping registration and proxy ratio. During the restoration process, both the adjacency weight and the time prior facet are considered, thus forming the node domain backhaul representation. If the backhaul target node is in a shared channel, the backhaul gain is reduced according to the sharing constraint, and the reduction is recorded in the shared table. If insufficient proxy coverage is detected during the backhaul, temporary neighborhood expansion is triggered. Neighborhood expansion uses the patch edge set in the topology log, and the expansion path and expansion ratio are written to the topology log for subsequent window reuse. After the backhaul is completed, the process proceeds to the feedforward transform. Semantically, the feedforward transform refers to performing position-by-position nonlinear mapping and channel scaling on the node-domain backhaul representation. It employs a feedforward network (FFN). The feedforward transform includes an expansion layer and a compression layer. The expansion layer is responsible for channel expansion and point-state activation in the time dimension, while the compression layer is responsible for channel convergence and forming residual superposition with the original node-level representation. The residual superposition ratio is read from the quality score as its initial value and smoothly updated over time within a window. If residual drift is detected to be higher than a threshold within a consecutive window, a drift event is registered in the event log, and the expansion layer channel ratio is reduced in the next window. The feedforward transform outputs the node-domain feedforward representation while retaining the residual ratio registration, which is directly referenced in the residual ratio injection stage of S400.

[0050] After the node domain feedforward representation is formed, normalization is performed. Semantically, normalization refers to constructing stable statistics in the channel and node dimensions and using these to correct the representation for scaling and bias. Layer Normalization (LN) is used for normalization; normalization parameters are migrated from the previous stable window and lightly updated within the current window, with the update rate and stability recorded in the operation log. When the input comes from a deweighted channel, the normalization bias term is reduced along with the quality score, and the reduction coefficient is recorded in the metadata buffer. After the normalization process, a node domain normalized representation is formed. At this point, node aggregation, adaptive estimation of the number of agents, attention initialization, node-to-agent attention calculation and aggregation, agent-to-node backpropagation, feedforward transformation, and normalization complete a closed-loop operation. This closed-loop operation is shared in inference and training modes. Triggering conditions include data arrival events, topology update events, and holiday template change events. Trigger records are stored in the event log and propagated downstream.

[0051] The output of this step is a two-stage attention modeling structure for nodes and agents. The node-agent two-stage attention modeling structure includes fields such as agent set, agent proportion, agent source mapping, node-to-agent mapping registration, agent aggregation representation, node domain backpropagation representation, node domain feedforward representation, residual proportion registration, and node domain normalization representation. The agent set, agent proportion, and agent source mapping are entered into the adjacent layer output pairing and feature concatenation of S300, serving as the agent-side source during cross-layer pairing. The node-to-agent mapping registration and agent aggregation representation are read by the rank configuration registration of S300, providing source facets for low-rank bottleneck compression and reconstruction. The node domain feedforward representation and node domain normalization representation, along with the residual proportion registration, are entered into the time dimension aggregation, routing weight initialization, temperature parameter registration, expert selection, feedforward calculation, residual proportion injection, readout shaping, horizon-weighted robust loss construction, and sparsity and utilization constraint injection operations of S400, serving as input facets for the end-stage hybrid expert steady-state decoding structure. Simultaneously, the shared table, bias log, and quality log of this step can be retrieved in the gated weight registration and gated residual superposition stages of S300, serving as prior references for the gated weights.

[0052] In summary, the technical effects of this step are as follows: Through two-stage attention modeling of the node domain and the proxy domain, the spatiotemporal information from the standardized embedding and temporal convolutional enhancement structure is converged at the proxy layer and backpropagated at the node layer, forming a stable and adjustable attention surface and feedforward surface. This provides a directly schedulable input structure for subsequent cross-layer low-rank bottleneck and gated residual fusion as well as end-point hybrid expert steady-state decoding.

[0053] S300: Obtain the two-stage attention modeling structure of nodes and agents, perform adjacent layer output pairing, feature splicing and rank configuration registration processing, perform low-rank bottleneck compression, reconstruction and gate weight registration processing, perform gate residual superposition and normalization and balancing processing, and generate a cross-layer low-rank bottleneck and gate residual fusion structure. The node-agent two-stage attention modeling structure is used as input for this step. This structure includes the agent set, agent percentage, agent source mapping, node-to-agent mapping registration, agent aggregation representation, node domain backpropagation representation, node domain feedforward representation, residual percentage registration, and node domain normalization representation. Alignment indexes and temporal semantic alignment sequences are read from the normalized embedding in S100, and shared constraints, bias percentages, and window scores are read from the shared table, bias log, and quality log in S200. Specifically, adjacent layer output pairing is first performed. Semantically, adjacent layer output pairing refers to matching the pairable channels of the preceding and following layers within the same window one by one. The processing order is: reading the alignment index, checking the node dimension and time dimension piece by piece, triggering candidate mapping when gaps exist, retrieving the nearest agent from the node-to-agent mapping registration, filling the gap according to the adjacent relationship weight, and recording the filling path. If a pairing conflict is detected, the marked node is frozen according to the shared table, and the reason for freezing and the time period are recorded in the event log. Subsequently, feature splicing is performed. Semantically, feature splicing refers to establishing a splicing order for paired channels according to the time dimension, node dimension, and proxy dimension, following the order of node domain first and proxy domain second. Node domain feedforward representation and node domain normalized representation are used first, followed by proxy aggregation representation. The three components are written to the splicing cache after homogenization in the channel dimension. If the window score is low, the splicing cache enters a deweighting path, and the deweighting factor and source window are written to the quality log for subsequent steps to read.

[0054] Further, a rank configuration registration is performed. Semantically, this registration refers to registering compressible dimensions and candidate rank sources for each concatenated channel. The candidate sources include temporally semantically aligned sequences, node-to-proxy mapping registrations, and channel proportions recorded in the shared table. The processing flow involves calculating channel stability and mutual information characterization, generating a rank score, and then writing the score, source characterization, and window identifier into the rank configuration registration table. When channel stability fluctuates within a continuous window, a suppression flag is added to the registration table. This suppression flag triggers channel folding during subsequent low-rank bottleneck compression and reconstruction stages. The process then proceeds to low-rank bottleneck compression and reconstruction. Semantically, this involves decomposing the concatenated cache using the rank scores provided by the rank configuration registration table, performing principal component truncation, and subspace reconstruction to form two branches: bottleneck representation and reconstruction representation. Channel decomposition employs an orthogonal decomposition paradigm, principal component truncation follows the suppression flags and source characterization weights in the registration table, and subspace reconstruction establishes reconstruction weight stacks in both the node and time dimensions. These weight stacks are derived from proxy source mapping and adjacency weights. If insufficient proxy coverage is detected, temporary subspace expansion is initiated, the expansion path is written to the topology log, and the temporary path is cleaned up and the recovery time is recorded after expansion. Understandably, low-rank bottleneck compression and reconstruction outputs three data items: bottleneck representation, reconstruction representation, and rank configuration snapshot. The rank configuration snapshot, along with the rank configuration registration table, serves as the basis for decision-making in subsequent gating weight registration and gating residual overlay in this step.

[0055] Subsequently, gating weight registration is performed. Semantically, gating weight registration refers to assigning gating channel opening and closing ratios for different source representations and forming a traceable register. The processing order is as follows: reading channel statistics of bottleneck representations and reconstructed representations, reading shared constraints in the shared table and weighting factors in the quality log, and reading historical proportions in the residual proportion register. After the three types of information are normalized and mapped at the same scale, initial gating values ​​are generated. Then, the initial gating values ​​are reduced or relaxed according to the suppression flag in the rank configuration snapshot. The reduced or relaxed results are written into the gating weight register. If cross-window gating oscillations occur, an oscillation flag is marked in the register, and a smoothing strategy is triggered. The smoothing strategy generates a smoothing proportion through a time neighborhood sliding window. The smoothing proportion is written back to the gating weight register for the next round of reading and writing. After registration, gated residual overlay is performed. Semantically, gated residual overlay refers to weighting and overlaying the bottleneck representation, reconstruction representation, and node domain feedforward representation according to the gate weight registration and residual ratio registration, position by position, to generate a gated residual fusion channel. The overlay process first performs short-window mixing in the time dimension, and then performs neighborhood diffusion in the node dimension. The diffusion radius of the neighborhood diffusion is given by the adjacent relationship weight. When the diffusion radius triggers the upper limit, the diffusion process enters a limiting state and is written to the quality log. Subsequently, normalization is performed. Semantically, normalization refers to establishing channel statistics for the gated residual fusion channel and performing scale correction and offset correction. The parameters are migrated from the previous stable window and lightly updated in the current window. The update rate and stability are written to the operation and maintenance log. If the current window is in a deweighted path, the offset term is reduced according to the window score, and the reduction ratio is written to the metadata buffer. After normalization is completed, balancing is performed. Semantically, balancing refers to balancing the energy level and channel proportion of the fused channel and the original splicing buffer. The method is to construct balancing statistics and suppress channels with excessive proportion, and compensate channels with insufficient proportion. The balancing results are written into the balancing statistics record, and the gating proportion in the register is synchronously corrected to complete one closed loop.

[0056] After completing the above processing, the output product of this step is formed: a cross-layer low-rank bottleneck and gated residual fusion structure. This cross-layer low-rank bottleneck and gated residual fusion structure includes fields such as bottleneck representation, reconstruction representation, gated residual fusion channel, rank configuration registration table, rank configuration snapshot, gated weight registration, balancing statistical records, and normalized parameter set. The bottleneck representation and reconstruction representation serve as the basic channel in the time-dimensional aggregation of S400; the gated residual fusion channel serves as the routing facet in the routing weight initialization and temperature parameter registration of S400; the rank configuration registration table and rank configuration snapshot serve as the decision-making basis in the expert selection and feedforward calculation of S400; and the gated weight registration and balancing statistical records are used in the S400 process. The residual ratio of 00 is injected and readout shaping as a proportion reference. The normalized parameter set runs through the robust loss construction and sparsity and utilization constraint injection operations of S400, forming the input facet of the end-end hybrid expert steady-state decoding structure. At the same time, the gate weight registration and balancing statistics are written back to the bias log of S200 for subsequent windows to read during attention initialization and channel allocation stages, thus completing the closed-loop connection with the previous steps. The logs throughout the entire process include event logs, quality logs, and operation and maintenance logs. These three types of logs provide traceable records and anomaly localization basis when running across windows. The technical effect of this step can be summarized as follows: cross-layer interconnection channels are constructed by pairing adjacent layer outputs and concatenating features; a compact subspace is established by rank configuration registration and low-rank bottleneck compression and reconstruction; and the channel proportion and energy level are stabilized by gate weight registration, gated residual superposition, normalization, and balancing, providing a structured and adjustable fusion representation for subsequent end-end decoding.

[0057] S400, based on a cross-layer low-rank bottleneck and gated residual fusion structure, performs time-dimensional aggregation and routing weight initialization, registers temperature parameters, selects experts and performs feedforward calculations, performs residual ratio injection, readout shaping and horizon-weighted robust loss construction, and sparsity and utilization constraint injection, generating a terminal hybrid expert steady-state decoding structure. The cross-layer low-rank bottleneck and gated residual fusion structure is used as the input for this step. This structure includes bottleneck representation, reconstruction representation, gated residual fusion channel, rank configuration registration table, rank configuration snapshot, gated weight registration, balancing statistics record, and normalized parameter set. Residual proportion registration, shared table, and bias log are read from the node and agent two-stage attention modeling structure. Temporal semantic alignment sequence and alignment index are read from the normalized embedding. Specifically, temporal aggregation is first performed. Semantically, temporal aggregation refers to constructing a unified time axis and performing segmented weighted superposition of multi-scale temporal components from the bottleneck representation, reconstruction representation, and gated residual fusion channel. The processing order is to complete piecewise alignment according to the alignment index, and then generate aggregation weights based on the periodic and event characteristics in the temporal semantic alignment sequence. The weights are smoothly updated within the short window according to the window position. When an over-proportion channel from the balancing statistics record is detected, temporal aggregation puts the channel into a limiting state and writes it to the quality log. If there are missing test backfill traces in the window, endpoint fade-in / fade-out is enabled at the aggregation endpoint to suppress boundary fluctuations. The result of time dimension aggregation is aggregated time channel and aggregated record. The aggregated time channel is used as the input facet in the subsequent route weight initialization in this step, and the aggregated record and the original window identifier are stored together in the metadata buffer for traceability.

[0058] Further, route weight initialization and temperature parameter registration are performed. Semantically, route weight initialization refers to generating the initial distribution of expert route scores and injecting prior information from previous steps. This is achieved by reading the channel proportion from the gating weight registration and the historical proportion from the residual proportion registration, combining this with the short-range and medium-range change rates of the aggregated time channels to form the initial route profile. When a shared node is marked in the shared table, the initial route profile is weighted lower at the corresponding position and recorded in the bias log. Semantically, temperature parameter registration refers to setting a temperature scale that controls the sharpness of the route distribution. The initial temperature value is taken from the energy level balancing amount in the balancing statistics record and is slowly adjusted according to the window score. The adjustment step size and direction are written to the operation and maintenance log. If route oscillations are detected in consecutive windows, the temperature scale is increased and an oscillation flag is registered in the event log. These two items form the two fields: the route weight initialization table and the route temperature registration. Both, along with the aggregated time channels, enter the expert selection stage. Simultaneously, the route temperature registration is read as an adjustment factor in the subsequent sparsity and utilization constraint injection operations of this step.

[0059] Subsequently, expert selection and feedforward computation are performed. Expert selection semantically refers to selecting a small number of active experts from the candidate expert set for each time slice and node location based on the routing weight initialization table and routing temperature registration. Simultaneously, a sparse mask and expert percentage list are generated. To avoid overloading individual experts, historical occupancy is read from the bias log, and load balancing correction is implemented within the same time window, with the correction magnitude recorded in utilization statistics. The full English name for Mixture of Experts (MoE) is used in this implementation, where the expert unit is composed of a combination of a feedforward network sublayer and a normalization sublayer. Each selected expert receives local segments from the aggregated time channel and the gated residual fusion channel, performs position-by-position nonlinear mapping and channel scaling, and outputs the expert response. If a time slice does not reach the minimum coverage, a backup expert is triggered, and the replacement path and source are written to the event log. After the feedforward calculation is completed, the expert response is backfilled into the original time axis and node axis according to the sparse mask to form the expert hybrid response, and at the same time, the expert call details and utilization statistics are generated. The expert hybrid response and the node domain feedforward representation are aligned in the channel dimension to prepare for residual proportional injection.

[0060] After obtaining the expert mixed response, residual scaling injection and readout shaping are performed. Semantically, residual scaling injection refers to the position-wise weighted superposition of the expert mixed response and prior representation based on the residual scaling registration, forming the final fused response. This process reads the initial scaling from the residual scaling registration and performs a smooth update in conjunction with the current window score. When a cross-window drift event is detected, the step size of the smooth update is automatically reduced, and the reduction period is recorded in the operation log. Semantically, readout shaping refers to mapping the final fused response to the target prediction dimension and restoring the structure required by the external interface. The processing order is to establish readout channel statistics and perform scale correction and offset correction. The parameters are derived from the migration values ​​of the normalized parameter set and are lightly updated in the current window. If the balancing statistics record shows energy level bias, readout shaping compensates for the bias term and writes the compensation amount into the readout shaping record. After readout shaping is completed, two fields are obtained: the terminal readout sequence and the readout shaping record. The terminal readout sequence is directly accessible to the external interface in training or inference mode, and the readout shaping record, along with the window identifier, is written to the metadata buffer for retrieval by the evaluation module.

[0061] Subsequently, the construction of a horizon-weighted robust loss and the injection of sparsity and utilization constraints are performed. Semantically, the construction of the horizon-weighted robust loss refers to assigning different weights to multiple predicted horizons and constructing a segmented penalty term insensitive to anomalies. The processing order is as follows: read the horizon configuration in the semantically aligned sequence at reading time, generate horizon weights, then compare the deviation between the terminal read sequence and the reference sequence for each horizon and apply a segmented penalty. The truncation factor of the penalty is adaptively adjusted according to the window score. The penalties for all horizons are accumulated to form the robust loss, and the weights and truncation factors are written into the loss construction register. Semantically, the injection of sparsity and utilization constraints refers to adjusting the number of expert activations and the call distribution through sparse regularization and load balancing regularization. Specifically, it involves reading sparse mask statistics to obtain the activation density and applying a sparse penalty to windows exceeding the target density, while simultaneously reading utilization statistics to calculate the call variance between experts and applying a balancing penalty when the variance is too large. When the routing temperature registration is in an upward state, the coefficient of the sparse penalty decreases accordingly, and the coefficient of the balancing penalty increases accordingly. The updated coefficients are written into the constraint injection register. In training mode, the robust loss and the two types of constraints are used together to backpropagate gradients. In inference mode, they are only used for behavior monitoring and written to the monitoring log. After constraint injection, three fields are generated: loss construction registration, sparse constraint term registration, and utilization constraint term registration. Among them, the utilization constraint term registration and expert call details are written back to the preceding gating weight registration to provide prior facets for the gating initialization of subsequent windows.

[0062] After completing the above processing, the output product of this step, the terminal hybrid expert steady-state decoding structure, is formed. The terminal hybrid expert steady-state decoding structure includes fields such as terminal readout sequence, aggregation time channel, routing weight initialization table, routing temperature registration, expert hybrid response, expert call details, utilization statistics, residual ratio injection record, readout shaping record, loss construction registration, sparse constraint term registration, and utilization constraint term registration. Among them, the terminal readout sequence and readout shaping record are delivered to the external interface, the aggregation time channel, routing temperature registration, expert call details, and utilization statistics are reused in subsequent consecutive windows in this step, the loss construction registration and the two types of constraint terms are read by the optimization routine in training mode, and the gating-related registration and constraint registration are written in reverse to the bias log of the node and agent two-stage attention modeling structure for attention initialization and channel allocation reading in the next window, so as to maintain stable connection with the previous steps in cross-window operation.

[0063] In summary, the technical effects of this step are as follows: a stable and controllable expert route is constructed through temporal aggregation and routing temperature adjustment; a division of labor and cooperation on the spatiotemporal scale is formed through expert selection and residual ratio injection; and the training and inference behavior is stabilized and a structured terminal readout sequence is output through the joint injection of horizon-weighted robust loss and sparsity and utilization constraints.

[0064] like Figure 3As shown, Figure 3 This is a software architecture diagram provided for an embodiment of this application. The data acquisition subsystem is responsible for receiving input data such as raw historical multi-node traffic sequences; the model calculation subsystem carries the core algorithm of this patented method and is used to execute the modeling and calculation process from S100 to S400; the storage and management subsystem is responsible for handling the persistence and scheduling of intermediate results (such as node-to-agent mapping registration, quality logs, etc.); and the application service interface provides the output and service of prediction results. This system architecture supports centralized deployment in the cloud or distributed deployment on local edge nodes, and has high flexibility and practicality.

[0065] like Figure 4 As shown, Figure 4 This is a preferred internal structure diagram provided in an embodiment of this application. The fusion process begins by concatenating the output features of adjacent layers (corresponding to 301 in the diagram); then, bottleneck compression is performed through a low-rank weight matrix (302) to reduce dimensionality; nonlinearity is then introduced through a nonlinear activation function (303); next, feature reconstruction is performed by reconstructing the weight matrix (304); finally, the reconstructed features are residually superimposed with the original input (or features of skip connections), and after layer normalization (305), the robust fused feature F* is output. This structure intuitively demonstrates the key role of low-rank bottlenecks and residual connections in ensuring efficient information flow and fusion.

[0066] like Figure 5 As shown, Figure 5 This is a preferred internal structure diagram of the end-to-end hybrid expert stabilization decoding module provided in this application embodiment. The decoding module mainly includes an input gating unit (401), expert sub-networks (402), a weighted fusion unit (403), and a residual scaling and normalization unit (404). Input features are first routed and allocated by the input gating unit (401) to control the activation of different expert sub-networks (402); the outputs of each expert network are integrated in the weighted fusion unit (403); finally, the fusion result is processed by the residual scaling and normalization unit (404) to stabilize the training process and output the final prediction result Y*. This module effectively improves the robustness of long-term prediction through expert division of labor and temperature control mechanisms.

[0067] In a preferred embodiment, the traffic flow prediction system of the present invention comprises the following main modules: The system includes a data preprocessing module, a spatiotemporal embedding construction module, an encoder layer (including a cross-layer low-rank fusion module), a decoder layer (including a stable hybrid expert mechanism), and a prediction output and loss optimization module.

[0068] Input is historical traffic sequence ,in, Enter historical traffic sequences; Indicates batch size, Indicates time steps, Indicates the number of sensor nodes, This represents the number of input channels (e.g., speed, flow rate, etc.). The output is the prediction result for several future time steps. ,in, : Predicted output for several future time steps.

[0069] Furthermore, the data preprocessing and embedding module Data standardization: Input traffic data is linearly normalized or Z-score standardized to eliminate dimensional differences; missing values ​​can be filled in by linear interpolation or moving average.

[0070] Temporal and spatial embedding: Time index encoding (such as periodic features like hours and days of the week) is used and mapped to a time embedding vector. The node's geographical topology is determined by the adjacency matrix. Encoding as spatial embedding These two features are concatenated with the input features to form an enhanced input:

[0071] in, For the enhanced input tensor; For time embedding vectors; This is the spatial embedding vector. Further, the encoder structure and cross-layer low-rank fusion module: the encoding part adopts a multi-layer stacked Transformer-like structure. Each encoder layer contains three core steps: (1) Temporal Convolution Enhancement Module To compensate for the shortcomings of the original self-attention in local dependency modeling, a one-dimensional convolution (TCN) is added before the input of each layer to enhance it:

[0072] in, The feature tensor after temporal convolution enhancement corresponds to time step. ; For the current time step Input features; This is a one-dimensional convolution operation used to extract local temporal patterns; This represents the size of the convolution kernel or the size of a local time window.

[0073] Achieve dynamic smoothing of local time windows.

[0074] (2) Spatial Agent Attention Module To reduce the computational complexity of spatiotemporal attention, this invention designs an adaptive proxy mechanism. In each layer, node features... Mapped to several proxy nodes :

[0075] in, :Proxy node representation; :Proxy index; : Node index; These represent the query, key, and value vectors, respectively. , , ,in, These are learnable weights; Feature dimension; Indicates temperature parameter; :node Input features; T represents transpose; Then, reverse aggregation from Proxy to Node is performed:

[0076] in, For nodes Updated features; temperature parameters It is used to control the smoothness of attention distribution.

[0077] Number of agents Determined adaptively by the latest time-series features: ,in

[0078] in, Represents the sigmoid function; It should refer to the characteristics of the latest time step; It is a multilayer perceptron; and These are the minimum and maximum numbers of agents; It is the gate value; Indicates rounding down; To find the average value function; This mechanism reduces computational overhead while preserving important spatial interaction information.

[0079] (3) Cross-layer low-rank fusion module: To address the fragmentation problem between multi-layer representations, this invention introduces a cross-layer fusion mechanism. Let the first... Layer and First The layer outputs are respectively , First, perform splicing or summation:

[0080] in, : splicing characteristics; , They represent the first Layer and first The output feature tensor of the layer.

[0081] Furthermore, compression and reconstruction are achieved through low-rank bottlenecks:

[0082]

[0083] in, :low rank; Weight matrix; Weight matrix; Hyperbolic tangent activation function; : The reconstructed feature tensor.

[0084] Then introduce a gated residual structure:

[0085] in, : The output features after fusion; Layer normalization; Gating weight matrix; The dropout method randomly filters out some neurons to prevent overfitting. This structure achieves lossy compression and feature compensation of interlayer information, reduces parameter redundancy, and improves representation consistency.

[0086] Furthermore, the decoder is coupled with a stable hybrid expert mechanism: The decoding part uses a single-layer or shallow-layer structure, and only decodes the final fused representation. Make predictions.

[0087] This invention introduces a Stable Mixture-of-Experts (MoE) network at the decoding end to improve long-term prediction stability.

[0088] For input Calculation output:

[0089] in: :Predicted output; Number of expert subnetworks; Temperature parameter; : Residual proportionality coefficient (value range: 0.3 to 0.6); : Each expert's feedforward subnetwork; Gated networks; The softmax function; Refers to the decoder input; This represents the expert subnetwork index.

[0090] By employing temperature regulation and residual scaling, gradient oscillations and error accumulation in multi-step prediction can be effectively suppressed.

[0091] Furthermore, regarding the loss function and training strategy: This invention employs a horizon-weighted robust loss function:

[0092] in: Total loss function; Predict the length of the horizon; :Prediction step index; Weighting coefficients for different predicted horizons; : Proxy sparse regularization terms; Agent utilization equilibrium term; :Smoothing coefficient of Huber loss; :Proxy sparse regularization Weighting coefficients; Agent utilization equilibrium item The weighting coefficients.

[0093] Used to constrain the sparsity of the agent attention matrix:

[0094] in, : Regularization weights in; Along the batch and time Average of dimensions; Agent attention matrix.

[0095] During training, the usage variance and average value of agent nodes are monitored to prevent agent collapse or over-concentration.

[0096] To verify the effectiveness of the present invention, in a specific embodiment, the publicly available traffic flow dataset PEMS08 was used for the experiment. The experimental setup was consistent with the existing TLAST model, with both the input and prediction lengths being 12 time steps, and the mean absolute error (MAE) was used as the core evaluation metric.

[0097] The experimental results are shown in Appendix Table 1:

[0098] Specifically, to verify the effectiveness of the proposed method, a comparative experiment was conducted on the public traffic flow dataset PEMS08. The results, as shown in Table 1 (Table 1 shows the experimental results), demonstrate that the proposed method significantly improves model efficiency and robustness while maintaining prediction accuracy (average MAE of 13.1779, slightly better than the baseline model's 13.1807). Specifically, due to the introduction of a low-rank bottleneck and cross-layer fusion module, the number of model parameters is reduced by approximately 18%, and the average training and inference time is shortened by approximately 10%. More importantly, as shown in Table 2 (Table 2 compares the prediction results of the proposed method and the baseline model on the PEMS08 dataset), at longer prediction horizons (steps 7 to 12), the prediction error of the proposed method exhibits a smoother convergence trend, with its error volatility (measured by standard deviation σ(MAE)) reduced by approximately 6%, maintaining a low level of 0.025~0.029 over a long period, significantly better than the volatility range of the conventional Transformer model (approximately 0.035~0.04). This fully verifies that the present invention, through cross-layer low-rank feature fusion and gated residual mechanism, has significant effects in reducing model parameter sensitivity and improving long-term prediction stability, and has clear engineering application value.

[0099] Table 2 Comparison of prediction results of the present invention and the baseline model on the PEMS08 dataset.

[0100] Regarding feasibility and industrial application prospects, this invention is implemented based on standard deep learning frameworks such as PyTorch. Its core modules (such as adaptive agent attention, cross-layer low-rank fusion LRF, and steady-state hybrid expert decoding MoE) are all lightweight and pluggable, exhibiting good engineering adaptability. The method can be effectively applied to short-term traffic flow prediction, traffic signal timing optimization, emergency traffic event response, and travel information services in various scenarios such as highways, urban road networks, and park IoT. Experimental results show that it outperforms existing baselines in both accuracy and stability of multi-step prediction, demonstrating clear industrial application value and promising prospects for widespread adoption.

Claims

1. A cross-layer low-rank fusion spatiotemporal traffic flow prediction modeling method, characterized in that, include: The system acquires historical multi-node traffic sequences, time labels, and road topology, performs standardization, missing data completion, and label encoding, and performs temporal embedding, spatial embedding construction, and temporal convolutional enhancement processing to generate standardized embedding and temporal convolutional enhancement structures. Based on the standardized embedding and temporal convolutional enhancement structure, node aggregation, agent number adaptive estimation and attention initialization are performed. Node-to-agent attention calculation and aggregation, agent-to-node backpropagation and feedforward transformation, and normalization are performed to generate a two-stage attention modeling structure for nodes and agents. The two-stage attention modeling structure of nodes and agents is obtained. Adjacent layer output pairing, feature splicing and rank configuration registration are performed. Low-rank bottleneck compression, reconstruction and gating weight registration are performed. Gating residual superposition, normalization and balancing are performed to generate a cross-layer low-rank bottleneck and gating residual fusion structure. Based on the cross-layer low-rank bottleneck and gated residual fusion structure, time-dimensional aggregation and routing weight initialization are performed, temperature parameter registration, expert selection and feedforward calculation are performed, residual ratio injection, readout shaping and horizon-weighted robust loss construction, sparsity and utilization constraint injection are performed, and a terminal hybrid expert steady-state decoding structure is generated.

2. The method according to claim 1, characterized in that, Historical multi-node traffic sequences, time tags, and road topology include: The historical multi-node traffic sequence is a multi-channel time series composed of speed, flow, occupancy and related observations continuously reported by road monitoring terminals, including node identifiers, sampling times and observation types; the time tags include intraday location, intraweek location, holiday attributes, weather conditions and event trigger attributes, forming tag fields and tag vocabulary; the road topology describes the connectivity, direction, hierarchy and adjacency of road segments, and includes road segment length, number of lanes, speed limit, signal control phase and historical coupling degree.

3. The method according to claim 1, characterized in that, The process of standardization, test completion, and label encoding also includes: Standardization processing includes extracting the central trend and dispersion amplitude based on historical distribution statistics to generate a standardized parameter set and performing dimensionless mapping and scale unification; Missing test completion includes prioritizing backfilling based on the same node and the same tier, and generating an aligned index by interpolating adjacent nodes in the topology.

4. The method according to claim 1, characterized in that, The temporal convolution enhancement process also includes: The temporal convolution enhancement process includes constructing a one-dimensional convolution stack using a temporal convolutional network, performing slice-by-slice filtering for each node and each time slice, and switching to a shared channel and registering a shared node flag when a node is detected to be invalid for an extended period.

5. The method according to claim 1, characterized in that, The process of performing node aggregation, adaptive estimation of agent number, and attention initialization also includes: The node aggregation includes generating an aggregation weight field based on topological adjacency weights, quality scores, and shared node tags, and enabling a buffered connection strategy at the endpoints of the sliding window. The adaptive estimation of the number of agents includes generating an agent size score based on the structural diversity index, the time fluctuation index and the quality score, and using a smooth interpolation strategy to generate a soft gating ratio. The attention initialization includes reading multi-scale temporal channels from the temporal convolutional enhancement structure to construct a temporal prior surface, reading adjacency weights from spatial embedding to construct a topological prior surface, and assigning temporal bias and topological bias weight ratios to each head of the multi-head attention.

6. The method according to claim 1, characterized in that, The process of performing attention calculation and aggregation from the execution node to the agent, and backhaul and feedforward transformation from the agent to the node, also includes: The node-to-agent attention calculation and aggregation includes aligning node-level queries with agent-level key values ​​and sending them into a multi-head channel, and implementing gate control suppression for abnormally high correlation. The agent-to-node backhaul includes restoration based on the node-to-agent mapping registration and agent ratio, taking into account the adjacent relationship weight and time prior surface. When the agent coverage is insufficient, a temporary neighborhood expansion is triggered. The feedforward transformation includes performing nonlinear mapping from the feedforward network and reading the initial value of the residual superposition ratio from the quality score and updating it smoothly within the window.

7. The method according to claim 1, characterized in that, The process of pairing adjacent layer outputs, concatenating features, and registering rank configuration also includes: The adjacent layer output pairing includes checking the node dimension and time dimension piece by piece based on the alignment index. When there is a gap, the nearest agent is retrieved from the node to agent mapping registration and the gap is filled based on the adjacent relationship weight. The feature splicing adopts the splicing order of node field first and proxy field second and performs channel homogenization. When the window score is low, it enters the weight reduction path. The rank configuration registration includes generating rank scores based on channel stability and mutual information characterization, and registering compressible dimensions and candidate rank sources.

8. The method according to claim 1, characterized in that, The process of performing low-rank bottleneck compression, reconstruction, and gating weight registration also includes: The low-rank bottleneck compression and reconstruction includes channel decomposition using an orthogonal decomposition paradigm, principal component truncation following the suppression flag in the rank configuration registration, and reconstruction weight stacks based on proxy source mapping and adjacency relation weights established in the node dimension and time dimension, respectively. The gating weight registration includes generating initial gating values ​​based on bottleneck representation, channel statistics of reconstruction representation, shared constraints, weight reduction factors in quality logs, and residual ratio registration, and then reducing or relaxing them according to the suppression flag in the rank configuration snapshot.

9. The method according to claim 1, characterized in that, The process of performing gated residual superposition, normalization, and balancing also includes: The gated residual superposition includes weighting the bottleneck representation, reconstruction representation and node domain feedforward representation according to the gate weight and residual ratio, first performing short window mixing in the time dimension and then performing neighborhood diffusion in the node dimension. The normalization includes migrating parameters from the previous stable window and updating them lightly, and the balancing process includes balancing the energy level and channel proportion of the fused channel and the original splicing buffer.

10. The method according to claim 1, characterized in that, The process of temperature parameter registration, expert selection, and feedforward calculation also includes: The temperature parameter registration includes setting the temperature scale, with the initial temperature value taken from the balance statistics record and adjusted according to the window score; The expert selection process includes selecting active experts from the candidate expert set based on the route weight initialization table and route temperature registration, and generating a sparse mask and an expert percentage list. The feedforward computation includes performing nonlinear mapping using a hybrid expert structure.

Citation Information

Cited By

  • Analysis scheduling method and system based on multi-label routing

    CN122027553A

  • An analysis scheduling method and system based on multi-label routing

    CN122027553B