Cross-region adaptive joint transfer learning method for urban spatio-temporal prediction
Patent Information
- Application Number
- CN202610930877.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-26
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]本发明为克服上述的不足之处,目的在于提供用于城市时空预测的跨区域自适应联合迁移学习方法,在捕捉复杂的时空依赖关系的同时,缓解由边界导致的数据稀疏问题,兼顾迁移泛化能力、隐私安全防护的联邦迁移学习框架,提高城市时空预测的准确性
[0059]本发明的有益效果在于:利用LLM强大的语义推理能力挖掘隐性区域关联,有效缓解行政边界导致的数据稀疏与依赖断裂问题;同时设计了基于参数解耦的SS-MoE机制,在保护数据隐私的前提下实现跨区域知识的选择性迁移,显著抑制了分布差异引发的负迁移,模型的跨区域泛化能力与预测鲁棒性显著提高。
Smart Images

Figure CN122819362A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of urban spatiotemporal prediction, and more particularly to a cross-regional adaptive joint transfer learning method for urban spatiotemporal prediction. Background Technology
[0002] Urban spatiotemporal prediction is crucial for smart city management. Traditional static planning and manual scheduling cannot adapt to the dynamic evolution and changes of cities. As cities develop, dynamic, intelligent, and proactive prediction is needed to achieve dynamic management, risk warning, and refined governance. Urban spatiotemporal prediction relies on massive, multi-source, and high-sampling-density spatiotemporal data. However, existing regulations strictly limit the leakage of personal information, and data fragmentation caused by administrative boundaries severely hinders effective model training, resulting in data-driven modeling distortion. Cross-regional and cross-city data support amplifies privacy leaks and data security risks, while data anonymization significantly damages spatiotemporal characteristics, affecting prediction accuracy.
[0003] In order to overcome data limitations and make full use of the knowledge of data-rich regions while protecting data privacy, Federated Transfer Learning (FTL) has become a promising solution for cross-regional urban prediction tasks. However, traditional learning methods face two challenges: (1) Difficulty in acquiring knowledge within the region: The lack of boundaries weakens the spatiotemporal dependence and leads to data sparsity and imbalance in the distribution of samples across the entire region, increasing the difficulty of unified modeling; (2) Cross-regional knowledge transfer bias: The distribution difference between the source region and the target region will cause negative transfer, which will significantly reduce the generalization ability of the pre-trained model, thereby reducing the robustness of the model and the prediction accuracy. Summary of the Invention
[0004] To overcome the aforementioned shortcomings, this invention aims to provide a cross-regional adaptive joint transfer learning method for urban spatiotemporal prediction. This method captures complex spatiotemporal dependencies while mitigating data sparsity caused by boundaries. It also provides a federated transfer learning framework that balances transfer generalization ability and privacy protection, thereby improving the accuracy of urban spatiotemporal prediction.
[0005] This invention achieves the above objective through the following scheme: a cross-regional adaptive joint transfer learning method for urban spatiotemporal prediction, comprising:
[0006] (1) Construct a spatiotemporal model within the local framework to capture spatiotemporal dependencies and output cross-modal enhanced spatiotemporal features. The model includes:
[0007] (1.1) Spatiotemporal network coding branch used to extract spatiotemporal topology and fuse multi-scale dynamic dependencies;
[0008] (1.2) Used to introduce general semantic knowledge and parse cross-regional implicit associations in large language models to drive encoding branches;
[0009] (1.3) Cross-modal alignment mechanism for eliminating modal gap and realizing deep interaction between semantic space and feature space;
[0010] (2) A federated transfer learning framework is proposed to achieve cross-regional knowledge sharing under privacy protection, including:
[0011] (2.1) A federated parameter sharing mechanism for cross-regional security collaborative training and knowledge aggregation;
[0012] (2.2) Shared sparse hybrid expert module for filtering irrelevant source domain noise and suppressing negative migration across regions;
[0013] (3) The spatiotemporal features enhanced by cross-modality are fed into the shared sparse hybrid expert module, the expert weighted prediction results are dynamically integrated, and the final prediction results are generated by combining the local prediction.
[0014] Preferably, step (1.1) includes:
[0015] (1.1.1) Input is a spatiotemporal graph ,in For a set of nodes, Let edge sets represent the connection relationships. An adjacency matrix is used to describe the physical connections between nodes. The node feature matrix is used to extract spatial features. With time characteristics ;
[0016] (1.1.2) For spatial features, neighbor information is aggregated through a graph convolutional network, and then dynamic relationships are captured using a graph attention network:
[0017]
[0018]
[0019] in, Essentially, it is a node feature matrix. High-dimensional representation after linear projection For degree matrix, Let be the learnable weight matrix of the l-th layer. GAT stands for Graph Attention Network, which is the activation function.
[0020] (1.1.3) For time-related characteristics, gated loop units are used to preserve causality, and attention mechanisms are used to focus on key moments:
[0021]
[0022]
[0023] in, This represents a gated loop unit, and Attn() represents the attention mechanism;
[0024] (1.1.4) After linearly aligning the spatial and temporal features and combining them with structural priors, hybrid spatiotemporal features are generated. :
[0025]
[0026] .
[0027] Preferably, step (1.2) includes: converting the node feature matrix... and adjacency matrix Mapped to a natural language prompt, deep semantics are extracted using a large language model (LLM) with frozen parameters, and the last token is used as the representative:
[0028]
[0029]
[0030] in, This represents the final feature that drives the encoding branch of the large language model. This represents the mapping function for Prompt.
[0031] Preferably, step (1.3) includes:
[0032] (1.3.1) Based on hybrid spatiotemporal features For Query, LLM semantics Using key and value, cross-modal fusion features are obtained through multi-head attention fusion:
[0033]
[0034] in, This indicates multi-head attention fusion;
[0035] (1.3.2) Concatenate the original features to prevent information loss, and output the local prediction through the projection layer:
[0036]
[0037]
[0038] in, This represents the projection function.
[0039] Preferably, step (2.1) includes:
[0040] (2.1.1) Each client only uploads shared parameters. The server performs a weighted average to generate global parameters:
[0041]
[0042] in, These are globally shared parameters generated by server aggregation. Indicates the first The region in the first Shared module parameters of the wheel, This represents the total number of regions participating in the training.
[0043] (2.1.2) Subsequently, client-side adaptive updates are performed, with the client updating based on the cosine similarity with the global parameters. Make personalized adjustments:
[0044]
[0045]
[0046] in, These are the updated locally shared parameters. It is the cosine similarity weight factor, used to measure the similarity between local parameters and global parameters; the more similar they are, the greater the weight.
[0047] Preferably, step (2.2) includes:
[0048] (2.2.1) When making predictions in the target area, route selection is performed, expert weights are calculated, and the Top-k are selected:
[0049]
[0050]
[0051] in, It is a gating network. It is the distribution of expert weights. These are the k highest weights after normalization;
[0052] (2.2.2) Perform expert ensemble and weighted summation of the outputs of the activated experts:
[0053]
[0054] in, It is the total number of experts. It is the first The fusion and embedding output of individual experts It is the result of expert weighted prediction.
[0055] Preferably, step (3) includes: fusing dynamic MoE output with local personalized prediction:
[0056]
[0057] in, It is the weighting coefficient that balances expert output and local output. This is the final prediction result.
[0058] A structure for implementing the above method includes: an encoder, a cross-modal alignment module, and a hybrid expert decoder; the encoder includes a spatiotemporal network coding branch module and a large language model-driven coding branch module, the cross-modal alignment module uses an attention mechanism, and the hybrid expert decoder adopts a shared sparse hybrid expert architecture and uses gated routing; the spatiotemporal network coding branch module and the large language model-driven coding branch module are respectively connected to the cross-modal alignment module, and the cross-modal alignment module is connected to the hybrid expert decoder; the spatiotemporal graph is input into the spatiotemporal network coding branch module and the large language model-driven coding branch module to obtain cross-modal enhanced spatiotemporal features and local predictions, the cross-modal enhanced spatiotemporal features are input into the hybrid expert decoder to dynamically integrate expert weighted prediction results, and combined with local predictions to generate the final prediction result.
[0059] The beneficial effects of this invention are as follows: by utilizing the powerful semantic reasoning capabilities of LLM to mine implicit regional associations, the data sparsity and dependency breakage problems caused by administrative boundaries are effectively alleviated; at the same time, an SS-MoE mechanism based on parameter decoupling is designed to achieve selective transfer of cross-regional knowledge while protecting data privacy, significantly suppressing negative transfer caused by distribution differences, and significantly improving the model's cross-regional generalization ability and prediction robustness. Attached Figure Description
[0060] Figure 1 This is a schematic diagram of the steps of the method of the present invention;
[0061] Figure 2 This is a schematic diagram of the spatiotemporal model structure constructed within the local framework in the method of this invention;
[0062] Figure 3 This is a schematic diagram of a federated transfer learning framework process in the method of this invention. Detailed Implementation
[0063] The present invention will be further described below with reference to specific implementation examples, but the scope of protection of the present invention is not limited thereto:
[0064] Example: Figure 1 As shown, a cross-regional adaptive joint transfer learning method for urban spatiotemporal prediction is illustrated in one embodiment using ride-hailing demand datasets from cities A and B. The dataset from city A includes 524 sensor / region nodes, 1120 edges, a time interval of 10 minutes, and 17280 time steps; the dataset from city B includes 627 sensor / region nodes, 4845 edges, a time interval of 10 minutes, and 17280 time steps. The METIS subgraph partitioning algorithm is used to divide the dataset into 5 non-overlapping client regions, with 4 regions serving as source regions and 1 region as the target region. During prediction, the time window length T=12, meaning that the ride-hailing demand for the next time slice is predicted using data from the past 120 minutes. The specific method steps include:
[0065] (1) Construct a spatiotemporal model within the local framework to capture spatiotemporal dependencies, and output cross-modal enhanced spatiotemporal features, such as Figure 2 As shown, the model includes:
[0066] (1.1) A spatiotemporal network coding branch used to extract spatiotemporal topology and fuse multi-scale dynamic dependencies. Specifically, it includes the following steps:
[0067] (1.1.1) Input is a spatiotemporal graph ,in For a set of nodes, Let edge sets represent the connection relationships. An adjacency matrix is used to describe the physical connections between nodes. The node feature matrix is used to extract spatial features. With time characteristics ;
[0068] (1.1.2) For spatial features, neighbor information is aggregated through a Graph Convolutional Network (GCN), and then dynamic relationships are captured using a Graph Attention Network (GAT):
[0069]
[0070]
[0071] in, Essentially, it is a node feature matrix. High-dimensional representation after linear projection For degree matrix, Let be the learnable weight matrix of the l-th layer. For activation functions;
[0072] (1.1.3) For temporal features, a Gated Recurrent Unit (GRU) is used to preserve causality, and an attention mechanism is used to focus on key moments:
[0073]
[0074]
[0075] in, This represents a gated loop unit, and Attn() represents the attention mechanism;
[0076] (1.1.4) After linearly aligning the spatial and temporal features and combining them with structural priors, hybrid spatiotemporal features are generated. :
[0077]
[0078] .
[0079] (1.2) Used to introduce general semantic knowledge and parse cross-regional implicit associations. Large language model-driven encoding branch: The original data and adjacency matrix Mapped to a natural language prompt, deep semantics are extracted using the LLM with frozen parameters, and the last token is used as the representative:
[0080]
[0081]
[0082] in, This represents the final characteristic of the LLM branch. This represents the mapping function for Prompt.
[0083] (1.3) Cross-modal alignment mechanism for eliminating modal gap and enabling deep interaction between semantic space and feature space.
[0084] (1.3.1) Based on hybrid spatiotemporal features For Query, LLM semantics Using key and value, cross-modal fusion features are obtained through multi-head attention (MHA) fusion. :
[0085] .
[0086] (1.3.2) Concatenate the original features to prevent information loss, and output the local prediction through the projection layer:
[0087]
[0088]
[0089] in, This represents the projection function.
[0090] (2) A federated transfer learning framework is proposed to achieve cross-regional knowledge sharing under privacy protection, such as... Figure 3 As shown, it includes:
[0091] (2.1) A federated parameter sharing mechanism for cross-regional secure collaborative training and knowledge aggregation. Specifically, it includes the following steps:
[0092] (2.1.1) Each client only uploads shared parameters. The server performs a weighted average to generate global parameters:
[0093]
[0094] in, These are globally shared parameters generated by server aggregation. Indicates the first The region in the first Shared module parameters of the wheel, This represents the total number of regions participating in the training. In one embodiment, during the t-th round of federated training, the shared parameters uploaded by the four region clients can be simplified as follows: =0.56、 =0.61、 =0.59、 =0.63, then the server aggregation yields: =(0.56+0.61+0.59+0.63) / 4=0.5975.
[0095] (2.1.2) Subsequently, client-side adaptive updates are performed, with the client updating based on the cosine similarity with the global parameters. Make personalized adjustments:
[0096]
[0097]
[0098] in, These are the updated locally shared parameters. This is the cosine similarity weighting factor, used to measure the similarity between local and global parameters; the greater the similarity, the higher the weight. Locally shared parameters for the target region. =0.63, which is a globally shared parameter. Cosine similarity = 0.5975 If the value is 0.8, then the adaptive update is: =0.63+0.8×(0.5975-0.63)=0.604.
[0099] (2.2) A shared sparse hybrid expert module for filtering irrelevant source domain noise and suppressing negative migration across regions. Specifically, it includes the following steps:
[0100] (2.2.1) When making predictions in the target area, route selection is performed, expert weights are calculated, and the Top-k are selected:
[0101]
[0102]
[0103] in, It is a gating network. It is the distribution of expert weights. These are the k highest weights after normalization;
[0104] (2.2.2) Perform expert ensemble and weighted summation of the outputs of the activated experts:
[0105]
[0106] in, It is the total number of experts. It is the first The fusion and embedding output of individual experts It is the result of expert weighted prediction.
[0107] (3) The spatiotemporal features enhanced by cross-modality are fed into the shared sparse hybrid expert module to dynamically integrate multi-source expert knowledge and generate the final prediction results. .
[0108] The final prediction result integrates the dynamic MoE output with the local prediction from step (1):
[0109]
[0110] in, It is the weighting coefficient that balances the expert output and the local output.
[0111] In one embodiment, the method of the present invention is validated on datasets of two typical urban spatiotemporal prediction scenarios:
[0112] Scenario 1: Demand Forecast for Cross-Regional Ride-Hailing Services within City A
[0113] Source Region: The rich dataset of City A is divided into multiple regions. The spatiotemporal data of each region is highly dynamic and irregular, and contains complete traffic flow data for three consecutive months.
[0114] Target area: City A dataset is a single area with severely sparse data, containing only 3 days of traffic flow data.
[0115] Scenario 2: Cross-city ride-hailing demand forecast between cities A and B
[0116] Source Region: The rich dataset of City A is divided into multiple regions. The spatiotemporal data of each region is highly dynamic and irregular, and contains complete traffic flow data for three consecutive months.
[0117] Target area: City B dataset is a single area with severely sparse data, containing only 3 days of traffic flow data.
[0118] To verify the prediction accuracy and transfer stability of the model in this invention, a comparison was made with existing mainstream spatiotemporal prediction models in Scenario 1 (sparse data scenario of shared bicycles). The evaluation metrics used were the mean absolute error (MAE) and root mean square error (RMSE) over 60 minutes; smaller values indicate more accurate predictions. The experimental comparison results are shown in Table 1 below:
[0119] Traditional federal model FedAvg 3.487 5.160 This invention reduces costs by an average of 9.7%. Traditional cross-regional model pFedCTP 3.329 4.953 This invention reduces costs by an average of 5.4%. This invention model CRAFT 3.110 4.718 Excellent performance
[0120] Table 1
[0121] As shown in Table 1, when the target region faces severe data sparsity, the traditional federated model suffers from large errors due to overfitting; the traditional cross-regional model exhibits limited performance improvement due to "negative migration" caused by the difference in distribution between the source and target domains. In contrast, this invention utilizes a large model to introduce semantic common sense and filters out geographical noise through the SS-MoE mechanism, achieving significantly superior prediction accuracy.
[0122] To further verify the "negative migration resistance" (i.e., migration stability) during cross-city migration, the fluctuation of prediction error of each model in the target domain was tested, and the results are shown in Table 2:
[0123] Traditional cross-regional model ST-GFSL 2.966 3.996 This invention reduces costs by an average of 6.8%. Traditional cross-regional model pFedCTP 2.677 3.851 This invention reduces costs by an average of 0.6%. This invention model CRAFT 2.661 3.831 Excellent performance
[0124] Table 2
[0125] The results show that as the differences in urban characteristics and environment between the two cities widen, the prediction error of traditional cross-regional migration models increases (poor stability and severe negative migration); while the model of this invention, due to the design of a federated adaptive personalized update mechanism based on parameter decoupling and sparse gating routing, exhibits extremely strong migration stability and cross-regional robustness.
[0126] A structure for implementing the above method includes: an encoder, a cross-modal alignment module, and a hybrid expert decoder; the encoder includes a spatiotemporal network coding branch module and a large language model-driven coding branch module; the cross-modal alignment module uses an attention mechanism; the hybrid expert decoder adopts a shared sparse hybrid expert architecture and uses gated routing; the spatiotemporal network coding branch module and the large language model-driven coding branch module are respectively connected to the cross-modal alignment module, and the cross-modal alignment module is connected to the hybrid expert decoder. The spatiotemporal graph... As feature inputs, they are fed into the spatiotemporal network coding branch module and the large language model-driven coding branch module, respectively. In the spatiotemporal network coding branch module, the spatiotemporal graph first obtains spatial features through a graph convolutional network and a graph attention network. Then the spacetime diagram Temporal features are also obtained through gated loop units and attention mechanisms. Spatial features and time characteristics After linear alignment is achieved through linear layers and linear fusion layers, combined with structural priors, hybrid spatiotemporal features are generated. In the large language model-driven encoding branch module, the features of the spatiotemporal graph are processed by natural language prompts and the large language model outputs the final features of the LLM branch. With hybrid spatiotemporal features For Query, LLM semantics Using key and value, multi-head attention fusion is used to obtain cross-modal enhanced spatiotemporal features. and local forecasts Spatiotemporal features of cross-modal enhancement Input hybrid expert decoder dynamically integrates expert weighted prediction results Combined with local forecasts Generate the final prediction results .
[0127] The above description describes specific embodiments of the present invention and the technical principles employed. Any changes made in accordance with the concept of the present invention that do not exceed the spirit of the specification and drawings should still fall within the protection scope of the present invention.
Claims
1. A cross-regional adaptive joint transfer learning method for urban spatiotemporal prediction, characterized in that... include: (1) Construct a spatiotemporal model within the local framework to capture spatiotemporal dependencies, and output cross-modal enhanced spatiotemporal features and local predictions. The model includes: (1.1) Spatiotemporal network coding branch used to extract spatiotemporal topology and fuse multi-scale dynamic dependencies; (1.2) Used to introduce general semantic knowledge and parse cross-regional implicit associations in large language models to drive encoding branches; (1.3) Cross-modal mechanisms for bridging the modal gap and enabling deep interaction between the semantic space and the feature space; (2) A federated transfer learning framework is proposed, including: (2.1) A federated parameter sharing mechanism for cross-regional security collaborative training and knowledge aggregation; (2.2) Shared sparse hybrid expert module for filtering irrelevant source domain noise and suppressing negative migration across regions; (3) The spatiotemporal features enhanced by cross-modality are fed into the shared sparse hybrid expert module, the expert weighted prediction results are dynamically integrated, and the final prediction results are generated by combining the local prediction.
2. The cross-regional adaptive joint transfer learning method for urban spatiotemporal prediction according to claim 1, characterized in that, Step (1.1) includes: (1.1.1) Spatiotemporal diagram Input spatiotemporal model, where For a set of nodes, Let edge sets represent the connection relationships. An adjacency matrix is used to describe the physical connections between nodes. For the node feature matrix, extract spatial features respectively. With time characteristics ; (1.1.2) After linearly aligning the spatial and temporal features and combining them with structural priors, hybrid spatiotemporal features are generated. : , 。 3. The cross-regional adaptive joint transfer learning method for urban spatiotemporal prediction according to claim 2, characterized in that, In step (1.1.1), the extraction of spatial features includes the following steps: aggregating neighbor information through a graph convolutional network, and then capturing dynamic relationships using a graph attention network: , , in, Essentially, it is a node feature matrix. High-dimensional representation after linear projection For degree matrix, Let be the learnable weight matrix of the l-th layer. GAT stands for Graph Attention Network, which is the activation function.
4. The cross-regional adaptive joint transfer learning method for urban spatiotemporal prediction according to claim 2, characterized in that, In step (1.1.1), the extraction of temporal features includes the following steps: using a gated recurrent unit to preserve causality, and using an attention mechanism to focus on key moments: , , in, This represents a gated loop unit, and Attn() represents the attention mechanism.
5. The cross-regional adaptive joint transfer learning method for urban spatiotemporal prediction according to claim 1, characterized in that, Step (1.2) includes: converting the node feature matrix and adjacency matrix Mapped to a natural language prompt, deep semantics are extracted using a large language model (LLM) with frozen parameters, and the last token is used as the representative: , , in, This represents the final feature that drives the encoding branch of the large language model. This represents the mapping function for Prompt.
6. The cross-regional adaptive joint transfer learning method for urban spatiotemporal prediction according to claim 1, characterized in that, Step (1.3) includes: (1.3.1) Based on hybrid spatiotemporal features For Query, LLM semantics Using multi-head attention fusion as the key and value, cross-modal enhanced spatiotemporal features are obtained: , in, This indicates multi-head attention fusion; (1.3.2) Concatenate the original features to prevent information loss, and output the local prediction through the projection layer. : , , in, This represents the projection function.
7. The cross-regional adaptive joint transfer learning method for urban spatiotemporal prediction according to claim 1, characterized in that, Step (2.1) includes: (2.1.1) Each client only uploads shared parameters. The server performs a weighted average to generate global parameters: , in, These are globally shared parameters generated by server aggregation. Indicates the first The region in the first Shared module parameters of the wheel, This represents the total number of regions participating in the training. (2.1.2) Subsequently, client-side adaptive updates are performed, with the client updating based on the cosine similarity with the global parameters. Make personalized adjustments: , , in, These are the updated locally shared parameters. It is the cosine similarity weight factor, used to measure the similarity between local parameters and global parameters; the more similar they are, the greater the weight.
8. The cross-regional adaptive joint transfer learning method for urban spatiotemporal prediction according to claim 1, characterized in that, Step (2.2) includes: (2.2.1) When making predictions in the target area, route selection is performed, expert weights are calculated, and the Top-k are selected: , , in, It is a gating network. It is the distribution of expert weights. These are the k highest weights after normalization; (2.2.2) Perform expert ensemble and weighted summation of the outputs of the activated experts: , in, It is the total number of experts. It is the first The fusion and embedding output of individual experts It is the result of expert weighted prediction.
9. The cross-regional adaptive joint transfer learning method for urban spatiotemporal prediction according to claim 1, characterized in that, Step (3) includes: fusing dynamic MoE output with local prediction: , in, It is the weighting coefficient that balances expert output and local output. This is the final prediction result.
10. A structure for implementing the above method, characterized in that... include: The system comprises an encoder, a cross-modal alignment module, and a hybrid expert decoder. The encoder includes a spatiotemporal network coding branch module and a large language model-driven coding branch module. The cross-modal alignment module uses an attention mechanism, and the hybrid expert decoder adopts a shared sparse hybrid expert architecture with gated routing. The spatiotemporal network coding branch module and the large language model-driven coding branch module are respectively connected to the cross-modal alignment module, which is connected to the hybrid expert decoder. The spatiotemporal graph is input into the spatiotemporal network coding branch module and the large language model-driven coding branch module to obtain cross-modal enhanced spatiotemporal features and local predictions. The cross-modal enhanced spatiotemporal features are input into the hybrid expert decoder to dynamically integrate expert weighted prediction results and combine them with local predictions to generate the final prediction result.