A method for learning dynamic representations of urban areas based on contextual representation learning.
By constructing a continuous dynamic graph and a multi-level information extractor, a dynamic representation vector is built for urban areas, which solves the problem of insufficient urban area characterization in the existing technology, and can quickly respond to the task needs of changing environments, reduce resource consumption and calculation costs, and provide accurate downstream task support.
Patent Information
- Application Number
- CN202510805132.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-06-17
AI Technical Summary
The existing technology cannot effectively capture urban regional representations dynamically, resulting in insufficient accuracy and generalization in changing urban downstream tasks, and requires separate fine-tuning for each task, consuming a large amount of resources.
By constructing a continuous dynamic graph, dynamic representation vectors are constructed for urban areas using short-time, medium-time and long-time information extractors, and combined with context representation learning, it is directly applied to multiple downstream tasks to reduce dependence on label data.
It realizes multi-dimensional and multi-level modeling of urban areas, can quickly respond to task requirements in changing environments, reduce computing costs and time consumption, and provide accurate market forecasting and personalized recommendation services.
Smart Images

Figure CN120317401B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of urban function and spatiotemporal data mining, and in particular relates to a method for learning dynamic representations of urban areas based on contextual representation learning. Background Art
[0002] Large-scale group mobility data plays a crucial role in the study of contemporary urban function. By recording the movement trajectories and behavioral patterns of people in urban spaces, this data provides a solid analytical foundation for a number of downstream urban tasks, such as regional traffic flow prediction, regional housing price prediction, and regional crime rate prediction. Because cities are composed of a variety of distinct areas, including commercial and residential areas, understanding urban function requires effectively learning high-quality regional representations.
[0003] In recent years, many researchers have modeled and learned urban region representations based on large-scale group mobility data. By obtaining a universal urban region representation, they aim to solve various downstream urban tasks. Consequently, the requirements for the accuracy and generalization of urban region representation modeling are increasing. Currently, existing methods primarily rely on using inter-regional mobility events as a measure of the strength of relationships between regions. However, this process ignores the actual duration of mobility events and fails to model urban region representations at different time scales, resulting in a significant loss of group mobility information. In real-world downstream task scenarios, due to the dynamic nature of mobility activities, downstream tasks related to urban functions are also constantly changing. Traditional methods cannot meet the requirements of fine-grained and dynamic tasks. Furthermore, existing methods require further fine-tuning of the learned urban region representation for each downstream task, which consumes additional time and resources. Therefore, how to dynamically capture and learn urban region representations and directly apply these representations to various downstream tasks without fine-tuning has become an urgent problem. Summary of the Invention
[0004] To solve the above technical problems, the present invention provides a method for learning dynamic representations of urban areas based on contextual representation learning, comprising the following steps:
[0005] Step S1: Construct a continuous dynamic graph based on public transportation data. The continuous dynamic graph is a collection of movement data sorted by time. ;
[0006] Step S2: Construct an urban area dynamic representation vector for each urban area , construct short-time-information feature extractor, medium-time-information fusion extractor and long-time-information update extractor, based on and Extract basic dynamic feature vector, medium time granularity urban area dynamic representation vector and current city dynamic representation vector respectively ;
[0007] Step S3: Based on and Constructing a historical region representation sequence and a current representation, and training the short-time information feature extractor, the medium-time information fusion extractor, and the long-time information update extractor based on the current representation and a plurality of pre-training tasks;
[0008] Step S4: Use the trained short-time information feature extractor, medium-time information fusion extractor and long-time information update extractor to obtain the urban area dynamic representation, and combine the urban area dynamic representation with context representation learning to apply it to multiple downstream tasks.
[0009] Beneficial effects:
[0010] 1. This paper proposes a dynamic learning method for urban region representation. Using a multi-level spatiotemporal information extractor, this method extracts spatiotemporal information at different time granularities from raw large-scale group mobility data, enabling multi-dimensional and multi-level modeling of cities. To address the problem that previous methods cannot be directly applied to dynamic urban downstream tasks, such as predicting multi-day urban regional traffic flow, this paper proposes a dynamic urban region representation encoding framework. By maintaining a dynamic urban region memory and continuously updating it with each detailed travel data, this framework enables dynamic modeling of regional representations, providing more accurate support for market forecasting, traffic flow analysis, and other applications.
[0011] 2. The present invention proposes to predict the city's downstream tasks based on contextual representation learning. The model can be automatically adjusted according to the input task information and historical data, which greatly reduces the dependence on a large amount of labeled data and demonstrates a strong multi-tasking capability, because it does not need to be fine-tuned separately for each new task, which enables it to quickly respond to various task requirements in a changing environment. The present invention can easily cope with multiple related tasks such as price forecasting and user demand forecasting without the need for special training for each task. In the scenario of resource scheduling, the present invention can optimize resource allocation in real time according to dynamically changing needs, and the personalized recommendation system can quickly provide users with accurate recommendation services through contextual information. Through this efficient task switching and information extraction mechanism, the present invention can exert significant performance advantages in practical applications and greatly reduce the time and computing costs of enterprises in dealing with complex tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 This is a flow chart of a method for learning dynamic representations of urban areas based on contextual representation learning according to the present invention;
[0013] Figure 2Schematic diagram of downstream task application based on context representation learning in an embodiment of the present invention. DETAILED DESCRIPTION
[0014] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only intended to illustrate the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.
[0015] Example 1
[0016] like Figure 1 As shown, an embodiment of the present invention provides a method for learning dynamic representations of urban areas based on contextual representation learning, comprising the following steps:
[0017] Step S1: Construct a continuous dynamic graph based on public transportation data. The continuous dynamic graph is a collection of mobile data sorted by time. ;
[0018] Step S2: Construct an urban area dynamic representation vector for each urban area , construct short-time-information feature extractor, medium-time-information fusion extractor and long-time-information update extractor, based on and Extract basic dynamic feature vector, medium time granularity urban area dynamic representation vector and current city dynamic representation vector respectively ;
[0019] Step S3: Based on and Construct a historical region representation sequence and a current representation. Based on the current representation and combined with multiple pre-training tasks, train a short-term information feature extractor, a medium-term information fusion extractor, and a long-term information update extractor.
[0020] Step S4: Use the trained short-time information feature extractor, medium-time information fusion extractor, and long-time information update extractor to obtain the dynamic representation of the urban area, and combine the dynamic representation of the urban area with context representation learning to apply it to multiple downstream tasks.
[0021] In one embodiment, the above step S1: constructs a continuous dynamic graph based on public transportation data, which is a collection of a series of movement data sorted by time. , specifically including:
[0022] Step S11: Eliminate public transportation data with missing information according to simple rules, and construct mobile data based on this ,in represents the i-th individual mobile data record, Respectively represent the starting longitude, starting latitude, ending longitude, ending latitude, departure time, and arrival time;
[0023] Step S12: Sort all mobile data in ascending order according to departure time, and divide the processed mobile data into several travel data blocks according to fine-grained time. , which is calculated as follows:
[0024] (1)
[0025] in, is the current time, Is the current travel data block The start time, The travel time belongs to and The collection of all mobile data between;
[0026] Step S13: Divide the city map into several areas according to the current city planning basis, and use the mapping function to The starting longitude and latitude and the ending longitude and latitude in the data are mapped to the regional serial number, and the regional serial number is divided according to the starting longitude and latitude in the data to obtain the processed large-scale group movement data. ,in, represents the i-th mobile data after processing, They represent the starting area, arrival area and departure time respectively.
[0027] In one embodiment, the above step S2: constructs a city area dynamic representation vector for each city area , construct short-time-information feature extractor, medium-time-information fusion extractor and long-time-information update extractor, based on and Extract basic dynamic feature vector, medium time granularity urban area dynamic representation vector and current city dynamic representation vector respectively , specifically including:
[0028] Step S21: Use the city area encoder to generate the Maintain a dynamic representation vector of a city area , used to store the dynamic changes of area i;
[0029] Step S22: Construct a short-term information feature extractor to encode fine-grained motion data into basic dynamic feature vectors , the specific steps are as follows:
[0030] First, divide the travel data time into hours and calculate the time used for coding area The starting time period is arrive Internal fine-grained The basic dynamic eigenvector of , the formula is as follows:
[0031] (2)
[0032] (3)
[0033] (4)
[0034] in, and are trainable parameters; and All represent intermediate variables; It will and Functions stitched together; and Representing regions and region Dynamic representation vector of ; and Represented in The departure area is and the arrival area is All travel events;
[0035] Then, the time information of the travel data is encoded according to the time attenuation coefficient, and the information is preliminarily aggregated according to the departure area and arrival area of the trip to obtain the basic dynamic feature vector of the time period. ;
[0036] Step S23: Construct a mid-time-information fusion extractor to perform data enhancement on the basic dynamic feature vector to obtain the enhanced mid-time granularity urban area dynamic representation vector :
[0037] For the basic dynamic feature vector Perform cross-time attention aggregation, and the calculation formula is as follows:
[0038] (5)
[0039] (6)
[0040] in, and are all trainable parameters, is a normalization function used to reduce the internal offset of the input data, is the dynamic representation vector of urban areas with medium temporal granularity obtained after cross-temporal aggregation data enhancement;
[0041] Step S24: Construct a long-term information update extractor for Perform selective learning and storage to generate the current city dynamic representation vector :
[0042] according to The city history dynamic representation vector and the learnable parameters are used to dynamically update the city dynamic representation. The formula is as follows:
[0043] (7)
[0044] in, are trainable parameters, is the dynamic representation vector of urban history, is the element-wise product operation;
[0045] Learnable parameters are introduced to control the flow and update of urban area representation information, ensuring that the model can capture urban area characteristics for a long time.
[0046] In one embodiment, the above step S3: based on and Construct a historical region representation sequence and a current representation. Based on the current representation, train a short-term information feature extractor, a medium-term information fusion extractor, and a long-term information update extractor in combination with various pre-training tasks. Specifically, the following steps are involved:
[0047] Step S31: Dynamic representation vector of urban area with medium time granularity Input the long-term information update extractor and get the current city dynamic representation vector and urban area dynamic representation vector Combined, we obtain a historical regional representation sequence;
[0048] Step S32: extract the first C representations from the historical region representation sequence in chronological order, generate the current representation through the region representation attention fusion mechanism, and add the current representation to the historical region representation sequence;
[0049] Step S33: Design pre-training tasks for adjusting the parameters of the extractor, including: regional input and output flow prediction tasks and inter-regional flow data prediction tasks, as shown in the following formula:
[0050] (8)
[0051] (9)
[0052] (10)
[0053] in, 、 、 They represent the input flow data, output flow data, and OD transfer data from area i to area j at time t, respectively; and Respectively represent the selection from The moment begins time interval;
[0054] Step S34: Pre-train the short-time information feature extractor, the medium-time information fusion extractor, and the long-time information update extractor through the current representation, construct the loss function L, and back-propagate to update the model parameters:
[0055] (11)
[0056] (12)
[0057] Among them, m represents the number of pre-training tasks, n represents the number of urban areas, Represent the actual value and predicted value of pre-training task k in region i, respectively. represents the parameter used to adjust the weight of the pre-training task; K represents the total number of tasks;
[0058] Finally, the trained short-time information feature extractor, medium-time information fusion extractor and long-time information update extractor are obtained.
[0059] In one embodiment, step S4 above: using the trained short-term information feature extractor, medium-term information fusion extractor, and long-term information update extractor to obtain a dynamic representation of the urban area, and combining the dynamic representation of the urban area with contextual representation learning to apply it to multiple downstream tasks, specifically includes:
[0060] Step S41: Initialize the historical region representation sequence Used to store the representation of urban historical areas and use the urban area encoder to obtain the representation of urban area i at time t Storage Access ;
[0061] Step S42: Characterize the sequence from history Extract the historical region representation of the previous C days, and take the task data corresponding to the previous C days as the support node representation for prediction according to the current task:
[0062] (13)
[0063] Among them, k represents that the sample data belongs to the kth downstream task, represents the dynamic representation vector of region i at time t, represents the historical representation vector of the downstream task k at time t;
[0064] Step S43: Based on the attention mechanism, first construct a query vector based on the current moment representation of the city area:
[0065] (14)
[0066] in, The weight matrix is a learnable parameter;
[0067] Step S44: Construct a support vector library based on the historical representation vectors of the urban area and the historical data of the specific downstream task , and calculate the association weights between the current city region representation and the historical representation and data based on the query vector :
[0068] (15)
[0069] (16)
[0070] in, Indicates that region i is The key-value vector at the moment, Are two different weight matrices, which are learnable parameters;
[0071] Step S45: Pass Perform weighted aggregation and linear transformation on the kth downstream task data of the previous c days to obtain the target task data for prediction:
[0072] (17)
[0073] in, Represents the predicted value of the kth downstream task that needs to predict region i, Represents learnable parameters. MLP is a linear prediction layer that adjusts parameters during pre-training, and the parameters do not change during downstream task reasoning.
[0074] The trained three extractors mentioned above are combined with context representation learning to perform multiple urban area downstream tasks, such as crime rate prediction, housing price prediction, and regional traffic prediction.
[0075] Figure 2 A schematic diagram of downstream task applications based on contextual representation learning is shown.
[0076] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for learning dynamic representations of urban areas based on contextual representation learning, characterized in that: include: Step S1: Construct a continuous dynamic graph based on public transportation data. The continuous dynamic graph is a collection of movement data sorted by time. ; Step S2: Construct an urban area dynamic representation vector for each urban area , construct short-time-information feature extractor, medium-time-information fusion extractor and long-time-information update extractor, based on and Extract basic dynamic feature vector, medium time granularity urban area dynamic representation vector and current city dynamic representation vector respectively , specifically including: Step S21: Use the city area encoder to generate the Maintain a dynamic representation vector of a city area , used for storage area Dynamic changes; Step S22: Construct a short-term information feature extractor to encode fine-grained motion data into basic dynamic feature vectors , the specific steps are as follows: First, divide the travel data time into hours and calculate the time used for coding area The starting time period is arrive Internal fine-grained The basic dynamic eigenvector of , the formula is as follows: (2) (3) (4) in, and are trainable parameters; and All represent intermediate variables; It will and Functions stitched together; and Representing regions and region Dynamic representation vector of ; and Represented in The departure area is and the arrival area is All travel events; Then, the time information of the travel data is encoded according to the time attenuation coefficient, and the information is preliminarily aggregated according to the departure area and arrival area of the trip to obtain the basic dynamic feature vector of the time period. ; Step S23: Construct a mid-time-information fusion extractor to perform data enhancement on the basic dynamic feature vector to obtain an enhanced mid-time granularity urban area dynamic representation vector : For the basic dynamic feature vector Perform cross-time attention aggregation, and the calculation formula is as follows: (5) (6) in, and are all trainable parameters, is the normalization function, A dynamic representation vector for urban areas at medium time granularity; Step S24: Construct a long-term information update extractor for Perform selective learning and storage to generate the current city dynamic representation vector : according to The city history dynamic representation vector and the learnable parameters are used to dynamically update the city dynamic representation. The formula is as follows: (7) in, are trainable parameters, is the dynamic representation vector of urban history, is the element-wise product operation; Step S3: Based on and Constructing a historical region representation sequence and a current representation, wherein the current representation is a dynamic representation of the urban region at the current moment, and training the short-term information feature extractor, the medium-term information fusion extractor, and the long-term information update extractor based on the current representation and in combination with multiple pre-training tasks; Step S4: Use the trained short-time information feature extractor, medium-time information fusion extractor and long-time information update extractor to obtain the urban area dynamic representation, and combine the urban area dynamic representation with context representation learning to apply it to multiple downstream tasks.
2. The urban area dynamic representation learning method based on contextual representation learning according to claim 1 is characterized in that: Step S1: Construct a continuous dynamic graph based on public transportation data. The continuous dynamic graph is a collection of movement data sorted by time. , specifically including: Step S11: Eliminate public transportation data with missing information according to simple rules, and construct mobile data based on this ,in represents the i-th individual mobile data record, Respectively represent the starting longitude, starting latitude, ending longitude, ending latitude, departure time, and arrival time; Step S12: Sort all mobile data in ascending order according to departure time, and divide the processed mobile data into several travel data blocks according to fine-grained time. , which is calculated as follows: (1) in, is the current time, Is the current travel data block The start time, The travel time belongs to and The collection of all mobile data between; Step S13: Divide the city map into several areas according to the current city planning basis, and use the mapping function to The starting longitude and latitude and the ending longitude and latitude in the data are mapped to the regional serial number, and the regional serial number is divided according to the starting longitude and latitude in the data to obtain the processed large-scale group movement data. ,in, represents the i-th mobile data after processing, They represent the starting area, arrival area and departure time respectively.
3. The urban area dynamic representation learning method based on contextual representation learning according to claim 2 is characterized in that: Step S3: Based on and Constructing a historical region representation sequence and a current representation, and training the short-term information feature extractor, the medium-term information fusion extractor, and the long-term information update extractor based on the current representation and multiple pre-training tasks, specifically including: Step S31: Dynamic representation vector of the urban area with medium time granularity Input the long-term information update extractor and obtain the current city dynamic representation vector Dynamic representation vector of the urban area Combined, we obtain a historical regional representation sequence; Step S32: extract the first C representations from the historical region representation sequence in chronological order, generate a current representation through the region representation attention fusion mechanism, and add the current representation to the historical region representation sequence; Step S33: Design pre-training tasks for adjusting the parameters of the extractor, including: regional input and output flow prediction tasks and inter-regional flow data prediction tasks, as shown in the following formula: (8) (9) (10) in, 、 、 They represent the input flow data, output flow data and OD transfer data from area i to area j at time t, respectively. and Respectively represent the selection from The moment begins time interval; Step S34: Pre-train the short-time information feature extractor, the medium-time information fusion extractor, and the long-time information update extractor through the current representation, construct a loss function L, and back-propagate to update model parameters: (11) (12) Among them, m represents the number of pre-training tasks, n represents the number of urban areas, They represent the actual value and predicted value of the model pre-training task k in region i, represents the parameter used to adjust the weight of the pre-training task; K represents the total number of tasks; Finally, the trained short-time information feature extractor, medium-time information fusion extractor and long-time information update extractor are obtained.
4. The urban area dynamic representation learning method based on contextual representation learning according to claim 3 is characterized in that: Step S4: using the trained short-term information feature extractor, medium-term information fusion extractor, and long-term information update extractor to obtain a dynamic representation of the urban area, and combining the dynamic representation of the urban area with context representation learning to apply it to multiple downstream tasks, specifically including: Step S41: Initialize the historical region representation sequence Used to store the representation of urban historical areas and use the urban area encoder to obtain the representation of urban area i at time t Storage Access ; Step S42: From the historical representation sequence Extract the historical region representation of the previous C days, and take the task data corresponding to the previous C days as the support node representation for prediction according to the current task: (13) Among them, k represents that the sample data belongs to the kth downstream task, represents the dynamic representation vector of region i at time t, represents the historical representation vector of the downstream task k at time t; Step S43: Based on the attention mechanism, first construct a query vector based on the current moment representation of the city area: (14) in, The weight matrix is a learnable parameter; Step S44: Construct a support vector library based on the historical representation vector of the urban area and the historical data of the specific downstream task, and calculate the association weight of the current urban area representation and the historical representation and data based on the query vector. : (15) (16) in, Indicates that region i is The key-value vector at the moment, Are two different weight matrices, which are learnable parameters, and T represents the transpose operation of the matrix; Step S45: Pass To the front The k-th downstream task data of the day is weighted aggregated and linearly changed to obtain the target task data for prediction: (17) in, Represents the value of the kth downstream task that needs to predict the area 𝑖, Represents learnable parameters. MLP is a linear prediction layer that adjusts parameters during pre-training, and the parameters do not change during downstream task reasoning.
Citation Information
Patent Citations
Urban short-term traffic flow prediction method, system and equipment
CN118097982A
Multi-granularity graph learning system for cross-city logistics demand prediction
CN118627675A