Explanatable multi-modal city region representation method
By extracting multimodal features and generating representation vectors using graph structure and variational graph autoencoder, the problem of relying on labeled data in the prior art is solved, and high-precision and interpretable urban area representation is achieved, which is suitable for large-scale urban data sets.
Patent Information
- Application Number
- CN202510257148.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-29
AI Technical Summary
The existing urban area characterization methods rely on a large amount of labeled data, making it difficult to fully capture the complex functions and characteristics of urban areas, and have limited characterization capabilities, especially poor performance in aspects other than traffic flow.
By extracting POI features, remote sensing image features, OD inflow features and parking lot features as node features, combining OD flow interaction features and distance features as edge features, using graph structure to describe regional correlation, multi-head attention mechanism and variational graph autoencoder VGAE generate representation vectors, and using decoder to reconstruct features, and finally interpreting the graph substructure and node features that contribute the most to the representation vector through the interpretation module.
It can automatically extract potential feature representations without a large amount of label data, improve the accuracy and interpretability of urban area representations, comprehensively capture the characteristics of urban areas, and support urban planning and management.
Smart Images

Figure CN120388296A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of urban area representation, and more particularly, to an interpretable multi-modal urban area representation method. Background Art
[0002] Graph neural network (GNN) is a neural network model capable of processing graph-structured data, and it has achieved remarkable results in urban area representation learning. However, in practical applications, applying graph neural network to urban area representation learning has the following drawbacks:
[0003] Most of the existing methods for representing urban areas rely on supervised learning and require a large amount of labeled data to train the model. However, in the functional representation of urban areas, the functions and characteristics of urban areas are complex and variable, and it is difficult to accurately describe them with simple labels. Moreover, the representation methods are often designed to solve a specific task. For example, the representation method is particularly good at predicting traffic flow in urban areas, but has no representation ability or poor representation ability in other aspects. Summary of the Invention
[0004] The purpose of this application is to provide an interpretable multi-modal urban area representation method, which can improve the diversity and accuracy of feature representation.
[0005] This application is implemented as follows:
[0006] This application provides an interpretable multi-modal urban area representation method, including the following steps:
[0007] Divide the target city into multiple regions, each region as a node, and the edges formed by connecting the nodes as the relationships between the regions;
[0008] Extract POI features, remote sensing image features, OD inflow features, and parking lot features as node features to represent the functional attributes and traffic mobility of each region, and extract OD flow interaction features and distance features as edge features to represent the traffic flow and spatial proximity between regions;
[0009] Describe the correlation between regions through a graph structure, capture the dependencies between different dimensions within a region through a multi-head attention mechanism to obtain the processed node features and edge features, and input the processed node features and edge features into a variational graph autoencoder (VGAE) for processing to generate a representation vector;
[0010] Use a decoder to reconstruct the node features and edge features from the representation vector;
[0011] Use an interpretation module to interpret the graph sub-structures and node features that contribute the most to each dimension of the representation vector.
[0012] Compared with the prior art, the present application has at least the following advantages or beneficial effects:
[0013] The present application provides an interpretable multi-modal urban area representation method. By extracting POI features, remote sensing image features, OD inflow features, and parking lot features as node features to characterize the functional attributes and traffic mobility of each area, and extracting OD flow interaction features and distance features as edge features to characterize the traffic flow and spatial proximity between areas, it can capture the characteristics of urban areas more comprehensively and improve the accuracy of area representation. The correlation between areas is described by a graph structure, and the multi-head attention mechanism is used to capture the dependencies between different dimensions within the area to obtain the processed node features and edge features, effectively capturing the complex dependencies within and between urban areas. By inputting the processed node features and edge features into a variational graph autoencoder (VGAE) for processing to generate representation vectors, latent feature representations are automatically extracted from multi-source data without the need for a large amount of labeled data, and it is applicable to large-scale urban datasets. By using a decoder to reconstruct the node features and edge features from the representation vectors, while processing both area feature reconstruction and inter-area relationship reconstruction, the complexity of urban areas can be better captured, providing more comprehensive support for urban planning and management. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0015] Figure 1 It is a flowchart of an interpretable multi-modal urban area representation method of the present application;
[0016] Figure 2 It is a schematic structural diagram of the EMVGAE model. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Usually, the components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations.
[0018] The following will describe in detail some embodiments of the present application in conjunction with the drawings. Without conflict, the various embodiments and the various features in the embodiments below can be combined with each other.
[0019] Embodiment
[0020] An embodiment of the present application provides an interpretable multimodal urban area representation method, which can enhance the representation ability of urban areas.
[0021] Please refer to Figure 1 , the interpretable multimodal urban area representation method includes the following steps:
[0022] S1: Divide the target city into multiple regions, with each region as a node, and the edges formed by connecting the nodes as the relationships between the regions;
[0023] Specifically, the target city is preferably divided into grids of 1 km × 1 km, with each grid as a region and also representing a node in the graph, and the relationships between the grids are represented by edges.
[0024] S2: Extract POI features, remote sensing image features, OD inflow features, and parking lot features as node features to characterize the functional attributes and traffic mobility of each region, and extract OD flow interaction features and distance features as edge features to characterize the traffic flow and spatial proximity between regions;
[0025] Specifically, POI features, remote sensing image features, OD inflow features, and parking lot features are extracted from POI data, remote sensing image data, OD inflow data, and parking lot data as node features to characterize the functional attributes and traffic mobility of each region. OD flow interaction features and distance features are extracted from OD flow interaction data and distance data as edge features to characterize the traffic flow and spatial proximity between regions. POI data, remote sensing image data, OD inflow data, parking lot data, OD flow interaction data, and distance data are all high-dimensional input data, and these data generate low-dimensional representation vectors after a series of subsequent processes. One of the objectives of the present invention is to extract high-dimensional input data to generate low-dimensional representation vectors, which can improve the efficiency of data processing and analysis, facilitate data visualization, and optimize data storage and transmission. POI (Point of Interest) features refer to specific locations on the map, such as restaurants, hospitals, schools, etc. Remote sensing images are images of the Earth's surface captured from the air or space. By processing and analyzing these images, features such as land use type, vegetation cover, building density, etc. can be extracted. OD (Origin-Destination) inflow features refer to the attractiveness of a certain region as a destination, that is, how much human or vehicle flow flows into this region from other regions, which reflects the traffic demand and attractiveness of this region. Parking lot features include information such as the number of parking spaces, usage frequency, parking fees, etc. These features can reflect the traffic capacity, parking demand, and possible traffic congestion conditions of the region. OD flow interaction features refer to the exchange of human or vehicle flow between regions, reflecting the traffic mobility and interdependence between regions. The distance feature refers to the physical distance or traffic time between regions, which is an important basis for evaluating the spatial proximity between regions. POI data, remote sensing image feature data, and OD inflow data are used as input features of nodes; these features can comprehensively describe the functional attributes (such as business districts, residential areas, industrial areas, etc.) and traffic mobility of each region. While OD inflow data and map data are used as input features of edges. These data capture the interaction relationships between regions and characterize the traffic flow and spatial proximity between regions. This step is a method of comprehensively using multiple data sources to characterize the functional attributes of a city or region, traffic mobility, and the interaction relationships between regions. Through these feature extractions, more accurate and comprehensive data support can be provided for fields such as urban planning, traffic management, and economic analysis. It can also capture the characteristics of urban regions more comprehensively and improve the accuracy of regional representation.
[0026] S3: Describe the correlation between regions through a graph structure, capture the dependencies between different dimensions within a region through a multi-head attention mechanism to obtain processed node features and edge features, and input the processed node features and edge features into a variational graph autoencoder VGAE for processing to generate representation vectors;
[0027] Specifically, the graph structure is used as a data representation method, where nodes represent different regions, and edges represent certain relationships or correlations between these regions. The correlations between target urban regions are described by the graph structure, where multiple data sources are used to reflect different regional attributes. The multi-head attention mechanism is a deep learning technique that allows the model to simultaneously focus on multiple different parts when processing input data. In this process, the multi-head attention mechanism is used to capture the dependencies between different dimensions (POI, remote sensing images, OD inflows, parking lots, OD flow interactions, and distances) within a region. These dependencies are used to update or process the original node features and edge features, making them better reflect the complexity within the region and the interactions between dimensions. The variational graph autoencoder (VGAE) is a model that combines graph neural networks and variational autoencoders, and it can learn the latent representation of graph-structured data. In this step, the processed node features and edge features are passed as input data to the VGAE. The VGAE converts these input data into a representation vector in a low-dimensional latent space. These representation vectors capture the global and local features of the graph-structured data and can be used for subsequent analysis tasks. This step automatically extracts the latent feature representation from multi-source data without the need for a large amount of labeled data and is applicable to large-scale urban datasets.
[0028] S4: Use the decoder to reconstruct the node features and edge features from the representation vector;
[0029] Specifically, the decoder is used to reconstruct the feature vector of the node or the feature of the edge from the low-dimensional representation vector. That is, the decoder recovers the features of the nodes and edges in the original graph from the representation vector. For node features, the decoder outputs a vector with the same dimension as the original node feature vector. For edge features, if the graph data contains explicit edge features, the decoder also reconstructs these features.
[0030] S5: Use the interpretation module to interpret the graph substructures and node features that contribute the most to each dimension of the representation vector.
[0031] Specifically, the interpretation module is used to reveal which graph substructures and node features contribute the most to each dimension of the representation vector. This helps us better understand how the representation vector extracts and utilizes information from the original graph data, thereby improving interpretability and transparency.
[0032] In some embodiments of the present invention, extracting POI features includes:
[0033] Select multiple types of points of interest, extract the category and location information of the POIs, and calculate the proportion of each POI category in each area based on the category and location information of the POIs through longitude and latitude to obtain a vector (1, n); for each area, obtain a vector with multiple values, and then calculate the POU embedding vector obtained from the contribution degrees of different POIs according to TF-IDF as: P = (P1, P2, P3,..., P n ), where P is the contribution degree vector of the area POIs; P1,..., P n respectively represent the proportion of the contribution degrees of POIs with different labels.
[0034] Specifically, the distributions of different types of POIs reflect the characteristics of the city and the main activity types, and directly affect the regional function distribution. In this embodiment, 14 types of points of interest are selected, the category and location information of the POIs are extracted, and the proportion of each POI category in each area is calculated through longitude and latitude to obtain a vector (1, n). For each area, obtain a vector with 14 values, and then calculate the POU embedding vector obtained from the contribution degrees of different POIs according to TF-IDF as: P = (P1, P2, P3,..., P n ), where P is the contribution degree vector of the area POIs. P1,..., P n respectively represent the proportion of the contribution degrees of POIs with different labels.
[0035] In some embodiments of the present invention, the extraction of OD inflow characteristics includes:
[0036] Collect taxi OD data for a preset duration, where the preset duration includes one month, and one month is divided into weekdays and weekends; the data includes the boarding time, alighting time, boarding longitude and latitude, and alighting longitude and latitude; divide a day into multiple time periods, divide a day from 6 am to 6 pm into 6 time periods, so there are a total of 6 × num day OD volume intervals, and num day represents the total number of days in a month. Extract the longitude and latitude of the alighting location to calculate the landing area, and extract the alighting time to calculate the time slice at the time of landing; finally, obtain the regional inflow time series vector (1, m) of each area, where m is the OD volume interval in the above-mentioned one month, that is, the dimension; in the OD inflow vector T(0, 0, 24, 14,..., 0), 24 indicates that the OD volume flowing into this area during the time period from 10 am to 12 pm on weekdays is 24; define the OD inflow vector T as follows: T = (T1, T2, T3,..., T m ), where T is the OD inflow vector of the area, and T1,..., T m respectively represent the OD inflow distributions in different time dimensions.
[0037] Specifically, the inflow OD feature is the inflow of taxi OD in a region, which reflects the characteristics of the region to a certain extent. In this embodiment, one month of taxi OD data is collected and divided into weekdays and weekends. The data includes the pick-up time, drop-off time, pick-up longitude and latitude, and drop-off longitude and latitude. The day is divided into 6 time periods from 6 am to 6 pm, so there are a total of 6×num day OD volume intervals, where num day represents the total number of days in a month. The longitude and latitude of the drop-off location are extracted to calculate the landing area, and the drop-off time is extracted to calculate the time slice at the time of landing. Finally, the regional inflow time series vector (1, m) of each region is obtained, where m is the OD volume interval of the above month, that is, the dimension. For example, in the OD inflow vector T(0, 0, 24, 14,..., 0), 24 indicates that the OD volume flowing into this region during the period from 10 am to 12 pm on a weekday is 24. The OD inflow vector T is defined as follows: T = (T1, T2, T3,..., T m ), where T is the OD inflow vector of the region, and T1,..., T m respectively represent the OD inflow distributions in different time dimensions.
[0038] In some embodiments of the present invention, extracting remote sensing image features includes:
[0039] Using ResNet-50 to extract feature vectors from remote sensing images, and the output of the convolutional layer is obtained as follows: Y = W*X + b, where X is the input remote sensing image matrix, Y is the new feature obtained after convolution, W is the weight tensor, and b is the offset term; * represents the convolution operation; then y is defined as the residual block: y = F(x) + x, where F(x) is the residual function; finally, calculate the global average pooling: where Z is the feature after the residual connection; the shape of the feature map Z is (C, H, W), where C is the number of channels, and H and W are the height and width respectively; the resulting feature vector is obtained: S = (S1, S2, S3,..., S k ), where S1,..., S k represent the features of different dimensions of the remote sensing image of the region.
[0040] Specifically, in urban representation learning, remote sensing images are important spatial data sources, and their feature vectors can reveal the functional similarities and differences of different urban areas. Deep learning techniques, especially convolutional neural networks (CNNs), have shown powerful capabilities in extracting high-dimensional feature vectors from remote sensing images. Therefore, ResNet-50 is used to extract feature vectors from remote sensing images, and the output of the convolutional layer is as follows: Y = W * X + b, where X is the input remote sensing image matrix, Y is the new feature obtained after convolution, W is the weight tensor, and b is the bias term. * represents the convolution operation. Then, y is defined as the residual block: y = F(x) + x, where F(x) is the residual function. Finally, global average pooling is calculated: where Z is the feature after residual connection. The shape of the feature map Z is (C, H, W), where C is the number of channels, and H and W are the height and width, respectively. The resulting feature vector is obtained: S = (S1, S2, S3,..., S k ), where S1,..., S k represent the features of remote sensing images in different dimensions of the region.
[0041] In some embodiments of the present invention, extracting parking lot features includes:
[0042] Construct a parking lot quantity vector L(1,1). The parking lot has only one-dimensional quantity dimension in the regional vector, which reflects the normalized parking record quantity dimension of the area and represents the relative scale of parking activities in each area; combine the feature vectors of POI, OD inflow, remote sensing images, and parking lots into a feature vector to represent a region: V = (P, T, S, L), where P is the POI feature vector, T is the OD inflow feature vector, S is the remote sensing image feature vector, and L is the parking lot feature vector.
[0043] Specifically, the distribution of parking lots provides insights into travel patterns because the accessibility of parking lots affects local traffic flow, land use, and economic activities. In the present invention, a parking lot quantity vector L(1,1) is constructed for each region, that is, the parking lot has only one-dimensional quantity dimension in the regional vector, which reflects the normalized parking record quantity dimension of the area and represents the relative scale of parking activities in each area. Combine the feature vectors of POI, OD inflow, remote sensing images, and parking lots into a feature vector to represent a region: V = Concatenate(P, T, S, L), where P is the POI feature vector, T is the regional inflow feature vector, S is the remote sensing image feature vector, and L is the parking lot feature vector. This comprehensive representation effectively captures the spatial, semantic, and mobility features of each region, providing a solid foundation for further analysis.
[0044] In some embodiments of the present invention, extracting OD flow interaction features includes:
[0045] Construct an OD interaction matrix to reflect the functional similarity of each region at different times. The OD interaction matrix tracks the taxi flow between regions at 6-hour intervals; for each trip, the algorithm calculates the pick-up and drop-off grid positions using longitude and latitude, and then assigns the trip to the correct time interval; since the time series of OD interactions is three-dimensional, represent the matrix as I, where I[i,j,k] represents the number of OD interactions between region i and region j within a certain time k; the OD interaction matrix is constructed as follows: I = (I1, I2, I3,..., I m ), where I1, I2, I3,..., I m is the interaction matrix of OD interactions between regions within a certain time interval in a month.
[0046] Specifically, this embodiment considers the spatio-temporal OD flow interaction and distance correlation as the relationships between regions. These relationships are estimated as the edges of a graph, capturing the interactions and dependencies that contribute to the overall urban dynamics. The correlation between urban regions is often revealed by the patterns of vehicle traffic interactions. These traffic exchanges reflect the functional and spatial relationships between regions. Construct an OD interaction matrix on the time series to reflect the functional similarity of each region at different times. The OD interaction matrix tracks the taxi flow between regions at 6-hour intervals (from 6 am to 6 pm). For each trip, the algorithm calculates the pick-up and drop-off grid positions using longitude and latitude, and then assigns the trip to the correct time interval. Since the time series of OD interactions is three-dimensional, represent the matrix as I, where I[i,j,k] represents the number of OD interactions between region i and region j within a certain time k. The OD interaction matrix is constructed as follows: I = (I1, I2, I3,..., I m ), where I1, I2, I3,..., I m is the interaction matrix of OD interactions between regions within a certain time interval in a month.
[0047] In some embodiments of the present invention, extracting distance features includes:
[0048] Obtain a distance matrix using a distance decay formula: where d i,j is the Euclidean distance between two regions i and j, and β is the decay factor; concatenate the distance matrix on the OD interaction sequence matrix to form edge features of multiple time scales and spaces as follows: Q = Concatenate(I, D), where I is the taxi OD interaction sequence matrix and D is the inter-region distance matrix.
[0049] Specifically, considering the influence of distance factors on the itinerary, the Euclidean distance between the regional centers is calculated, and the distance matrix is obtained using the distance decay formula: where d i,j is the Euclidean distance between two regions i and j, and β is the decay factor. Concatenate the distance matrix on the OD interaction sequence matrix to form edge features in multiple time scales and spaces as follows: Q = Concatenate(I, D), where I is the taxi OD interaction sequence matrix and D is the inter-regional distance matrix. This combination captures both the movement flow between regions and the spatial distance between regions. By including spatial and time information, it ensures the accurate capture of the interactions between regions, enabling a more precise analysis of urban mobility and its potential structural relationships.
[0050] In some embodiments of the present invention, the step of describing the correlation between regions through a graph structure, capturing the dependencies between different dimensions within a region through a multi-head attention mechanism to obtain processed node features and edge features, and inputting the processed node features and edge features into a variational graph autoencoder VGAE for processing to generate representation vectors specifically includes:
[0051] Represent the graph structure as G = (V, E, W), where V represents the nodes in the graph, E represents the edges in the graph, and W represents the weights of the edges; the node features are composed of the concatenation of P, T, S, and L feature vectors, denoted as V = (P, T, S, L); each node contains F = n + m + k + 1 features, where n, m, k, and l respectively represent the dimension numbers of the above-mentioned POI, OD inflow, remote sensing image, and parking lot. The dimension of the node feature matrix X is X ∈ R N×F , where N represents the number of nodes in the graph and F represents the node feature dimension; the edge features are composed of the concatenation of OD flow interaction features and the distance matrix, denoted as E = (I, D); assume the graph has M edges, and each edge contains Y = m + 1 features, making the dimension of the edge feature matrix Q e be Q e ∈ R M×Y ; Extract node features and edge features respectively;
[0052] Use the multi-head attention mechanism to capture the interrelationships between different dimensions within a region, and assign different attention weights to enhance the expressive ability of the representation; Generate query vector Q, key vector K, and value vector V from the node feature matrix X through linear transformations respectively: Q = XW Q , K = XW K , V = XW V , where W Q , W K , W Vis a learnable weight matrix; then calculate the single-head attention weights, calculate the similarity between the query Q and the key K using the dot product, and transform it into attention weights through softmax: In the formula, QK represents the feature attention correlation within all regions, is the scaling factor; to capture the interrelationships between different dimensions within the region, perform the calculation of multi-head attention, use h independent attention heads, with different weight matrices for each head, concatenate the outputs of all attention heads and pass them through a linear transformation: MultiHead(Q, K, V) = Concat(head1, head2,..., head h )W O , where W O is the weight of the output linear transformation; finally, perform a residual connection between the output of the multi-head attention and the input, and obtain the feature matrix updated by the attention mechanism through layer normalization: Z = LayerNorm(X + MultiHead(Q, K, V));
[0053] Calculate the edge weights reflecting the spatial movement dynamics to effectively simulate the relationships between regions; these weights come from the edge feature Q i,j , which includes OD flow interaction features and distance features; define the edge weights as: w i,j = σ(W E *MLP(Q i,j ))), where Q i,j is the original edge feature vector between node i and node j, σ is the sigmoid function, and W E is the trainable weight matrix; the obtained edge weights w i,j combine both long-distance interactions and short-distance adjacency relationships, effectively capturing the comprehensive spatial relationships within the graph;
[0054] To effectively model the influence of adjacent nodes, a graph attention mechanism is adopted, which ensures that the model captures local and global interactions by focusing on the most relevant nodes among adjacent nodes; the attention coefficient is calculated as: where is the attention score between node i and node j at layer l, capturing the semantic correlation between the two region features, a (l)T is the learnable attention vector at layer l, T represents the vector transpose, and W (l) is the weight matrix at layer l, and represent the node features of node i and node j at layer l respectively, and || represents concatenation; the final attention coefficient α i,j combines the learned attention score e i,j and the edge weight w i,j , as follows:
[0055] Denote the neighbor nodes of i. This combination enables the model to consider both node-level correlations and edge-level spatial dynamics, enabling it to capture more subtle inter-regional interactions. Using the calculated attention coefficient α i,j , update the node features as follows: Denote node i itself and its neighbor nodes; after applying the attention mechanism, the latent representation of the node is obtained: Z = X (l+1) , where Z ∈ R N*d , N represents the number of nodes in the graph, d represents the node feature dimension, and Z represents the latent representation matrix;
[0056] For the variational graph autoencoder VGAE, the output of the feature vector is further mapped to a mean vector and a log standard deviation vector as follows: μ = ZW μ + b μ , logσ 2 = ZW σ + b σ , where W μ and b μ are the weight matrix and bias term used in the mean calculation, and W σ and b σ are the weight matrix and bias term used in the log standard deviation calculation; μ, logσ 2 are the mean and log standard deviation vectors of the output respectively;
[0057] To facilitate the calculation of the gradient, the reparameterization technique is introduced to sample the 16-dimensional latent variable z to obtain the final representation vector: Z = μ + σ ⊙ ∈, where μ represents the mean vector, used to determine the central position of the latent variable; σ represents the standard deviation vector, used to control the distribution range of the latent variable; ⊙ is the Hadamard product, indicating element-wise multiplication between vectors; ∈ represents a random noise vector that follows the standard normal distribution , used to introduce randomness, thereby generating the latent variable; Z is the representation vector, that is, the final sampling result combining the mean μ, the standard deviation σ, and the random noise ∈;
[0058] In some embodiments of the present invention, the specific steps of using the decoder to reconstruct the node features and edge features from the representation vector include:
[0059] For the reconstruction of node features, pass the low-dimensional representation vector z through multiple stacked MLP layers and pass the relu activation function to construct the original regional features: X' = f MLP (z), where X' is the reconstructed node feature matrix;
[0060] For the reconstruction of edge features, the representation vectors z of the regions with interaction relationshipsi and z j Combine them, and construct the original regional interaction feature: E through multiple stacked MLP layers with the relu activation function i,j′ = f MLP (z i , z j ), where E i,j′ is the reconstructed edge feature matrix;
[0061] Its loss function consists of three parts: the KL divergence term, the node information reconstruction error term, and the edge information reconstruction error term, aiming to effectively reconstruct the features and correlations of complex multi-scale spatio-temporal urban regions; the total loss is as follows:
[0062]
[0063] where X and Q e represent the node and edge features on the graph, p(z) is the prior distribution of the latent variable z, is the variational posterior distribution, q(z|X, Q e ) is the probability distribution of the latent variable z given the node features and edge features, p(X|z) and p(Q e |z) are the conditional probability distributions of the node features and edge features given the latent variable; the specific form is:
[0064]
[0065] where σ i represents the latent standard deviation vector of node i, μ i represents the latent mean vector of node i, N represents the total number of nodes on the graph, M represents the total number of edges on the graph, X i is the original feature of node i, X i′ is the reconstructed feature of node i, E i,j represents the original edge feature between nodes i and j, E i,j′ represents the reconstructed edge feature of i, j.
[0066] In some embodiments of the present invention, it further includes using an interpretation module to interpret the representation vector, specifically:
[0067] Use GNNExplainer to reveal the graph substructure and node features that contribute the most to each dimension of the representation vector z. Specifically, input the pre-trained E-MVGAE model, graph structure G, and the learned representation vector z into the interpretation module; in the multi-objective optimization framework, regard each dimension z ij of the representation vector z as an independent objective; specifically, the objective is to identify the subgraph and the subset of node features which respectively represent GSi is a subgraph of G, X Si is a characteristic subset of X, and they make important contributions to the prediction of z ij ; we obtain:
[0068]
[0069] where H(z ij ) is the entropy of z ij , and H(z ij |G = G Si , X = X Si ) is the conditional entropy given the subgraph G Si and the characteristic subset X Si ;
[0070] Since H(z ij ) is a constant, maximizing the mutual information is equivalent to minimizing the conditional entropy:
[0071]
[0072] In this expression, H(z ij |G = G Si , X = X Si ) represents the uncertainty of z Si given the selected subgraph G Si and the node characteristics X ij ; minimizing this conditional entropy ensures that the selected subgraph and characteristics are highly predictive of the specific embedding dimension z ij ; P(z ij |G = G Si , X = X Si is the probability of observing z ij given the selected explanatory components, represents the expectation of the conditional distribution of z ij ; ensuring that the selected subgraph and characteristics capture the most critical patterns and relationships underlying z ij ;
[0073] To identify the most important dimensions in the learned representation vector z, the mutual information MI(z ij ,(G Si , X Si )) between each dimension z ij and all possible subgraphs G Si and node characteristics X Si is calculated; this process quantifies the contribution of each subgraph and characteristic subset to explaining a specific embedding dimension; the dimensions are sorted according to the calculated mutual information values, and the dimension with the highest mutual information value is selected as the most important dimension; to achieve this, two key mechanisms are used; first, a binary feature mask F ∈ {0,1} is appliedd to determine the key features z of each dimension ij , the selected features are represented as:
[0074] X Si = X ⊙ F ij , where ⊙ represents element-wise multiplication; use the soft adjacency mask A S ∈ [0, 1] n×n to select the subgraph G Si , ensuring that only the most influential nodes and edges are retained; then, jointly optimize these mechanisms to maximize the mutual information: By designing this optimization process, focus on the most relevant structures and feature-based components in each dimension of the embedding. By maximizing the mutual information, ensure that the selected subgraphs and features contain the most predictive patterns and relationships. This approach provides a detailed and interpretable representation of the embedding space, revealing the different contributions of features and subgraphs to each dimension.
[0075] In the present invention, an interpretable multi-modal urban area representation method can be implemented through the EMVGAE model, please refer to Figure 2 , Figure 2 which is the structural schematic diagram of the EMVGAE model. The model includes an EMVGAE encoder, a dual-task decoder, and an interpretation module.
[0076] EMVGAE encoder: responsible for processing high-dimensional input data to generate low-dimensional representation vectors, and uses the graph attention network GAT and the multi-head attention mechanism to learn node features. The graph attention network GAT can automatically focus on key regions, and by calculating the attention scores between regions to weight the relationships between different regions, it can effectively capture the dependencies between regions. The multi-head attention mechanism can capture the dependencies between different dimensions within a region. Then, further process the high-dimensional features of the nodes through the variational graph autoencoder (VGAE) to generate low-dimensional representation vectors. This process can enhance the model's compression ability for regional features by learning the latent variable distribution while retaining important information.
[0077] Dual-task Decoder: The dual-task decoder is responsible for reconstructing node features and edge features from the low-dimensional representation vectors. The dual-task decoder performs two tasks simultaneously: one is to reconstruct regional features, and the other is to reconstruct the relationships between regions. Its first task is to reconstruct the original features of the region (such as POIs, traffic flow, parking lot distribution, etc.) from the low-dimensional representation vectors. The low-dimensional vectors are decoded through a multi-layer perceptron (MLP) to generate predicted node features. The second task of the decoder is to reconstruct the relationships (edge features) between regions from the low-dimensional representation vectors, that is, to predict the traffic flow, interaction intensity, etc. between regions. The decoder concatenates the low-dimensional representations between node pairs and inputs them into the MLP for reconstruction. The decoder is not only responsible for the recovery of node features, but also ensures that the representation vector of each node can reflect the relative relationships between regions (such as traffic flow or spatial connections between regions) by learning the relationships and interactions between nodes. This dual-task mechanism enables the model to not only accurately represent the functional characteristics of each region, but also capture the interaction relationships between different regions. Specifically, one part of the decoder is used to recover the original regional features, and the other part focuses on reconstructing the connection strength between nodes, thereby realizing the simultaneous modeling of regional characteristics and interaction patterns.
[0078] Explanation Module: To improve the transparency and interpretability of the model, the GNNExplainer module is introduced. This module analyzes the low-dimensional representation vectors in detail and clearly indicates which specific input elements play a key role in the final result. GNNExplainer identifies important features and nodes by perturbing the input data and observing the changes in the model output, thereby revealing the specific contributions of different input features to the representation vectors. For example, for the representation vector of this region, the scenic spots and tourism dimensions in the POIs have the greatest impact on the final representation, indicating that this region is mainly composed of tourist attractions.
[0079] For those skilled in the art, it is obvious that this application is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of this application, this application can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of this application is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in this application. Any reference signs in the claims should not be regarded as limiting the claimed rights.
Claims
1. An interpretable multi-modal urban area representation method, characterized in that, It includes the following steps: Divide the target city into multiple regions, with each region serving as a node, and the edges formed by connecting the nodes representing the relationships between the regions; Extract POI features, remote sensing image features, OD inflow features, and parking lot features as node features to characterize the functional attributes and traffic mobility of each region, and extract OD flow interaction features and distance features as edge features to characterize the traffic flow and spatial proximity between regions; Describe the correlation between regions through a graph structure, capture the dependencies between different dimensions within the regions through a multi-head attention mechanism to obtain the processed node features and edge features, and input the processed node features and edge features into a variational graph autoencoder (VGAE) for processing to generate representation vectors; Use a decoder to reconstruct the node features and edge features from the representation vectors; Use an interpretation module to interpret the graph sub-structures and node features that contribute the most to each dimension of the representation vectors.
2. The interpretable multi-modal urban area representation method according to claim 1, wherein The extraction of POI features includes: Select multiple types of points of interest, extract the category and location information of the POI, and calculate the proportion of each POI category in each area based on the category and location information of the POI through longitude and latitude to obtain a vector (1, n); for each area, obtain a vector with multiple values, and then according to TF-IDF, the POU embedding vector calculated from the contribution degrees of different POIs is: P = (P1, P2, P3,..., P n ), where P is the contribution degree vector of the area POI; P1,..., P n respectively represent the proportion of the contribution degrees of POIs with different labels.
3. The method for interpretable multi-modal urban area representation according to claim 2, wherein The extraction of OD inflow features includes: Collect taxi OD data for a preset duration, where the preset duration includes one month, and one month is divided into weekdays and weekends; the data includes pick-up time, drop-off time, pick-up longitude and latitude, and drop-off longitude and latitude; extract the longitude and latitude of the drop-off location to calculate the landing area, and extract the drop-off time to calculate the time slice at the landing point; finally, obtain the regional inflow time series vector (1, m) for each region, where m is the OD volume interval for the above-mentioned one month, i.e., the dimension; in the OD inflow vector T(0, 0, 24, 14,..., 0), 24 indicates that the OD volume flowing into this region during the period from 10 am to 12 pm on a weekday is 24; define the OD inflow vector T as follows: T = (T1, T2, T3,..., T m ), where T is the OD inflow vector for the region, and T1,..., T m respectively represent the OD inflow distributions in different time dimensions.
4. The method for interpretable multi-modal urban area representation according to claim 3, wherein The extraction of remote sensing image features includes: Use ResNet-50 to extract feature vectors from remote sensing images. The output of the convolutional layer is as follows: Y = W * X + b, where X is the input remote sensing image matrix, Y is the new feature obtained after convolution, W is the weight tensor, and b is the bias term; * represents the convolution operation; then define y as the residual block: y = F(x) + x, where F(x) is the residual function; finally, calculate the global average pooling: where Z is the feature after residual connection; the shape of the feature map Z is (C, H, W), where C is the number of channels, and H and W are the height and width respectively; obtain the resulting feature vector: S = (S1, S2, S3,..., S k ), where S1,..., S k represent the features of remote sensing images in different dimensions of this area.
5. The method for interpretable multi-modal urban area representation according to claim 4, wherein The extraction of parking lot features includes: Construct a parking lot quantity vector L(1,1). The parking lot has only one-dimensional quantity dimension in the regional vector, which reflects the normalized parking record quantity dimension of the area and represents the relative scale of parking activities in each area; Combine the feature vectors of POI, OD inflow, remote sensing image, and parking lot into a feature vector to represent a region: V = (P, T, S, L), where P is the POI feature vector, T is the OD inflow feature vector, S is the remote sensing image feature vector, and L is the parking lot feature vector.
6. The method for interpretable multi-modal urban area representation according to claim 5, wherein The extraction of OD flow interaction features includes: Construct an OD interaction matrix to reflect the functional similarity of each region at different times. The OD interaction matrix tracks the taxi flow between regions at 6-hour intervals; for each trip, the algorithm calculates the pick-up and drop-off grid positions using longitude and latitude, and then assigns the trip to the correct time interval; since the time series of OD interactions is three-dimensional, the matrix is represented as I, where I[i, j, k] represents the number of OD interactions between region i and region j within a certain time k; the OD interaction matrix is constructed as follows: I = (I1, I2, I3,..., I m ), where I1, I2, I3,..., I m is the interaction matrix of OD interactions between regions within a certain time interval in a month.
7. The method for interpretable multi-modal urban area representation according to claim 6, wherein The extraction of distance features includes: The distance matrix is obtained using the distance decay formula: where d i,j is the Euclidean distance between two regions i and j, and β is the decay factor; the distance matrix is concatenated on the OD interaction sequence matrix to form edge features in multiple time scales and spaces as follows: Q = Concatenate(I, D), where I is the taxi OD interaction sequence matrix and D is the inter-region distance matrix.
8. The method for interpretable multi-modal urban area representation according to claim 7, wherein The step of describing the correlation between regions through a graph structure, capturing the dependencies between different dimensions within the regions through a multi-head attention mechanism to obtain the processed node features and edge features, and inputting the processed node features and edge features into a variational graph autoencoder (VGAE) for processing to generate representation vectors specifically includes: The graph structure is represented as G = (V, E, W), where V represents the nodes in the graph, E represents the edges in the graph, and W represents the weights of the edges; the node features are composed of the concatenation of the P, T, S, and L feature vectors, denoted as V = (P, T, S, L); each node contains F = n + m + k + 1 features, where n, m, k, and l represent the dimensionality numbers of the above-mentioned POI, OD inflow, remote sensing image, and parking lot respectively, and the dimensionality of the node feature matrix X is X ∈ R N×F , where N represents the number of nodes in the graph and F represents the node feature dimension; the edge features are composed of the concatenation of the OD flow interaction features and the distance matrix, denoted as E = (I, D); assuming the graph has M edges, each edge contains Y = m + 1 features, and the dimensionality of the edge feature matrix Q e is Q e ∈ R M×Y ; extract the node features and edge features respectively; The multi-head attention mechanism is used to capture the interrelationships between different dimensions within the region, and different attention weights are assigned to enhance the expressive power of the representation. The node feature matrix X is respectively linearly transformed to generate a query vector Q, a key vector K, and a value vector V: Q = XW Q , K = XW K , V = XW V , where W Q , W K , W V are learnable weight matrices. Then, the single-head attention weights are calculated. The similarity between the query Q and the key K is calculated using the dot product and transformed into attention weights through softmax: In the formula, QK represents the feature attention correlation within all regions, is the scaling factor. To capture the interrelationships between different dimensions within the region, multi-head attention is calculated. h independent attention heads are used, and the weight matrix of each head is different. The outputs of all attention heads are concatenated and passed through a linear transformation: MultiHead(Q, K, V) = Concat(head1, head2,..., head h )W O , where W O is the weight of the output linear transformation. Finally, the output of the multi-head attention is connected with the input through a residual connection, and the feature matrix updated by the attention mechanism is obtained through layer normalization: Z = LayerNorm(X + MultiHead(Q, K, V)); Calculate the edge weights reflecting the dynamics of spatial movement to effectively simulate the relationships between regions; these weights are derived from edge features which include OD flow interaction features and distance features; define the edge weights as: where is the original edge feature vector between node i and node j, σ is the sigmoid function, and W E is the trainable weight matrix; the obtained edge weight w i,j combines both long-distance interactions and short-distance adjacency relationships, effectively capturing the comprehensive spatial relationships within the graph; To effectively model the influence of neighboring nodes, a graph attention mechanism is adopted, which ensures that the model captures local and global interactions by focusing on the most relevant nodes among neighboring nodes; the attention coefficient is calculated as: where is the attention score between node i and node j at layer l, capturing the semantic correlation between the features of two regions, a (l)T is the learnable attention vector at layer l, T represents vector transpose, W (l) is the weight matrix at layer l, and represent the node features of node i and node j at layer l respectively, || represents concatenation; the final attention coefficient α i,j combines the learned attention score e i,j and the edge weight w i,j , as follows: $k\in N(i)$ represents the neighbor nodes of $i$. This combination enables the model to take into account both node-level correlations and edge-level spatial dynamics, allowing it to capture more subtle inter-regional interactions; using the calculated attention coefficient $\alpha$ i,j , the node features are updated as follows: represents node $i$ itself and its neighbor nodes; after applying the attention mechanism, the latent representation of the node is obtained: $Z = X$ (l+1) , where $Z\in\mathbb{R}$ N*d , $N$ represents the number of nodes in the graph, $d$ represents the dimension of node features, and $Z$ represents the latent representation matrix; For the variational graph autoencoder VGAE, the output of the feature vector is further mapped to a mean vector and a log standard deviation vector as follows: μ = ZW μ + b μ , logσ 2 = ZW σ + b σ , where W μ and b μ are the weight matrix and bias term used in mean calculation, and W σ and b σ are the weight matrix and bias term used in log standard deviation calculation; μ, logσ 2 are the mean and log standard deviation vectors of the output respectively; To facilitate the calculation of gradients, the reparameterization technique is introduced to sample the 16-dimensional latent variable z to obtain the final representation vector: z = μ + σ ⊙ ∈, where μ represents the mean vector, which is used to determine the central position of the latent variable; σ represents the standard deviation vector, which is used to control the distribution range of the latent variable; ⊙ is the Hadamard product, indicating element-wise multiplication between vectors; ∈ represents a random noise vector that follows the standard normal distribution and is used to introduce randomness to generate the latent variable; Z is the representation vector, that is, the final sampling result combining the mean μ, the standard deviation σ, and the random noise ∈.
9. The method for interpretable multi-modal urban area representation according to claim 8, wherein The specific steps of using a decoder to reconstruct the node features and edge features from the representation vectors include: For the reconstruction of node features, the low-dimensional representation vector z is passed through multiple stacked MLP layers and the relu activation function is passed to construct the original region features: X′ = f MLP (z), where X′ is the reconstructed node feature matrix; For the reconstruction of edge features, the representation vectors z of the regions with interaction relationships i and z j are combined, and the original regional interaction features: E are constructed through multiple stacked MLP layers with the relu activation function i,j′ = f MLP (z i , z j ), where E i,j′ is the reconstructed edge feature matrix; Its loss function consists of three parts: a KL divergence term, a node information reconstruction error term, and an edge information reconstruction error term, aiming to effectively reconstruct the features and correlations of complex multi-scale spatio-temporal urban areas; The total loss is as follows: where X and Q e represent node and edge features on the graph, p(z) is the prior distribution of the latent variable z, is the variational posterior distribution, q(z|X, Q e ) is the probability distribution of the latent variable z given the node features and edge features, p(X|z) and p(Q e |z) are the conditional probability distributions of the node features and edge features given the latent variable; the specific forms are: where σ i represents the potential standard deviation vector of node i, μ i represents the potential mean vector of node i, N represents the total number of nodes on the graph, M represents the total number of edges on the graph, X i is the original feature of node i, X i′ is the feature of node i after reconstruction, E i,j represents the original edge feature between nodes i and j, E i,j′ represents the feature of edge i, j after reconstruction.
10. A method for interpretable multi-modal urban area representation according to claim 9, characterized in that, The specific steps for the usage explanation module to explain the graph substructures and node features with the greatest contribution to each dimension of the representation vector include: Use GNNExplainer to reveal the graph substructures and node features that contribute the most to each dimension of the representation vector z. Specifically, input the pre-trained E-MVGAE model, the graph structure G, and the learned representation vector z into the explanation module; in the multi-objective optimization framework, treat each dimension z ij of the representation vector z as an independent objective; specifically, the objective is to identify the subgraph and the subset of node features which respectively represent that G Si is a subgraph of G, X Si is a subset of the features of X, and they make important contributions to the prediction of z ij ; obtain: where H(z ij ) is the entropy of z ij , and H(z ij |G = G Si , X = X Si ) is the conditional entropy given the subgraph G Si and the feature subset X Si ; Since H(z ij ) is a constant, maximizing the mutual information is equivalent to minimizing the conditional entropy: In this expression, H(z ij |G = G Si , X = X Si ) represents the uncertainty of the given selected subgraph G Si and the node feature X Si , z ij ; minimizing this conditional entropy ensures that the selected subgraph and features are highly predictive for a specific embedding dimension z ij ; P(z ij |G = G Si , X = X Si is the probability of observing z ij given the selected explanatory components, denotes the expectation of the conditional distribution over z ij ; ensuring that the selected subgraph and characteristics capture the most critical patterns and relationships underlying z ij ; To identify the most important dimensions in the learned representation vector z, the mutual information MI(z ij with all possible subgraphs G Si and node features X Si is calculated for each dimension z ij ,(G Si ,X Si )); this process quantifies the contribution of each subgraph and feature subset to explaining a specific embedding dimension; the dimensions are sorted according to the calculated mutual information values, and the dimension with the highest mutual information value is selected as the most important dimension; to achieve this, two key mechanisms are used; first, a binary feature mask F ∈ {0,1} d is applied to determine the key features z ij for each dimension, and the selected features are represented as: X Si = X ⊙ F ij , where ⊙ represents element-wise multiplication; using the soft adjacency mask A S ∈ [0, 1] n×n to select the subgraph G Si , ensuring that only the most influential nodes and edges are retained; then, jointly optimize these mechanisms to maximize the mutual information: