An unsupervised region representation learning method and system for learning city region embeddings

By integrating graph and hypergraph comparative learning into an unsupervised region representation learning method, the limitations of traditional urban region analysis methods in processing multi-source heterogeneous data are overcome, a more accurate and adaptable urban region embedding is achieved, and the efficiency of urban planning and management is improved.

CN119416854BActive Publication Date: 2025-10-17CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411451723.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-17
Publication Date
2025-10-17
Estimated Expiration
2044-10-17

AI Technical Summary

Technical Problem

Traditional urban area analysis methods are difficult to effectively capture the complex relationships and dynamic changes in multi-source heterogeneous urban data, lack universality, and cannot adapt to the characteristics of different cities and regions, resulting in insufficient accuracy and practicality of analysis results.

Method used

An unsupervised region representation learning method is adopted to learn urban region embeddings from multimodal data through integrated graph and hypergraph comparative learning, construct a regional hybrid graph network to capture pairwise and high-order relationships, and optimize model parameters through cross-module comparative learning to generate highly adaptable region embeddings.

Benefits of technology

It improves the accuracy and robustness of urban area analysis, can adapt to different urban environments and task requirements, provides a more comprehensive regional representation, and enhances the model's ability to capture complex interactions and data adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119416854B_ABST
    Figure CN119416854B_ABST
Patent Text Reader

Abstract

The application relates to an unsupervised region representation learning method and system for learning city region embedding, and belongs to the technical field of city computing and intelligent systems. The method specifically comprises the following steps: S1, constructing a region mixed graph network, and utilizing multi-modal data to capture the pair and group relationships between regions; S2, performing graph and hypergraph contrast learning, learning discriminative representations of region nodes and high-order relationships between regions from graph and hypergraph structures through parallel contrast learning modules; S3, cross-module contrast learning, promoting information exchange between graph and hypergraph node representations through a controller component, and enhancing the learning ability of the model; S4, optimizing model parameters, generating more effective region embedding by jointly optimizing graph loss, hypergraph loss and cross-module loss, and supporting various downstream tasks. The technical scheme provided by the application improves the accuracy and efficiency of city region analysis and enhances the ability of the model to capture complex interactions between regions.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of urban computing and intelligent systems, and relates to an unsupervised regional representation learning method and system for learning urban regional embeddings. BACKGROUND

[0002] With the acceleration of urbanization and the development of information technology, the explosive growth of urban data provides unprecedented opportunities and challenges for urban planning, traffic management, public safety and other fields. Urban regional analysis has become increasingly important in urban planning, traffic management, public safety and other fields. Traditional urban regional analysis methods are often limited to evaluating the characteristics and patterns of individual regions independently, ignoring the complex interactions and dynamic relationships between regions, which greatly limits the accuracy and practicality of the analysis results.

[0003] The complex relationships between urban regions are reflected in multiple aspects. First, the geographical proximity between different regions leads to close connections in terms of human flow, material flow, information flow, etc. For example, the human flow attraction of a business center is not only affected by its own business atmosphere and facility perfection, but also closely related to the surrounding traffic convenience, residential population density and other factors. Similarly, the housing prices of residential areas are also influenced by multiple factors such as surrounding educational resources, medical services, green environment, etc.

[0004] Secondly, the functional diversity and spatial heterogeneity of urban regions further increase the difficulty of analysis. Different regions may carry multiple functions such as residence, business, industry, culture and entertainment, each with its own unique operating rules and influencing factors. Therefore, when conducting urban regional analysis, it is necessary to consider factors such as regional function, land use status, population density, economic development level, etc. to ensure the comprehensiveness and accuracy of the analysis results.

[0005] However, traditional urban regional analysis methods have some shortcomings in dealing with these complex problems. On the one hand, the information provided by a single data source is often one-sided and cannot fully reflect the true situation of the region; on the other hand, simple feature engineering and statistical models are difficult to capture the complex relationships and dynamic changes between regions. Therefore, researchers have begun to seek new methods to cope with these challenges.

[0006] In recent years, the rise of deep learning technology has provided new ideas for urban regional analysis. Deep learning models have strong feature extraction and pattern recognition capabilities, and can automatically learn useful information from large-scale data without much human intervention. By applying deep learning technology to the field of urban regional analysis, more complex and accurate models can be built to capture the complex relationships and dynamic changes between regions.

[0007] However, applying deep learning techniques to urban area analysis also faces many challenges. First, how to effectively collect and integrate multi-source heterogeneous urban data is a key problem. Different data sources differ in format, precision, timeliness, etc. Proper data preprocessing and fusion techniques are needed to ensure data consistency and availability. Second, how to design a reasonable model structure and parameters to adapt to the characteristics of different cities and regions is also an important problem. Different cities differ in geographical environment, cultural background, economic development, etc. Model optimization and adjustment need to be made according to specific circumstances.

[0008] In summary, there is an urgent need for a universal method that can capture the complex relationships of urban areas for effective urban area analysis. This method should be able to handle large-scale, multi-modal urban data and automatically learn useful information from the data to guide decision-making in urban planning, traffic management, public safety, and other fields. At the same time, the method needs to be able to adapt to different city and regional characteristics, providing accurate and robust regional embedding results. SUMMARY

[0009] Therefore, the purpose of the present application is to provide an unsupervised regional representation learning method and system for learning urban area embedding, which learns comprehensive regional embedding from multi-modal data by integrating graph and hypergraph contrast learning, can capture pairwise and high-order relationships between regions, and automatically adapt to different urban environments and task requirements. The technical solution provided by the present application can realize a deeper understanding of urban areas and provide more scientific and effective support for urban area analysis.

[0010] To achieve the above purpose, the present application provides the following technical solutions:

[0011] An unsupervised regional representation learning method for learning urban area embedding, the method comprising the following steps:

[0012] S1, constructing a regional mixed graph network model, using multi-modal data to capture pairwise and group relationships between regions;

[0013] S2, performing graph and hypergraph contrast learning, through parallel contrast learning modules, respectively learning discriminative representations of regional nodes and high-order relationships between regions from graph and hypergraph structures;

[0014] S3, cross-module contrast learning, through the controller component to facilitate information exchange between graph and hypergraph node representations, enhancing the learning ability of the model;

[0015] S4, optimizing model parameters, through joint optimization of graph loss, hypergraph loss and cross-module loss, generating more effective regional embedding to support various downstream tasks.

[0016] Further, in step S1, the street view context embedding step and the pair edge construction step are included, and the specific steps are as follows:

[0017] S11, street view context embedding: using street view image data, extracting visual features in the region through a visual encoder, and encoding these features into the initial node embedding of the region; using Inception-V3 architecture and importing pre-trained weights as the encoder of the street view context embedding; for a given region r i a set of street view images in the final visual embedding where f η represents the visual encoder, a represents the number of randomly sampled street view images in r i ; the embedding of all regions is represented as These embeddings capture the physical environment and visual information in the region, providing rich contextual information for subsequent graph learning; Inception-V3 is a deep convolutional neural network developed by Google, which is an improved version of the earlier version of the Inception series. Inception-V3 is widely used in image classification, target detection and other tasks.

[0018] Google Inception Net won first place in the 2014 ImageNet Large Scale Visual Recognition Competition (ILSVRC). The network wins with structural innovation, replacing the fully connected layer with a global average pooling layer, greatly reducing the parameter quantity. This network model is generally referred to as Inception V1. In Inception V2, the Batch Normalization method is introduced, which speeds up the convergence of training. In the Inception V3 model, by splitting the two-dimensional convolutional layer into two one-dimensional convolutional layers, not only the parameter quantity is reduced, but also the overfitting phenomenon is alleviated. Inception-V3 is composed of an initial feature extraction module (Stem) and multiple Inception modules, each module contains parallel 1x1, 3x3, 5x5 convolution and pooling operations to capture features of different scales. The convolution factor decomposition technique decomposes larger convolution kernels into multiple small convolution kernels to reduce the amount of calculation. The network uses global average pooling instead of fully connected layers at the end, and finally outputs the classification result through the softmax layer. The entire architecture design focuses on efficiency and parameter optimization, suitable for large-scale image classification tasks.

[0019] S12, pair edge construction: based on the geographical proximity between regions and human mobility data, construct pair edges in the mixed graph; based on human mobility to construct a matrix representing the normalized number of trips between regions r i and r j , based on the geographical neighbor adjacency matrix A s ∈{0,1} n×n ; in the initial edge matrix A′ f , the most relevant k1 edges (k1=10) of each node are reserved by KNN algorithm, and the updated edge matrix A f is obtained, then the fusion adjacency matrix A=A s o A f is obtained by element-wise logical OR operation, the value of element a ij in matrix A represents whether there is a link edge between regions r i and r j , which represents the direct interaction and spatial relationship between regions, and the graph of the region is obtained

[0020] S13, super-edge construction: to capture the high-order relationship between regions, such as common POI distribution pattern or similar human flow pattern, construct super-edge to combine multiple region nodes together; first, based on human mobility, select the closest k2 (k2=10) regions from the source and destination of each region to form a super-edge; all these super-edge sets constitute the super-edge association matrix based on mobility data Secondly, by calculating the cosine similarity between POI features, and selecting the most similar k3 (k3=10) nodes to form a super-edge, similarly construct the association matrix Then, the final super-edge association matrix H=H f ||H p is generated by connecting operation to merge H f and H p , where m=m1+m2; obtain the supergraph The introduction of super-edge makes the mixed graph capable of representing more complex relationships between regions.

[0021] S14, mixed graph construction: while completing the street view context embedding and edge construction process, combine the edge set and super-edge set obtained in the previous steps to obtain a mixed graph based on region as the input of the framework, the mixed graph is represented as

[0022] Further, in step S2, including graph / supergraph enhancement, graph / supergraph representation and graph / supergraph contrast loss function 3 steps, specifically including:

[0023] S21, graph / hypergraph augmentation: use data augmentation techniques to perturb the graph and hypergraph structure, generate different views to simulate and alleviate the noise problem in graph and hypergraph; for graph structure, by sampling a random mask matrix from Bernoulli distribution, with probability p m randomly mask some dimensions of node feature matrix X, with probability p r randomly delete some edges in G , to generate two augmented graph views and Similarly, for hypergraph structure, with probability p' m randomly mask node features in different dimensions, with probability p' r randomly delete some hyperedges in H , to generate two augmented hypergraph views and Since the nodes in graph and hypergraph essentially represent regions, feature masking is the same for both, that is,

[0024] S22, graph / hypergraph representation: use graph convolutional network GCN and hypergraph neural network HGNN as encoder to learn the representation of nodes and hyperedges from the augmented graph and hypergraph; graph representation learning focuses on updating the embedding of each node by aggregating neighbor node information, which is and input into the graph convolutional network GCN encoder with shared parameters to obtain the node embedding matrix Y1 and Y2 of the two views, and then project the node embedding matrix through a two-layer MLP projection head to a smaller latent space to obtain the projected embedding matrix and Similarly, hypergraph representation learning is performed on the generated two augmented views through the hypergraph neural network (HGNN) encoder; through an iterative manner, the node and hyperedge degree matrix, as well as the trainable weight matrix, are used to update the embedding of nodes and hyperedges; finally, the node embedding matrix P1, P2 and hyperedge embedding matrix Q1, Q2 of the two views are obtained; in order to improve the effect of contrastive learning, a two-layer MLP is used as a projection head to project the node and hyperedge embedding to a smaller latent space, respectively, to obtain the node and hyperedge embedding matrix of the corresponding view and

[0025] S23, graph / hypergraph contrastive loss: design a contrastive loss function to guide the model to learn to distinguish the representations of the same region between different views, and to close the representations of similar regions in different views; for graph contrastive learning, use the InfoNCE loss function in contrastive learning to calculate the contrastive loss at the node level; for any node v i ​The embedding in the first view The corresponding embedding in the second view as an anchor point The node-level contrastive loss is calculated by maximizing the similarity between positive samples (calculated by the cosine similarity function) and minimizing the similarity between the anchor point and negative samples, i.e. Where τ v is a temperature parameter representing the penalty of difficult negative samples, and s(·) is a cosine similarity function used to measure the similarity between two representations; Similarly, for hypergraph contrastive learning, node-level, hyperedge-level and membership-level contrastive losses are designed to capture high-order relationships in hypergraphs and promote the learning of representations with high-order information; After obtaining the projected node and hyperedge embedding matrices, node-level, hyperedge-level and membership-level contrastive losses are calculated respectively; The node-level contrastive loss is calculated by maximizing the similarity of the same node's representation in different views while minimizing the similarity between different nodes, i.e. Where The hyperedge-level contrastive loss is calculated by maximizing the similarity of the same hyperedge's representation in different views, i.e. Where And the membership-level contrastive loss considers the association between nodes and hyperedges to enhance the representation of nodes and hyperedges, i.e. Where Finally, the three levels of contrastive loss are weighted and summed to obtain the total hypergraph contrastive loss

[0026] Further, in step S3, by introducing a cross-module contrastive learning mechanism, the quality and expressiveness of the region embedding are further optimized; Specifically, it includes:

[0027] S31, information fusion: In order to integrate the two sets of region embedding representations obtained in the parallel running graph contrastive learning module and hypergraph contrastive learning module, a weight-based fusion mechanism is adopted to perform weighted 2 fusion on the node embedding representations Z v and Z n from different modules to obtain the fused representation Z f = αZ v +(1-α)Z n , where α ∈ (0, 1) is a weight parameter used to balance the contributions of different modules;

[0028] S32, contrastive loss calculation: Using the fused node embedding representation Z f , the fusion contrastive loss is defined by calculating the contrastive loss between the same node pairs Then the contrastive loss of all nodes is integrated to obtain the overall loss function of cross-module contrastive learning to optimize the final node representation.

[0029] Further, in step S4, a comprehensive regional embedding learning process is implemented by jointly optimizing the loss functions of all contrastive learning modules to obtain the final regional embedding; specifically including:

[0030] The overall loss is defined as Where λ1 and λ2 are parameters measuring the importance of After model training, the final regional representation E is obtained In each training cycle, the graph and hypergraph enhancement operations are performed, first generating two enhanced graphs and hypergraphs from and Then input them into the corresponding shared graph encoder f θ (·) and hypergraph encoder to learn the relevant representation; subsequently, different projection heads (p φ (·), p ψ (·) and ) map these embeddings to the latent space and calculate the contrastive loss of the graph and hypergraph modules respectively; next, the controller component is used to fuse the node representations of the graph and hypergraph modules, thereby generating the cross-module contrastive loss; finally, the model is trained by optimizing the overall objective loss; after training, the well-trained model will output the regional representation E, which can be applied to various downstream tasks.

[0031] The present application also provides an unsupervised regional representation learning system for learning city regional embeddings.

[0032] The present application has the following advantages:

[0033] The present application provides a framework named HyperRegion, which is specially designed for urban computing tasks to learn comprehensive regional embeddings from multi-modal data in an unsupervised manner. The framework can capture pair-wise and high-order relationships between regions by integrating graph and hypergraph contrastive learning, thus providing more comprehensive regional representations. The HyperRegion framework promotes information exchange and collaboration between different modules by performing graph contrastive learning and hypergraph contrastive learning in parallel, and introducing cross-module contrastive learning, which optimizes the model's ability to capture complex interactions between regions, greatly improving the quality and robustness of regional embeddings. The framework not only improves the accuracy of regional representations, but also provides highly adaptive solutions for different urban regions and downstream tasks. In addition, the design of HyperRegion takes into account the noise and incompleteness of data, and through the self-supervised contrastive learning paradigm, it enhances the robustness and generalization ability of the learned representations. Overall, the present application greatly improves the efficiency and effectiveness of urban region analysis, and has important significance for the development of future urban computing and intelligent systems.

[0034] Additional advantages, objects, and features of the application will be set forth in part by the description that follows, and in part will become apparent to those skilled in the art upon examination of same, or can be learned by practice of the application. The objects and other advantages of the application can be realized and attained by the BRIEF DESCRIPTION OF DRAWINGS

[0035] In order to make the purposes, technical solutions and advantages of the present application clearer, the preferred detailed description of the present application will be combined with the drawings as follows, wherein:

[0036] Figure 1 The system block diagram of the present application scheme is shown in the figure.

[0037] Figure 2 The flowchart of the present application scheme is shown in the figure. DETAILED DESCRIPTION

[0038] The technical solutions of the present application will be described in detail below in combination with the drawings.

[0039] The present application proposes an unsupervised regional representation learning method and system for learning city region embeddings, which proposes a framework named HyperRegion. The framework first utilizes multi-modal data and graph neural networks to capture key features within and outside the region, which are considered crucial for determining regional embeddings. Next, HyperRegion constructs a regional mixed graph network by considering human mobility, geographical proximity, point of interest distribution, and visual semantics. In this embodiment, the proposed HyperRegion framework is evaluated in different city environments (e.g., different city regions and multi-modal data sets) and various downstream tasks (such as crime prediction, check-in prediction, and land use classification). The empirical results show that HyperRegion shows significant effectiveness and efficiency in performance, effectively solving the limitations of existing methods in processing complex urban data.

[0040] Figure 1 The system block diagram of the present application scheme, Figure 2 The flowchart of the present application scheme. In the technical scheme of the present application, the 7 concepts involved are as follows:

[0041] The first concept: point of interest (POI): a point of interest is a place in the city with a specific function or significance, such as restaurants, shops, schools, and hospitals, etc. The characteristics of these places are the services or activities they provide, which have an important influence on the daily life of urban residents. In the HyperRegion framework, the distribution and characteristics of POIs are used to enrich the regional embedding, helping to understand the function and attractiveness of the region.

[0042] The second concept: human mobility: human mobility refers to the movement patterns and behaviors of people in the city, including commuting, leisure, and other daily activities. This mobility not only reflects the interaction between city regions, but also affects the city dynamics and regional vitality. In the HyperRegion framework, human mobility data is used to capture the dynamic connections between regions, providing a basis for city traffic planning and regional development analysis.

[0043] The third concept: geographical neighbor: geographical neighbor describes the spatial proximity relationship between regions, which is based on the proximity of geographical location. The concept of geographical neighbor is very important when constructing the graph structure, as it directly affects the way of connecting the edges between regions. In the HyperRegion framework, the geographical neighbor relationship is represented by pairwise edges, which helps to simulate the local spatial relationship and interaction between regions.

[0044] Concept 4: Street View Images: Street view images provide visual information of urban areas, including physical environmental features such as buildings, roads, public spaces, etc. These image data can be used to extract visual features of the area, such as architectural style, greenery level, and pedestrian density, etc. In the HyperRegion framework, street view image data can be used to initialize the visual embedding of the area nodes or further enrich the multi-modal feature representation of the area through a visual encoder.

[0045] Concept 5: Area Embedding: It involves converting areas in the city into low-dimensional vector space. These embedding vectors can preserve the geographical and functional information of the areas, as well as the relationships between areas, so that complex urban data can be effectively processed by machine learning algorithms. In area embedding, the spatial and semantic relationships between areas are usually maintained, for example, if two areas are geographically adjacent or functionally similar, their vector representations should also be close to each other. In the HyperRegion framework, these embedding vectors are used to determine the characteristics of the urban areas and provide support for urban analysis and prediction tasks. By learning these embeddings, the system can better understand and predict the role of each area in the city's functions and activities, thus making more accurate area function identification and prediction.

[0046] Concept 6: Graph Contrastive Learning: Graph contrastive learning is a self-supervised learning method that learns discriminative embeddings by maximizing the consistency between graph representations under different views. In the HyperRegion framework, the graph contrastive learning module generates different graph views by applying graph augmentation strategies and designs a contrastive loss function to guide the model learning. This method can improve the robustness and discriminability of the learned area embeddings, especially in the presence of data noise and incompleteness.

[0047] Concept 7: Hypergraph Contrastive Learning: Hypergraph contrastive learning is an extension of graph contrastive learning, which naturally models high-order relationships between multiple entities by introducing hyperedges. In the HyperRegion framework, the hypergraph contrastive learning module not only considers the pairwise relationships between areas, but also captures the group relationships between areas, thus providing a more comprehensive description of area interactions. This method helps to learn area representations containing high-order information, enhancing the model's understanding of complex urban phenomena.

[0048] As shown in Figure 1 and Figure 2 , the system framework of the present application mainly includes four modules: mixed graph construction module, graph contrastive learning module, hypergraph contrastive learning module, cross-module contrastive module, Figure 1 As shown in the system block diagram of the present application, the mixed graph construction module: this module aims to integrate multi-modal data sources and construct a regional mixed graph network, which includes the following steps:

[0049] Step 1: Utilize street view image data to extract visual features X within the region through a visual encoder, and encode these features as initial node embeddings for the region. These embeddings capture the physical environment and visual information within the region, providing rich contextual information for subsequent graph learning.

[0050] Step 2: Based on the geographical proximity between regions and human mobility data, construct pairwise edges in the mixed graph. In the initial edge matrix A' constructed based on human mobility f , the most relevant k1 edges for each node are retained through the KNN algorithm, resulting in an updated edge matrix A f . Subsequently, the adjacency construction matrix A s is obtained through element-wise logical OR operation, and these edges represent direct interactions and spatial relationships between regions, resulting in the region graph G

[0051] Step 3: To capture high-order relationships between regions, such as common POI distribution patterns or similar human mobility patterns, construct hyperedges to combine multiple region nodes. First, based on human mobility, select the closest k2 regions to each region from both source and destination perspectives to form hyperedges, constructing a hyperedge association matrix H f based on mobility data. Second, calculate the cosine similarity between POI features and select the most similar k3 nodes to form hyperedges, similarly constructing the association matrix H p . Then, combine H f and H p through connection operations to generate the final hyperedge association matrix H, resulting in the hypergraph H The introduction of hyperedges enables the mixed graph to represent more complex relationships between regions.

[0052] Step 4: Through the graph and hypergraph structures, integrate intra-regional attributes and inter-regional relationships to construct a region mixed graph containing graph and hypergraph structures Each node v represents a region, and the connections between regions are represented by the connection matrices A, H of edges and hyperedges, providing input for subsequent contrastive learning.

[0053] Graph contrastive learning module: The purpose of this module is to learn discriminative representations of region nodes through graph structures, including the following steps:

[0054] Step 1: Apply graph augmentation strategies such as node feature masking and edge removal to generate two different graph views and

[0055] Step 2: Use a shared graph neural network-based encoder f θ(·) Learn node representations from these two augmented views and map them to a low-dimensional latent space using projection heads p φ (·) Map node representations to a low-dimensional latent space.

[0056] Step Three: Compute contrastive loss By optimizing this loss function, the model is encouraged to learn robust and discriminative node representations that can distinguish different views.

[0057] Hypergraph Contrastive Learning Module: This module aims to learn high-order relationships between regions through hypergraph structures, including the following steps:

[0058] Step One: Apply hypergraph augmentation strategies such as node feature masking and hyperedge removal to generate two different hypergraph views and

[0059] Step Two: Use a shared hypergraph neural network-based encoder to form node and hyperedge representations and map them to a low-dimensional latent space using projection heads p ψ (·) and .

[0060] Step Three: Design a three-level contrastive loss mechanism based on the InfoNCE loss function in contrastive learning, including node-level, hyperedge-level, and membership contrastive loss. Finally, the three levels of contrastive loss are weighted and summed to obtain the total hypergraph contrastive loss By optimizing these loss functions, the model is encouraged to learn representations with high-order information from hypergraph structures.

[0061] Cross-Module Contrastive Module: To enhance information exchange and complementarity between different modules, this module fuses the node representations learned by the graph and hypergraph through a controller component, including the following steps:

[0062] Step One: Use a weighted fusion mechanism to fuse the node representations learned by the graph and hypergraph to form a fused representation Z f .

[0063] Step Two: Compute the fused contrastive loss This loss considers the representations of the same node in different modules, encouraging the model to integrate common and complementary information.

[0064] Step Three: Jointly optimize all contrastive losses, including graph contrastive loss hypergraph contrastive loss and cross-module contrastive loss to generate more effective regional embeddings.

[0065] To sum up, the application proposes a framework named HyperRegion, which firstly integrates multi-modal data using a region feature fusion module, then learns local and global features of the region through a graph contrast learning module and a hypergraph contrast learning module. Finally, the cross-modal contrast module is used to enhance the information exchange and complementarity of the representation between different modules. Extensive experiments on different city computing tasks (such as check-in prediction and crime prediction) show that HyperRegion has significant effectiveness and superiority in performance, effectively solving the limitations of existing methods in regional representation learning.

[0066] Finally, it should be pointed out that the above embodiments are only used to illustrate the technical solutions of the application and are not limiting. Although the application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the application can be modified without departing from the spirit and scope of the technical solutions, and all should be covered in the scope of the claims of the application.

Claims

1. An unsupervised region representation learning method for learning urban region embeddings, characterized by: The method comprises the following steps: S1. Build a regional hybrid graph network model to capture pairwise and group relationships between regions using multimodal data. S2, perform graph and hypergraph contrastive learning, through parallel contrastive learning modules, learn the discriminative representation of regional nodes and the high-order relations between regions from the graph and hypergraph structures respectively; S3, cross-module comparative learning, promotes information exchange between graph and hypergraph node representations through the controller component, enhancing the learning ability of the model; S4. Optimize model parameters to generate more effective region embeddings by jointly optimizing graph loss, hypergraph loss, and cross-module loss to support various downstream tasks; the overall loss is defined as Where λ1 and λ2 are the measures of The parameters of importance; the final region representation is obtained after model training; In step S1, including the street view context embedding step and the paired edge construction step, the specific steps are as follows: S11. Street view context embedding: Using street view image data, the visual encoder extracts the visual features in the region and encodes these features into the initial node embedding of the region; the Inception-V3 architecture is used, and the pre-trained weights are imported as the encoder of the street view context embedding; for a given region r i A set of street view images within The final visual embedding of where f η represents the visual encoder, a represents r i The number of randomly sampled street view images in ; the embedding of all regions is represented as S12. Paired Edge Construction: Construct paired edges in a hybrid graph based on geographic proximity between regions and human mobility data; construct a matrix based on human mobility. Represents region r i and r j The normalized number of trips between them is used to construct the matrix A based on the adjacency of geographical neighbors. s ∈{0,1} n×n ; In the initial edge matrix A' f In the example, the KNN algorithm is used to retain the most relevant k1 edges of each node, and the updated edge matrix A is obtained. f , then, the fused adjacency matrix is ​​obtained by element-by-element logical OR operation Element a in matrix A ij The value represents the area r i and r j Are there links between them? These edges represent the direct interactions and spatial relationships between regions, and the regional graph is obtained. S13. Hyperedge construction: To capture high-order relationships between regions, we construct hyperedges that group multiple regional nodes together. First, based on human mobility, we select the k2 regions closest to each region from the perspectives of the region’s source and destination to form hyperedges. All these hyperedge sets constitute a hyperedge association matrix based on mobility data. Secondly, by calculating the cosine similarity between POI features and selecting the most similar k3 nodes to form a hyperedge, the association matrix is ​​constructed similarly Then merge H by the join operation f and H p Generate the final hyperedge correlation matrix H = H f ||H p ,in m=m1+m2; get hypergraph S14, hybrid graph construction: While completing the street view context embedding and edge construction process, combine the edge sets obtained in the previous steps and hyperedge sets Get based on region The mixed graph of is taken as the input of the framework, and the mixed graph is represented as In step S2, there are three steps: graph / hypergraph enhancement, graph / hypergraph representation, and graph / hypergraph contrast loss function, specifically including: S21. Graph / Hypergraph Enhancement: Use data augmentation techniques to perturb graph and hypergraph structures and generate different views to simulate and alleviate the noise problem in graphs and hypergraphs. For graph structures, a random masking matrix is ​​sampled from the Bernoulli distribution with probability p. m Randomly shield some dimensions of the node feature matrix X with probability p r Randomly delete graph Some edges in the original graph Generate enhanced two-graph view and Similarly, for the hypergraph structure, with probability p' m Randomly mask node features in different dimensions with probability p' r Random deletion For some hyperedges in , we get two enhanced hypergraph views and Since nodes in graphs and hypergraphs essentially represent regions, feature masking is the same for both, i.e. S22. Graph / Hypergraph Representation: Using graph convolutional networks (GCNs) and hypergraph neural networks (HGNNs) as encoders, we learn the representations of nodes and hyperedges from enhanced graphs and hypergraphs. Graph representation learning focuses on updating the embedding of each node by aggregating neighboring node information. and Input into the shared parameter graph convolutional network GCN encoder to obtain the node embedding matrices Y1 and Y2 of the two views, and then project the node embedding matrix into a smaller potential space through a two-layer MLP projection head to obtain the projected embedding matrix and Similarly, hypergraph representation learning is performed on the two generated enhanced views by using a hypergraph neural network encoder to learn node and hyperedge representations. The embeddings of nodes and hyperedges are updated iteratively using the degree matrices of nodes and hyperedges and a trainable weight matrix. Finally, the node embedding matrices P1 and P2 and the hyperedge embedding matrices Q1 and Q2 of the two views are obtained. Two layers of MLP are used as projection heads to project the embeddings of nodes and hyperedges into smaller latent spaces, respectively, to obtain the embedding matrices of nodes and hyperedges of the corresponding views. and S23. Graph / hypergraph contrast loss: Design a contrast loss function to guide the model to learn to distinguish the representation of the same region between different views and to bring the representation of similar regions in different views closer. For graph contrast learning, use the InfoNCE loss function in contrast learning to calculate the node-level contrast loss. For any node v i , embed it in the first view As an anchor, the corresponding embedding in the second view As a positive sample, other nodes are used as negative samples. The node-level contrast loss is calculated by maximizing the similarity between positive samples and minimizing the similarity between anchor points and negative samples. Right now in τ v represents the temperature parameter for the hard negative sample penalty, s(·) is the cosine similarity function used to measure the similarity between two representations, represents the number of nodes; similarly, for hypergraph contrastive learning, node-level, hyperedge-level, and membership-level contrastive losses are designed to capture high-order relationships in the hypergraph and promote learning representations with high-order information; after obtaining the projected node and hyperedge embedding matrices, the node-level, hyperedge-level, and membership-level contrastive losses are calculated respectively; the node-level contrastive loss is calculated by maximizing the number of nodes with the same v i Representation in different views and The similarity between different nodes is calculated by minimizing the similarity between them, that is, in Hyperedge-level contrastive loss is achieved by maximizing the same hyperedge e i Representation in different views and The similarity is calculated, that is, in τ e is the temperature parameter, represents the number of hyperedges; while the contrast loss at the affiliation level considers the association between nodes and hyperedges and enhances the representation of nodes and hyperedges, i.e. in τ m is the temperature parameter, Representative node v i The number of hyperedges it belongs to, is the discriminator implemented by a bilinear network; finally, the weighted summation of the three levels of contrast loss is used to obtain the total hypergraph contrast loss In step S3, the quality and expressiveness of region embeddings are further optimized by introducing a cross-module comparative learning mechanism. Specifically, S31, Information Fusion: Using a weight-based fusion mechanism, nodes from different modules are embedded into the representation Z v and Z n Perform weighted fusion to obtain the fusion representation Z f =αZ v +(1-α)Z n , that is, respectively and Where α∈(0,1) is a weight parameter used to balance the contribution of different modules; S32, contrast loss calculation: introduce fusion contrast loss, use the fused node embedding representation Z f ,Will and The same node pairs in the region r are regarded as positive pairs, otherwise as negative pairs; i As an anchor node, the fusion contrast loss is defined by calculating the contrast loss between the same node pairs Then the contrast losses of all nodes are combined to obtain the overall loss function of cross-module contrast learning where τ c For l c The corresponding temperature parameters are used to optimize the final node representation.

2. The unsupervised region representation learning method for learning urban region embedding according to claim 1, characterized in that: In step S4, a comprehensive region embedding learning process is implemented to obtain the final region embedding by jointly optimizing the loss functions of all contrastive learning modules; specifically, it includes: The total loss is defined as Where λ1 and λ2 are the measures of Parameters of importance; the final region representation is obtained after model training In each training cycle, the graph and hypergraph enhancement operations are performed, first from and Generate two enhanced graphs and hypergraphs, which are then fed into the corresponding shared graph encoders f θ (·) and hypergraph encoder to learn related representations; subsequently, different projection heads (p φ (·), p ψ (·)and ) maps these embeddings into the latent space and calculates the contrastive loss of the graph and hypergraph modules separately; next, a controller component is used to fuse the node representations of the graph and hypergraph modules, resulting in a cross-module contrastive loss; finally, the model is trained by optimizing the overall objective loss; after training, the well-trained model will output the region representation E, which can be applied to various downstream tasks.

3. An unsupervised region representation learning system for learning urban region embeddings, characterized by: The system adopts the method according to any one of claims 1 to 2.