Structure perception multi-view city representation learning method and system with coordination fusion and alignment

By employing a structure-aware multi-view city representation learning method, we have solved the problems of misleading positive and negative sample selection across views and conflicting optimization objectives in existing technologies, generating higher-quality region embeddings and improving the stability and generalization performance of the model.

CN121564630APending Publication Date: 2026-02-24FUZHOU UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511723302.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-21
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing urban embedding methods lack structure awareness, leading to misleading selection of positive and negative samples across views, failing to effectively capture potential semantic associations, and conflicting optimization objectives between fusion and contrastive learning, resulting in unstable gradient updates and suboptimal embedding structures.

Method used

A structure-aware multi-view city representation learning method is adopted. The sparse autoencoder module extracts specific view representations to enhance the interaction between views. The cross-view transformer module optimizes the fusion process and enhances consistency through the structure-aware multi-view contrast learning module. The soft Lagrange constraint strategy is combined to coordinate gradient conflicts.

Benefits of technology

It improves the discriminative power and semantic consistency of region embeddings, alleviates gradient conflicts, enhances the stability and generalization performance of the model, and generates higher-quality consensus representations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121564630A_ABST
    Figure CN121564630A_ABST
Patent Text Reader

Abstract

The invention provides a structure perception multi-view city representation learning method and system with coordinated fusion and alignment, and a structure perception multi-view city representation learning model with coordinated fusion and alignment is established on the basis of target city POI and taxi travel data. The method further comprises the steps that multi-view data are constructed, original data are reconstructed through a sparse auto-encoder module, and specific representation of each view is extracted; executing an enhanced specific view representation module to enhance the expression ability of the specific view representation by modeling beneficial inter-view interaction; a cross-view converter module is adopted, and a multi-view fusion process is optimized by means of cross-view converter module region similarity; executing a structure perception multi-view contrast learning module, and enhancing the consistency between the consensus representation and the specific view representation; and optimizing model parameters, executing a soft Lagrange constraint-based training strategy, and solving the problem of a suboptimal solution caused by gradient conflicts in joint learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention proposes a structure-aware multi-view city representation learning method and system with coordinated fusion and alignment, which relates to the field of spatial information technology. Background Technology

[0002] Embedding methods in current urban applications can be broadly categorized into two types: single-view methods and multi-view methods. Single-view methods typically rely on a single type of urban data to construct a regional representation, such as human trajectory data, point of interest (POI) data, or remote sensing imagery data. These methods are computationally efficient but neglect the complementarity of different data sources, making it difficult to fully represent regional semantics.

[0003] The shortcomings of existing technologies are mainly reflected in the lack of structure awareness in contrastive learning: In existing unified integration frameworks, when contrastive learning constructs cross-view positive and negative sample pairs, positive samples are different view representations of the same area, while negative samples are randomly selected from other areas. However, in real urban environments, areas in different spatial locations may have similar functional structures or semantic roles. For example, commercial areas with similar POI composition, population flow patterns, or socioeconomic characteristics may play similar roles in the urban system. Treating such semantically similar areas as negative samples may generate misleading supervisory signals, hindering the model from recognizing shared features of geographically distant but functionally similar areas. Therefore, contrastive learning that ignores structural similarity often fails to capture potential semantic associations, reducing the compactness and discriminativeness of the representations.

[0004] The shortcomings of existing technologies also lie in the optimization conflict between fusion and alignment: Feature fusion typically enhances semantic consistency by minimizing the differences in shared semantic representations between different views, while contrastive learning improves discriminative power and generalization ability by maximizing the semantic differences between city representations. These two objectives inherently conflict in their optimization directions. Without proper coordination, the fusion module may cause the representations of all views to converge excessively, while the contrastive module may push the representations of structurally different regions further apart. This situation leads to gradient update conflicts, resulting in unstable convergence and suboptimal embedding structures. Therefore, effectively coordinating these conflicting gradients within a unified fusion and alignment framework is the core challenge in obtaining robust and highly generalizable region embeddings. Summary of the Invention

[0005] In view of this, to fill the gaps and deficiencies in existing technologies, this invention proposes a structure-aware multi-view city representation learning method and system with coordinated fusion and alignment. This invention simultaneously enhances structure awareness capabilities and resolves the optimization conflict between the fusion target and the comparison target in multi-view representation learning. To overcome the problem of missing structure awareness, this invention incorporates a structure-guided comparison module, integrating functional similarity into the negative sampling process to avoid treating semantically similar regions as negative samples, thereby improving semantic consistency and cross-view alignment. To alleviate the optimization conflict between the fusion module and the comparison module, this method employs a collaborative training strategy to coordinate gradient conflicts, ensuring more stable convergence and obtaining higher-quality embeddings.

[0006] This invention proposes a structure-aware multi-view city representation learning method and system with coordinated fusion and alignment, including the following:

[0007] This invention proposes a structure-aware multi-view city representation learning method with coordinated fusion and alignment. The method is characterized by establishing a structure-aware multi-view representation learning model with coordinated fusion and alignment based on target city POI and taxi travel data. This method includes the following:

[0008] Step S1: Construct multi-view data;

[0009] Step S2: Use a sparse autoencoder module to reconstruct the original data and extract a specific representation for each view;

[0010] Step S3: Execute the Enhance Specific View Representation module to enhance the expressiveness of specific view representations by modeling beneficial inter-view interactions;

[0011] Step S4: Employ a cross-view transformer module, including optimizing the multi-view fusion process by leveraging the region similarity of the cross-view transformer module to generate an initial consensus representation;

[0012] Step S5: Execute the structure-aware multi-view contrastive learning module to enhance the consistency between the consensus representation and the specific view representation, and generate the final consensus representation;

[0013] Step S6: Optimize model parameters and implement a training strategy based on soft Lagrangian constraints to solve the suboptimal solution problem caused by gradient conflict in joint learning.

[0014] Further, step S1 includes the following:

[0015] Step S11: Use POI to represent the semantic features of the region, and define the semantic features of the region as P, including the following:

[0016] P = {p1, p2, p3, ..., p} n},p i ∈R c

[0017] Where C is the number of POI categories, p i p is the semantic feature representation of region i. i Each dimension is represented by the number of POIs of a specific category within the region, and the dimension size is C.

[0018] POI stands for Point of Interest.

[0019] Step S12: Define regional interaction characteristics, including defining taxi travel in the target city as an interaction behavior between regions;

[0020] The regional interaction features are divided into outflow features and inflow features. The outflow feature is defined as S, which includes the following:

[0021] S = {s1, s2, s3, ..., s} n},s i ∈R n

[0022] Where n is the number of regions, s i The number of taxi trips from region i to other regions within a specific time period, s i Each dimension represents the number of times a user travels from region i to different regions.

[0023] Specifically, defining taxi travel in the target city as an inter-regional interaction includes: treating people taking taxis as an inter-regional interaction.

[0024] The inflow characteristic is defined as D, which includes the following:

[0025] D = (d1, d2, d3, ..., d n ),d i ∈R n

[0026] Where d i Each dimension represents the number of times region i is reached from different regions, R n Represents each inflow feature vector d i It is an n-dimensional vector.

[0027] Further, step S2 includes the following:

[0028] Step S21: Reconstruct the original data using a sparse autoencoder module and extract a specific representation for each view, including the following:

[0029] Learn and extract representative embedding representations from the original view;

[0030] Introducing reconstruction loss L r : Utilizing a reconstruction loss L that combines reconstruction error and sparsity regularization r This is used to improve the model's generalization performance and its ability to capture regional features.

[0031] Further, step S3 includes the following:

[0032] Step S31: After obtaining the initial view-specific embedding representation in the sparse autoencoder module, the view-specific features of each view are enhanced by integrating complementary information from other views: the features of the current view and other views are concatenated along the feature dimension to form a joint representation H; then, the joint representation is transformed by a nonlinear function and multiplied element-wise with the original view-specific features to generate the enhanced representation Z. v :

[0033] Z v =H v ⊙M(H)

[0034] Where ⊙ represents element-wise multiplication, H v Let H be a specific view representation of the v-th view; M() is a nonlinear transformation function that projects the joint representation H onto the view with respect to H. v Spaces with the same dimensions ensure consistency in the feature space.

[0035] Further, step S4 includes the following:

[0036] Step S41: Use the learnable matrix W Q W K and W V Embed the features of different views into Z v Project them onto the query, key, and value spaces respectively; then establish pairwise relationships between regions through an attention mechanism, including the following:

[0037] Z = concat(Z) 1 Z 2 ,…,Z v )

[0038] Q = Z × W Q

[0039] K = Z × W K

[0040] V = Z × W V

[0041]

[0042] Where Q represents the query space, K represents the key space, and V represents the value space;

[0043] Here, softmax() represents the normalization exponential function; concat() represents the function that concatenates two or more arrays or strings;

[0044] Where O represents the structure-aware similarity matrix, O ij Representing region r i With r j The structural relationship between them; V j Representing region r j The corresponding value vector; scaling factor d k The dimension is consistent with the key vector to stabilize attention calculations and reduce redundancy caused by multi-view stitching; a fully connected layer is used to... Further optimization yields the initial consensus representation.

[0045] The fully connected layer belongs to the cross-view fusion transformer module. For region r i The intermediate fusion representation aggregated through the cross-view attention mechanism, and the consensus representation. It is the final output of the cross-view fusion transformer module;

[0046] Step S42: Introduce a consistency loss function to ensure that the consensus representation can effectively capture the shared semantic information among the views, where the difference between the consensus representation and the specific representations of each view is minimized; consistency loss L con The definition includes the following:

[0047]

[0048] Where D(·,·) represents the similarity measure between two representations; by minimizing the consistency loss L con This allows for the preservation of key view-specific features while promoting semantic alignment between views.

[0049] Further, step S5 includes the following:

[0050] Step S51: Introduce the structure-aware multi-view contrast learning module. The structure-aware multi-view contrast learning module uses the cross-view fusion transformer module to calculate the structural similarity matrix O. ij And quantize the region r i With r j Structural similarity between them;

[0051] Step S52: Based on the similarity matrix O ij Design a dynamic weighting mechanism, where when O ij When it increases, exp(-αOij The term ) effectively suppresses or eliminates the influence of these regions as negative samples, avoiding the generation of misleading supervisory signals; conversely, when O ij When the value is small, the region is considered a reliable negative sample;

[0052] Step S53: Introduce a dynamic temperature scaling strategy to adjust temperature parameters based on structural similarity. This suppresses the influence of similar regions during similarity calculation, reduces interference with contrastive learning, promotes more discriminative embedding learning, and yields the final consensus representation. Finally, the inter-view contrast loss of the v-th view is obtained. The definition includes the following:

[0053]

[0054] τ ij =τ(1+βO) ij )

[0055] in, For region r i The consensus expressed; and They are regions r i and r j Enhanced view-specific representation under the v-th view; τ is the reference temperature parameter; τ ij For region r i and r j The adaptive temperature parameters between them; α and β are the dynamic weighting intensity and temperature adjustment intensity, respectively; D'(·,·) is the embedding similarity quantization function under temperature control;

[0056] Step S54: The total contrast loss of all views is determined by analyzing each view. Summing yields:

[0057]

[0058] in, L represents the v-th view. c This represents the total contrast loss across all views.

[0059] Further, step S6 includes the following:

[0060] Step S61: The training strategy based on soft Lagrangian constraints includes a multi-stage collaborative optimization strategy, which includes the following:

[0061] Each training iteration is decomposed into a series of subtasks: first, the sparse autoencoder module is optimized, then the augmented specific view representation and cross-view transformer modules are jointly optimized, and finally the structure-aware multi-view contrast learning module is optimized. This order strictly follows the internal information flow of the framework, with the output of the previous stage serving as the input of the next stage, reflecting the dependency between the stages. In each stage, only the parameters of that stage are updated, while all other parameters are frozen.

[0062] Step S62: The training strategy based on soft Lagrange constraints includes establishing a soft Lagrange constraint mechanism, which includes the following:

[0063] A soft consistency regularization term is introduced between adjacent optimization stages to construct a cross-stage collaborative objective; the soft consistency regularization mechanism simulates the Lagrange multiplier in a soft form; the final loss is defined as:

[0064] L con-final =L con +λ1·R(Z,H)

[0065]

[0066] Where R(·,·) represents the use of cosine similarity to measure the similarity between adjacent stage embeddings; hyperparameters λ1 and λ2 control the strength of soft constraints and adjust the degree of cooperation between stages; the soft Lagrangian constraint mechanism adopts a delayed activation strategy, activating the soft constraints only after a preset number of iterations;

[0067] Where L con-final For the corrected eventual consistency loss; L c-final The final loss;

[0068] Consensus representation and These are the final outputs of the cross-view fusion transformer module and the structure-aware multi-view contrastive learning module, respectively. The cross-view fusion transformer module fuses complementary information from all views and, through inter-view attention mechanisms and contrastive learning, forms a shared feature representation that is effective for downstream tasks.

[0069] Step S63: The training strategy based on soft Lagrangian constraints includes establishing a state-aware constraint propagation mechanism, wherein the state-aware constraint propagation mechanism dynamically adjusts the strength of the soft constraints, including the following:

[0070]

[0071] in, and Let ω and ω represent the second norms of the functional gradients of the reconstruction loss and the consistency loss, respectively. The hyperparameters controlling the strength and sensitivity of the baseline constraints are: when the gradient increases in the previous stage, the constraint strength λ will decrease, providing learning space for the model and avoiding error propagation; the training strategy based on soft Lagrangian constraints achieves the robustness of the model to noise and optimization variance through a state-aware mechanism.

[0072] According to a second aspect of the present invention, a structure-aware multi-view city representation learning system with coordinated fusion and alignment is proposed, characterized in that it is implemented by a structure-aware multi-view city representation learning method with coordinated fusion and alignment as described in any one of the present invention, wherein the structure-aware multi-view city representation learning method with coordinated fusion and alignment includes:

[0073] The sparse autoencoder module is used to reconstruct the original data and extract a specific representation for each view;

[0074] Enhance the expressive power of specific view representation modules by modeling beneficial inter-view interactions;

[0075] The cross-view transformer module is used to optimize the multi-view fusion process by leveraging regional similarity to generate an initial consensus representation that has been robustly enhanced.

[0076] The structure-aware multi-view contrastive learning module is used to enhance the consistency between the consensus representation and the specific view representation, and to generate the final consensus representation;

[0077] A training strategy module based on soft Lagrangian constraints is used to solve the suboptimal solution problem caused by gradient conflict in joint learning.

[0078] According to a third aspect of the present invention, the present invention proposes a structure-aware multi-view city representation learning system with coordinated fusion and alignment, comprising an electronic device, wherein the electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, when the processor executes the computer program, it implements a structure-aware multi-view city representation learning method with coordinated fusion and alignment as described in any one of the present invention.

[0079] According to a fourth aspect of the present invention, the present invention proposes a structure-aware multi-view city representation learning system with coordinated fusion and alignment, comprising a computer-readable storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, it implements a structure-aware multi-view city representation learning method with coordinated fusion and alignment as described in any one of the present invention.

[0080] The present invention has the following advantages:

[0081] (1) This invention proposes a contrastive learning mechanism that explicitly incorporates the potential structural relationships between regions into the negative sample selection process. By avoiding semantically misleading supervision signals, it enhances the compactness of structurally similar region representations and improves the discriminative ability of the embedding process; (2) This invention designs a new optimization strategy that combines module training with soft Lagrange regularization. This strategy effectively coordinates the fusion and contrastive learning objectives, alleviates gradient conflicts, and improves the stability and efficiency of joint optimization; (3) The transformer-based fusion module designed in this invention extracts consensus representations by making full use of the complementary information of similar regions, effectively reducing inconsistencies and redundancy between different views. Attached Figure Description

[0082] Figure 1 This is a schematic diagram of the steps of the present invention.

[0083] Figure 2 This is a schematic diagram of the technical framework of the structure-aware multi-view city representation learning method with coordinated fusion and alignment of the present invention.

[0084] Figure 3 This is a schematic diagram comparing the ablation test performance results of the present invention. Detailed Implementation

[0085] The technical solution of the present invention will now be described in detail with reference to the accompanying drawings.

[0086] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0087] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments of the present invention; as used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise; furthermore, it should be understood that when the terms “comprising” and / or “including” are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0088] like Figures 1 to 3 As shown, this invention proposes a structure-aware multi-view city representation learning method and system with coordinated fusion and alignment, characterized by the following:

[0089] This invention proposes a structure-aware multi-view city representation learning method with coordinated fusion and alignment. The method is characterized by establishing a structure-aware multi-view representation learning model with coordinated fusion and alignment based on target city POI and taxi travel data. This method includes the following:

[0090] Step S1: Construct multi-view data;

[0091] Step S2: Use a sparse autoencoder module to reconstruct the original data and extract a specific representation for each view;

[0092] Step S3: Execute the Enhance Specific View Representation module to enhance the expressiveness of specific view representations by modeling beneficial inter-view interactions;

[0093] Step S4: Employ a cross-view transformer module, including optimizing the multi-view fusion process by leveraging the region similarity of the cross-view transformer module to generate an initial consensus representation;

[0094] Step S5: Execute the structure-aware multi-view contrastive learning module to enhance the consistency between the consensus representation and the specific view representation, and generate the final consensus representation;

[0095] Step S6: Optimize model parameters and implement a training strategy based on soft Lagrangian constraints to solve the suboptimal solution problem caused by gradient conflict in joint learning.

[0096] In this invention, the sparse autoencoder module is represented by VR, the augmented specific view representation module is represented by EVSR, the cross-view transformer module is represented by CVTF, and the structure-aware multi-view contrast learning module is represented by MVSCL.

[0097] Further, step S1 includes the following:

[0098] Step S11: Use POI to represent the semantic features of the region, and define the semantic features of the region as P, including the following:

[0099] P = {p1, p2, p3, ..., p} n},p i ∈R c

[0100] Where C is the number of POI categories, p i p is the semantic feature representation of region i. i Each dimension is represented by the number of POIs of a specific category within the region, and the dimension size is C;

[0101] Where POI represents Point of Interest;

[0102] Step S12: Define regional interaction characteristics, including defining taxi travel in the target city as an interaction behavior between regions;

[0103] The regional interaction features are divided into outflow features and inflow features. The outflow feature is defined as S, which includes the following:

[0104] S = {s1, s2, s3, ..., s} n},S i ∈R n

[0105] Where n is the number of regions, s i The number of taxi trips from region i to other regions within a specific time period, s i Each dimension represents the number of times a user travels from region i to different regions.

[0106] The inflow characteristic is defined as D, which includes the following:

[0107] D = (d1, d2, d3, ..., d n ),d i ∈R n

[0108] Where d i Each dimension represents the number of times region i is reached from different regions, R n Represents each inflow feature vector d i It is an n-dimensional vector.

[0109] Further, step S2 includes the following:

[0110] Step S21: Reconstruct the original data using a sparse autoencoder module and extract a specific representation for each view, including the following:

[0111] Learn and extract representative embedding representations from the original view;

[0112] Introducing reconstruction loss L r : Utilizing a reconstruction loss L that combines reconstruction error and sparsity regularization r This is used to improve the model's generalization performance and its ability to capture regional features.

[0113] Further, step S3 includes the following:

[0114] Step S31: After obtaining the initial view-specific embedding representation in the sparse autoencoder module, the view-specific features of each view are enhanced by integrating complementary information from other views: the features of the current view and other views are concatenated along the feature dimension to form a joint representation H; then, the joint representation is transformed by a nonlinear function and multiplied element-wise with the original view-specific features to generate the enhanced representation Z. v :

[0115] Z v =H v ⊙M(H)

[0116] Where ⊙ represents element-wise multiplication, H v Let H be a specific view representation of the v-th view; M() is a nonlinear transformation function that projects the joint representation H onto the view with respect to H. v Spaces with the same dimensions ensure consistency in the feature space.

[0117] Further, step S4 includes the following:

[0118] In one embodiment of the present invention, a multi-view structure-aware contrastive learning module is proposed, which optimizes the negative sampling strategy by combining the structural relationships between regions. Its core idea is to avoid treating structurally similar regions as negative samples, thereby eliminating misleading supervisory signals.

[0119] Step S41: Use the learnable matrix W Q W k and W V Embed the features of different views into Z v Project them onto the query, key, and value spaces respectively; then establish pairwise relationships between regions through an attention mechanism, including the following:

[0120] Z = concat(Z) 1 Z 2 ,…,Z v )

[0121] Q = Z × W Q

[0122] K = Z × W K

[0123] V = Z × W V

[0124]

[0125]

[0126] Where Q represents the query space, K represents the key space, and V represents the value space;

[0127] Here, softmax() represents the normalization exponential function; concat() represents the function that concatenates two or more arrays or strings;

[0128] Where O represents the structure-aware similarity matrix, O ij Representing region r i With r j The structural relationship between them; V j Representing region r j The corresponding value vector; scaling factor d k The dimension is consistent with the key vector to stabilize attention calculations and reduce redundancy caused by multi-view stitching; a fully connected layer is used to... Further optimization yields the initial consensus representation.

[0129] The fully connected layer belongs to the cross-view fusion transformer module. For region r i The intermediate fusion representation aggregated through the cross-view attention mechanism, and the consensus representation. It is the final output of the cross-view fusion transformer module;

[0130] Step S42: Introduce a consistency loss function to ensure that the consensus representation can effectively capture the shared semantic information among the views, where the difference between the consensus representation and the specific representations of each view is minimized; consistency loss L con The definition includes the following:

[0131]

[0132] Where D(·,·) represents the similarity measure between two representations; by minimizing the consistency loss L com This allows for the preservation of key view-specific features while promoting semantic alignment between views.

[0133] Further, step S5 includes the following:

[0134] Step S51: Introduce the structure-aware multi-view contrast learning module. The structure-aware multi-view contrast learning module uses the cross-view fusion transformer module to calculate the structural similarity matrix O. ij And quantize the region r i With r j Structural similarity between them;

[0135] Step S52: Based on the similarity matrix O ij Design a dynamic weighting mechanism, where when O ij When it increases, exp(-αO ij The term ) effectively suppresses or eliminates the influence of these regions as negative samples, avoiding the generation of misleading supervisory signals; conversely, when Oij When the size is small, the region is considered a reliable negative sample.

[0136] In one embodiment of the present invention, when O ij When the value is large (indicating high structural similarity), exp(-αO) ij The term ) will effectively suppress or eliminate the influence of these regions as negative samples, avoiding the generation of misleading supervisory signals. Conversely, when O ij When the size is small, the region is considered a reliable negative sample.

[0137] Step S53: Introduce a dynamic temperature scaling strategy to adjust temperature parameters based on structural similarity. This suppresses the influence of similar regions during similarity calculation, reduces interference with contrastive learning, promotes more discriminative embedding learning, and yields the final consensus representation. Finally, the inter-view contrast loss of the v-th view is obtained. The definition includes the following:

[0138]

[0139] τ ij =τ(1+βO) ij )

[0140] in, For region r i The consensus expressed; and They are regions r i and r j Enhanced view-specific representation under the v-th view; τ is the reference temperature parameter; τ ij For region r i and r j The adaptive temperature parameters between them are used; α and β are the dynamic weighting intensity and temperature adjustment intensity, respectively; D'(·,·) is the embedding similarity quantization function under temperature control; the final consensus representation is obtained.

[0141] Step S54: The total contrast loss of all views is determined by analyzing each view. Summing yields:

[0142]

[0143] in, L represents the v-th view. c This represents the total contrast loss across all views.

[0144] Further, step S6 includes the following:

[0145] In one embodiment of the present invention, in step S6, a training strategy based on soft Lagrangian constraints is used to effectively solve the suboptimal solution problem caused by gradient conflicts in joint learning. Each training iteration consists of the following sequential sub-stages: optimization of learnable parameters in the sparse autoencoder module, followed by the augmentation of specific view representations and cross-view transformer modules, and finally the structure-aware multi-view contrastive learning module. In each stage, only the parameters of that stage are updated, while all other parameters are frozen. This alternating optimization ensures that the parameters of each module gradually approach local optima, and parameter freezing prevents gradient conflicts.

[0146] Step S61: The training strategy based on soft Lagrangian constraints includes a multi-stage collaborative optimization strategy, which includes the following:

[0147] Each training iteration is decomposed into a series of subtasks: first, the sparse autoencoder module is optimized, then the augmented specific view representation and cross-view transformer modules are jointly optimized, and finally the structure-aware multi-view contrast learning module is optimized. This order strictly follows the internal information flow of the framework, with the output of the previous stage serving as the input of the next stage, reflecting the dependency between the stages. In each stage, only the parameters of that stage are updated, while all other parameters are frozen.

[0148] Step S62: The training strategy based on soft Lagrange constraints includes establishing a soft Lagrange constraint mechanism, which includes the following:

[0149] A soft consistency regularization term is introduced between adjacent optimization stages to construct a cross-stage collaborative objective; the soft consistency regularization mechanism simulates the Lagrange multiplier in a soft form; the final loss is defined as:

[0150] L con-final =L con +λ1·R(Z,H)

[0151]

[0152] Where R(·,·) represents the use of cosine similarity to measure the similarity between adjacent stage embeddings; hyperparameters λ1 and λ2 control the strength of soft constraints and adjust the degree of cooperation between stages; the soft Lagrangian constraint mechanism adopts a delayed activation strategy, activating the soft constraints only after a preset number of iterations;

[0153] Where L con-final For the corrected eventual consistency loss; L c-final The final loss;

[0154] Consensus representation and These are the final outputs of the cross-view fusion transformer module and the structure-aware multi-view contrastive learning module, respectively. The cross-view fusion transformer module fuses complementary information from all views and, through inter-view attention mechanisms and contrastive learning, forms a shared feature representation that is effective for downstream tasks.

[0155] Step S63: The training strategy based on soft Lagrangian constraints includes establishing a state-aware constraint propagation mechanism, wherein the state-aware constraint propagation mechanism dynamically adjusts the strength of the soft constraints, including the following:

[0156]

[0157] in, and Let ω and ω represent the second norms of the functional gradients of the reconstruction loss and the consistency loss, respectively. The hyperparameters controlling the strength and sensitivity of the baseline constraints are: when the gradient increases in the previous stage, the constraint strength λ will decrease, providing learning space for the model and avoiding error propagation; the training strategy based on soft Lagrangian constraints achieves the robustness of the model to noise and optimization variance through a state-aware mechanism.

[0158] like Figure 2 As shown, according to a second aspect of the present invention, the present invention proposes a structure-aware multi-view city representation learning system with coordinated fusion and alignment, characterized in that it is implemented by a structure-aware multi-view city representation learning method with coordinated fusion and alignment as described in any one of the present invention, wherein the structure-aware multi-view city representation learning method with coordinated fusion and alignment includes:

[0159] The sparse autoencoder module is used to reconstruct the original data and extract a specific representation for each view;

[0160] Enhance the expressive power of specific view representation modules by modeling beneficial inter-view interactions;

[0161] The cross-view transformer module is used to optimize the multi-view fusion process by leveraging regional similarity to generate a robust consensus representation.

[0162] The structure-aware multi-view contrastive learning module is used to enhance the consistency between the consensus representation and the specific view representation, and to generate the final consensus representation;

[0163] A training strategy module based on soft Lagrangian constraints is used to solve the suboptimal solution problem caused by gradient conflict in joint learning.

[0164] According to a third aspect of the present invention, the present invention proposes a structure-aware multi-view city representation learning system with coordinated fusion and alignment, comprising an electronic device, wherein the electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, when the processor executes the computer program, it implements a structure-aware multi-view city representation learning method with coordinated fusion and alignment as described in any one of the present invention.

[0165] According to a fourth aspect of the present invention, the present invention proposes a structure-aware multi-view city representation learning system with coordinated fusion and alignment, comprising a computer-readable storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, it implements a structure-aware multi-view city representation learning method with coordinated fusion and alignment as described in any one of the present invention.

[0166] In addition to the above, the present invention also has related embodiments, including the following:

[0167] This invention uses urban datasets from region A and region B to learn regional embedding representations using multi-view fusion and structure-aware contrastive learning methods. The learned regional embedding representations are then applied to two types of downstream tasks to verify the reliability of the model. These two types of downstream tasks include prediction tasks (regional attractiveness prediction and regional service demand prediction) and classification tasks (land use classification).

[0168] This invention selected seven advanced models in the field for performance comparison, and the results are shown in Tables 1 and 2.

[0169] In Table 1, for the regional attractiveness prediction, OURS is the optimal model, and ReCP is the suboptimal model. In Table 1, for the land use classification, OURS is the optimal model, and HREP is the suboptimal model.

[0170] In Table 2, for the regional attractiveness prediction category, OURS is the optimal model, and ReCP is the suboptimal model. In Table 2, for the land use classification category, OURS is the optimal model, and ReCP is the suboptimal model.

[0171]

[0172]

[0173] Table 1 Comparison of model performance in region A

[0174]

[0175] Table 2 Comparison of model performance in region B

[0176] The prediction task of this invention includes the following:

[0177] like Figure 3 As shown, regional popularity prediction and regional service demand prediction are common downstream tasks. This invention uses the number of check-ins and service requests in each region as the region's popularity and service demand, respectively. Using region representations learned by different methods as input, a Lasso regression model is used to train the prediction task, and all experimental results are obtained through 5-fold cross-validation. Tables 1 and 2 show the evaluation metrics results of the prediction task, including MAE, RMSE, and R. 2 The following conclusions can be drawn from the results:

[0178] The significant advantage of the model proposed in this invention is that, as can be seen from the results of the three indicators, the model proposed in this invention exhibits the best performance. For example, as... Figure 3 As shown, compared with the suboptimal model ReCP, R 2 These figures represent increases of 10.52% and 13.16%, respectively.

[0179] The multi-view method exhibits excellent performance: MVURE, MGFN, HREP, ReCP, and the method of this invention all fuse multi-view information through attention mechanisms and contrastive learning strategies, demonstrating excellent predictive performance and further verifying the importance of multi-view learning.

[0180] Single-view and inefficient fusion methods perform poorly: ZE-Mob, MV-PN, and ReMVC have relatively poor performance, for reasons including:

[0181] 1) ZE-Mob: This method uses only a single view for interaction modeling, making it sensitive to noise and reducing the accuracy of the embedding representation. Furthermore, it relies heavily on local co-occurrence information while ignoring global information, thus weakening the learning effect of inter-region correlations.

[0182] 2) MV-PN and ReMVC: Although both methods utilize multi-view information for region representation learning, MV-PN fails to fully capture deep-level interaction features by fusing multi-view information through a simple method; while ReMVC fails to effectively extract semantic information before information interaction, affecting the effect of subsequent comparative learning and ultimately reducing the overall model performance.

[0183] This invention proposes a classification task, including the following:

[0184] Another downstream task is land use classification. This invention uses K-means to cluster the region embeddings into different categories to evaluate the effectiveness of region representation learning. Table 1 shows the relevant metrics for classification performance, including NMI, ARI, and F-measure. The following are the main findings summarized from the experimental results of this invention:

[0185] The model of this invention achieves optimal classification performance: By optimizing multi-view fusion and utilizing region similarity modeling to ensure the differences in features between views, the method of this invention extracts a more robust consensus representation, thus achieving the best classification performance. This result highlights the advantages of the method of this invention.

[0186] Significant Improvement in Contrastive Learning Strategy: Compared to existing contrastive learning models, the method of this invention enhances the consistency between view-specific representations and consensus representations while further strengthening the similarity of embedded representations of similar regions through structure-aware contrastive learning. This strategy effectively alleviates the problem of inconsistent representations of similar regions in previous studies, thereby significantly improving classification performance.

[0187] While fusion-based models offer strong performance, they also have limitations: Although fusion-based models such as MVURE and MGFN demonstrate strong classification capabilities, redundant view-specific information may dominate during feature fusion, affecting the classification results. This limitation highlights the advantages of the method in this invention in optimizing features between views.

[0188] Next, the present invention designed five variant models to evaluate the effectiveness of each module in the framework and analyze their specific impact on overall performance.

[0189] The descriptions of the various variant models are as follows:

[0190] MVFSAC-w / o-VR: Replaces the sparse autoencoder module (VR) with a multilayer perceptron (MLP).

[0191] MVFSAC-w / o-CVTF: Replaces the Cross-View Transformer Module (CVTF) for feature fusion with simple feature concatenation.

[0192] MVFSAC-w / o-EVSR: Removes Enhanced Specific View Representation Modules (EVSR) from the model.

[0193] MVFSAC-w / o-MVCLS: Removes the structure-aware multi-view contrastive learning module (MVSCL) from the model.

[0194] MVFSAC-w-CL: Replaces the structure-aware multi-view contrastive learning module with the standard contrastive learning (CL) module.

[0195] The present invention applied the above-mentioned variant in the two aforementioned experiments, and the results are as follows: Figure 3 As shown.

[0196] in Figure 3 (a) represents the regional attractiveness prediction. Figure 3 (b) indicates land use classification.

[0197] Overall, the proposed model significantly outperforms the five variant models, demonstrating the key contributions of each module to the overall performance. A detailed analysis follows:

[0198] Effectiveness of the sparse autoencoder module: Replacing the sparse autoencoder module with a multilayer perceptron reduced model performance by 4.27% to 8.72%. This demonstrates that the sparse autoencoder effectively captures key features and improves robustness to noise in the input data through reconstruction loss. This characteristic makes the model more stable when dealing with noisy data or outliers, thereby improving the reliability of the region embedding representation.

[0199] Effectiveness of the cross-view transformer module: Replacing the cross-view transformer module with simple feature concatenation resulted in a greater drop in model performance, ranging from 11.45% to 14.12%, particularly poor performance in the region popularity prediction task. This is because simple concatenation fails to fully exploit the correlations between different views, leading to the accumulation of private information, which masks consensus features, thereby weakening the model's effectiveness and classification accuracy.

[0200] Enhancing the effectiveness of the view-specific representation module: This module improves the quality of view-specific representations by leveraging the complementarity among multiple views. By enriching the view-specific representations, this module enhances the effectiveness of the similarity between the consensus representation and the view-specific representation, thereby amplifying the role of contrastive learning. This mechanism effectively improves the model's performance in downstream tasks.

[0201] Effectiveness of the Structure-Aware Multi-View Contrast Learning Module: In the land use classification task, the variant with the Structure-Aware Multi-View Contrast Learning Module removed (MVFSAC-w / o-MVCLS) performed the worst, highlighting the crucial role of the MVFSAC module in improving model performance. The MVFSAC module guides the model to enhance representational similarity between similar regions by introducing structural relationship constraints while maintaining consistency between consensus representations and view-specific representations. This mechanism is highly aligned with the core objectives of downstream tasks. Furthermore, replacing the MVFSAC variant with the standard contrast learning module (MVFSAC-w-CL) further validates the effectiveness of MVFSAC in handling these tasks.

[0202] The above are preferred embodiments of the present invention. Any changes made to the technical solution of the present invention that do not exceed the scope of the technical solution of the present invention shall fall within the protection scope of the present invention.

Claims

1. A structure-aware multi-view city representation learning method with coordinated fusion and alignment, characterized in that, The aforementioned structure-aware multi-view city representation learning method with coordinated fusion and alignment establishes a structure-aware multi-view representation learning model with coordinated fusion and alignment based on target city POI and taxi travel data. This method also includes the following: Step S1: Construct multi-view data; Step S2: Use a sparse autoencoder module to reconstruct the original data and extract a specific representation for each view; Step S3: Execute the Enhance Specific View Representation module to enhance the expressiveness of specific view representations by modeling beneficial inter-view interactions; Step S4: Employ a cross-view transformer module, including optimizing the multi-view fusion process by leveraging the region similarity of the cross-view transformer module to generate an initial consensus representation; Step S5: Execute the structure-aware multi-view contrastive learning module to enhance the consistency between the consensus representation and the specific view representation, and generate the final consensus representation; Step S6: Optimize model parameters and implement a training strategy based on soft Lagrangian constraints to solve the suboptimal solution problem caused by gradient conflict in joint learning.

2. The structure-aware multi-view city representation learning method with coordinated fusion and alignment as described in claim 1, characterized in that, Step S1 includes the following: Step S11: Use POI to represent the semantic features of the region, and define the semantic features of the region as P, including the following: P={p1,p2,p3,...,p n },p i ∈R c Where C is the number of POI categories, p i p is the semantic feature representation of region i. i Each dimension is represented by the number of POIs of a specific category within the region, and the dimension size is C; Where POI represents Point of Interest; Step S12: Define regional interaction characteristics, including defining taxi travel in the target city as an interaction behavior between regions; The regional interaction features are divided into outflow features and inflow features. The outflow feature is defined as S, which includes the following: S={s1,s2,s3,…,s n },s i ∈R n Where n is the number of regions, s i The number of taxi trips from region i to other regions within a specific time period, s i Each dimension represents the number of times a user travels from region i to different regions. The inflow characteristic is defined as D, which includes the following: D=(d1,d2,d3,…,d n )d i ∈R n Where d i Each dimension represents the number of times region i is reached from different regions, R n Represents each inflow feature vector d i It is an n-dimensional vector.

3. The structure-aware multi-view city representation learning method with coordinated fusion and alignment as described in claim 2, characterized in that, Step S2 includes the following: Step S21: Reconstruct the original data using a sparse autoencoder module and extract a specific representation for each view, including the following: Learn and extract representative embedding representations from the original view; Introducing reconstruction loss L r : Utilizing a reconstruction loss L that combines reconstruction error and sparsity regularization r This is used to improve the model's generalization performance and its ability to capture regional features.

4. The structure-aware multi-view city representation learning method with coordinated fusion and alignment as described in claim 3, characterized in that, Step S3 includes the following: Step S31: After obtaining the initial view-specific embedding representation in the sparse autoencoder module, the view-specific features of each view are enhanced by integrating complementary information from other views: the features of the current view and other views are concatenated along the feature dimension to form a joint representation H; then, the joint representation is transformed by a nonlinear function and multiplied element-wise with the original view-specific features to generate the enhanced representation Z. v : Z v =H v ⊙M(H) Where ⊙ represents element-wise multiplication, H v Let H be a specific view representation of the v-th view; M() is a nonlinear transformation function that projects the joint representation H onto the view with respect to H. v Spaces with the same dimensions ensure consistency in the feature space.

5. The structure-aware multi-view city representation learning method with coordinated fusion and alignment as described in claim 4, characterized in that, Step S4 includes the following: Step S41: Use the learnable matrix W Q W K and W V Embed the features of different views into Z v Project them onto the query, key, and value spaces respectively; then establish pairwise relationships between regions through an attention mechanism, including the following: Z=concat(Z 1 ,Z 2 ,…,Z v ) Q=Z×W Q K=Z×W K H=Z×W K Where Q represents the query space, K represents the key space, and V represents the value space; Here, softmax() represents the normalization exponential function; concat() represents the function that concatenates two or more arrays or strings; Where O represents the structure-aware similarity matrix, O ij Representing region r i With r j The structural relationship between them; V j Representing region r j The corresponding value vector; scaling factor d k The dimension is consistent with the key vector to stabilize attention computation; to reduce redundancy caused by multi-view stitching, a fully connected layer is used. Further optimization yields the initial consensus representation. The fully connected layer belongs to the cross-view fusion transformer module. For region r i The intermediate fusion representation aggregated through the cross-view attention mechanism, and the consensus representation. It is the final output of the cross-view fusion transformer module; Step S42: Introduce a consistency loss function to ensure that the consensus representation can effectively capture the shared semantic information among the views, where the difference between the consensus representation and the specific representations of each view is minimized; consistency loss L con The definition includes the following: Where D(·,·) represents the similarity measure between two representations; by minimizing the consistency loss L con This allows for the preservation of key view-specific features while promoting semantic alignment between views.

6. The structure-aware multi-view city representation learning method with coordinated fusion and alignment according to claim 5, characterized in that, Step S5 includes the following: Step S51: Introduce the structure-aware multi-view contrast learning module. The structure-aware multi-view contrast learning module uses the cross-view fusion transformer module to calculate the structural similarity matrix O. ij And quantize the region r i With r j Structural similarity between them; Step S52: Based on the similarity matrix O ij Design a dynamic weighting mechanism, where when O ij When it increases, exp(-αQ) ij The term ) effectively suppresses or eliminates the influence of these regions as negative samples, avoiding the generation of misleading supervisory signals; conversely, when O ij When the value is small, the region is considered a reliable negative sample; Step S53: Introduce a dynamic temperature scaling strategy to adjust temperature parameters based on structural similarity. This suppresses the influence of similar regions during similarity calculation, reduces interference with contrastive learning, promotes more discriminative embedding learning, and yields the final consensus representation. Finally, the inter-view contrast loss of the v-th view is obtained. The definition includes the following: t ij =τ(1+βO ij ) in, For region r i The consensus expressed; and They are regions r i and r j Enhanced view-specific representation under the v-th view; τ is the reference temperature parameter; τ ij For region r i and r j The adaptive temperature parameters between; α and β are the dynamic weighting intensity and temperature adjustment intensity, respectively; D'(·,·) is the embedding similarity quantization function under temperature control; Step S54: The total contrast loss of all views is obtained by analyzing each view. Summing yields: in, L represents the v-th view. c This represents the total contrast loss across all views.

7. The structure-aware multi-view city representation learning method with coordinated fusion and alignment as described in claim 6, characterized in that, Step S6 includes the following: Step S61: The training strategy based on soft Lagrangian constraints includes a multi-stage collaborative optimization strategy, which includes the following: Each training iteration is decomposed into a series of subtasks: first, optimize the sparse autoencoder module, then jointly optimize the augmented specific view representation and cross-view transformer modules, and finally optimize the structure-aware multi-view contrast learning module. This order strictly follows the internal information flow of the framework, with the output of the previous stage serving as the input of the next stage, reflecting the dependency between the stages. In each stage, only the parameters of that stage are updated, while all other parameters are frozen. Step S62: The training strategy based on soft Lagrange constraints includes establishing a soft Lagrange constraint mechanism, which includes the following: A soft consistency regularization term is introduced between adjacent optimization stages to construct a cross-stage collaborative objective; the soft consistency regularization mechanism simulates the Lagrange multiplier in a soft form; the final loss is defined as: Lcon - final=Lcom+λ1·R(Z,H) Where R(·,·) represents the use of cosine similarity to measure the similarity between adjacent stage embeddings; hyperparameters λ1 and λ2 control the strength of soft constraints and adjust the degree of cooperation between stages; the soft Lagrangian constraint mechanism adopts a delayed activation strategy, activating the soft constraints only after a preset number of iterations; Where L con-final For the corrected eventual consistency loss; L c-final For the ultimate loss; Consensus representation and These are the final outputs of the cross-view fusion transformer module and the structure-aware multi-view contrastive learning module, respectively. The cross-view fusion transformer module fuses complementary information from all views and, through inter-view attention mechanisms and contrastive learning, forms a shared feature representation that is effective for downstream tasks. Step S63: The training strategy based on soft Lagrangian constraints includes establishing a state-aware constraint propagation mechanism. The state-aware constraint propagation mechanism dynamically adjusts the strength of soft constraints, including the following: in, and Let ω and ω represent the second norms of the functional gradients of the reconstruction loss and the consistency loss, respectively. The hyperparameters controlling the strength and sensitivity of the baseline constraints are: when the gradient increases in the previous stage, the constraint strength λ will decrease, providing learning space for the model and avoiding error propagation; the training strategy based on soft Lagrangian constraints achieves the robustness of the model to noise and optimization variance through a state-aware mechanism.

8. A structure-aware multi-view city representation learning system with coordinated fusion and alignment, characterized in that, This is achieved through a structure-aware multi-view city representation learning method with coordinated fusion and alignment as described in any one of claims 1 to 7, wherein the structure-aware multi-view city representation learning method with coordinated fusion and alignment includes: The sparse autoencoder module is used to reconstruct the original data and extract a specific representation for each view; Enhance the expressive power of specific view representation modules by modeling beneficial inter-view interactions; The cross-view transformer module is used to optimize the multi-view fusion process by leveraging regional similarity to generate a robust consensus representation. The structure-aware multi-view contrastive learning module is used to enhance the consistency between consensus representations and specific view representations; A training strategy module based on soft Lagrangian constraints is used to solve the suboptimal solution problem caused by gradient conflict in joint learning.

9. A structure-aware multi-view city representation learning system with coordinated fusion and alignment, comprising an electronic device, wherein the electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements a structure-aware multi-view city representation learning method with coordinated fusion and alignment as described in any one of claims 1 to 7.

10. A structure-aware multi-view city representation learning system with coordinated fusion and alignment, comprising a computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements a structure-aware multi-view city representation learning method with coordinated fusion and alignment as described in any one of claims 1 to 7.

Citation Information

Cited By

  • A method for representing urban areas based on multi-view joint and structure-aware contrast

    CN122289943A

  • A multi-view urban area embedding method based on spatial function consistency

    CN122336073A

  • A multi-view urban area embedding method based on spatial function consistency

    CN122336073B