Spatial Heterogeneity Adaptive Learning Method and System for OD Flow Inference

By using a spatial heterogeneity adaptive learning system and leveraging the joint optimization of a heterogeneity representation encoder and multiple expert networks, the problem of limited accuracy in existing OD flow inference methods in complex regions is solved, and more accurate spatial interaction relationship modeling is achieved.

CN122491464APending Publication Date: 2026-07-31INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS
Filing Date
2026-03-25
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing OD flow inference methods are difficult to accurately reflect the differences in spatial interaction relationships in urban areas with multiple centers and significant functional differences, and their accuracy is limited due to reliance on external prior knowledge or parameter sharing.

Method used

A spatial heterogeneity adaptive learning system is adopted, which uses a heterogeneity representation encoder and multiple parallel expert networks to perform joint optimization using a multi-objective loss function, adaptively characterizing spatial differences and performing differential modeling.

Benefits of technology

It improves the accuracy and stability of OD flow inference, can accurately reflect the differences in spatial interaction relationships in complex regions, avoids error accumulation, and enhances the model's ability to express local interaction patterns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122491464A_ABST
    Figure CN122491464A_ABST
Patent Text Reader

Abstract

This application provides a spatial heterogeneity adaptive learning method and system for OD traffic inference, belonging to the fields of artificial intelligence and geospatial data analysis. The system includes: a spatial heterogeneity adaptive representation learning module that receives the original input features of OD pairs and maps these features into heterogeneous representation vectors using a heterogeneity representation encoder; a heterogeneity-aware inference module connected to the spatial heterogeneity adaptive representation learning module, comprising a probabilistic gating network and multiple parallel expert networks; the probabilistic gating network is used to generate soft route weights; each expert network is used to generate preliminary traffic prediction values; the heterogeneity-aware inference module performs a weighted summation of the preliminary traffic prediction values ​​from each expert network based on the soft route weights to obtain the final OD traffic prediction value; and an end-to-end joint optimization module uses a multi-objective loss function to simultaneously perform joint optimization and training of the network parameters. This system can improve the applicability and generalization ability of OD traffic inference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of artificial intelligence and geospatial data analysis technology, and in particular to a spatial heterogeneity adaptive learning method and system for OD flow inference. Background Technology

[0002] Origin–Destination (OD) flow data is a type of geographic information data used to describe the flow relationship of human activities between different spatial units. By depicting the correspondence between the origin (O) and destination (D) of movement behavior, it reflects the strength of connections and spatial interaction patterns between regions. This type of data plays an important role in applications such as urban spatial structure analysis, regional accessibility assessment, spatial interaction modeling, and urban functional area identification.

[0003] With the development of mobile sensing and positioning technologies, the means of acquiring data related to human travel activities are constantly enriching. However, in practical applications, due to limitations such as data collection methods, sample coverage, differences in data standards, and privacy protection and compliance requirements, it is difficult to directly obtain comprehensive and structurally complete OD traffic data. Therefore, inferring and supplementing OD traffic based on limited observation data has become an important research and application direction in related technical fields.

[0004] Existing methods for OD flow inference include various technical approaches, such as traditional models based on spatial interaction assumptions and data-driven models based on machine learning. Most of these existing technologies adopt a unified global modeling strategy, that is, to describe the OD flow relationship in the entire study area through fixed functional forms or shared model parameters. In inference, these methods usually assume that the factors affecting OD flow and their mechanisms of action are consistent in different spatial areas (i.e., the spatial stationarity assumption). They are applicable to certain scenarios with relatively simple spatial structures. However, in urban areas with multiple centers and significant functional differences, this globally unified modeling approach often leads to the smoothing or suppression of local interaction patterns, making it difficult to accurately reflect the differences in spatial interaction relationships between different areas. This can easily lead to large deviations in the inference results in complex areas.

[0005] To address the aforementioned issues, some existing technologies have begun to incorporate spatial heterogeneity factors into the OD flow inference process, considering the differences in the relationship between OD flow and its driving factors across different spatial regions. Related methods typically construct spatial heterogeneity features and use them as model inputs to improve inference accuracy. However, in practical applications, the following shortcomings remain: Firstly, some methods rely on pre-defined spatial structure features or external prior knowledge, such as fixed network indicators or predefined spatial partitions, when characterizing spatial heterogeneity. Their effectiveness is highly dependent on the rationality of prior assumptions, making it difficult to adapt to complex and changing spatial environments, and lacking the ability to automatically learn potential spatial structures and interaction patterns from observational data. Secondly, while some methods introduce spatial heterogeneity features, their inference models still employ a parameter-sharing global optimization approach, imposing uniform parameter constraints across different regions or OD pairs. This limits the model's ability to express regional differences and specific spatial interaction patterns, making it difficult to fully leverage the role of spatial heterogeneity features. In addition, some existing technologies separate the learning process of spatial heterogeneity features from the OD flow inference process into multiple stages. This two-stage processing method leads to a misalignment between the feature learning target and the flow inference target, which can easily cause error accumulation during the inference process and affect the overall inference effect. Summary of the Invention

[0006] This invention addresses the problem that existing OD flow inference technologies struggle to simultaneously balance the ability to represent spatial heterogeneity and the accuracy of data-driven inference. It proposes a technical solution that can adaptively characterize spatial differences under limited observation data conditions and perform differentiated modeling of different spatial interaction modes.

[0007] To achieve the above objectives, this application provides the following technical solution: This application provides a spatial heterogeneity adaptive learning system for OD flow inference, the system being used to process OD pair datasets consisting of origin O and destination D, the system comprising: The spatial heterogeneity adaptive representation learning module is configured to receive the original input features of OD pairs, map the original input features into heterogeneous representation vectors using a heterogeneous representation encoder, and perform L2 normalization on the heterogeneous representation vectors; during the model training phase, the heterogeneous representation encoder is optimized based on the contrastive learning objective. The heterogeneity-aware inference module, connected to the spatial heterogeneity adaptive representation learning module, includes a probabilistic gating network and multiple parallel expert networks. The probabilistic gating network is configured to receive normalized heterogeneity representation vectors and generate soft routing weights corresponding to each expert network. Each expert network is configured to receive the original input features and output a corresponding traffic prediction value. The heterogeneity-aware inference module is further configured to perform a weighted summation of the traffic prediction values ​​output by each expert network based on the soft routing weights to obtain the final OD traffic prediction value. The end-to-end joint optimization module is configured to utilize a multi-objective loss function to jointly optimize and train the network parameters of the spatial heterogeneity adaptive representation learning module and the heterogeneity-aware inference module.

[0008] Preferably, the spatial heterogeneity adaptive representation learning module includes a heterogeneity representation encoder, which is configured to perform the following operations: Receive the original input features of the OD pair, which include the node attribute features of the start point O and the end point D and the spatial relationship features between the start point O and the end point D; The original input features are mapped to a low-dimensional embedding space using a multilayer perceptron (MLP), thereby generating heterogeneous representation vectors for distinguishing different spatial interaction modes.

[0009] Preferably, the spatial heterogeneity adaptive representation learning module constructs a contrastive learning objective based on OD flow values, and the contrastive learning objective is configured as follows: Pseudo-labels for comparative learning are generated based on the actual flow values ​​of the OD pairs; Based on the pseudo-labels, the heterogeneity representation learning process is constrained by a supervised contrastive loss function, so that the representations of OD pairs with similar flow rates are close to each other in the embedding space, while the representations of OD pairs with large flow rate differences are far apart in the embedding space.

[0010] Preferably, the expression for the supervised contrastive loss function is as follows: , In the formula, This indicates that there is a supervised comparison of losses. This represents the set of sample indices for the current training batch. It is any sample in a certain training batch. Indicates sample The positive sample set, samples yes Any positive sample in, It is the set of all samples, including both positive and negative samples. yes Any sample in, , , Representing samples respectively ,sample ,sample Normalized heterogeneity representation vector, Represents the positive sample set Number of samples included This indicates that the cosine similarity between two vectors is calculated. It's a hyperparameter.

[0011] Preferably, the probability gating network is further configured as follows: A linear transformation layer is configured to receive the heterogeneity representation vector and generate a corresponding linear transformation. dimensional original fraction vector; The Softmax activation layer, connected to the linear transformation layer, is configured to activate the linear transformation layer. The original score vector is normalized to generate soft route weights representing the probabilities of each OD pair being assigned to each expert network. The expert network is a deep residual neural network structure.

[0012] Preferably, the heterogeneity-aware inference module further includes a load balancing constraint unit, which is configured to calculate the expert load balancing loss based on the soft route weights generated by the probabilistic gating network. The expert load balancing loss is constructed as a measure of the difference between the average weights assigned to all expert networks within a training batch and the ideal uniform distribution.

[0013] Preferably, the multi-objective loss function is a weighted sum of the main task regression loss, the supervised comparison loss, and the expert load balancing loss; The regression loss for the main task uses the Hubel loss function to measure the deviation between the predicted OD flow value and the actual flow value.

[0014] This embodiment provides a spatial heterogeneity adaptive learning method for OD flow inference, the method being executed by the system described in any of the above embodiments, including: Obtain an OD pair dataset consisting of a starting point O and an ending point D. The OD pair dataset contains the original input features. The spatial heterogeneity adaptive representation learning module receives the original input features of the OD pair, maps the original input features into a heterogeneity representation vector using the heterogeneity representation encoder, and performs L2 normalization on the heterogeneity representation vector; during the model training phase, the heterogeneity representation encoder is optimized based on the contrastive learning objective. The heterogeneity-aware inference module includes a probabilistic gating network and multiple parallel expert networks. The probabilistic gating network is configured to receive a normalized heterogeneity representation vector and generate soft routing weights corresponding to each expert network. Each expert network is configured to receive the original input features and output a corresponding traffic prediction value. The heterogeneity-aware inference module is further configured to perform a weighted summation of the traffic prediction values ​​output by each expert network based on the soft routing weights to obtain the final OD traffic prediction value. The end-to-end joint optimization module calculates the regression loss of the main task, the supervised comparison loss, and the expert load balancing loss based on the weight distribution of soft routers, and performs joint optimization on all network parameters.

[0015] This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the steps of the method described in the above embodiment.

[0016] Beneficial effects: The technical solution provided in this application, by setting up a spatial heterogeneity adaptive representation learning module, can automatically learn potential spatial interaction patterns from the original input data of OD pairs, achieving adaptive mining of spatial heterogeneity features without the need for manually setting spatial prior assumptions. Through the cooperation of a probabilistic gating network and multiple parallel expert networks in the heterogeneity-aware inference module, the smoothing or suppression of local interaction patterns caused by global parameter sharing is avoided, more accurately reflecting the differences in spatial interaction relationships between different areas in multi-center, functionally differentiated urban areas. The network parameters are jointly optimized and trained using a multi-objective loss function, ensuring that the learned heterogeneity representation vectors can directly serve the traffic inference task, preventing error propagation and accumulation at different stages, and improving the stability and accuracy of the overall inference effect. Attached Figure Description

[0017] Figure 1 This is a logic diagram of a spatial heterogeneity adaptive learning system for OD flow inference provided according to some embodiments of this application. Detailed Implementation

[0018] The purpose of this invention is to overcome the limitations of existing OD flow inference techniques, such as insufficient spatial heterogeneity characterization, over-reliance on external prior knowledge, and accuracy constraints caused by global sharing of inference model parameters. This invention proposes an OD flow inference technique that can adaptively model spatial heterogeneity under limited observation data. Specifically, this embodiment provides a spatial heterogeneity adaptive learning method and system for OD flow inference. Based on a joint learning framework of contrastive representation learning and expert perception (ALSH-OD model), it adaptively characterizes the interaction differences between different spatial regions and obtains high-precision OD flow inference results even with only partial observation flow data.

[0019] The embodiments of this application will now be described with reference to the accompanying drawings.

[0020] This embodiment provides a spatial heterogeneity adaptive learning system for OD flow inference. This system is used to process OD pair datasets consisting of origin O and destination D, such as... Figure 1 As shown, the system includes: The spatial heterogeneity adaptive representation learning module is configured to receive the original input features of OD pairs, map the original input features into heterogeneous representation vectors using the heterogeneity representation encoder, and perform L2 normalization on the heterogeneous representation vectors; during the model training phase, the heterogeneity representation encoder is optimized based on the contrastive learning objective. The heterogeneity-aware inference module, connected to the spatial heterogeneity adaptive representation learning module, includes a probabilistic gating network and multiple parallel expert networks. The probabilistic gating network is configured to receive normalized heterogeneity representation vectors and generate soft routing weights corresponding to each expert network. Each expert network is configured to receive raw input features and output corresponding traffic prediction values. The heterogeneity-aware inference module is also configured to perform a weighted summation of the traffic prediction values ​​output by each expert network based on the soft routing weights to obtain the final OD traffic prediction value. The end-to-end joint optimization module is configured to utilize a multi-objective loss function to jointly optimize and train the network parameters (i.e., all network parameters) of the spatial heterogeneity adaptive representation learning module and the heterogeneity-aware inference module.

[0021] The technical solution of this embodiment will be described in detail below.

[0022] I. Problem Definition.

[0023] In traffic inference tasks, traffic refers to the quantity or intensity of moving entities (such as people, vehicles, data packets, goods, etc.) that start from a specific starting point and arrive at a specific destination within a specific time period.

[0024] Origin and destination form an ordered origin-destination pair (OD pair). OD pairs have interactive directionality; the flow direction from origin O to destination D and the flow direction from destination D to origin O are physically considered two independent OD pairs. The OD flow inference task aims to infer the flow of a given set of OD pairs (all OD pairs constitute the OD dataset, denoted as: Learn a mapping function Each OD pair From the eigenvector The description includes node attributes of the start and end points and the spatial relationship between them (such as distance, interaction frequency, traffic intensity, etc.). The goal of modeling is to ensure that the generated inferred values... Able to accurately approximate the actual flow rate .

[0025] The spatial heterogeneity adaptive learning system for OD traffic inference provided in this embodiment includes three core modules working in concert: (1) Adaptive Heterogeneity Representation Learning module: Through a contrastive learning objective guided by a downstream inference task, it adaptively learns heterogeneous representation vectors with strong discriminativeness from the original input features to effectively distinguish different spatial interaction patterns; (2) Heterogeneity-Aware Soft Routing Mechanism (also known as the Heterogeneity-Aware Inference Module): Based on the above heterogeneous representation vectors, it generates soft routing weights and dynamically assigns each OD pair to multiple expert networks; each expert independently models the local interaction pattern they are responsible for and generates the final traffic prediction value through probability weighted aggregation; (3) End-to-End Joint Optimization strategy. Strategy (also known as the end-to-end joint optimization module): Designs a unified multi-objective loss function to deeply integrate heterogeneous representation learning, expert routing, and traffic prediction tasks, achieving end-to-end training throughout the entire process and ensuring that all components are collaboratively optimized under a unified objective.

[0026] II. Spatial Heterogeneity Adaptive Representation Learning Module.

[0027] Characterizing spatial heterogeneity is a core challenge in geographic flow modeling. Traditional methods often rely on external prior knowledge to characterize spatial heterogeneity, or decouple the representation learning stage from the downstream flow inference task. This results in learned representations that, while reflecting some spatial interaction patterns, may not necessarily serve the optimization of inference accuracy. To overcome these limitations, this embodiment provides an adaptive representation learning module for spatial heterogeneity. By introducing contrastive learning, the OD flow magnitude is used as an intrinsic supervisory signal to guide the model to directly learn heterogeneous representations beneficial to the inference task from OD data. This design does not require pre-defined spatial partitions, avoids the disconnect between representation learning and prediction objectives, and helps the model autonomously identify potential interaction structures closely related to the flow generation mechanism.

[0028] Specifically, the spatial heterogeneity adaptive representation learning module comprises two components: (i) a heterogeneous representation encoder, used to map the original OD features of the OD pair dataset to a low-dimensional embedding space. Further, this heterogeneous representation encoder is configured to perform the following operations: receive the original input features of the OD pairs, which include node attribute features of the starting point O and the ending point D, and spatial relationship features between the starting point O and the ending point D; map the original input features to the low-dimensional embedding space using a multilayer perceptron (MLP), thereby generating heterogeneous representation vectors for distinguishing different spatial interaction modes. (ii) a contrastive learning objective based on OD flow size, used to guide the structural optimization of the embedding space, enabling it to semantically distinguish different flow patterns (i.e., spatial interaction modes), thereby supporting more accurate inference. Furthermore, the spatial heterogeneity adaptive representation learning module constructs a contrastive learning objective based on OD flow values. The contrastive learning objective is configured as follows: generate pseudo-labels for contrastive learning based on the real flow values ​​of OD pairs; based on the pseudo-labels, constrain the heterogeneity representation learning process through a supervised contrastive loss function, so that the representations of OD pairs with similar flow values ​​are close to each other in the embedding space, and the representations of OD pairs with large flow differences are far apart in the embedding space.

[0029] The heterogeneity representation encoder and contrastive learning objective are described in detail below.

[0030] (1) Heterogeneity representation encoder.

[0031] The heterogeneity-aware routing module includes a heterogeneity representation encoder. Its input is a feature vector containing OD node attribute features and spatial interaction relationships (such as distance, interaction frequency, traffic intensity, etc.). The output is a heterogeneous representation vector. This process involves transforming the high-dimensional original OD input vector... Mapping to a semantically meaningful low-dimensional embedding space, it can be represented as ,in This representation vector represents the learnable parameters of the encoder. The aim is to encode abstract features that can distinguish different spatial interaction patterns, providing highly discriminative inputs for subsequent gating networks.

[0032] Spatial interaction patterns refer to a set of interactive relationships between origin-destination (OD) pairs in geographic space, exhibiting similar flow generation mechanisms or being dependent on specific spatial structures. Specifically, at the physical attribute level, this term manifests as a specific combination of characteristics exhibited by a set of OD pairs in terms of spatial impedance (e.g., geographical distance), interaction intensity (e.g., flow rate), and socio-economic attributes of origin and destination (e.g., regional GDP, industrial structure). Typical examples include "short-distance, high-frequency intra-city cluster commuter flows" driven by regional economic synergy and "long-distance bulk cargo inter-regional transportation flows" influenced by national logistics network architecture. Simultaneously, at the model representation level, this term refers to the fact that the aforementioned OD pairs are mapped to similar vector representations in the hidden feature space of a deep learning model. Because these OD pairs follow similar nonlinear mapping rules from input features to flow output, they tend to be activated and processed by the same expert network in the model, thus representing differentiated flow characteristics and interaction mechanisms under different geographic environments from a data-driven perspective.

[0033] Structurally, heterogeneity characterizes the encoder. This is a multilayer perceptron (MLP) consisting of two fully connected layers. To improve the stability of the training process and accelerate convergence, a BatchNormalization operation is introduced after each fully connected layer, and a Gaussian error linear unit (GELU) is used as the activation function to achieve a continuous and smooth nonlinear transformation. Encoder The final output heterogeneity representation vector pass Norm normalization ( To unify the length, subsequent comparative learning can focus entirely on vector direction, thereby learning more fundamental pattern similarity.

[0034] (2) Comparative learning objectives of representation learning.

[0035] To ensure the encoder The generated heterogeneous representation vectors are effective for downstream inference tasks, avoiding decoupling between representation learning and inference tasks. The encoder's learning process is guided by a contrastive learning objective. The core idea of ​​this objective is that OD samples with similar flow rates attract each other, while OD samples with vastly different flow rates repel each other. This is achieved by utilizing the task's own supervisory signal, namely the actual OD flow rate values. This guides the construction of the encoder embedding space.

[0036] Specifically, the system first takes its continuous real flow values (After logarithmic transformation) it is discretized using a quantile binning strategy to... Each category, for each sample (i.e., the corresponding OD pair) is assigned a data-driven pseudo-label. Based on this pseudo-label, for a sample... All other samples within the same batch that have the same pseudo-label Together they constitute its positive sample set All pseudo-labels that differ from these are considered negative samples. The contrastive learning objective is achieved through a supervised contrastive loss function. The implementation, its mathematical expression is as follows: (1) in, It is for the sample The number of positive samples, This represents the set of sample indices for the current training batch. represent Any positive sample in The heterogeneous representation vector, It is the set of all samples, including both positive and negative samples. yes Any sample in, Indicates two processes Cosine similarity between representation vectors normalized by norm. It is a sample The representation vector (heterogeneous representation vector). It is a sample The representation vector, It is a temperature hyperparameter used to adjust the model's sensitivity to distinguishing between positive and negative sample pairs. (.) indicates exponentiation.

[0037] This loss function effectively aligns representation learning with the final regression task objective, ensuring that accurate routing decisions can be made based on heterogeneous information that is highly relevant to the task.

[0038] III. Heterogeneity-aware soft routing mechanism.

[0039] After obtaining embedding vectors (i.e., heterogeneous representation vectors) that can effectively represent different spatial interaction patterns, the key issue is how to construct an inference architecture to fully utilize this heterogeneous information and perform differentiated modeling of different spatial interaction patterns. Even when heterogeneous features are used as input, the parameter sharing mechanism of traditional global models still weakens spatial heterogeneous information, thereby inhibiting the modeling of local specific interaction patterns.

[0040] To overcome this limitation, this embodiment introduces a Mixture of Experts (MoE) architecture in the heterogeneity-aware inference module, applying a divide-and-conquer strategy to OD traffic inference. The Mixture of Experts architecture includes... Each expert network (referred to as "expert") focuses on learning a specific local spatial interaction pattern (such as long-distance bulk flows or high-frequency short-distance flows within urban agglomerations), thereby improving the model's ability to express complex spatial heterogeneity. To achieve this goal, this embodiment designs a heterogeneity-aware soft routing mechanism (also known as a heterogeneity-aware routing module), consisting of four collaborative components: (i) a probabilistic gating network that receives the normalized heterogeneity representation vector and dynamically generates the assigned weights (also known as gating weights or soft routing weights) for each expert based on the heterogeneity representation vector; (ii) a set of parallel expert networks, each of which receives the original input features and outputs the corresponding flow prediction value, that is, each expert network independently infers the original input of the OD data; (iii) a weighted fusion module, based on... The soft route weights sum the traffic prediction values ​​output by each expert network to obtain the final OD traffic prediction value. That is, the outputs of each expert are aggregated according to the gating weights to form the final prediction; (iv) A load balancing constraint unit calculates the expert load balancing loss according to the soft route weights generated by the probabilistic gating network to prevent route collapse (i.e., a few experts dominate the allocation) and ensure that all experts are effectively activated and utilized during training. This mechanism not only supports adaptive modeling of different spatial interaction modes, but also ensures the diversity and stability of expert division of labor through explicit route supervision and balancing constraints.

[0041] Furthermore, a probabilistic gating network is configured to receive the normalized heterogeneity representation vector and generate soft routing weights for each OD pair corresponding to each of the multiple expert networks based on the heterogeneity representation; multiple expert networks are configured to receive the original input features and output corresponding traffic prediction values; a weighted fusion module, connected to the probabilistic gating network and the multiple expert networks, is configured to perform weighted aggregation (weighted summation) on the traffic prediction values ​​output by the multiple expert networks according to the soft routing weights to generate the final OD traffic prediction value.

[0042] The following section provides a detailed explanation of the implementation of the probabilistic gating network, the expert network with multiple parallel settings, and the weighted fusion module.

[0043] (1) Probabilistic gating network.

[0044] To obtain the normalized heterogeneity representation vector that characterizes the heterogeneity of the sample space. After that, probabilistic gating networks Transform it into a The soft route weights of individual experts, the probabilistic gating network It consists of a linear transformation layer followed by a Softmax activation function, and its parameter set is denoted as: .

[0045] First, the input representation vector A linear transformation layer is used to generate a... A logit vector of dimension (i.e.) (original fraction vector) The calculation process is as follows: (2) in, and These are probabilistic gating networks Weight matrix and bias vector; logit vector Logit refers to the unnormalized raw score vector generated before the output layer of the gated network. Each element value in this vector (called the logit value) is a real number with an unbounded range. Its magnitude directly reflects the raw evidence support or unnormalized confidence of the model for the corresponding class based on the input features. A larger positive logit value usually indicates that the model favors that class, while a negative value indicates that it does not favor that class. The logit vector itself is not a probability distribution because the sum of its elements is not necessarily 1 and its values ​​are unbounded.

[0046] Subsequently, to transform the logit vector This is transformed into an effective probability distribution, and then normalized using the Softmax function to obtain the final result. Dimensional expert assignment probability vector The vector The first in element That is, to represent a sample Assigned to the The probability of each expert (i.e., the soft route weight) is calculated as follows: (3) The above calculation process ensures that the vector satisfy final vector These factors constitute the weighting coefficients for all experts, directly guiding the aggregation process of subsequent expert inferences.

[0047] (2) The structure of the expert network.

[0048] In the scheme of this application, OD flow inference is derived from... An independent network of experts Responsibility. The total number of expert networks. is a hyperparameter whose value determines the granularity of the model's representation of heterogeneous patterns. These expert networks are structurally identical but possess their own independent parameter sets, allowing each expert to focus on learning specific nonlinear mapping relationships under different spatial interaction patterns. A network of experts analyzed the input samples. The inference process is represented as ,in It is the first The expert on the first The inferred value of a sample, It is the first The original input features of each expert It is the first Learnable parameters for each expert.

[0049] Each expert network is a deep residual neural network structure. Deep residual networks can effectively capture complex nonlinear relationships. Specifically, each expert network consists of a linear input layer, It consists of a residual block and a linear output layer. The number of residual blocks... `<parameter>` is a hyperparameter used to determine the depth and non-linear fitting capability of each expert network. Each residual block contains two fully connected layers, batch normalization, a GELU activation function, and a Dropout layer. This structure alleviates gradient vanishing and improves the model's generalization ability. After obtaining deep features from each residual block, a linear output layer maps the deep features to the expert's preliminary flow prediction values. It should also be noted that the output of each expert network is a flow prediction value in logarithmic space, i.e. .

[0050] (3) Weighted fusion module (used for weighted output aggregation).

[0051] Model on samples The final OD flow prediction value, i.e. the final inferred flow. By all Independent inferences from individual experts The weighted aggregation is obtained by performing weighted aggregation, and the aggregation weights are the expert allocation probability vectors generated by the aforementioned probabilistic gating network. This process can be described as a weighted summation: (4) The probability-based aggregation mechanism is the core of the end-to-end optimization solution provided in this embodiment. Through the weighted aggregation mechanism, it is ensured that the final loss function can simultaneously guide the probabilistic gating network and each expert network, thereby achieving collaborative learning of all modules.

[0052] (4) Expert load balancing loss.

[0053] To ensure the effectiveness of the hybrid expert architecture, constraints need to be imposed on the soft routing mechanism of the probabilistic gating network to prevent the vast majority of samples from being concentrated in a few dominant experts, i.e., routing collapse. Routing collapse causes the remaining experts to degenerate due to lack of effective training, thereby weakening the model's diversity and overall representational ability. Therefore, this application introduces an expert load balancing loss. This explicitly promotes load balancing among experts, ensuring that all experts are fully activated and utilized during training.

[0054] Expert load balancing losses The core idea is to encourage that, within a training batch, the average probability (average weight) assigned to all experts should be as uniform as possible.

[0055] First, the model calculates within a batch (size: Each expert Average load undertaken (Average weight): (5) Then, a vector is constructed from the average load of all experts. By calculating vectors With an ideal uniform distribution vector The mean squared error (MSE) between the two values ​​is used to construct the load balancing loss. (6) By minimizing This prompts probabilistic gating networks to explore the possibility of allocating samples to different experts, ensuring that all experts can participate in the learning process. This not only maintains the diversity of expert functions but also promotes the division of labor and cooperation among experts, fully realizing the potential of the MoE model.

[0056] IV. Joint Optimization Strategy.

[0057] To achieve coordinated optimization of the three core components—heterogeneous representation learning, expert routing, and traffic inference—this embodiment employs an end-to-end training strategy, jointly optimizing all modules through a unified multi-objective loss function. The overall objective function is as follows: The system integrates three losses: main task regression loss, supervised comparison loss, and expert load balancing loss, which correspond to traffic inference accuracy, representation effectiveness, and routing balance, respectively, and are presented as follows: (7) in, The regression loss is used for the main task and measures the model's inference bias regarding OD flow. It is the supervised contrastive loss defined by formula (1), used to guide the learning of heterogeneous representations; This is the expert load balancing loss defined in formula (6), used to promote balanced route allocation. Hyperparameter and These are used to adjust the weights of the two auxiliary loss terms in the overall optimization objective.

[0058] Given that OD flow data typically exhibits a significant long-tail distribution, the traditional mean squared error (MSE) loss assigns excessive weight to the error of high-flow samples, causing the model to overfit a small number of high-intensity flows during the optimization process, while under-modeling the majority of low-intensity OD flows.

[0059] To enhance the model's robustness to extreme values, this embodiment uses the Huber loss function as... This loss function cleverly combines the advantages of MSE and MAE: when the absolute value of the inference error does not exceed a threshold When the error exceeds a certain threshold, a quadratic form is used to ensure gradient smoothness; when the error exceeds a certain threshold... When the flow rate is high, it transforms into a linear function, thus effectively suppressing the dominant role of high-flow samples in the overall loss.

[0060] The loss function is in a size of The calculation is performed on a batch basis, and its specific form is as follows: (8) in, This is the actual flow rate value after logarithmic transformation. These are the inferred values ​​corresponding to the model. It is an adjustable hyperparameter that controls the model's sensitivity to extreme bias samples.

[0061] This embodiment minimizes... This allows the model to maintain high fitting accuracy to typical flow patterns while effectively mitigating the optimization bias caused by the long-tailed distribution.

[0062] V. Model Training and Inference.

[0063] The model constructed in this application involves the collaborative optimization of three modules during training: heterogeneous representation learning, expert routing, and traffic inference. To clearly demonstrate this joint learning paradigm, the complete end-to-end training process is summarized in pseudocode, as shown in Table 1. Table 1 is as follows: Table 1. Pseudocode of the model training and inference framework

[0064] Table 1 shows a detailed description of the entire process from inputting OD feature data, dynamically allocating samples to expert subnetworks via a probabilistic gating network, performing differential inference in the expert network, and finally using the joint loss function for gradient backpropagation and parameter update.

[0065] The technical solution provided in this embodiment will be further explained below with reference to experiments.

[0066] In this embodiment, two representative OD flow datasets are selected as experimental data: intercity freight flow at the national scale (China Intercity Freight Dataset) and commuter flow at the city scale (New York City Commuter Dataset). These two datasets are preprocessed and used as input to the spatial heterogeneity adaptive learning system for OD flow inference (hereinafter referred to as the ALSH-OD model) provided in this embodiment. The spatial heterogeneity adaptive representation learning module acquires heterogeneity representations for the two datasets. Multiple independent expert networks in the heterogeneity-aware inference module perform parallel preliminary inferences on the OD data input. Then, based on the soft routing weights of each expert network, a weighted aggregation is performed to generate the final OD flow prediction value. Selecting two datasets at different scales—national and city—for experiments allows for the evaluation of the system's ability to model spatial heterogeneity and its generalization performance at different spatial scales.

[0067] Meanwhile, seven existing representative models were selected as baseline models and compared with the system provided in this embodiment: Theoretical models: gravity model (GM) and radiation model (RM).

[0068] Traditional machine learning models: Random Forest (RF) and Gradient Boosting Regression Tree (GBRT).

[0069] Deep learning models: Deep Gravity (DG), Physics-Guided Graph Neural Network (PG-MFG), and Hybrid Model Based on Community Detection (HMCG-LGBM).

[0070] This embodiment uses four indicators to comprehensively evaluate inference performance: Root Mean Square Error (RMSE): Used to evaluate the model's fitting accuracy to high-flow samples; Mean Absolute Error (MAE): Reflects the model's average inference bias on the overall sample; Commuter Common Component (CPC): Quantifies the degree of overlap between inferred and actual flow in spatial distribution; Spearman Correlation Coefficient (SCC): Evaluates the rank correlation between the inferred result and the actual value.

[0071] Specifically, the ALSH-OD model provided in this embodiment is implemented based on the PyTorch framework. An end-to-end joint optimization strategy is adopted, and the total loss function is set as a multi-objective loss function, which includes the main task regression loss (Huber Loss), supervised contrastive loss, and expert load balancing loss. The heterogeneous representation encoder is updated synchronously through the backpropagation algorithm. The network parameters of the heterogeneity perception inference module (including the probabilistic gating network and various expert networks) were determined. The specific experimental procedure is as follows: Data collection: Acquire intercity freight flow data and New York City commuter flow data at a national scale.

[0072] Data preprocessing: Logarithmic transformation is performed on the raw OD flow data to mitigate the influence of long-tail distribution, and node attribute features are standardized to form the China Intercity Freight Dataset and the New York City Commuter Dataset. Each record in the dataset represents an OD pair sample consisting of a starting point O and an ending point D. Each OD pair has its original OD features, which are used as raw input features and are fed into the ALSH-OD model for processing.

[0073] Heterogeneous representation learning: The heterogeneous encoder maps the input OD features into low-dimensional embedding vectors. Under the constraint of supervised contrastive loss, samples with similar flow patterns are brought closer to each other in the feature space, and heterogeneous representation vectors are output.

[0074] Soft route allocation: The probabilistic gating network dynamically generates probability weights (soft route weights) for each OD sample and assigns them to each expert network based on the learned heterogeneity representation.

[0075] Differentiated inference and aggregation: Each expert network performs flow prediction independently, and the final OD flow prediction result is obtained by weighting the outputs of each expert according to probability weights.

[0076] The experimental results are shown in Table 2, which is as follows: Table 2. Overall performance of model accuracy

[0077] In the table, * indicates that the p-value is less than 0.001.

[0078] As shown in Table 2, the experimental results demonstrate that the ALSH-OD model proposed in this embodiment achieves the best inference accuracy on both datasets. On the China Intercity Freight Dataset, ALSH-OD achieves an RMSE of 425.465 and a MAE of 75.740. Compared to the suboptimal model (HMCG-LGBM), the RMSE is reduced by 4.18%; compared to the commonly used Deep Gravity (DG) and PG-MFG models, the RMSE is significantly reduced by 36.8% and 32.6%, respectively. This result illustrates the significant advantage of the ALSH-OD model provided in this embodiment in handling large-scale, highly heterogeneous regional interaction tasks. On the New York City Commuter Dataset, ALSH-OD also performs excellently, with an RMSE reduced to 4.074, outperforming the strong GBRT model (RMSE=4.325), a reduction of approximately 5.8%. Simultaneously, it achieves an optimal CPC value of 0.786, indicating that the ALSH-OD model can accurately reconstruct the fine-scale urban commuter structure.

[0079] Therefore, ALSH-OD achieves consistent performance improvements on two real-world datasets with significant spatial heterogeneity: intercity freight in China and commuting in New York City. On the intercity freight dataset in China, ALSH-OD's RMSE is 32.6% lower than the current state-of-the-art baseline model, indicating its stronger ability to characterize complex spatial interaction patterns in tasks involving inter-regional flow inference with significant spatial heterogeneity. Further analysis shows that, despite not providing any prior labels regarding flow type or regional function during training, the multiple expert subnetworks automatically learned by the model exhibit semantically clear division of labor in spatial interaction patterns: some experts focus on high-intensity cross-regional freight flows, while others capture intra-city commuting, medium-sized inter-regional connections, or sparse long-distance interactions, respectively. This data-driven expert division of labor structure is highly consistent with known national logistics networks, urban agglomeration development patterns, and regional economic connections, indicating that the model can automatically distinguish different types of geographical interaction processes based on the inherent structure of the data. This characteristic not only verifies the model's ability to capture real-world spatial interaction mechanisms but also provides an interpretable analytical basis for applications such as urban functional area identification, refined traffic demand prediction, and regional collaborative development assessment.

[0080] In summary, this embodiment proposes a spatial heterogeneity adaptive learning system for OD flow inference, based on a joint learning framework of contrastive representation learning and expert perception (ALSH-OD), which adaptively discovers and models spatial heterogeneity from observation data. This approach theoretically expands the methodological path for spatial heterogeneity modeling: by introducing task-driven supervised contrastive learning, ALSH-OD reconstructs spatial heterogeneity into an intrinsic, end-to-end optimizable, learnable representation of OD flow data, thus avoiding dependence on external prior knowledge (such as distance decay assumptions in gravity models, network centrality indices, or community detection results). Simultaneously, leveraging the heterogeneity-aware soft routing mechanism provided by the spatial heterogeneity adaptive representation learning module, the model can perform data-driven soft partitioning of OD pairs based on interaction semantics and allocate them to multiple parameter-independent expert subnetworks. This allows different experts to focus on fitting specific types of local interaction patterns. This design not only alleviates the averaging bias of the global model on spatial heterogeneity but also achieves joint optimization of representation learning and flow inference, avoiding the target mismatch problem in traditional two-stage methods. This provides a feasible solution for constructing task-consistent, geographically interpretable intelligent modeling.

[0081] Based on the same inventive concept, this embodiment provides a spatial heterogeneity adaptive learning method for OD flow inference, which is executed by the system provided in any of the above embodiments, including: Obtain a dataset of OD pairs consisting of a starting point O and an ending point D, wherein the OD pairs contain the original input features; Spatial heterogeneity adaptive representation learning steps: Based on the original input features, through contrastive learning guided by the downstream OD flow inference task, output heterogeneous representations that can distinguish different spatial interaction modes; The system receives the original input features of the OD pair, maps the original input features into a heterogeneous representation vector using a heterogeneous representation encoder, and then processes the heterogeneous representation vector. Normalization processing; during the training phase, the heterogeneity representation encoder is optimized based on the contrastive learning objective; The inference steps of heterogeneity perception include: generating soft route weights based on the heterogeneity representation vector; multiple parallel expert networks receiving the original input features and outputting preliminary traffic prediction values; and weighting and summing the preliminary traffic prediction values ​​of each expert network according to the soft route weights to obtain the final OD traffic prediction value. End-to-end joint optimization steps: Calculate the main task regression loss, supervised comparison loss, and expert load balancing loss based on soft router weight distribution, and jointly optimize the network parameters.

[0082] In the above technical solution, the spatial heterogeneity adaptive representation learning step is used to adaptively discover potential spatial structure differences from observation data, avoiding dependence on preset spatial rules or external prior knowledge; the heterogeneity perception inference step is used to break through the limitations of traditional global parameter sharing models, enabling different regions or different OD pairs to correspond to different modeling methods, thereby improving the model's ability to express region-specific spatial interaction relationships; the end-to-end joint optimization step ensures that the learned (spatial) heterogeneity representation serves the OD flow inference target, reduces error propagation during the phased modeling process, and improves the accuracy and stability of the inference results.

[0083] The above-described solutions in this application can effectively characterize the spatial distribution characteristics of OD flow in complex spatial environments with multiple centers and heterogeneous functions, thereby improving the applicability and generalization ability of the OD flow inference method in different application scenarios.

[0084] This embodiment also provides a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method provided in the foregoing embodiments.

[0085] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A spatial heterogeneity adaptive learning system for OD flow inference, the system being used to process a dataset of OD pairs consisting of origin O and destination D, the system comprising: A spatial heterogeneity adaptive representation learning module is configured to receive the original input features of OD pairs, map the original input features into heterogeneous representation vectors using a heterogeneity representation encoder, and perform L2 normalization on the heterogeneous representation vectors. During the model training phase, the heterogeneity representation encoder is optimized based on the contrastive learning objective; The heterogeneity-aware inference module, which is connected to the spatial heterogeneity adaptive representation learning module, includes a probabilistic gating network and multiple parallel expert networks. The probabilistic gating network is configured to receive the normalized heterogeneity representation vector and generate soft routing weights corresponding to each expert network; each expert network is configured to receive the original input features and output the corresponding traffic prediction value; the heterogeneity-aware inference module is further configured to perform a weighted summation of the traffic prediction values ​​output by each expert network based on the soft routing weights to obtain the final OD traffic prediction value. The end-to-end joint optimization module is configured to utilize a multi-objective loss function to jointly optimize and train the network parameters of the spatial heterogeneity adaptive representation learning module and the heterogeneity-aware inference module.

2. The system according to claim 1, characterized in that, The spatial heterogeneity adaptive representation learning module includes a heterogeneity representation encoder, which is configured to perform the following operations: Receive the original input features of the OD pair, which include the node attribute features of the start point O and the end point D and the spatial relationship features between the start point O and the end point D; The original input features are mapped to a low-dimensional embedding space using a multilayer perceptron (MLP), thereby generating heterogeneous representation vectors for distinguishing different spatial interaction modes.

3. The system according to claim 1 or 2, characterized in that, The spatial heterogeneity adaptive representation learning module constructs a contrastive learning objective based on OD flow values, and the contrastive learning objective is configured as follows: Pseudo-labels for comparative learning are generated based on the actual flow values ​​of the OD pairs; Based on the pseudo-labels, the heterogeneity representation learning process is constrained by a supervised contrastive loss function, so that the representations of OD pairs with similar flow rates are close to each other in the embedding space, while the representations of OD pairs with large flow rate differences are far apart in the embedding space.

4. The system according to claim 3, characterized in that, The expression for the supervised contrastive loss function is as follows: , In the formula, This indicates that there is a supervised comparison of losses. This represents the set of sample indices for the current training batch. It is any sample in a certain training batch. Indicates sample The positive sample set, samples yes Any positive sample in, It is the set of all samples, including both positive and negative samples, where sample a is... Any sample in, , , Representing samples respectively ,sample ,sample Normalized heterogeneity representation vector, Represents the positive sample set Number of samples included This indicates that the cosine similarity between two vectors is calculated. It's a hyperparameter.

5. The system according to claim 1, characterized in that, The probability gating network is further configured as follows: A linear transformation layer is configured to receive the heterogeneity representation vector and generate a corresponding linear transformation. dimensional original fraction vector; The Softmax activation layer, connected to the linear transformation layer, is configured to activate the linear transformation layer. The original score vector is normalized to generate soft route weights representing the probabilities of each OD pair being assigned to each expert network. The expert network is a deep residual neural network structure.

6. The system according to claim 5, characterized in that, The heterogeneity-aware inference module further includes a load balancing constraint unit, which is configured to calculate the expert load balancing loss based on the soft route weights generated by the probabilistic gating network. The expert load balancing loss is constructed as a measure of the difference between the average weights assigned to all expert networks within a training batch and the ideal uniform distribution.

7. The system according to claim 6, characterized in that, The multi-objective loss function is a weighted sum of the main task regression loss, the supervised comparison loss, and the expert load balancing loss; The regression loss for the main task uses the Hubel loss function to measure the deviation between the predicted OD flow value and the actual flow value.

8. A spatial heterogeneity adaptive learning method for OD flow inference, characterized in that, The method is performed by the system described in any one of claims 1 to 7, and includes: Obtain a dataset of OD pairs consisting of a starting point O and an ending point D, wherein the OD pairs contain the original input features; The spatial heterogeneity adaptive representation learning module receives the original input features of the OD pair, maps the original input features into a heterogeneity representation vector using the heterogeneity representation encoder, and performs L2 normalization on the heterogeneity representation vector; during the model training phase, the heterogeneity representation encoder is optimized based on the contrastive learning objective. The heterogeneity-aware inference module includes a probabilistic gating network and multiple parallel expert networks. The probabilistic gating network is configured to receive a normalized heterogeneity representation vector and generate soft routing weights corresponding to each expert network. Each expert network is configured to receive the original input features and output a corresponding traffic prediction value. The heterogeneity-aware inference module is further configured to perform a weighted summation of the traffic prediction values ​​output by each expert network based on the soft routing weights to obtain the final OD traffic prediction value. The end-to-end joint optimization module calculates the main task regression loss, supervised comparison loss, and expert load balancing loss based on soft router weight distribution, and performs joint optimization of network parameters.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method of claim 8.