AOI Matching Method Based on Spatial Context Awareness and Adaptive Multimodal Feature Fusion
Patent Information
- Application Number
- CN202610758512.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-29
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2046-05-29
AI Technical Summary
[0005]有鉴于此,本申请实施例提供一种基于空间上下文感知与自适应多模态特征融合的AOI匹配方法,该方法可以解决名称缺失、轮廓粗糙等特殊数据和复杂异构场景下AOI匹配性能不佳的问题,提高AOI的匹配精度
[0011]在本申请实施例中,对来自两个不同地图平台的AOI分别进行几何特征提取和文本特征提取,得到各AOI各自对应的几何嵌入特征和文本嵌入特征;通过将来自两个不同地图平台的AOI及其邻域POI建模为包含多种节点类型与关系的异构空间图,并引入异构图Transformer模型对异构空间图进行上下文聚合,能够有效融合POI语义与AOI拓扑关系,弥补了单一特征表达能力的不足;进一步,基于每个AOI的几何嵌入特征、文本嵌入特征和空间上下文嵌入特征构建全局空间上下文向量,并通过上下文感知的GRU门控融合机制进行自适应去噪,可以得到高质量的多模态精炼特征;之后,通过源内注意力机制融合同一地图平台中AOI的多模态精炼特征,得到两种地图平台各自对应的源内嵌入特征,并通过跨源注意力机制进行双向交叉注意力计算,得到两种地图平台各自对应的源间嵌入特征,实现了多模态精炼特征的深度融合和跨源语义对齐;最后,基于两种地图平台各自对应的源间嵌入特征,以及候选AOI对的空间嵌入向量,确定候选AOI对的匹配概率,可以保证AOI的匹配结果符合地理空间一致性,实现精准匹配。该方法可以解决名称缺失、轮廓粗糙等特殊数据和复杂异构场景下AOI匹配性能不佳的问题,提高AOI的匹配精度。
Smart Images

Figure CN122310147B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of geographic information science and artificial intelligence technology, specifically to an AOI matching method based on spatial context awareness and adaptive multimodal feature fusion. Background Technology
[0002] Areas of Interest (AOIs), as the basic unit of digital urban spatial cognition and understanding, have become a core task in the fusion of multi-source geospatial data due to their accurate matching. Especially when integrating heterogeneous data sources such as commercial map platforms (e.g., Baidu Maps) and open community maps (e.g., Open Street Maps), AOI matching aims to align semantically or functionally equivalent AOIs across different systems. This is of great significance for building a unified urban information model, supporting key smart city applications (e.g., public facility planning, commercial site selection analysis, urban computing), and improving the accuracy of location services and the efficiency of emergency response.
[0003] However, multi-source geographic data have inherent differences in collection standards, modeling granularity, and semantic systems, posing three major challenges to AOI matching: (1) Geometric heterogeneity. Differences in coordinate systems and collection accuracy within the same map platform lead to significant inconsistencies in the shape, scale, or boundaries of the same entity, resulting in nonlinear offsets; (2) Semantic heterogeneity. Incompatibility in classification systems, naming language differences, and varying descriptive granularity among different data sources create semantic gaps; (3) Spatial distribution heterogeneity. Data coverage, update frequency, and density are highly uneven across different regions and platforms. These heterogeneities restrict the deep integration of multi-source geographic data and exacerbate the difficulty of matching.
[0004] Existing AOI matching methods integrate geometry and semantics, which significantly improves AOI matching performance. However, they still treat AOI as an isolated object. Furthermore, multimodal fusion often adopts static splicing or fixed weighting strategies, resulting in poor AOI matching performance when faced with special data such as missing names and rough outlines, as well as complex heterogeneous scenarios. The matching accuracy of AOI is difficult to guarantee. Summary of the Invention
[0005] In view of this, embodiments of this application provide an AOI matching method based on spatial context awareness and adaptive multimodal feature fusion. This method can solve the problem of poor AOI matching performance under special data such as missing names and coarse contours and complex heterogeneous scenarios, and improve the matching accuracy of AOI.
[0006] The AOI matching method based on spatial context awareness and adaptive multimodal feature fusion provided in this application includes: Geometric and textual features are extracted from AOIs from two different map platforms to obtain their respective geometric and textual embedding features. Each AOI is used as a graph node, and Points of Interest (POIs) are introduced as auxiliary graph nodes to construct a heterogeneous spatial graph. The edges of the heterogeneous spatial graph include containment edges from POIs to AOIs within the same map platform, adjacency edges between AOIs within the same map platform that satisfy the first spatial proximity condition, and candidate edges between candidate AOI pairs across map platforms. Candidate AOI pairs are two AOIs from different map platforms that satisfy the second spatial proximity condition. The heterogeneous graph Transformer model is used to perform context aggregation and feature updates on the heterogeneous spatial graph to obtain the spatial context embedding features of each AOI. A global spatial context vector is constructed based on the geometric embedding features, text embedding features, and spatial context embedding features of each AOI. Adaptive denoising is then performed using a context-aware GRU gated fusion mechanism to obtain multimodal refined features for each AOI. The multimodal refined features include geometric refined features, text refined features, and spatial context refined features. The multimodal refined features of AOIs from the same map platform are fused using an intra-source attention mechanism to obtain the intra-source embedding features corresponding to each of the two map platforms. Bidirectional cross-attention calculation is then performed using a cross-source attention mechanism to obtain the inter-source embedding features corresponding to each of the two map platforms. Based on the inter-source embedding features corresponding to each of the two map platforms and the spatial embedding vectors of the candidate AOI pairs, the matching probability of the candidate AOI pairs is determined.
[0007] Optionally, context aggregation and feature updates are performed on the heterogeneous spatial graphs to obtain the spatial context embedding features for each AOI, including: Based on the node type and the corresponding linear transformation matrix, the geometric embedding features and text embedding features of the target node in the heterogeneous spatial graph are linearly transformed to obtain the first query vector. The geometric embedding features and text embedding features of the source nodes adjacent to the target node in the heterogeneous spatial graph are also linearly transformed to obtain the first key vector and the first value vector. The target node can be any AOI in the heterogeneous spatial graph, and the linear transformation matrices corresponding to the first key vector and the first value vector are different. Based on the first key vector and the first query vector, and combined with the attention weight matrix associated with the edge type, the attention score is calculated. Based on the attention score, the first value vector is weighted, aggregated, and normalized to obtain the spatial context embedding features corresponding to the target node.
[0008] Optionally, a global spatial context vector is constructed based on the geometric embedding features, text embedding features, and spatial context embedding features of each AOI, and adaptive denoising is performed through a context-aware GRU gated fusion mechanism to obtain multimodal refined features for each AOI, including: The geometric embedding features, text embedding features, and spatial context embedding features of the reference AOI are averaged element-wise to obtain the global spatial context vector of the reference AOI; the reference AOI can be any one of multiple AOIs. The global spatial context vector of the reference AOI is concatenated with the features of a specific modality to obtain the corresponding concatenated vector; the features of the specific modality can be any one of the geometric embedding features, text embedding features, and spatial context embedding features of the reference AOI. The concatenated vector is input into the reset gate of the GRU to obtain the corresponding reset vector. The features of the specific modality and the corresponding reset vector are convolved and concatenated with the global spatial context vector of the reference AOI. The concatenated features are reconstructed using the tanh activation function to obtain the corresponding reconstructed features. The concatenated vector is input into the update gate of the GRU to obtain the corresponding update vector. The update vector is used to perform weighted fusion of the features of the specific modality and the reconstructed features to obtain the refined features of the specific modality of the reference AOI.
[0009] Optionally, the two different map platforms include a first map platform and a second map platform; bidirectional cross-attention calculation is performed through a cross-source attention mechanism to obtain the inter-source embedding features corresponding to each of the two map platforms, including: The source-internal embedding features of the first map platform are linearly transformed to obtain the second query vector. The source-internal embedding features of the second map platform are linearly transformed to obtain the second key vector and the second value vector. The first cross-attention feature is calculated. The first cross-attention feature is residually fused with the source-internal embedding features of the first map platform to obtain the source-internal embedding features of the first map platform. The source-embedded features of the second map platform are linearly transformed to obtain the third query vector. The source-embedded features of the first map platform are linearly transformed to obtain the third key vector and the third value vector. The second cross-attention feature is calculated. The second cross-attention feature is residually fused with the source-embedded features of the second map platform to obtain the source-inter-embedded features of the second map platform.
[0010] Optionally, based on the inter-source embedding features corresponding to the two map platforms and the spatial embedding vectors of the candidate AOI pairs, the matching probability of the candidate AOI pairs is determined, including: The source embedding features corresponding to the two map platforms are concatenated to obtain aligned depth features. Spatial metric vectors are constructed based on at least the intersection-union ratio, centroid distance, and area ratio of candidate AOI pairs, and linear transformation is performed to obtain spatial embedding vectors of candidate AOI pairs. The aligned depth features are concatenated with the spatial embedding vectors and input into a classifier composed of multilayer perceptrons. The matching probability of candidate AOI pairs is calculated by using the sigmoid activation function.
[0011] In this embodiment, geometric and textual features are extracted from AOIs from two different map platforms to obtain their respective geometric and textual embedding features. By modeling the AOIs from the two different map platforms and their neighboring POIs as a heterogeneous spatial graph containing multiple node types and relationships, and introducing a heterogeneous graph Transformer model to perform contextual aggregation on the heterogeneous spatial graph, the semantics of POIs and the topological relationships of AOIs can be effectively integrated, making up for the shortcomings of single feature representation. Furthermore, a global spatial context vector is constructed based on the geometric embedding features, textual embedding features, and spatial context embedding features of each AOI, and contextual awareness is used to... The method employs a known GRU-gated fusion mechanism for adaptive denoising, yielding high-quality multimodal refined features. Subsequently, an intra-source attention mechanism is used to fuse the multimodal refined features of AOIs from the same map platform, resulting in intra-source embedding features for each of the two map platforms. A cross-source attention mechanism is then used for bidirectional cross-attention calculation to obtain inter-source embedding features for each of the two map platforms, achieving deep fusion of multimodal refined features and cross-source semantic alignment. Finally, based on the inter-source embedding features for each of the two map platforms and the spatial embedding vectors of candidate AOI pairs, the matching probability of the candidate AOI pairs is determined, ensuring that the AOI matching results conform to geospatial consistency and achieving accurate matching. This method can address the poor AOI matching performance issues caused by special data such as missing names and coarse contours, as well as complex heterogeneous scenarios, thereby improving AOI matching accuracy. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort, wherein: Figure 1 A flowchart illustrating an AOI matching method based on spatial context awareness and adaptive multimodal feature fusion provided in this application embodiment; Figure 2 A schematic diagram illustrating a workflow for calculating modal refinement features using GRU, as provided in this application; Figure 3 This application provides a schematic diagram illustrating the process of calculating the first reference cross-attention feature corresponding to Baidu Maps. Figure 4 A schematic diagram of the processing flow of an AOI matching method based on spatial context awareness and adaptive multimodal feature fusion provided in this application; Figure 5 This is a schematic diagram of the experimental results of a module ablation provided in this application. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. The following embodiments are used to illustrate this application, but are not intended to limit the scope of this application. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0014] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0015] It should be noted that the terms "first, second, and third" used in the embodiments of the present invention are only used to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, and third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of the present invention described herein can be implemented in an order other than that illustrated or described herein.
[0016] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which the embodiments of this application pertain. It should also be understood that terms such as those defined in general dictionaries should be understood to have a meaning consistent with their meaning in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.
[0017] This application provides an AOI matching method based on spatial context awareness and adaptive multimodal feature fusion, such as... Figure 1 The diagram shown is a flowchart illustrating an AOI matching method based on spatial context awareness and adaptive multimodal feature fusion provided in this application embodiment. The method includes: S101. Perform geometric feature extraction and text feature extraction on AOIs from two different map platforms to obtain the geometric embedding features and text embedding features corresponding to each AOI.
[0018] In some embodiments, the two different map platforms can be two map platforms that need to match their AOIs (matching the AOI in one map platform with the AOI in another map platform). These two map platforms can be Baidu Maps, Open Street Map (OSM), or Amap. In this application, Baidu Maps and OSM will be used as the two map platforms in the following description.
[0019] In some embodiments, for each AOI from two different map platforms, multiple dimensions of geometric features of the AOI can be extracted in the EPSG:3857 coordinate system. These geometric features include 44 dimensions such as the perimeter, area, aspect ratio, compactness, shape index, center coordinates, adjacency relationship with other AOIs, and inclusion relationship of the AOI. Then, based on these geometric features, a corresponding geometric feature vector is constructed, and a linear transformation is performed on the geometric feature vector to obtain a geometric embedding feature of a specific dimension. The specific dimension is a unified dimension selected by researchers after comprehensive consideration of empirical rules, computational resource constraints, and experimental verification. This application uses a specific dimension of 128 for illustration. The geometric feature vector can be transformed using formula (1). A linear transformation is performed to obtain 128-dimensional geometric embedding features. : (1); in, Let be a learnable parameter, representing the geometric projection matrix. , Indicates geometric feature bias. , The 44-dimensional geometric feature vector is linearly combined and mapped to a 128-dimensional space. The feature information useful for AOI matching tasks is extracted by weighted summation. ReLU represents the activation function.
[0020] In some embodiments, for each AOI from two different map platforms, its textual attribute features, such as the semantic information of the AOI's name, address, and category, can be extracted, and these textual semantic information are concatenated into a string. Then, the concatenated string is input into the BERT-Base-Chinese model for pre-training encoding. After encoding, textual features representing global semantics are extracted, and the textual features are linearly transformed to obtain text embedding features of a specific dimension. For example, the textual features extracted by the BERT-Base-Chinese model can be processed using formula (2). A linear transformation is performed to obtain 128-dimensional text embedding features. : (2); in, This represents the string obtained by concatenating semantic information from the text. This represents the text features extracted by the BERT-Base-Chinese model. Here are the learnable parameters, representing the text projection matrix. , Indicates text feature bias. .
[0021] S102. Construct a heterogeneous spatial graph by using each AOI as a graph node and introducing points of interest (POIs) as auxiliary graph nodes.
[0022] It should be noted that the edges of the heterogeneous spatial graph include: containing edges pointing from POIs to AOIs within the same map platform, adjacent edges between AOIs within the same map platform that satisfy the first spatial proximity condition, and candidate edges between candidate AOI pairs across map platforms; a candidate AOI pair is two AOIs from different map platforms that satisfy the second spatial proximity condition. Specifically, AOIs satisfying the first spatial proximity condition can be edges between AOIs within the same map platform whose centroid distance is less than a first distance threshold. Two AOIs satisfying the second spatial proximity condition can be two AOIs from different map platforms whose spatial distance is within a certain range (e.g., centroid distance less than the second distance threshold, and the second distance threshold greater than the first distance threshold), or two AOIs from different map platforms whose spatial distance is within a certain range and whose AOI names are similar and whose categories are the same.
[0023] In some embodiments, edges are included to aggregate functional information within an AOI (such as a residential community containing multiple shops); adjacent edges are used to capture the external neighborhood structure of the AOI; candidate edges refer to edges that connect candidate AOI pairs (potential matching AOI pairs) across map platforms. Candidate edges serve as a bridge for cross-source information interaction, enabling the model to compare feature differences from heterogeneous data sources.
[0024] In some embodiments, a heterogeneous spatial graph can represent the spatial relationships between AOIs on different map platforms, and the node types in the graph include primary nodes. , and auxiliary nodes and , , These represent the area AOI entities of Baidu Maps and OSM, respectively. and These represent the regional POI entities of Baidu and OSM, respectively.
[0025] S103. Use the heterogeneous graph Transformer model to perform context aggregation and feature update on the heterogeneous spatial graph to obtain the spatial context embedding features of each AOI.
[0026] In some embodiments, during the process of context aggregation and feature update of heterogeneous spatial graphs using the Heterogeneous Graph Transformer (HGT) model, the HGT model aggregates information of different types of nodes (POIs, adjacent AOIs) in the neighborhood of each AOI node in the heterogeneous spatial graph through a type-aware message passing mechanism, thereby updating its own features to context-embedded features containing spatial topology.
[0027] S104. Construct a global spatial context vector based on the geometric embedding features, text embedding features, and spatial context embedding features of each AOI, and perform adaptive denoising through a context-aware GRU gated fusion mechanism to obtain the multimodal refined features of each AOI.
[0028] Among them, multimodal refined features include geometric refined features, textual refined features, and spatial context refined features.
[0029] In some embodiments, the geometric embedding features, text embedding features, and spatial context embedding features of each AOI can be element-wise averaged to obtain the global spatial context vector. Then, for each AOI, the geometric embedding features are concatenated with the global spatial context vector, and the concatenated vector is adaptively denoised and reconstructed using a context-aware GRU-gated fusion mechanism to obtain the geometrically refined features of the corresponding AOI; the text embedding features are concatenated with the global spatial context vector, and the concatenated vector is adaptively denoised and reconstructed using a context-aware GRU-gated fusion mechanism to obtain the textually refined features of the corresponding AOI; the spatial context embedding features are concatenated with the global spatial context vector, and the concatenated vector is adaptively denoised and reconstructed using a context-aware GRU-gated fusion mechanism to obtain the spatial context refined features of the corresponding AOI.
[0030] S105. By fusing the multimodal refined features of AOIs in the same map platform through the intra-source attention mechanism, the intra-source embedding features corresponding to the two map platforms are obtained. Then, by performing bidirectional cross-attention calculation through the cross-source attention mechanism, the inter-source embedding features corresponding to the two map platforms are obtained.
[0031] In some embodiments, an intra-source attention mechanism (such as a multi-head self-attention mechanism) can be used to interact with the multimodal refined features of each AOI in the same map platform to obtain a feature sequence after interaction between multiple AOI modalities. This feature sequence is then averaged along the modal dimension to obtain intra-source embedded features. Subsequently, the intra-source embedded features from different map platforms are linearly transformed into query vectors, key vectors, and value vectors. A cross-source attention mechanism is then used to perform bidirectional cross-attention calculation on the query vectors, key vectors, and value vectors to obtain the inter-source embedded features corresponding to each of the two map platforms.
[0032] S106. Based on the inter-source embedding features corresponding to the two map platforms and the spatial embedding vectors of the candidate AOI pairs, determine the matching probability of the candidate AOI pairs.
[0033] In some embodiments, for each candidate AOI pair, a spatial metric vector containing key indicators such as crossover ratio, centroid distance (after smoothing), and adjacency judgment is calculated based on its original coordinates, and a corresponding spatial embedding vector is obtained by linear projection. Finally, the matching probability of the candidate AOI pair is calculated based on the spatial embedding vector of the candidate AOI pair and the source embedding features corresponding to the two map platforms.
[0034] In this embodiment, geometric and textual features are extracted from AOIs from two different map platforms to obtain the geometric embedding features and textual embedding features corresponding to each AOI. By modeling the AOIs from the two different map platforms and their neighboring POIs as a heterogeneous spatial graph containing multiple node types and relationships, and introducing the HGT model to perform context aggregation on the heterogeneous spatial graph, the semantics of POIs and the topological relationships of AOIs can be effectively integrated, making up for the shortcomings of single feature representation. Furthermore, a global spatial context vector is constructed based on the geometric embedding features, textual embedding features, and spatial context embedding features of each AOI, and a context-aware GRU gate is used. An adaptive denoising mechanism using a controlled fusion method yields high-quality multimodal refined features. Subsequently, an intra-source attention mechanism is used to fuse the multimodal refined features of AOIs from the same map platform, resulting in intra-source embedding features for each of the two map platforms. A cross-source attention mechanism is then used for bidirectional cross-attention calculation to obtain inter-source embedding features for each of the two map platforms, achieving deep fusion of multimodal refined features and cross-source semantic alignment. Finally, based on the inter-source embedding features for each of the two map platforms and the spatial embedding vectors of candidate AOI pairs, the matching probability of the candidate AOI pairs is determined, ensuring that the AOI matching results conform to geospatial consistency and achieving accurate matching. This method can address the poor AOI matching performance issues caused by special data such as missing names and coarse contours, as well as complex heterogeneous scenarios, thereby improving AOI matching accuracy.
[0035] In some embodiments of this application, the step S103, "performing context aggregation and feature updates on the heterogeneous spatial graph to obtain the spatial context embedding features of each AOI," can be achieved through the following steps: Based on the node type and the corresponding linear transformation matrix, the geometric embedding features and text embedding features of the target node in the heterogeneous spatial graph are linearly transformed to obtain the first query vector. The geometric embedding features and text embedding features of the source nodes adjacent to the target node in the heterogeneous spatial graph are also linearly transformed to obtain the first key vector and the first value vector. Based on the first key vector and the first query vector, and combined with the attention weight matrix associated with the edge type, an attention score is calculated. The first value vector is then weighted, aggregated, and normalized based on the attention score to obtain the spatial context embedding features corresponding to the target node. The first query vector, the first key vector, and the first value vector have the same dimension.
[0036] In some embodiments, the target node is any AOI in the heterogeneous spatial graph, that is, all AOIs in the heterogeneous spatial graph will be used as the target node to calculate its corresponding spatial context embedding feature once. The source nodes adjacent to the target node may include other AOIs and / or POIs.
[0037] In some embodiments, the node types include AOI and POI types. The linear transformation matrices differ between different node types. Each node type includes a first linear transformation matrix for calculating the query vector, a second linear transformation matrix for calculating the key vector, and a third linear transformation matrix for calculating the value vector. The first, second, and third linear transformation matrices for nodes of the same type are different. The first, second, and third linear transformation matrices for two different node types are also different. Each linear transformation matrix can be determined during the HGT model training process.
[0038] In some embodiments, the linear transformation matrices corresponding to the first key vector and the first value vector are different. That is, a first linear transformation can be performed on the geometric embedding features and text embedding features of the source node to obtain the first key vector, and a second linear transformation can be performed on the geometric embedding features and text embedding features of the source node to obtain the first value vector. The difference between the first linear transformation and the second linear transformation is that the linear transformation matrices used are different.
[0039] In some embodiments, an attention score can be calculated based on a first key vector, a first query vector, and an attention weight matrix associated with the edge type. For example, the attention score can be calculated according to formula (3). : (3); in, , These represent the first and second hexagonal lines of the heterogeneous graph Transformer model. The key vector of each attention head (corresponding to the first key vector, based on the source node) (obtained by calculating geometric embedding features and text embedding features) and query vector (corresponding to the first query vector, based on the target node) (Calculated from geometric embedding features and text embedding features). Indicates transpose. These are learnable parameters representing the attention weight matrix associated with the edge type; that is, different edge types (including edges, adjacent edges, and candidate edges) correspond to different attention weight matrices. The dimension of the input features of the HGT model (total dimension, 128 in this application) represents the total dimension. These respectively represent the number of attention heads (in this application) ), This represents the normalized exponential function.
[0040] In some embodiments, the source nodes corresponding to the target node may include multiple ones, so the first key vector and the first value vector may also include multiple ones. For each source node of the target node, multiple attention scores can be obtained by formula (1). Then, the multiple attention scores and multiple first value vectors can be weighted and summed (that is, the attention scores and the corresponding first value vectors are multiplied and added together), and the result of the weighted summation is normalized to obtain the spatial context embedding features.
[0041] In some embodiments, the HGT model includes stacked... One HGT layer, target node In the Spatial context embedding features of layers By aggregating all its neighbors (source node) The first value vector of ) To perform the update, as shown in formula (4): (4); in, This represents the set of source nodes corresponding to the target node. Represents the target node In the The contextual embedding features of the layer. Formula (4) can be understood as first based on attention scores. The aggregated message is obtained by weighted summation, and then passed through a linear layer specific to the target node type. Then, after activation function Finally, it is combined with the context embedding features of the previous layer. Add them together to get the target node. In the Spatial context embedding features of layers In addition, it is also necessary to... Perform layer normalization to obtain the target node The final spatial context embedding features.
[0042] It should be noted that by stacking L HGT layers, each AOI can aggregate information from its multi-hop neighborhood, so that the final output context embedding features not only encode the features of the AOI itself, but also the structural information of its surrounding nodes, thus providing data support for subsequent accurate matching.
[0043] Understandably, based on the node type and the corresponding linear transformation matrix, the geometric and textual embedding features of the target node are linearly transformed into the first query vector, and the geometric and textual embedding features of the source node are linearly transformed into the first key vector and the first value vector. Based on the first key vector and the first query vector, and combined with the attention weight matrix associated with the edge type, an attention score is calculated. Based on the attention score, the first value vector is weighted, aggregated, and normalized, thereby capturing the structural information of the POI and AOI around the target node. This ensures that the final spatial context embedding features not only include the features of the AOI itself, but also the structural information of its surrounding nodes, providing rich data support for subsequent accurate matching.
[0044] In some embodiments of this application, step S104 can be implemented through the following steps: averaging the geometric embedding features, text embedding features, and spatial context embedding features of the reference AOI element-wise to obtain the global spatial context vector of the reference AOI; concatenating the global spatial context vector of the reference AOI with the features of a specific modality to obtain the corresponding concatenated vector; inputting the concatenated vector into the reset gate of the GRU to obtain the corresponding reset vector; convolving the features of the specific modality with the corresponding reset vector and concatenating them with the global spatial context vector of the reference AOI; reconstructing the concatenated features using the tanh activation function to obtain the corresponding reconstructed features; inputting the concatenated vector into the update gate of the GRU to obtain the corresponding update vector; and using the update vector to perform weighted fusion of the specific modality features and the reconstructed features to obtain the refined features of the specific modality of the reference AOI.
[0045] In some embodiments, the reference AOI is any one of multiple AOIs, meaning that for each AOI, its multimodal refined features need to be calculated. The geometric embedding features, text embedding features, and spatial context embedding features of the reference AOI have the same dimension. The elements at corresponding positions in the features (vectors) of the three modalities of the reference AOI can be averaged to generate the global spatial context vector of the reference AOI.
[0046] In some embodiments, a specific modality includes any one of three modalities: geometric, textual, and spatial context. Correspondingly, the features of the specific modality are any one of the geometric embedding features, textual embedding features, and spatial context embedding features of the reference AOI. Below, using the geometric modality as the specific modality, the calculation of the refined features of the geometric modality of the reference AOI is explained.
[0047] First, the global spatial context vector and geometric embedding features of the reference AOI are concatenated. Then, the concatenated vector is input into the reset gate of the GRU to obtain the corresponding reset vector, and simultaneously input into the update gate of the GRU to obtain the corresponding update vector. The geometric embedding features and the reset vector are convolved, and then concatenated with the global spatial context vector of the reference AOI. The concatenated features are then reconstructed using the tanh activation function to obtain the corresponding reconstructed features. Finally, the updated vector is used to perform a weighted fusion of the geometric embedding features and the reconstructed features to obtain the refined features of the geometric modality of the reference AOI. It should be noted that the calculation method for the refined features of other modalities is similar to this process, and will not be elaborated here.
[0048] For example, such as Figure 2 The diagram illustrates a workflow for calculating modal refinement features using GRU, as provided in this application. First, geometric embedding features are based on the reference AOI. Text embedding features and spatial context embedding features Constructing global space context vectors Global space context vector Used to guide feature denoising and reconstruction; then, the reset gate is based on features of a specific modality. ( Represents a specific mode, and global space context vector The concatenated vectors are used to learn element-wise weights to identify and suppress noise dimensions (such as missing semantic attributes in the map platform or coordinates affected by measurement errors), generating a denoised reset vector. Reset vector The calculation method is expressed by formula (5): (5); in, This represents the activation function. Indicates splicing, It resets the gate for a specific mode. The learnable weight matrix, It resets the gate for a specific mode. The bias.
[0049] After obtaining the reset vector Then, the tanh activation function is introduced to generate reconstructed features corresponding to specific modes. As shown in formula (6): (6); in, It is the tanh activation function with respect to a specific mode. The learnable weight matrix, It is the tanh activation function with respect to a specific mode. The bias, This represents element-wise multiplication. This step involves fusing the filtered reset vector. With global space context vector To reconstruct feature representations, it is possible to potentially fill in missing attributes (such as inferring category information based on spatial neighbors).
[0050] Furthermore, the update gate is based on features of a specific modality. and global space context vector The concatenated vectors are used to generate the update vector. This process is represented by formula (7): (7); in, It is an update gate with respect to a specific modality. The learnable weight matrix, It is an update gate with respect to a specific modality. The bias.
[0051] Finally, using the update vector Features of a specific mode and reconstructed features By fusing, a specific mode can be obtained. Refined characteristics This process is represented by formula (8): (8); The significance of formula (8) lies in the fact that it can update the vector Control the flow of information, if the characteristics of a specific modality reliable, If the features of a specific modality Noisy or missing Reconstruction features enhanced by context Replace it.
[0052] Understandably, by concatenating the global context vector of the reference AOI with the features of the specific modality to obtain a concatenated vector, and inputting it into the reset gate of the GRU to obtain the corresponding reset vector, the features of the specific modality and the corresponding reset vector are convolved and concatenated with the global context vector of the reference AOI. The concatenated features are then reconstructed using the tanh activation function to obtain the corresponding reconstructed features. The concatenated vector is then input into the update gate of the GRU to obtain the corresponding update vector. The update vector is then used to perform weighted fusion of the features of the specific modality and the reconstructed features to obtain the refined features of the specific modality of the reference AOI. This achieves dimensionality-level adaptive selection, which can suppress unreliable inputs while retaining effective features, providing accurate feature representations for subsequent matching tasks.
[0053] In some embodiments of this application, the two different map platforms include a first map platform and a second map platform; the step S106, "performing bidirectional cross-attention calculation through a cross-source attention mechanism to obtain the inter-source embedding features corresponding to the two map platforms," can be achieved through the following steps: The source-internal embedding features of the first map platform are linearly transformed to obtain the second query vector. The source-internal embedding features of the second map platform are linearly transformed to obtain the second key vector and the second value vector, and the first cross-attention feature is calculated. The first cross-attention feature is residually fused with the source-internal embedding features of the first map platform to obtain the source-internal embedding features of the first map platform. The source-internal embedding features of the second map platform are linearly transformed to obtain the third query vector. The source-internal embedding features of the first map platform are linearly transformed to obtain the third key vector and the third value vector, and the second cross-attention feature is calculated. The second cross-attention feature is residually fused with the source-internal embedding features of the second map platform to obtain the source-internal embedding features of the second map platform.
[0054] In some embodiments, after obtaining the multimodal refined features of each AOI, for the first map platform or the second map platform, the multimodal refined features of each AOI can be aggregated and layer-normalized to obtain a unified refined feature representation. This can be expressed by formula (9): (9); in, Representation layer normalization, These represent refined features for the geometric modality, text modality, and spatial context modality, respectively.
[0055] After achieving unified and refined characteristics Following this, multimodal refined features of AOIs from the same map platform are fused through an intra-source attention mechanism. This intra-source attention mechanism is implemented using a Transformer encoder, which employs a multi-head self-attention mechanism to achieve the fusion of intra-source multimodal refined features. For the first... Each attention point can be used to calculate the corresponding attention score using formula (10). : (10); in, and It is a learnable linear projection matrix. Indicates the first The query projection matrix corresponding to each attention head, Indicates the first The key projection matrix corresponding to each attention head. Indicates the first The projection matrix of the values corresponding to each attention head. This represents the number of feature dimensions for a single self-attention head.
[0056] The attention outputs of each self-attention head are then concatenated, linearly projected, residual connected, and layer normalized before being averaged and pooled to obtain the source-embedded features corresponding to the first map platform and the source-embedded features corresponding to the second map platform.
[0057] In some embodiments, the source-internal embedding features of the first map platform can be transformed into a second query vector using the linear projection matrix of the query vector; the source-internal embedding features of the second map platform can be transformed into a second key vector using the projection matrix corresponding to the key vector; and the source-internal embedding features of the second map platform can be transformed into a second value vector using the projection matrix corresponding to the value vector. Then, a first cross-attention feature is calculated based on the second query vector, the second key vector, and the second value vector. The first cross-attention feature is then residually fused with the source-internal embedding features of the first map platform to obtain the source-internal embedding features of the first map platform. Correspondingly, the source-internal embedding features of the second map platform can be transformed into a third query vector using the linear projection matrix of the query vector; the source-internal embedding features of the first map platform can be transformed into a third key vector using the projection matrix corresponding to the key vector; and the source-internal embedding features of the first map platform can be transformed into a third value vector using the projection matrix corresponding to the value vector. Then, a second cross-attention feature is calculated based on the third query vector, the third key vector, and the third value vector. The second cross-attention feature is then residually fused with the source-internal embedding features of the second map platform to obtain the source-internal embedding features of the second map platform.
[0058] The cross-source attention mechanism is implemented through a Transformer decoder, which uses a multi-head cross-attention mechanism to fuse refined features from multiple modalities within the source. The first reference cross-attention feature of Baidu Maps is calculated using a cross-attention head. The calculation process is represented by formula (11): (11); in, Let represent the learnable linear projection matrix, and let represent the th . The query projection matrix corresponding to each cross-attention head. Indicates the first The key projection matrix corresponding to each cross-attention head Indicates the first The projection matrix of the values corresponding to each cross-attention head. The feature dimension of a single cross-attention head. Represents the source embedding feature corresponding to OSM, symmetrically, for the i-th Cross-attention head, second reference cross-attention feature of OSM The calculation process is represented by formula (12): (12); in, This represents the source embedded features corresponding to Baidu Maps.
[0059] After calculating the first reference cross-attention features of Baidu Maps and the second reference cross-attention features of OSM using multiple cross-attention heads, the first reference cross-attention features can be concatenated to obtain the first cross-attention features of Baidu Maps. The first cross-attention features are then residually fused with the source-internal embedding features of Baidu Maps to obtain the source-internal embedding features of Baidu Maps. Simultaneously, the second reference cross-attention features are concatenated to obtain the second cross-attention features of OSM. The second cross-attention features are then fused with the source-internal embedding features of OSM to obtain the source-internal embedding features of OSM.
[0060] For example, when the first map platform is Baidu Maps and the second map platform is OSM, the process of calculating the first reference cross-attention feature corresponding to Baidu Maps is as follows: Figure 3 As shown. Source embedded features of Baidu Maps. Transform into a query vector through linear transformation Source embedding features of OSM Transform into a key vector through a linear transformation. Sum value vector Then query vector and key vector The attention weights are calculated by multiplying the transposes of the vectors and then normalized using a softmax layer. The normalized attention weights and value vectors are then combined. Multiplying these together yields the first reference cross-attention feature corresponding to Baidu Maps. .
[0061] It is understandable that by bidirectionally interacting with the source-embedded features of the first map platform and the source-embedded features of the second map platform, features with similar semantics but different distributions (such as the "shopping mall" label in the first map platform and the "retail" label in the second map platform) can be aligned. This can, to some extent, make up for the heterogeneity gap between data sources, so that the source-embedded features of the two map platforms can provide a reliable basis for subsequent AOI matching tasks and improve the matching accuracy of AOI.
[0062] In some embodiments of this application, step S106 can be implemented by the following steps: The source embedding features corresponding to the two map platforms are concatenated to obtain aligned depth features. Spatial metric vectors are constructed based on at least the intersection-union ratio, centroid distance, and area ratio of candidate AOI pairs, and linear transformation is performed to obtain spatial embedding vectors of candidate AOI pairs. The aligned depth features are concatenated with the spatial embedding vectors and input into a classifier composed of multilayer perceptrons. The matching probability of candidate AOI pairs is calculated by using the sigmoid activation function.
[0063] In some embodiments, for each candidate AOI pair, a spatial metric vector can be constructed based on the intersection-union ratio, centroid distance, and area ratio calculated in the EPSG:3857 coordinate system. This spatial metric vector is then projected into a spatial embedding vector through a fully connected layer. The spatial embedding vector is concatenated with aligned depth features, and a classifier and a sigmoid activation function are used to calculate the matching probability of the corresponding candidate AOI pair. Further, if the matching probability of a candidate AOI pair is less than a matching probability threshold (e.g., 0.7, where the threshold can be selected based on the F1 score on the validation set, choosing the threshold with the highest F1 score), the candidate AOI pairs are considered to be different geographic entities. Conversely, if the matching probability of a candidate AOI pair is greater than or equal to the matching probability threshold, the candidate AOI pairs can be considered to be the same geographic entity, thus allowing the data of the candidate AOI pairs to be fused.
[0064] For example, such as Figure 4 The diagram shown illustrates the processing flow of the AOI matching method based on spatial context awareness and adaptive multimodal feature fusion provided in this application. The flow includes: Step 1: For each AOI in Baidu Maps and OSM, calculate the corresponding geometric embedding features. Text embedding features and spatial context embedding features .
[0065] Step 2: Based on geometric embedding features Text embedding features and spatial context embedding features Constructing global space context vectors .
[0066] Step 3: Calculate the multimodal refined features for each AOI using the GRU's reset and update gates, and then fuse the multimodal refined features to obtain a unified refined feature. .
[0067] Step 4: Utilize the in-source attention mechanism and refined features corresponding to each AOI within Baidu Maps. The source embedding features of Baidu Maps were obtained. ; and through the source attention mechanism and the refined features corresponding to each AOI in OSM. Obtain the source embedding features of OSM .
[0068] Step 5: Utilize cross-source attention mechanism based on source-embedded features from Baidu Maps. Source embedding features of OSM Bidirectional cross-attention calculation is performed to obtain the inter-source embedding features of Baidu Maps. and the inter-source embedding features of OSM .
[0069] Step 6: Embed the source-to-source features of Baidu Maps Inter-source embedding features of OSM Aggregate and combine the spatial embedding vectors of candidate AOI pairs. (Based on the IoU, centroid distance, and area ratio of candidate AOI pairs, etc.), predict the matching probability of candidate AOI pairs.
[0070] The method provided in this application uses an end-to-end training approach. The model includes a geometric feature extraction layer (performing linear transformations on geometric feature vectors), a text feature projection layer (performing linear transformations on text features), HGT, a GRU-gated fusion module (context-aware GRU-gated fusion mechanism), an in-source attention layer (Transformer encoder layer), a cross-source attention layer (bidirectional cross-attention), a spatial constraint projection layer (constructing spatial metric vectors based on the intersection-union ratio, centroid distance, and area ratio of candidate AOI pairs and performing linear transformations), and a classifier. The network parameters of each module are obtained through joint training. The trained network parameters include at least the parameters in formula (1). and In formula (2) and In formula (3) In formula (5) and In formula (6) , In formula (7) and In formula (10) and In formula (11) .
[0071] The training data consists of AOI data from Baidu Maps and OSM in multiple cities, covering major categories such as commercial, residential, transportation, and public services. For each Baidu Maps AOI, spatial indexing is used to find OSM AOIs within 500 meters of its centroid, forming matching entity pairs. If two AOIs represent the same geographic entity, they are marked as positive samples. Otherwise, it is a negative sample. 70% of the data is used for training, 15% for validation, and 15% for testing, ensuring that data from the same city does not cross sets to avoid overfitting caused by spatial autocorrelation.
[0072] The batch size is set to 128, with each batch containing 128 entity pairs (each entity pair contains one Baidu AOI and one OSM AOI). During training, a batch of data is randomly selected from the training dataset and input into the geometric feature extraction layer, text feature projection layer, and HGT model in parallel for feature extraction. The output prediction probability is then processed by a GRU gated fusion module, an intra-source attention layer, a cross-source attention layer, a spatial constraint projection layer, and a classifier. The weighted cross-entropy loss between the predicted probability of the current batch and the true label is calculated, and the error gradient is backpropagated to all learnable network parameters using a chain rule. The maximum number of training epochs is set to 100, and an early stopping mechanism is introduced. At the end of each epoch, the model weights are frozen, and inference is performed on an independent validation set. The F1 score is calculated. If the validation F1 score does not improve within 10 consecutive epochs, training is terminated early to prevent overfitting.
[0073] In multi-source AOI matching tasks, the number of non-matching entity pairs (negative samples) in the real world is usually much larger than the number of matching entity pairs (positive samples), resulting in significant class imbalance in the data distribution. To prevent the model from being dominated by the majority class during gradient updates, this application uses a weighted cross-entropy loss function as the optimization objective function. For a given mini-batch sample set, the loss function... As shown in the following formula (13): (13); in, Indicates batch size. Indicates the first The true label of each entity pair; This represents the predicted probability output by the model. and This is the category penalty weight, whose value is inversely proportional to the ratio of positive to negative samples in the training set. In this application, it is taken as... , These represent the number of negative samples and the number of positive samples, respectively. By introducing these weights, the model's penalty for misclassification of minority class (positive samples) is dynamically increased, thereby optimizing the model's recall capability under highly imbalanced distributions.
[0074] The hyperparameter configuration for model training comprehensively considered network representation capacity, convergence stability, and hardware resources. All experiments were deployed on a single NVIDIA A100 (80GB) GPU, with a fixed random seed of 42 to ensure reproducibility. For optimization, the AdamW optimizer was chosen. Compared to the standard Adam, it decouples weight decay from gradient updates, effectively improving the generalization ability of deep multimodal networks containing Transformer structures. The initial learning rate was set to... Weight decay is set to This combination ensures smooth fine-tuning of the pre-trained language model while driving rapid convergence of the graph structure initialized from scratch. Furthermore, the batch size of the training data is set to 128, a value that fully utilizes the A100's memory throughput while ensuring the unbiasedness and smoothness of the mini-batch gradient estimation.
[0075] To comprehensively evaluate the effectiveness of the AOI matching method based on spatial context awareness and adaptive multimodal feature fusion proposed in this application, city-level AOI and POI data from two typical map platforms were selected: one is Baidu Maps, a commercial map service provider, and the other is OSM, driven by an open community. The data covers an area of approximately 1200 square kilometers in the central urban area of Shanghai, China (boundary box: 31.0°~31.5°N, 121.0°~122.0°E), an area with highly dense urban functions and complex spatial structure. Due to the significant differences between the two platforms in data collection methods, spatial accuracy, semantic labeling systems, and update frequencies, a typical multi-source heterogeneous spatial data fusion scenario is formed, providing an ideal experimental field for verifying the matching ability of the algorithm in a real and complex environment.
[0076] In August 2025, this experiment obtained a vector dataset containing 8195 AOIs and an auxiliary dataset containing 1,345,532 POIs through the Baidu Maps API, covering 12 primary categories including commercial, residential, transportation hubs, and green spaces. Simultaneously, data for the corresponding areas was downloaded from OSM and cleaned, ultimately yielding 22,564 AOIs and 35,900 POIs, demonstrating the richness and contribution of the open-source community in topological representation. Table 1 summarizes the key statistical characteristics of each dataset.
[0077] Table 1 Key statistical features of the dataset ; In the data preprocessing stage, the raw OSM data is first parsed: Polygon or Point geometric objects are constructed based on the node dictionary, and their geometric validity is verified using the Shapely library; simultaneously, entities with missing names or category labels are removed (statistics show that approximately 40% of AOIs in OSM have missing name information). For Baidu Maps data, the points_zh field is parsed to generate Polygons, and all coordinates are uniformly converted to the WGS84 geographic coordinate system to ensure cross-platform spatial consistency. Under a unified experimental setting, the AOI matching method based on spatial context awareness and adaptive multimodal feature fusion proposed in this application is systematically compared with five current mainstream AOI matching and geospatial entity alignment methods (CrossGeo, GAMM, STGEM, DeepMatcher, and MultiModal-EM). All models are trained and tested on the same Shanghai AOI dataset, using a consistent data partitioning ratio (training set:validation set:test set = 8:1:1) and preprocessing workflow to ensure the fairness and comparability of the experimental results. F1 score, precision, and recall are the evaluation metrics. The experimental results of each method on the test set are shown in Table 2.
[0078] Table 2 F1 Improvements Compared to SOTA ; As shown in Table 2, the method provided in this application outperforms the corresponding evaluation metrics of existing methods in all three metrics. The method achieved an F1 score of 0.9837 on the test set, representing an improvement of 5.22 to 5.92 percentage points compared to the baseline model STGEM (F1=0.9315). The method provided in this application achieves a precision of 98.49% and a recall of 98.26%, indicating accurate identification of true AOI matching pairs. In contrast, existing methods struggle to balance precision and recall. For example, while DeepMatcher and STGEM have high recall (greater than 98.7%), their precision only remains in the 87%–88% range, reflecting a tendency to over-predict positive examples, leading to many non-matching pairs being incorrectly identified as matches. GAMM and MultiModal-EM, while slightly superior in precision, sacrifice some recall. The method in this application significantly reduces both types of errors, demonstrating excellent matching performance.
[0079] Experimental results show that the AOI matching method based on spatial context awareness and adaptive multimodal feature fusion provided in this application not only achieves a breakthrough in F1 score in multi-source heterogeneous AOI matching tasks, but also achieves a balance between accuracy and recall. This method can improve the matching accuracy of AOI.
[0080] To thoroughly evaluate the generalization ability of the proposed framework under different geographical environments, this application further conducted cross-city migration experiments. The model was first trained on the Shanghai dataset, which has a complex urban structure and high AOI density, and then directly applied to the test sets of Chengdu and Beijing, two target cities, without any fine-tuning. This setup requires the model to not only cope with the significant differences in spatial layout and functional area distribution among different metropolitan areas, but also to overcome the heterogeneity brought about by cultural and regional factors such as place name naming habits and semantic expression methods. The cross-city generalization performance indicators are shown in Table 3.
[0081] Table 3. Cross-city generalization performance indicators ; As shown in Table 3, despite unavoidable domain shifts, the model maintains excellent matching performance in the target cities, achieving F1 scores of 0.9631 and 0.9251 in Chengdu and Beijing, respectively, only slightly lower than in Shanghai. This result demonstrates that the AOI matching method proposed in this application, based on spatial context awareness and adaptive multimodal feature fusion, does not overfit to the specific absolute coordinates or local keywords of Shanghai. Instead, it successfully captures a more universal matching logic, namely, modeling the relative spatial relationships between AOIs through a heterogeneous graph Transformer and achieving cross-platform semantic tag alignment using an attention mechanism.
[0082] To verify the effectiveness of the proposed AOI matching method based on spatial context awareness and adaptive multimodal feature fusion, and to deeply analyze the contribution of each component to the overall performance, ablation experiments were conducted from two dimensions: input feature composition and model architecture design. All experiments were performed on the aforementioned Shanghai municipal dataset, using a unified data partitioning and training process.
[0083] First, we examined the effects of three types of feature inputs: geometric features (Geo), semantic features (Text), and spatial topological features (HGT) based on heterogeneous graph Transformer modeling. The results of the feature-level ablation experiment are shown in Table 4.
[0084] Table 4. Results of Feature-Level Ablation Experiments ; As shown in Table 4, the model achieves optimal performance when all three features are input together. Removing any one feature significantly reduces performance. Using only geometric and textual features (with HGT disabled), the F1 score drops to 0.9591, indicating that ignoring high-order spatial structure information weakens the model's understanding of complex urban layouts. Retaining geometric and HGT features but removing textual features results in an F1 score of 0.9774, demonstrating that while geometric shapes and spatial relationships support strong matching capabilities, the lack of semantic alignment limits the upper limit of accuracy. Relying solely on textual features and HGT yields an F1 score of 0.9757, slightly lower than Geo+HGT, reflecting that geometric consistency remains a key criterion in cross-platform AOI matching.
[0085] The results above show that the three types of features are significantly complementary. Geometric features provide basic spatial constraints, text features support alignment at the category and name levels, while HGT captures the spatial context structure by modeling the relative positions and adjacency relationships between AOIs. Only when the three work together can high-precision matching be achieved.
[0086] Furthermore, key modules within the model were ablated to evaluate their design value. The module ablation experiment results are as follows: Figure 5 As shown.
[0087] like Figure 5 (Where the horizontal axis represents the F1 score and the vertical axis represents the modules within the ablation experiment model.) As shown, the initial baseline model, using only an encoder structure, achieved an F1 score of 0.9650, indicating that relying solely on the original feature representation is insufficient to fully capture the semantic and spatial relationships required for cross-platform matching. Building upon this, a GRU-gated fusion module was first introduced, improving the F1 score to 0.9742, an increase of 0.92%, demonstrating that this gating structure effectively adjusts the weights of multi-source feature fusion and enhances the model's ability to suppress noise. Subsequently, an intra-source attention layer was added, further improving the F1 score to 0.9786 (an increase of 0.44%), validating its positive role in enhancing the contextual consistency within a single-platform AOI. Further introduction of a cross-source attention layer resulted in an F1 score of 0.9812, an improvement of 0.26%, indicating that cross-platform semantic alignment highly depends on a bidirectional interaction mechanism, which can significantly improve the matching accuracy between heterogeneous data. Finally, introducing a spatial constraint projection layer into the complete model improved the F1 score to 0.9837 (a 0.25% increase), making it the largest contributor among all modules and highlighting the fundamental role of spatial consistency priors in geographic entity matching tasks. Furthermore, removing all attention mechanisms reduced the F1 score to 0.9658, still better than most baseline methods, but significantly lower than the complete model, further confirming the central role of attention mechanisms in modeling cross-source semantic interactions.
[0088] In summary, this ablation experiment clearly reveals the cumulative contribution path of each module in the model: from basic encoding to gated fusion, and then to the synergistic optimization of hierarchical attention mechanisms and spatial constraints, a high-performance AOI matching framework is jointly constructed. It should be noted that, depending on implementation needs, the steps described in the embodiments of this application can be broken down into more steps, or two or more steps or parts of steps can be combined into new steps to achieve the purpose of the embodiments of this application.
[0089] The methods described in the embodiments of this application can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code downloaded over a network that is originally stored in a remote recording medium or a non-transitory machine-readable medium and will be stored in a local recording medium. Thus, the methods described herein can be processed by software stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components (e.g., RAM, ROM, flash memory, etc.) capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods described herein. Furthermore, when a general-purpose computer accesses code used to implement the methods shown herein, the execution of the code transforms the general-purpose computer into a dedicated computer for executing the methods shown herein.
[0090] Those skilled in the art will recognize that the units and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this application.
[0091] The above embodiments are only used to illustrate the embodiments of this application, and are not intended to limit the embodiments of this application. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of this application. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of this application, and the patent protection scope of the embodiments of this application should be defined by the claims.
Claims
1. An AOI matching method based on spatial context awareness and adaptive multimodal feature fusion, characterized in that, include: Geometric and textual features were extracted from AOIs from two different map platforms to obtain the geometric and textual embedding features corresponding to each AOI. Each AOI is used as a graph node, and points of interest (POIs) are introduced as auxiliary graph nodes to construct a heterogeneous spatial graph. The edges of the heterogeneous spatial graph include the containing edges of POIs pointing to AOIs within the same map platform, the adjacency edges between AOIs within the same map platform that satisfy the first spatial proximity condition, and the edges between candidate AOI pairs across map platforms. Candidate AOI pairs are two AOIs from different map platforms that satisfy the second spatial proximity condition; The heterogeneous graph Transformer model is used to perform context aggregation on the heterogeneous spatial graph to obtain the spatial context embedding features of each AOI. A global spatial context vector is constructed based on the geometric embedding features, text embedding features, and spatial context embedding features of each AOI. Adaptive denoising is then performed using a context-aware GRU gated fusion mechanism to obtain multimodal refined features for each AOI. This process includes: element-wise averaging of the geometric embedding features, text embedding features, and spatial context embedding features of the reference AOI to obtain the global spatial context vector of the reference AOI; the reference AOI can be any one of multiple AOIs; and concatenating the global spatial context vector of the reference AOI with features of a specific modality to obtain the corresponding concatenated vector; the features of the specific modality are the geometric embedding features, text embedding features, and spatial context embedding features of the reference AOI. The feature can be any one of the embedded features and spatial context embedded features; the concatenated vector is input into the reset gate of the GRU to obtain the corresponding reset vector; the feature of the specific modality and the corresponding reset vector are convolved and concatenated with the global spatial context vector of the reference AOI; the concatenated feature is reconstructed through the tanh activation function to obtain the corresponding reconstructed feature; the concatenated vector is input into the update gate of the GRU to obtain the corresponding update vector; the feature of the specific modality and the reconstructed feature are weighted and fused using the update vector to obtain the refined feature of the specific modality of the reference AOI; among which, the multimodal refined features include geometric refined features, text refined features and spatial context refined features; By fusing multimodal refined features of AOIs from the same map platform using an attention mechanism, source-internal embedding features corresponding to each of the two map platforms are obtained. Furthermore, bidirectional cross-attention calculation is performed using the attention mechanism to obtain source-internal embedding features corresponding to each of the two map platforms. The two map platforms include a first map platform and a second map platform. The process of obtaining source-internal embedding features for each map platform through bidirectional cross-attention calculation includes: linearly transforming the source-internal embedding features of the first map platform to obtain a second query vector; linearly transforming the source-internal embedding features of the second map platform to obtain a second key vector and a second value vector, and calculating a first cross-attention feature; performing residual fusion of the first cross-attention feature and the source-internal embedding features of the first map platform to obtain the source-internal embedding feature of the first map platform; linearly transforming the source-internal embedding features of the second map platform to obtain a third query vector; linearly transforming the source-internal embedding features of the first map platform to obtain a third key vector and a third value vector, and calculating a second cross-attention feature; and performing residual fusion of the second cross-attention feature and the source-internal embedding features of the second map platform to obtain the source-internal embedding feature of the second map platform. Based on the inter-source embedding features corresponding to the two map platforms and the spatial embedding vectors of the candidate AOI pairs, the matching probability of the candidate AOI pairs is determined; wherein, the spatial embedding vectors are constructed based at least on the intersection-union ratio, centroid distance and area ratio of the candidate AOI pairs.
2. The method according to claim 1, characterized in that, Context aggregation is performed on the heterogeneous spatial graphs to obtain the spatial context embedding features of each AOI, including: Based on the node type and the corresponding linear transformation matrix, the geometric embedding features and text embedding features of the target node in the heterogeneous spatial graph are linearly transformed to obtain the first query vector. The geometric embedding features and text embedding features of the source nodes adjacent to the target node in the heterogeneous spatial graph are linearly transformed to obtain the first key vector and the first value vector. The target node is any AOI in the heterogeneous spatial graph, and the linear transformation matrices corresponding to the first key vector and the first value vector are different. The attention score is calculated based on the first key vector and the first query vector, combined with the attention weight matrix associated with the edge type. The first value vector is weighted, aggregated, and normalized based on the attention score to obtain the spatial context embedding features corresponding to the target node.
3. The method according to claim 1, characterized in that, Based on the inter-source embedding features corresponding to the two map platforms and the spatial embedding vectors of candidate AOI pairs, the matching probability of candidate AOI pairs is determined, including: The source embedding features corresponding to the two map platforms are concatenated to obtain the aligned depth features; A spatial metric vector is constructed based at least on the intersection-union ratio, centroid distance, and area ratio of candidate AOI pairs, and a linear transformation is performed to obtain the spatial embedding vector of the candidate AOI pairs. The aligned deep features are concatenated with the spatial embedding vector and input into a classifier composed of multilayer perceptrons. The matching probability of candidate AOI pairs is calculated by using the sigmoid activation function.
Citation Information
Patent Citations
Map interest point query method and device, equipment, storage medium and program product
CN114329244A
Visual language navigation method and system based on collaborative alignment and adaptive fusion
CN120427010A