Cross-modal geographic interest point matching method based on fine-grained geographic semantics and comparative learning

By employing a method based on fine-grained geographic semantics and contrastive learning, the problems of insufficient user fuzzy queries and multimodal associations in geographic information matching are addressed, achieving accurate matching of cross-modal geographic information and improving the efficiency and accuracy of geographic information processing.

CN122045321APending Publication Date: 2026-05-15SHANGHAI NORMAL UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI NORMAL UNIVERSITY
Filing Date
2026-01-28
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

In existing technologies, geographic information matching suffers from problems such as insufficient fine-grained correlation between user fuzzy queries and multimodal geographic data, and weak cross-modal alignment mechanisms, making it difficult to achieve accurate cross-modal geographic information matching.

Method used

We employ a method based on fine-grained geographic semantics and contrastive learning. We extract features through text encoders and geographic encoders, combine dual-mode negative sampling and InfoNCE contrastive learning strategies, and use unimodal prediction confidence to perform cross-modal contrastive learning to establish alignment relationships between text and geographic semantic features. We also introduce geographic context aggregation mechanisms and multimodal interaction to achieve alignment and fusion of multimodal features.

Benefits of technology

It improves the efficiency of geographic information processing, solves the problems of cross-modal representation mismatch and fuzzy matching of user queries, achieves accurate alignment of geographic-related representations under different modalities, and improves the accuracy and efficiency of geographic information matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122045321A_ABST
    Figure CN122045321A_ABST
Patent Text Reader

Abstract

The invention relates to a cross-modal geographic interest point matching method based on fine-grained geographic semantics and comparative learning, which comprises the following steps: constructing fine-grained geographic semantics representation through multi-scale feature extraction of two-dimensional geographic coordinates in combination with embedding of geographic spatial semantics; a dual-mode negative sampling InfoNCE loss function is introduced, and joint optimization is carried out on representation alignment of text-geographic modals in a pre-training stage; through multi-modal comparison optimization, the discrimination of the negative sample is adjusted and enhanced; according to the method, a cross-modal decoupling-interaction mechanism is put forward, under the supervision of single-modal prediction, single-modal representation is aligned with representation which is verified to be effective through single-modal prediction through a multi-modal comparison method, and decoupling reasoning and calibration of text semantics and geographic space relation data are achieved. Compared with the prior art, the method has the advantages that the geographic semantic modeling granularity is finer, the cross-modal alignment mechanism is more perfect and efficient, and the query-interest point matching capability of the pre-training language model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of geographic multimodal pre-trained models, and in particular to a cross-modal geographic point of interest matching method based on fine-grained geographic semantics and contrastive learning. Background Technology

[0002] In the digital age, geographic information has become a bridge connecting the physical and digital worlds. Geographic information matching enables the association calculation between user queries and points of interest in geospatial databases, serving as a fundamental support for location services, intelligent navigation, and spatial data analysis.

[0003] A typical geographic information matching process consists of three key stages: a feature extraction layer that extracts geometric and attribute features from the raw data, a similarity calculation layer that calculates the similarity between the two types of features, and a decision fusion layer that generates the final result through weighted fusion or machine learning.

[0004] Traditional geographic information matching methods revolve around geometric features, attribute features, and the fusion of the two: In geometric feature methods, coordinate and distance matching is achieved by comparing coordinates or calculating spatial distances; topological relationship matching utilizes constraints such as adjacency and inclusion to match road networks and administrative regions; shape feature matching extracts descriptors such as contours for parcel matching; attribute feature methods focus on non-spatial attributes, solving place name matching through string algorithms, or improving accuracy by combining multiple attribute categories and attribute combinations; fusion methods overcome the limitations of single features through weighted comprehensive scores or rule base reasoning.

[0005] Furthermore, with the rise of probabilistic and machine learning methods, probabilistic models such as Bayesian methods quantify matching probabilities, while traditional machine learning methods such as support vector machines rely on training samples to achieve automated matching. In terms of deep learning, word2vec combined with relevant models improves the matching effect of unstructured addresses. For example, in deep learning, some studies have improved the estimation accuracy and consistency with the true value by combining POI semantic embedding, which integrates spatial co-occurrence and classification semantics, with LSTM-attention-MLP models, and can also combine domain knowledge to optimize the spatial annotation of geographic images. Multimodal learning improves matching performance by fusing multi-source data.

[0006] Despite significant progress in existing research, the diversity, distribution, and heterogeneity of geospatial information make accurate matching of geographic information difficult. Current geographic information matching technologies face challenges such as insufficient fine-grained correlation between user fuzzy queries and multimodal geographic data, and weak cross-modal alignment mechanisms.

[0007] Chinese invention patent CN120670635A discloses a multimodal natural language understanding and generation system and method, including: constructing a cross-modal pre-training module, training a multimodal encoder, and establishing a cross-modal associative mapping space; performing mixed prompt fine-tuning and constructing a cloze test template; extracting user multi-turn dialogue intent representations based on an intent reasoning network and retrieving external knowledge bases for fine-grained reasoning; constructing a unified semantic representation framework, embedding text, images, and speech into a unified space, and generating multimodal intent-aware query vectors; and generating entity-level multimodal responses based on a key-value memory-based knowledge query module, optimizing the semantic understanding and generation capabilities of the dialogue model. This invention improves multimodal information understanding and generation capabilities, achieves deep association and understanding of image and text information, enhances downstream task adaptability, improves task completion accuracy and efficiency, and realizes unified semantic representation of multimodal information, providing support for information retrieval and utilization. However, standard MLPs still suffer from problems such as spectral bias due to approximating low-frequency functions, difficulty in capturing high-frequency details, insufficient fine-grained correlation of multimodal geographic data, cross-modal representation mismatch, weak alignment mechanisms, and fuzzy user queries that make it difficult to accurately match geographic information.

[0008] In summary, there is currently a lack of a cross-modal geographic interest point matching method based on fine-grained geographic semantics and contrastive learning to solve or partially solve the above problems. Summary of the Invention

[0009] The purpose of this invention is to overcome the shortcomings of the existing technology by providing a cross-modal geographic point of interest matching method based on fine-grained geographic semantics and contrastive learning. This method aims to solve or partially solve the problems of standard MLP, such as spectral bias due to approximate low-frequency functions, difficulty in capturing high-frequency details, insufficient fine-grained correlation of multimodal geographic data, cross-modal representation mismatch, weak alignment mechanism, and difficulty in accurately matching geographic information due to fuzzy user queries.

[0010] The objective of this invention can be achieved through the following technical solutions: This invention provides a cross-modal geographic interest point matching method based on fine-grained geographic semantics and contrastive learning, specifically including: S1. For the acquired text data and geographic data, extract text semantic features and geographic semantic features respectively through text encoder and geographic encoder; S2. Utilizing the aforementioned textual semantic features and geographic semantic features, cross-modal contrastive learning is conducted by combining a dual-mode negative sampling InfoNCE contrastive learning strategy and multimodal interaction to establish an alignment relationship between textual semantic features and geographic semantic features; S3. Use unimodal prediction confidence as a weak supervision signal to supervise cross-modal alignment and achieve pre-training of geographic multimodal data; S4. Through the pre-training of the aforementioned geographic multimodal model, text information and geographic information are aligned to achieve cross-modal geographic points of interest matching.

[0011] As a preferred technical solution, the geocoder uses a location coding method based on random Fourier features and a multilayer perceptron to capture features of different granularities and combine them to obtain high-dimensional features.

[0012] As a preferred technical solution, the frequency of the sigma value of the random Fourier features is changed by adjusting the exponential allocation strategy to process geographic data of different resolutions and capture the features of different granularities.

[0013] As a preferred technical solution, the geographic data is obtained by integrating multi-scale features and geographic context aggregation mechanisms. Geometric operations are performed on the absolute positions of the geographic context semantics of targets around the geographic location, and cluster centers are obtained to obtain the geographic data.

[0014] As a preferred technical solution, the InfoNCE contrastive learning strategy integrates a contrastive learning strategy into a geographic pre-trained model, using the InfoNCE loss function as the contrastive learning objective to enhance the pre-trained model's ability to distinguish between positive and negative sample pairs. The formula for the InfoNCE loss function is: In the formula, For the i-th original input sample, The total number of positive samples. The index for positive samples. For adaptive temperature coefficient, The total number of negative samples. Index for negative samples. Indicates the first The positive sample features corresponding to each sample Indicates the first Each negative sample feature This represents the cosine similarity function.

[0015] As a preferred technical solution, the dual-mode negative sampling strategy is used to enable the model to learn the distinction between positive and negative samples at the global level. It includes the selection of negative sample sampling method and the setting of the number of negative samples. The negative sample sampling method includes global decoupling sampling and local coupling sampling. The global decoupling sampling randomly selects negative samples from the full data without associating them with positive samples. The local coupling sampling assigns a dedicated negative sample set to each positive sample, enabling the model to distinguish between positive and negative samples in context.

[0016] As a preferred technical solution, the multimodal interaction is based on repeated Transformer sublayers, and cross-modal associations are learned through a dual-mask task of mask language model and mask generation model to achieve alignment and fusion of multimodal features.

[0017] As a preferred technical solution, the multimodal interaction is achieved by setting up a text classifier, a geographic data classifier, and a fusion modality classifier for unimodal prediction, calculating the cross-entropy loss of each classifier, and combining the multimodal contrast loss under weak supervision of unimodal prediction to form a learning objective.

[0018] As a preferred technical solution, any one of the text classifier, geographic data classifier, and fusion modality classifier adopts a structure combining a fully connected layer and a Softmax activation function.

[0019] As a preferred technical solution, any one of the text classifier, geographic data classifier, and fusion modality classifier is based on single-sample classification loss. Batch average loss To achieve training, the formulas are as follows: In the formula, x = t, g, or f represent the text classifier, the geographic data classifier, and the fusion modality classifier, respectively. For sample index, For geographic entity categories, It is the c-th position of the unique hot tag. For prediction samples of textual modalities, geographic modalities, or fused modalities Belongs to the The probability of geographic entities, =10 -10 , is the minimum value. This represents the number of samples.

[0020] Compared with the prior art, the present invention has at least one of the following beneficial effects: (1) This invention optimizes geographic data from multiple dimensions through a geographic encoder, and uses a location coding method based on random Fourier features and a multilayer perceptron to capture features of different granularities and combine them to obtain high-dimensional features. This solves the problems of insufficient fine-grained correlation of multimodal geographic data due to the spectral deviation of standard MLP caused by approximate low-frequency functions, difficulty in capturing high-frequency details, and lack of fine-grained correlation of multimodal geographic data. It realizes comparative learning between different modalities of GPS and GIS and improves the efficiency of geographic information processing.

[0021] (2) This invention integrates the contrastive learning strategy into the construction process of the geographic pre-trained model, adopts Info-NCE as the core learning objective, and combines dual-mode negative sampling and multimodal interaction to solve the problems of cross-modal representation mismatch and weak alignment mechanism, and realizes the optimization of the model to capture geospatial semantic association and cross-modal representation consistency.

[0022] (3) This invention introduces a single-peak supervised multimodal contrastive learning method in the cross-modal alignment learning of geographic semantic matching. By setting three classifiers for text, geographic data and fusion modality to make single-peak prediction, the modality validity is determined by using actual geographic entities as real labels. Then, the cross-entropy loss of each classifier is calculated separately. Combined with the multimodal contrastive loss under weak supervision of single-peak prediction, the learning objective is formed. This solves the problem that user queries are fuzzy and difficult to accurately match geographic information in the prior art. It realizes accurate alignment of geographic related representations under different modalities and improves the matching effect of cross-modal geographic information. Attached Figure Description

[0023] Figure 1 This is a schematic diagram illustrating the steps of a cross-modal geographic point of interest matching method based on fine-grained geographic semantics and contrastive learning. Figure 2 This is a schematic diagram of the network framework of X-DisGSIR in the embodiment. Detailed Implementation

[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0025] Example 1 To address the problems existing in the prior art, this embodiment provides a cross-modal geographic point of interest matching method based on fine-grained geographic semantics and contrastive learning. The method steps are as follows: Figure 1 As shown, it specifically includes: S1. For the acquired text data and geographic data, extract text semantic features and geographic semantic features respectively through text encoder and geographic encoder.

[0026] Textual and geographic data are processed separately using a text encoder and a specialized geoencoder for feature extraction. Specifically, the specialized geoencoder employs a location encoding technique using Random Fourier Features (RFF). Before processing by a feedforward network, this addresses the performance issues that arise in high-dimensional representations of low-dimensional GPS coordinates. Standard Multilayer Perceptrons (MLPs) suffer from spectral bias and difficulty capturing high-frequency details due to their approximate low-frequency function, which degrades performance in applications. To enrich location features, this method enhances hierarchical representation in the location encoder: capturing features of different granularities from coarse to fine and combining them to enrich high-dimensional features. Specifically, an exponential allocation strategy is used to adjust the sigma value of the RFF to change the frequency, thereby handling GPS coordinates at different resolutions.

[0027] Considering the strong correlation between geographic location and its surrounding targets, this method further introduces a geographic context aggregation mechanism. A series of geometric operations, such as mean, standard deviation, kurtosis, and skewness, are applied to the absolute positions of the surrounding targets' geographic context (GC). These statistics essentially cluster these targets, and the cluster centers are used to obtain the geographic targets. The final geographic representation, by integrating multi-scale features and contextual information, establishes a robust mapping from coordinates to semantics, laying a solid foundation for subsequent cross-modal alignment.

[0028] The extracted textual semantic features and geographic semantic features are then combined for cross-modal alignment and multimodal fusion learning.

[0029] S2. By utilizing textual semantic features and geographic semantic features, and combining dual-mode negative sampling, adaptive temperature-adjusted InfoNCE contrastive learning strategies, and multimodal interaction, cross-modal contrastive learning is conducted to establish the alignment relationship between textual semantic features and geographic semantic features.

[0030] This scheme employs the InfoNCE contrastive learning strategy and a classic multimodal interaction module for cross-modal contrastive learning. The InfoNCE contrastive learning strategy integrates contrastive learning into the geographic pre-trained model, using the Info Noise-contrastive estimation (InfoNCE) loss function as the contrastive learning objective, and employs a dual-mode negative sampling strategy. The Loss_InfoNCE optimizes the model's ability to distinguish between positive and negative sample pairs, and its formula is as follows: In the formula, For the i-th original input sample, The total number of positive samples. The index for positive samples. For adaptive temperature coefficient, The total number of negative samples. Index for negative samples. Indicates the first The positive features corresponding to each sample. Indicates the first Each negative sample feature This represents the cosine similarity function. The adaptive temperature coefficient is also included. This plays a key regulatory role, with a value ranging from 0.05 to 0.5, and a higher value... The value makes the model focus on difficult samples, while the lower value... The value enhances the distinction between simple samples.

[0031] To address the unique characteristics of geographic data, this invention proposes a dual-mode negative sampling strategy, encompassing two dimensions: negative sample selection and negative key quantity setting. Global unpaired sampling randomly selects negative samples from the entire dataset, unrelated to positive samples. For example, if the positive sample is an image of a coffee shop in Beijing with its corresponding address, the negative sample could be randomly selected from POI data such as supermarkets and parks across the country, such as the name of a subway station in Shanghai. This mode helps the model learn global-level positive-negative discrimination capabilities and is suitable for scenarios where the bimodal geographic target is treated as independent samples and positive-negative discrimination between samples is required. However, the negative samples may be unrelated to the query sample, potentially leading to redundant and irrelevant features being learned by the model. Locally coupled sampling (paired) assigns a dedicated set of negative samples to each positive sample, enabling the model to distinguish between positive and negative samples in specific contexts. For example, if the positive sample is the same as above, the negative sample could be selected from information about other coffee shops, such as an image or address of another coffee shop in Beijing. The image represents the same modality, while the address represents a cross-modality. In geographic semantic matching tasks, if the model needs to learn the matching relationship between modal pairs of specific geographic targets, the paired mode is more suitable. Specifically, other geographic targets can be selected as negative samples through same-modal or cross-modal comparisons.

[0032] The choice of negative sample sampling method should be considered in conjunction with task characteristics: if the focus is on intramodal contrast and the dataset size is large enough to support low-noise random sampling, the global decoupling sampling mode is simple and efficient, suitable for tasks requiring strong generalization ability such as retrieval; if it is necessary to enhance cross-modal matching ability, such as matching the geographic context of a geographic target with the text of other targets as negative examples, the local coupling sampling mode is more suitable for scenarios requiring fine-grained discrimination ability such as reranking. This invention finds that if the main focus is on intramodal contrastive learning, and the dataset is large enough to support random selection of negative samples without introducing too much noise, then the global decoupling sampling mode may be a simple and effective choice. If it is desired that the model can learn cross-modal matching ability and construct a set of negative samples paired with each positive sample, such as selecting the text of other geographic targets as negative samples for the geographic context of each geographic target, then the local coupling sampling mode is more suitable.

[0033] The setting of the number of negative samples refers to the selection of the number of negative samples k, which represents the total number of negative samples participating in this comparison. When setting it, it is necessary to balance the sufficiency of learning and the computational cost: too few samples will result in insufficient coverage of sample information, making it difficult for the model to learn effective distinguishing features and limiting performance; too many samples will significantly increase the computational cost, and performance will no longer improve due to diminishing marginal returns after exceeding the threshold. Setting the number of negative keys to 192 represents the optimal balance between performance, computational efficiency, and hardware adaptability. This value avoids the problems associated with small negative key numbers, such as less than 128, which lead to insufficient negative sample diversity, model overfitting to a limited number of negative examples, and weak contrastive signals. It provides sufficient negative contrast evidence for InfoNCE loss, allowing the model to effectively learn the feature discrimination of samples. At the same time, it avoids the huge memory consumption and computational overhead of extremely large negative key numbers, such as 65536 and 131072. It eliminates the need to maintain a complex dynamic feature queue, adapts to the batch size of regular training, and is compatible with the memory limitations of ordinary hardware, such as batch sizes of 64, 128, or 256. It enables fast similarity calculation and backpropagation during training, balancing model convergence speed and practical training efficiency. This value is the optimal choice for small-to-medium batch contrastive learning training scenarios, balancing feature learning quality and engineering feasibility.

[0034] This scheme employs a classic multimodal interaction module based on repeated Transformer sublayers, specifically consisting of residual connections (Add), layer normalization (Norm), and a fully connected feedforward network. Add+Norm is used for stable training, while Feedforward is used for non-linear transformation of features. The module learns cross-modal associations through dual-masking tasks: MLM (Masked Language Model) – randomly masks parts of the text, allowing the model to predict the masked content, similar to a self-supervised task in text modality; MGM (Masked Generative Model) – randomly masks parts of the visual modality, such as image patches, allowing the model to reconstruct the masked content, similar to a self-supervised task in visual modality. The module uses Loss_mm optimization to achieve the alignment and fusion of multimodal features.

[0035] S3. Utilize unimodal prediction confidence as a weak supervision signal to supervise cross-modal alignment and achieve pre-training of geographic multimodal data.

[0036] Furthermore, the multimodal fusion module performs unimodal prediction by setting three classifiers for text, geographic data, and fused modality, and determines modality validity by using actual geographic entities as the true labels; then, it calculates the cross-entropy loss of each classifier separately, and combines it with the multimodal contrast loss under weak supervision of unimodal prediction to form the learning objective.

[0037] All three classifiers employ a lightweight structure of fully connected layers and Softmax activation, ensuring focus on modality effectiveness rather than the capabilities of complex models.

[0038] Specifically, text classifier f t Represented as: Geographic classifier f g Represented as: Fusion classifier f f Represented as: In the formula, For input features, , For classifier parameters, and Represents the predicted samples for each modality Belongs to the The probability of a class, x = t, g, or f. .

[0039] All three classifiers employ multi-class cross-entropy loss. The single-sample classification loss and batch average loss for each classifier are as follows: In the formula, x = t, g, or f represent the text classifier, the geographic data classifier, and the fusion modality classifier, respectively. For sample index, For geographic entity categories, It is the c-th position of the unique hot tag. The c-th position of the unique hot tag has only a true category value of 1, while the rest are 0. For text / geographic / fusion modality prediction samples Belongs to the The probability of geographic entities, =10 -10 , is the minimum value, avoid When log(0) reaches an infinite value, Let be the number of samples. Therefore, the formula can be simplified to calculating only the negative logarithm of the predicted probability of the true class, such as if the true class is . ,but This more intuitively demonstrates that the more accurate the prediction, the smaller the loss.

[0040] To verify the effectiveness of this method, based on the above embodiments, this scheme was tested on the GeoGLUE (GeoGraphic Language Understanding Evaluation) dataset and compared in depth with four representative baseline models—covering the general pre-trained language models RoBERTa, ERNIE, StructBERT, and the geo-domain-specific pre-trained model (MGEO). The X-DisGSIR model structure is as follows: Figure 2 As shown in Table 1, this dataset, jointly released by Alibaba DAMO Academy's Natural Language Processing Group and Gaode Maps, is one of the authoritative benchmarks in the field of geographic semantic understanding. It includes over 230,000 real map log text queries containing colloquial questions and structured addresses, as well as millions of associated POIs, including latitude, longitude, and GIS information. Table 1 Experimental Results The experiment was conducted on the rerank test set of the GeoGLUE dataset, which contains 50,000 real geographic queries, each with 20-40 initial candidate POIs. The evaluation criteria included: Recall@k (k=5), a core metric in information retrieval, which measures the proportion of positive POIs among the top k candidates; MRR@k (k=5), which measures the average reciprocal of the ranking of positive POIs, reflecting ranking accuracy; Top1acc, the proportion of queries where the top candidate is a positive POI, directly reflecting the accuracy of the first recommendation; and NDCG@k (k=5), the normalized loss cumulative gain, which balances candidate relevance and ranking, aligning with user perception. The table results show that the X-DisGSIR model achieves significant breakthroughs in Recall@5 and MRR@5 metrics: compared to the geographic-specific baseline MGEO, Recall@5 is improved by 1.1% and MRR@5 by 0.5%; compared to the best-performing general-purpose model StructBERT, Recall@5 is improved by 1.16% and MRR@5 by 0.7%. These results demonstrate that the X-DisGSIR model proposed in this invention significantly outperforms the basic semantic matching of general-purpose pre-trained models and the spatial information modeling of traditional geographic-specific models in terms of the comprehensive ability of combining geographic text semantic understanding with accurate ranking of candidate POIs, providing a better technical path for improving query and POI matching efficiency in map search.

[0041] S4. By pre-training with geographic multimodal data, text information and geographic information are aligned to achieve cross-modal geographic points of interest matching.

[0042] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A cross-modal geographic point of interest matching method based on fine-grained geographic semantics and contrastive learning, characterized in that, The method specifically includes: S1. For the acquired text data and geographic data, extract text semantic features and geographic semantic features respectively through text encoder and geographic encoder; S2. Utilizing the aforementioned textual semantic features and geographic semantic features, cross-modal contrastive learning is conducted by combining a dual-mode negative sampling InfoNCE contrastive learning strategy and multimodal interaction to establish an alignment relationship between textual semantic features and geographic semantic features; S3. Use unimodal prediction confidence as a weak supervision signal to supervise cross-modal alignment and achieve pre-training of geographic multimodal data; S4. Through the pre-training of the aforementioned geographic multimodal model, text information and geographic information are aligned to achieve cross-modal geographic points of interest matching.

2. The method for cross-modal geographic point of interest matching based on fine-grained geographic semantics and contrastive learning according to claim 1, characterized in that, The geocoder employs a location coding method based on random Fourier features and a multilayer perceptron to capture features of different granularities and combine them to obtain high-dimensional features.

3. The cross-modal geographic point of interest matching method based on fine-grained geographic semantics and contrastive learning according to claim 2, characterized in that, By adjusting the sigma value of the random Fourier features using an exponential allocation strategy to change the frequency, geographic data of different resolutions can be processed to capture features of different granularities.

4. The cross-modal geographic point of interest matching method based on fine-grained geographic semantics and contrastive learning according to claim 1, characterized in that, The geographic data is obtained by integrating multi-scale features and geographic context aggregation mechanisms. Geometric operations are performed on the absolute geographic context semantic positions of targets around the geographic location, and cluster centers are obtained to obtain the geographic data.

5. The cross-modal geographic point of interest matching method based on fine-grained geographic semantics and contrastive learning according to claim 1, characterized in that, The InfoNCE contrastive learning strategy involves integrating a contrastive learning approach into a geographic pre-trained model, using the InfoNCE loss function as the contrastive learning objective, to enhance the pre-trained model's ability to distinguish between positive and negative sample pairs. The formula for the InfoNCE loss function is as follows: In the formula, For the i-th original input sample, The total number of positive samples. The index for positive samples. For adaptive temperature coefficient, The total number of negative samples. Index for negative samples. Indicates the first The positive sample features corresponding to each sample Indicates the first Each negative sample feature This represents the cosine similarity function.

6. The cross-modal geographic point of interest matching method based on fine-grained geographic semantics and contrastive learning according to claim 1, characterized in that, The dual-mode negative sampling strategy is used to enable the model to learn the distinction between positive and negative samples at the global level. It includes the selection of negative sample sampling method and the setting of the number of negative samples. The negative sample sampling method includes global decoupling sampling and local coupling sampling. The global decoupling sampling randomly selects negative samples from the full data without associating them with positive samples. The local coupling sampling assigns a dedicated negative sample set to each positive sample, enabling the model to distinguish between positive and negative samples in context.

7. The cross-modal geographic point of interest matching method based on fine-grained geographic semantics and contrastive learning according to claim 1, characterized in that, The multimodal interaction is based on repeated Transformer sublayers. It learns cross-modal associations through a dual-mask task of mask language model and mask generation model to achieve alignment and fusion of multimodal features.

8. The cross-modal geographic point of interest matching method based on fine-grained geographic semantics and contrastive learning according to claim 1, characterized in that, The multimodal interaction described herein uses a text classifier, a geographic data classifier, and a fusion modality classifier to perform unimodal prediction, calculates the cross-entropy loss of each classifier, and combines the multimodal contrast loss under weak supervision of unimodal prediction to form the learning objective.

9. A cross-modal geographic point of interest matching method based on fine-grained geographic semantics and contrastive learning according to claim 8, characterized in that, For any of the text classifier, geographic data classifier, and fusion modality classifier mentioned above, a structure combining a fully connected layer and a Softmax activation function is adopted.

10. A cross-modal geographic point of interest matching method based on fine-grained geographic semantics and contrastive learning according to claim 8, characterized in that, For any of the text classifier, geographic data classifier, and fusion modality classifier mentioned above, each is based on a single-sample classification loss. Batch average loss To achieve training, the formulas are as follows: In the formula, x = t, g, or f represent the text classifier, the geographic data classifier, and the fusion modality classifier, respectively. For sample index, For geographic entity categories, It is the c-th position of the unique hot tag. For prediction samples of textual modalities, geographic modalities, or fused modalities Belongs to the The probability of geographic entities, =10 -10 , is the minimum value. This represents the number of samples.