Geographic entity information processing method and device, storage medium and program product

By combining geographic semantic feature encoders, distance feature encoders, and context feature encoders, and using a neural network model for feature fusion, the problem of matching different data sources in the cell database is solved, achieving high-precision and high-robust geographic entity matching.

CN121580127APending Publication Date: 2026-02-27TAOBAO CHINA SOFTWARE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511773180.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

In the process of expanding the cell database, how to accurately and efficiently determine whether different cell records in the ontology and third-party databases point to the same geographic entity in the real world when there is a lack of unique identifiers? Existing rule-based matching schemes cannot effectively utilize the high-noise text features and spatial location information of addresses, resulting in poor robustness.

Method used

The system employs a geographic semantic feature encoder, a distance feature encoder, and a geographic context feature encoder to extract features from name and address semantics, spatial location, and surrounding geographic environment. It utilizes a neural network model for feature fusion and accurate judgment, and outputs a classification result through a classification head network to determine whether the references point to the same geographic entity.

Benefits of technology

It improves the accuracy and robustness of geographic entity matching, better handles differences from different data sources, and enhances the information coverage and accuracy of the community database.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121580127A_ABST
    Figure CN121580127A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a geographic entity information processing method and device, a storage medium and a program product. In the embodiment of the invention, for at least two pieces of geographic entity description information, a geographic semantic feature encoder, a distance feature encoder and a geographic context feature encoder are respectively utilized to encode names and address semantics, spatial positions and surrounding geographic environments; obtaining a geographic semantic similarity, a spatial distance feature and a geographic context similarity; and inputting the geographic semantic similarity, the spatial distance feature and the geographic context similarity into a classification head network, and outputting a classification result of whether the at least two pieces of geographic entity description information point to the same geographic entity or not. Therefore, by introducing the model to learn the incidence relation between different features, full utilization of multi-dimensional features such as geographic semantics, spatial positions and surrounding geographic environments is realized, feature fusion and accurate judgment are realized in the classification head network, and the accuracy and robustness of geographic entity matching are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, device, storage medium and program product for processing geographic entity information. Background Technology

[0002] In smart city scenarios such as intelligent security, smart property management, community life services, and personalized information recommendations, residential communities, as the basic spatial units of urban residents' lives, constitute the core geographical anchor points for various application services. Therefore, building a high-coverage, high-precision, and high-timeliness community database is a key infrastructure to support the realization of accurate perception, intelligent decision-making, and personalized services in the above scenarios.

[0003] To improve the quality of the community database, community records from third-party data sources such as map platforms, real estate agencies, and government systems (hereinafter referred to as "third-party databases") can be introduced as a foundation to expand the information and enhance the semantics of the ontology, thereby constructing a community database with high coverage, high accuracy, and high timeliness.

[0004] However, different data sources have their own implementations in terms of cell naming conventions, coordinate precision, boundary definitions, and update frequencies, resulting in significant differences between them. Therefore, accurately and efficiently determining whether different cell records in the ontology database and third-party databases point to the same geographic entity in the real world, in the absence of unique identifiers, is a major technical challenge faced by the aforementioned expanded cell database. Summary of the Invention

[0005] This application provides a geographic entity information processing method, device, storage medium, and program product to accurately identify geographic entities, effectively address matching difficulties caused by inconsistent description information across data sources, and improve the accuracy and robustness of geographic entity matching.

[0006] This application provides a method for processing geographic entity information, comprising: acquiring at least two geographic entity description information, wherein the geographic entity description information includes name information, address information, and spatial location information of the geographic entity; using a geographic semantic feature encoder, calculating the geographic semantic similarity between the at least two geographic entity description information based on the name information and address information in the at least two geographic entity description information; using a distance feature encoder, calculating the spatial distance feature between the at least two geographic entity description information based on the spatial location information in the at least two geographic entity description information; using a geographic context feature encoder, calculating the geographic context similarity between the at least two geographic entity description information based on the at least two geographic entity description information and their respective surrounding geographic environment information; inputting the geographic semantic similarity, spatial distance feature, and geographic context similarity into a classification head network, and outputting a classification result indicating whether the at least two geographic entity description information points to the same geographic entity.

[0007] This application also provides an electronic device, including a memory and a processor. The memory is used to store a computer program, and the processor is coupled to the memory to execute the computer program to implement the steps in the methods described above.

[0008] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to perform the steps in the methods described above.

[0009] This application also provides a computer program product, which includes a computer program / instructions that, when executed by a processor, enable the processor to implement the steps described in the above method embodiments.

[0010] In this embodiment, for at least two geographic entity descriptions, a geographic semantic feature encoder, a distance feature encoder, and a geographic context feature encoder are used to encode name and address semantics, spatial location, and surrounding geographic environment to obtain geographic semantic similarity, spatial distance features, and geographic context similarity. These are then input into a classification head network to output a classification result indicating whether the descriptions of at least two geographic entities point to the same geographic entity. Thus, by introducing a neural network model to learn the relationships between different features, multi-dimensional features such as geographic semantics, spatial location, and surrounding geographic environment are fully utilized. Feature fusion and accurate judgment are achieved in the classification head network, improving the accuracy and robustness of geographic entity matching. Attached Figure Description

[0011] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1a A flowchart illustrating a geographic entity information processing method provided for an exemplary embodiment of this application; Figure 1b A schematic diagram of the structure of a geographic entity information processing system provided for an exemplary embodiment of this application; Figure 2a A schematic diagram of the structure of a geographic entity information processing system provided as another exemplary embodiment of this application; Figure 2b A schematic diagram of the structure of a geographic context relationship graph provided as another exemplary embodiment of this application; Figure 3a A schematic diagram illustrating the process of a classification head network performing classification prediction, as provided in yet another exemplary embodiment of this application; Figure 3b A schematic diagram illustrating the process of a classification head network performing classification prediction, as provided in yet another exemplary embodiment of this application; Figure 4 A flowchart illustrating yet another exemplary embodiment of the present application provides a model training method. Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0012] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0013] It should be noted that, in the cases involving user information in the embodiments of this application, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse. In addition, the various models involved in this application (including but not limited to language models or large models) comply with relevant laws and standards.

[0014] Community Entity Alignment refers to the process of determining different data records that point to the same entity in the real world. For example, it determines whether two community records from different sources point to the same physical community in the real world. The two different sources can be, for example, an ontology library and a third-party library. This process can also be called Community Mounting.

[0015] When determining whether different community records point to the same physical community in the real world, a rule-based matching scheme is usually adopted. The rule-based matching scheme includes exact matching of community names, fuzzy matching of community names, character similarity matching, area matching, and parent-child community matching. In the rule-based matching scheme, a series of hard-coded rules are maintained for judgment, such as the names being exactly the same and the distance being less than X meters, the names being the same after removing specific characters, alias matching, etc. The following is a related introduction.

[0016] Exact matching of community names: If the community names are exactly the same and the distance between the communities is less than 1000 meters, it is determined that two different community records point to the same physical community.

[0017] Fuzzy matching of community names: Remove specific characters in the community name, such as removing characters like "Yuan|Community|Residential Area|Apartment|District|Government|Courtyard|Mansion|Park|Square|Garden", and if the matched community names are exactly the same and the distance is less than 200 meters; if the address strings are the same and the similarity of the community names is greater than 0.2, it is determined that two different community records point to the same physical community.

[0018] Character similarity matching: If the longest common substring similarity is greater than 0.7 and the distance is less than 50 meters, it is determined that two different community records point to the same physical community.

[0019] Area matching: If the polygon area occupied by the ontology library community on the map can cover the longitude and latitude of the third-party community, it is determined that two different community records point to the same physical community.

[0020] Parent-child community matching: If the ontology library community is a sub-community, its parent community has the same name as the third-party community and the distance is less than 1000 meters, it is determined that two different community records point to the same physical community.

[0021] In the above scheme of implementing community mounting with a series of hard-coded rules, the utilization rate of high-noise text features such as addresses is extremely low, and the rich spatial location information behind the descriptive information related to community records cannot be effectively utilized, resulting in insufficient feature utilization. Secondly, the fixed distance threshold cannot adapt to the differences in community scales in different cities and regions, resulting in poor robustness.

[0022] To address the aforementioned technical issues, in this embodiment, for at least two geographic entity descriptions, a geographic semantic feature encoder, a distance feature encoder, and a geographic context feature encoder are used to encode name and address semantics, spatial location, and surrounding geographic environment to obtain geographic semantic similarity, spatial distance features, and geographic context similarity. These are then input into a classification head network, which outputs a classification result indicating whether the descriptions of at least two geographic entities point to the same geographic entity. Thus, by introducing a model to learn the relationships between different features, multi-dimensional features such as geographic semantics, spatial location, and surrounding geographic environment are fully utilized. Feature fusion and accurate judgment are achieved in the classification head network, improving the accuracy and robustness of geographic entity matching.

[0023] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.

[0024] Figure 1a This is a schematic diagram illustrating the structure of a geographic entity information processing method provided as an exemplary embodiment of this application. For example... Figure 1a As shown, the method includes: S101: Obtain description information for at least two geographic entities, each including the name, address, and spatial location information of the geographic entity; S102: Using a geographic semantic feature encoder, calculate the geographic semantic similarity between at least two geographic entity descriptions based on the name and address information in the descriptions of at least two geographic entities; S103: Using a distance feature encoder, calculate the spatial distance feature between at least two geographic entity descriptions based on the spatial location information in the descriptions of at least two geographic entities; S104: Using a geographic context feature encoder, calculate the geographic context similarity between at least two geographic entity descriptions based on the geographic context information of at least two geographic entities and their surrounding geographic environment information. S105: Input geographic semantic similarity, spatial distance features and geographic context similarity into the classification head network, and output the classification result of whether at least two geographic entity descriptions point to the same geographic entity.

[0025] In this embodiment, a geographic entity refers to an entity that can be determined in the geographic space of the real world. The type of geographic entity is not limited. Geographic entities may include, but are not limited to: residential geographic entities such as communities, residential areas, and buildings; public facility geographic entities such as schools, subway stations, bus stops, and hospitals; commercial and service geographic entities such as business districts, shopping centers, commercial buildings, and office buildings; and geographic entities divided according to coverage areas such as street coverage areas, regional centers, districts, counties, and cities. In subsequent embodiments, a community is used as an example of a geographic entity, but this does not constitute a limitation on the embodiments of this application.

[0026] In this embodiment, it is necessary to determine whether at least two geographic entity description information points to the same geographic entity. Taking a geographic entity as a cell, the geographic entity description information is called cell description information (or simply cell entry). The at least two geographic entity description information can come from different databases or from the same database. In the case of different databases, for example, they can come from an ontology database and a third-party database respectively. Therefore, it is possible to determine whether geographic entity description information from different databases points to the same geographic entity. For example, when enhancing the ontology database through a third-party database, geographic entity description information pointing to the same geographic entity can be merged or deduplicated. This improves the information coverage of the ontology database and also increases the information accuracy of the enhanced ontology database, avoiding information redundancy. In the case of the same database, the specific implementation of the database is not limited; it can be an ontology database, a third-party database, or other databases. Deduplication processing can be performed on geographic entity description information from the same database to determine whether it points to the same geographic entity. Merging or deduplicating at least two geographic entity description information pointing to the same geographic entity can yield an ontology database or third-party database with higher information accuracy.

[0027] In this embodiment, the geographic entity description information includes name information, address information, and spatial location information. To fully utilize the geographic entity description information, this embodiment uses a neural network model to extract features from three dimensions: name and address semantics, spatial location information, and surrounding geographic environment. This fully utilizes the features extracted from different information dimensions, and the neural network model uses these extracted features to accurately identify geographic entities.

[0028] It should be noted that the implementation form of the neural network model used in this application embodiment is not limited, and it can be various deep learning-based neural network models. Optionally, the various neural network models used can be deep learning models with relatively small parameter scales or deep learning models with relatively large parameter scales. The large model is merely an example, and this application embodiment does not limit the number of model parameters supported by the deep learning model used, aiming to meet actual needs. The deep learning models involved in this application embodiment are not limited to a specific type, and different models can be selected according to implementation needs. For example, an artificial intelligence-based language model (LM) can be used, and of course, a multimodal model (MM) capable of processing multiple modal information can also be used.

[0029] In this embodiment, the neural network model used includes at least: a geographic semantic feature encoder, a distance feature encoder, a geographic context feature encoder, and a classification head network. These neural network models cooperate and collaborate to determine whether the description information of at least two geographic entities points to the same geographic entity. Each model contributes to improving the accuracy of the determination. The role of each model is explained in detail below.

[0030] In this embodiment, a geographic semantic feature encoder is used to calculate the geographic semantic similarity between at least two geographic entity description information based on the name information and address information in the description information of at least two geographic entities.

[0031] Among these, the name and address information in the geographic entity description information is relatively text-based. Therefore, for the name and address information in at least two geographic entity description information, such as... Figure 1b As shown, the input is fed into the geographic semantic feature encoder to calculate the semantic similarity between at least two geographic entity descriptions.

[0032] For example, taking a residential community as a geographic entity, geographic entity description information for two communities is obtained from different database sources. For each community's geographic entity description information, the corresponding name and address information are extracted. In some embodiments, to facilitate semantic similarity calculation, the extracted name and address information are formatted into a structured input sequence containing preset tags, for example, [CLS] Name: <Name Information> Address: <Address Information> [SEP]. Here, [CLS] indicates the start of the sequence and can also be considered a classification tag, representing the aggregated information of the entire sequence. [SEP] is used to separate different geographic entity description information and can be called a segmentation tag. <Name Information> and <Address Information> refer to the name and address information of the geographic entity, respectively.

[0033] like Figure 1b As shown, geographic entity description information A1 is abbreviated as A1, and its corresponding name information is Community A; its corresponding address information is XX Road XX No. 2. Geographic entity description information A2 is abbreviated as A2, and its corresponding name information is Community A, Zone 2; its corresponding address information is YY Road YY No. 2. Optionally, any of the above geographic entity description information can be formatted as: [CLS] Name: <Name Information> Address: <Address Information> [SEP]. For example, [CLS] Name: Community A; Address: XX Road XX No. 2 [SEP]. Wherein, [CLS] and [SEP] are two identifiers. [CLS] is used to indicate the beginning of the input sequence, and its corresponding final hidden state is used as the global semantic representation vector of the entire input sequence pair, which is convenient for the geographic semantic feature encoder to output the global semantic representation vector for subsequent similarity calculation or classification decision; [SEP] is used to separate the description information of different geographic entities in the input sequence, so that the geographic semantic feature encoder can distinguish the text boundaries of different geographic entity description information and effectively model the semantic similarity between different geographic entity description information.

[0034] In this embodiment, a geographic semantic feature encoder is used to semantically encode description information of at least two geographic entities and obtain the geographic semantic similarity between the description information of at least two geographic entities. For example, the geographic semantic feature encoder can extract the global semantic representation vector at the [CLS] location to calculate the geographic semantic similarity between the description information of at least two geographic entities. Alternatively, semantic representation vectors from all locations can be taken and weighted and fused to calculate the geographic semantic similarity between the description information of at least two geographic entities. The geographic semantic similarity between the description information of at least two geographic entities deeply fuses the name and address information of the description information of at least two geographic entities and captures their semantic similarity. In this embodiment, a distance feature encoder is used to calculate the spatial distance feature between the description information of at least two geographic entities based on the spatial location information in the description information of at least two geographic entities. In this embodiment, as... Figure 1b As shown, spatial location information from at least two geographic entity descriptions is used as input to a distance feature encoder, which calculates the spatial distance features between the descriptions of at least two geographic entities.

[0035] In this embodiment, as Figure 1b As shown, spatial location information refers to information describing the spatial location of a geographic entity. This can be implemented as latitude and longitude information, but is not limited to this. Besides latitude and longitude information, spatial location information can also be implemented as relative location information reflecting the spatial relationship between a geographic entity and surrounding roads, landmarks, or other reference points.

[0036] In an optional embodiment, taking the latitude and longitude of a geographic entity as an example, the Haversian distance or spherical cosine distance between the latitude and longitude of any two geographic entities can be calculated, and the spatial distance characteristics between any two geographic entities can be determined based on the calculation result.

[0037] Furthermore, if the descriptions of two address entities point to the same geographic entity in the real world, their surrounding geographic information will highly overlap. This surrounding geographic information can include, for example, points of interest (POIs) such as shops, schools, and bus stops; or, for example, road structure information such as road density, road type, and the location of transportation hubs—there are no limitations on this.

[0038] In this embodiment, based on the address information in the description information of each geographic entity, a map service interface is invoked to obtain the surrounding environment information within a preset distance range around the geographic entity described in the geographic entity description information. The preset distance range can be 300 meters, 500 meters, 1 kilometer, etc., and is not limited thereto; it can be flexibly set according to accuracy requirements or coverage requirements.

[0039] Furthermore, using a geographic context feature encoder, the geographic context similarity between at least two geographic entity descriptions is calculated based on the geographic context similarity between the descriptions of at least two geographic entities and their surrounding geographic environment information.

[0040] Specifically, such as Figure 1b As shown, the geographic context feature encoder can not only fuse the descriptive information of each geographic entity with its corresponding geographic environment information to obtain a feature vector that can represent the context of the geographic entity, but also perform feature interaction processing on at least two geographic context feature vectors to construct a feature representation that reflects the contextual relationship between at least two geographic entities.

[0041] The feature interaction processing methods may include, but are not limited to: concatenating, subtracting, or combining concatenation and subtraction of at least two geographic context feature vectors. Through feature interaction processing, a feature vector is generated that characterizes the degree of similarity between the geographic entities described by at least two geographic entity description information in terms of geographic environment, i.e., geographic context similarity.

[0042] In this embodiment, geographic semantic similarity, spatial distance features, and geographic context similarity are input into the classification head network, which performs classification processing and finally outputs a classification result indicating whether at least two geographic entity descriptions point to the same geographic entity.

[0043] In an optional embodiment, the classification head network is implemented as a general classification head network. It can then employ a feature fusion approach, concatenating geographic semantic similarity features, spatial distance features, and geographic context similarity features to form a fused feature vector. This vector is then input into the classification head network for classification, outputting a classification result indicating whether the descriptions of two geographic entities point to the same geographic entity. In this optional embodiment, the classification head network can adopt a multilayer perceptron structure. Using a general classification head network for classification processing offers advantages such as simplicity and high efficiency.

[0044] In another optional embodiment, a hierarchical fusion architecture classification head network is designed, comprising multiple sub-classification head networks and a total classification head network. Geographic semantic similarity features, spatial distance features, and geographic context similarity features can be input into different sub-classification head networks to obtain corresponding initial classification results. Further, the total classification head network fuses and predicts the outputs of the multiple sub-classification head networks to obtain a classification result, which indicates whether the description information of two geographic entities points to the same geographic entity. Even more optionally, during the fusion and prediction process of the total classification network's outputs from the multiple sub-classification head networks, the influence of geographic semantic similarity features, spatial distance features, and geographic context similarity features can be comprehensively considered. That is, the geographic semantic similarity features, spatial distance features, and geographic context similarity features, along with the initial classification results output from the multiple sub-classification head networks, are simultaneously input into the total classification head network for fusion and prediction to obtain the final classification result. Among them, multiple sub-classification head networks generate local discrimination results based on different semantic dimensions of geographic entity description information. The overall classification head network performs nonlinear fusion and confidence calibration on the local discrimination results and outputs a globally consistent classification decision. This can effectively overcome the defects of single semantic dimensions being susceptible to noise interference or information loss, and significantly improve the accuracy, robustness and scene adaptability of geographic entity matching.

[0045] In this embodiment, for at least two geographic entity descriptions, a geographic semantic feature encoder, a distance feature encoder, and a geographic context feature encoder are used to encode name and address semantics, spatial location, and surrounding geographic environment to obtain geographic semantic similarity, spatial distance features, and geographic context similarity. These are then input into a classification head network, which outputs a classification result indicating whether the descriptions of at least two geographic entities point to the same geographic entity. Thus, by introducing a model to learn the relationships between different features, multi-dimensional features such as geographic semantics, spatial location, and surrounding geographic environment are fully utilized. Feature fusion and accurate judgment are achieved in the classification head network, improving the accuracy and robustness of geographic entity matching.

[0046] In one optional embodiment, obtaining at least two geographic entity description information includes: combining geographic entity description information from the ontology library and a third-party library in pairs, and using the geographic entity description information in each combination as the aforementioned at least two geographic entity description information; or, for each geographic entity description information in the third-party library, based on a rule-based matching strategy and / or a text similarity matching strategy, recalling at least one similar geographic entity description information in the ontology library to form the aforementioned at least two geographic entity description information. The following describes these two implementation methods in detail: In an optional embodiment, the at least two geographic entity descriptions can be obtained by fully combining geographic entity descriptions existing in one or more databases (e.g., an ontology library and / or at least one third-party library). That is, a geographic entity description is selected from each database to form a combination, and the geographic entity descriptions in this combination constitute the at least two geographic entity descriptions to be processed in this embodiment. Taking ontology library A, third-party libraries B and C as an example, the geographic entity descriptions in ontology library A, third-party libraries B and C are fully combined. Each combination includes one geographic entity description from ontology library A, one geographic entity description from third-party library B, and one geographic entity description from third-party library C. The geographic entity descriptions in each combination serve as the at least two geographic entity descriptions to be processed in this embodiment.

[0047] In another optional embodiment, similar geographic entity description information is filtered (or recalled) using certain methods as at least two geographic entity description information to be processed in this embodiment. For example, based on any geographic entity description information in the ontology library, similar geographic entity description information can be recalled from various third-party libraries. Then, the original geographic entity description information and all the recalled similar geographic entity description information are used together as at least two geographic entity description information to be processed in this embodiment. As another example, based on any geographic entity description information in the ontology library, similar geographic entity description information can be recalled from each third-party library. Then, for each third-party library, the original geographic entity description information and the similar geographic entity description information recalled from that third-party library are used together as at least two geographic entity description information to be processed in this embodiment. For example, for any third-party library, based on any geographic entity description information in that third-party library, similar geographic entity description information is retrieved from the ontology library. Then, that geographic entity description information and all similar geographic entity description information retrieved from the ontology library are used together as at least two geographic entity description information to be processed in this embodiment. Alternatively, for any third-party library, based on any geographic entity description information in that third-party library, similar geographic entity description information is retrieved from the ontology library. Then, that geographic entity description information and one similar geographic entity description information retrieved from the ontology library are used together as at least two geographic entity description information to be processed in this embodiment.

[0048] In this embodiment, given the acquisition of at least two geographic entity descriptions, the results of rule-based matching can be categorized into three cases. A detailed description of rule-based matching can be found in the preceding embodiments and will not be repeated here. The first case is where it can be clearly determined that the entities point to the same geographic entity; the second case is where it can be clearly determined that the entities point to different geographic entities; and the third case is where it cannot be clearly determined, i.e., the similarity is high, making it unsuitable to directly determine whether they point to the same geographic entity. For the third case, the descriptions of at least two highly similar geographic entities are used as subsequent processing objects and processed according to the geographic entity information processing method provided in this embodiment to determine whether the descriptions of the at least two geographic entities point to the same geographic entity.

[0049] In the above embodiments, the description information of geographical entities with high similarity is locked by recall, and then the method provided in this application embodiment is used to identify whether they point to the same geographical entity, instead of identifying the combination of all geographical entity description information, which can reduce the overall computational burden.

[0050] In one optional embodiment, when calculating the geographic semantic similarity between at least two geographic entity descriptions based on name and address information in the geographic entity description information using a geographic semantic feature encoder, the method includes: performing regularization enhancement, word segmentation, and embedding processing on the name and address information in the at least two geographic entity descriptions to construct the input sequence of the geographic semantic feature encoder; such as Figure 2a As shown, the input sequence is fed into the encoding network of the geographic semantic feature encoder, and multi-layer bidirectional contextual semantic encoding based on an attention mechanism is performed to obtain geographic semantic features with at least two geographic entity descriptions; such as Figure 2a As shown, the geographic semantic features of at least two geographic entity description information are input into the first similarity network in the geographic semantic feature encoder for similarity calculation, so as to obtain the geographic semantic similarity between at least two geographic entity description information.

[0051] In this embodiment, regular expression enhancement refers to the normalization processing of unstructured descriptive information, including but not limited to: standardizing capitalization, removing invalid characters, replacing special symbols, and normalizing names or abbreviations, to improve text consistency between at least two geographic entity descriptive information. Word segmentation involves splitting name and address information according to a preset dictionary to form a sequence of multiple smaller text units, such as words. Further, embedding processing is performed to construct word embeddings, paragraph embeddings, and location information embeddings for each text unit, forming a structured input representation.

[0052] In one optional embodiment, when performing regular expression enhancement, word segmentation, and embedding processing on the name information and address information in at least two geographic entity description information to construct the input sequence of the geographic semantic feature encoder, the process includes: for the name information in any geographic entity description information, identifying discriminative target words in the name information based on regular expressions, adding enhancement tags to the target words to obtain enhanced name information; concatenating the enhanced name information and address information in at least two geographic entity description information into sentence pairs according to a preset structure; based on the enhancement tags, segmenting the sentence pairs with the goal of segmenting the target word into a single word to obtain a word segmentation sequence; in the word segmentation sequence, the target word is segmented into a single word; and performing word embedding, sentence embedding, and position embedding on the word segments in the word segmentation sequence to obtain the input sequence of the geographic semantic feature encoder.

[0053] In this embodiment, the distinctive text, i.e., target words, in the name information is identified and labeled based on regular expressions. In some embodiments, target words are implemented as key words such as directional words or construction batch words. For example, directional words include "East District"; construction batch words include "Phase I" and "Phase II". These are extracted using regular expression matching, and preset enhancement markers, such as [F_START] and [F_END], are added before and after them to enhance and highlight the directional words or construction batch words. By identifying distinctive target words in the name information and adding enhancement markers, the target words form independent segmentation units after word segmentation and obtain more prominent feature expressions during the embedding process. For example, the target words will be identified as a complete semantic unit in subsequent word segmentation processes, avoiding further splitting. In this way, when the geographic semantic feature encoder performs the self-attention mechanism, it will assign higher attention weights to this type of word segmentation based on special markers, thereby highlighting the differences in semantic expression of the marked content. This enables the model to more accurately distinguish the structural differences with key influence in the description information of different geographic entities, thereby improving the accuracy of geographic semantic similarity calculation.

[0054] In this embodiment, after enhancement, the processed name and address information from at least two geographic entity descriptions are concatenated according to a preset structure to form sentence pairs. When segmenting these sentence pairs, since the target word is already wrapped by the enhancement tag, the target word and the enhancement tag form a complete unit and are not split, thus allowing for centralized expression. In other words, the target word is identified as an independent and important segment. Furthermore, word embedding, sentence embedding, and positional embedding are performed on the segmented words in the segmentation sequence. Word embedding represents the specific semantic meaning of the segmented words; sentence embedding distinguishes the descriptive information of different geographic entities; and positional embedding characterizes the relative or absolute position of the segmented words in the segmentation sequence.

[0055] By performing regularization enhancement, word segmentation, and word embedding, sentence embedding, and location embedding mapping on geographic entity description information, the input sequence achieves structured representation and text normalization before entering the geographic semantic feature encoder. This processing method effectively reduces noise interference, format differences, and expression inconsistencies in name and address information from different data sources, improving the semantic consistency of the input sequence. This enables the geographic semantic feature encoder to extract geographic semantic features more stably and accurately. Consequently, it enhances the reliability of semantic modeling and provides a more accurate semantic basis for subsequent determinations of whether they refer to the same geographic entity.

[0056] Then, the input sequence is input into the encoding network of the geographic semantic feature encoder to perform multi-layer bidirectional contextual semantic encoding based on the attention mechanism to obtain geographic semantic features of at least two geographic entity description information; and the geographic semantic features of at least two geographic entity description information are used as input to the first similarity network of the geographic semantic feature encoder for similarity calculation.

[0057] In this embodiment, the implementation method of the target encoding network is not limited. For example, the target encoding network can be implemented using Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), Self-Attention structures, residual connection structures, or network architectures based on an Encoder-Decoder framework. In this embodiment, bidirectional contextual semantic encoding refers to extracting features from the input sequence in both forward and backward directions to obtain forward semantic features and backward semantic features. By simultaneously utilizing the forward and backward semantic features of the input sequence, the feature representation at any position can be combined with its preceding and following context to obtain a more complete semantic expression.

[0058] By performing multi-layer bidirectional contextual semantic encoding based on an attention mechanism in the encoding network, the semantic relationships within the input sequence and between different geographic entity descriptions can be fully explored. This mechanism can capture long-distance dependencies and automatically focus on text units that contribute to semantic understanding, thereby generating more complete and stable geographic semantic features. Therefore, even if different data sources describe the same geographic entity in significantly different ways, this embodiment can still perform semantic alignment through contextual information, improving the ability to discriminate geographic semantic similarity.

[0059] The first similarity network is a network used to determine semantic relationships based on the input geographic semantic features. The first similarity network can output a result representing the semantic similarity of the descriptive information of two geographic entities, i.e., geographic semantic similarity, based on the encoded semantic features.

[0060] For example, when the name and address information of two geographic entities are highly semantically consistent, the first similarity network will output a high geographic semantic similarity; conversely, when the name and address information of two geographic entities are significantly different semantically, the first similarity network will output a low geographic semantic similarity. This processing step converts geographic semantic features into comparable similarity indicators, thus obtaining the geographic semantic similarity between at least two geographic entity descriptions.

[0061] It should be understood that this embodiment does not limit the specific structural form of the first similarity network. Whether a classification network or a regression network is used, as long as it can output a result representing the semantic similarity between the description information of at least two geographic entities based on geographic semantic features, it falls within the protection scope of this application.

[0062] In some embodiments, at least two semantic features are concatenated, subtracted, multiplied, or any combination thereof are used as input features for similarity calculation. Subsequently, the constructed input features are fed into a first similarity network, which models the correlation between two geographic semantic features through one or more nonlinear mapping layers, and outputs a similarity result characterizing the semantic proximity between the two features. When the first similarity network is a classification network, its output can be the probability of "pointing to the same geographic entity" and "pointing to different geographic entities"; when the first similarity network is a regression network, its output can be a similarity score within a preset range as the geographic semantic similarity.

[0063] In this embodiment, we will use at least two geographic entity descriptions as an example, where each of the two geographic entity descriptions is a cell entry from a third-party library and an ontology library, respectively. For a cell entry from the third-party library τ... Numbering =1,2,3…τ), for cell entries from ontology β, with Numbering =1,2,3…β). Geographic semantic similarity features Used to characterize the third-party library The first community entry and the first in the ontology. The degree of similarity between the names and addresses of each community entry at the semantic level. Specifically, record... The geographic semantic features corresponding to the name and address information of the j-th cell entry in the third-party library. For the first in the ontology library The geographic semantic features corresponding to the name and address information of each community entry.

[0064] In an optional embodiment, in the geographic semantic feature encoder, the geographic semantic features corresponding to the name and address information of the j-th cell entry in the third-party library can be calculated first. The geographic semantic features corresponding to the name and address information of the i-th cell entry in the ontology. Then, the geographic semantic features are further calculated. and The similarity, i.e. the encoded representation corresponding to the [CLS] position, serves as a semantic similarity feature. This feature reflects the degree of semantic matching between the descriptive information of two geographic entities, i.e., geographic semantic similarity, which can be expressed as: .

[0065] Further, optionally, when using a geographic semantic feature encoder to obtain the geographic semantic features of any geographic description information, the following implementation methods can be adopted, but are not limited to: In one alternative implementation, the input sequence includes segmentation markers for segmenting descriptive information of at least two geographic entities, such as the SEP mentioned above. Based on this, the input sequence can be fed into the encoding network of a geographic semantic feature encoder, where the following operations are performed: Based on the segmentation markers, multi-level bidirectional self-attention calculations are performed on the name and address information of the same geographic entity description information to obtain the fused semantic features of the same geographic entity description information itself. Based on the segmentation markers, multi-layer bidirectional cross-attention calculations are performed on the name information of different geographic entity description information to obtain cross-semantic features in the name dimension. Based on the segmentation markers, multi-layer bidirectional cross-attention calculations are performed on the address information of different geographic entity description information to obtain cross-semantic features in the address dimension. Global attention is used to calculate the fusion semantic features of each geographic entity description information in the input sequence, as well as the cross semantic features of different geographic entity description information in the name dimension and the address dimension, so as to obtain the geographic semantic features of each geographic entity description information.

[0066] In the above embodiments, with the goal of obtaining more accurate geographic semantic features of geographic entity description information, an innovative breakthrough is made in the use of attention mechanisms. On the one hand, local attention calculation and global attention calculation are integrated. On the other hand, self-attention mechanism and cross-attention mechanism are integrated for local attention calculation. Specifically, self-attention calculation is performed on the name information and address information of the same geographic entity description information, cross-attention calculation is performed on the name information of different geographic entity description information, and cross-attention calculation is performed on the address information of different geographic entity description information. This multi-layered and multi-dimensional approach, which integrates global attention computation with local attention computation, and self-attention computation with cross-attention computation, can better combat noise interference such as missing fields and typos in the name and address information of the same geographic entity description information. For the name or address information of different geographic entities description information, it can achieve field alignment, capture fine-grained differences, and improve the model's semantic understanding ability. This enables the model to more accurately capture the geographic semantic features of each geographic entity description information. For example, it can more accurately understand the semantic relevance between "XX Meihuali Garden" and "Meili Garden", and accurately distinguish between highly similar names such as "Meili Garden South District" and "Meili Garden North District" that point to different geographic entities. This provides an accurate data foundation for subsequent identification of whether they are the same geographic entity based on geographic semantic features.

[0067] In an optional embodiment, the geographic semantic feature encoder can be implemented as a large model with reasoning and generation capabilities, such as a large language model. Taking the implementation of the geographic semantic feature encoder as a large language model as an example, this implementation introduces a chain of thought to guide the encoding network to complete the encoding objective according to the set execution steps and order. In this embodiment, the encoding objective is to generate geographic semantic features containing descriptions of at least two geographic entities.

[0068] Furthermore, in this optional embodiment, before performing multi-layer bidirectional contextual semantic encoding based on an attention mechanism, the method further includes: based on the information dimensions and encoding objectives included in the input sequence, invoking a target language model to generate a target thought chain. The information dimensions include name information dimensions and address information dimensions. The target language model can be a language model for language or text processing, such as, but not limited to, a large language model.

[0069] In an optional embodiment, the process of generating the target thought chain includes: constructing prompt information for calling the target language model based on the encoded target, wherein the prompt information is used to indicate to the target language model the respective information dimensions of the name information and address information in the input sequence and the semantic reasoning task to be completed; inputting the prompt information and the input sequence into the target language model, wherein the target language model generates the reasoning steps required to achieve the encoded target and their execution order based on natural language understanding and reasoning capabilities, and outputs the target thought chain.

[0070] The target thought chain includes the execution steps required to guide the model to achieve the encoding target, as well as the execution order between these steps. During operation, the encoding network processes data step-by-step according to the steps described in this target thought chain.

[0071] In this embodiment, the target thought chain includes a structural parsing step, a spatial reasoning step, and a fusion encoding step. Based on this, the target thought chain and the input sequence can be input into the encoding network of the geographic semantic feature encoder, where the following operations are performed: Guided by the structural parsing step, the name information in the input sequence is structurally parsed according to the preset information hierarchy to obtain the hierarchical information that can represent the region level in the name information, and the corresponding hierarchical labels are added to the name information to transform the name information into a structured sequence with hierarchical labels to obtain the first intermediate sequence. Guided by the spatial reasoning steps, the address information in the first intermediate sequence is validated and corrected using the address knowledge base to obtain the second intermediate sequence; the validation includes, but is not limited to, hierarchical consistency, regional affiliation, and address field integrity. Guided by the fusion encoding step, the name information with hierarchical information and the corrected address information in the second intermediate sequence are subjected to multi-layer bidirectional contextual semantic encoding based on an attention mechanism. This enables the encoding network to simultaneously capture the hierarchical structure in the name information and the semantic relationships in the address information, thereby obtaining the geographic semantic features of at least two geographic entity description information. For the specific implementation of the attention-based multi-layer bidirectional contextual semantic encoding, please refer to the aforementioned embodiments. The only difference is the input; the implementation principle can be compared analogously to the aforementioned embodiments and will not be repeated here.

[0072] In this embodiment, guided by the structural parsing step, the encoding network parses name information according to a preset information hierarchy model. For example, it identifies hierarchical relationships in the name, such as regional levels. In some embodiments, the regional level can be divided into three levels, from top to bottom: first-level region, second-level region, and third-level region. Each higher-level region contains multiple lower-level regions, forming a multi-level hierarchical relationship from top to bottom. Hierarchical information is added before and after the parsed words to obtain a first intermediate sequence. Further, guided by the spatial reasoning step, the encoding network compares the address information in the first intermediate sequence with an address knowledge base for rationality verification and correction. If the address information fails the rationality verification, it is corrected according to the knowledge base to obtain a second intermediate sequence. Then, guided by the fusion encoding step, the encoding network performs multi-layer bidirectional contextual semantic encoding on the second intermediate sequence based on an attention mechanism to generate geographic semantic features for each geographic entity.

[0073] In this embodiment, by introducing a target thought chain and executing structural parsing, spatial reasoning, and fusion coding steps in the encoding network, the geographic semantic feature encoder can have a clear processing path during the encoding process. Specifically, the structural parsing step can supplement hierarchical information into the name information, the spatial reasoning step can improve the rationality of the address information based on the address knowledge base, and the fusion coding step can jointly model the processed name and address information. This processing method can reduce the interference caused by non-standard address representations on feature extraction, making the generated geographic semantic features more accurate and reliable.

[0074] Furthermore, such as Figure 2a As shown, a distance feature encoder is used to calculate the spatial distance features between at least two geographic entity descriptions based on their spatial location information. For example, the spatial location information from the descriptions of at least two geographic entities is input into the distance feature encoder, which converts the spatial location information into spatial distance features. In this embodiment, the implementation method of converting spatial location information into spatial distance features using the distance encoder is not limited.

[0075] In an optional embodiment, the following operations are performed in the distance feature encoder: based on spatial location information, the distance information between at least two geographic entity descriptions is calculated; the distance information is mapped from the numerical space to the vector space using an embedding layer to obtain the spatial distance features between the at least two geographic entity descriptions. This method is relatively simple and efficient. In this embodiment, the method of calculating the distance information is not limited. For example, it can be the Haversian distance or spherical cosine distance between the latitude and longitude of two cells. However, the distance information calculated based on spatial location information is a continuous value, and the range of values ​​may be large (e.g., from several meters to several kilometers). Directly inputting it into the model will introduce scale differences of different orders of magnitude. Therefore, this application embodiment provides an improved implementation method. In this implementation method, continuous distance values ​​are converted into discrete distance intervals by a preset distance interval rule, thereby eliminating interference from different scales. Specifically, in the improved implementation method, the spatial location information in the descriptions of at least two geographic entities is input into the distance feature encoder, and the distance encoder converts the spatial location information into spatial distance features. Specifically, the following operations are performed in the distance feature encoder: based on spatial location information, the distance information between at least two geographic entity descriptions is calculated; according to a set distance variation rule, the distance information is mapped to a target distance interval, the target distance interval has an interval number, and the interval number represents the position of the target distance interval in a series of sequentially arranged distance intervals; the interval number of the target distance interval is mapped from the numerical space to the vector space using an embedding layer to obtain the spatial distance features between at least two geographic entity descriptions.

[0076] In the improved implementation described above, distance information is mapped to a corresponding distance interval, i.e., the target distance interval, according to the set distance interval rules, and an interval number is assigned to the target distance interval. The interval number represents the position of the target distance interval within a series of sequentially arranged distance intervals. Since the interval number is a discrete variable, such as... Figure 2aAs shown, this embodiment introduces an embedding layer into the distance feature encoder to map interval numbers from the numerical space to the vector space, thereby obtaining spatial distance features between descriptions of at least two geographic entities. In this embodiment, the interval numbers have stable interval semantics, and the spatial distance features formed by mapping the interval numbers through the embedding layer enable distance information to participate in modeling in vector form. Therefore, this embodiment can achieve stable and effective spatial distance discrimination, significantly enhancing the robustness of geographic entity matching. This embodiment does not limit the mapping method and its corresponding mapping structure through the embedding layer. For example, the interval numbers can be used as ordinal numerical inputs, and through linear transformation and activation functions, such as ReLU (Rectified Linear Unit), the interval numbers can be mapped from the numerical space to the vector space, outputting spatial distance features. Alternatively, a fully connected layer can be used to map the interval numbers, that is, the interval numbers can be used as inputs, and the spatial distance features can be output through the fully connected layer.

[0077] Further optional, such as Figure 2a As shown, the geographic context feature encoder includes a graph attention network and a second similarity network. In this optional embodiment, when calculating the geographic context similarity between at least two geographic entity descriptions based on their respective surrounding geographic environment information using the geographic context feature encoder, the process includes: constructing a geographic context relationship graph corresponding to each of the at least two geographic entity descriptions based on their respective surrounding geographic environment information; inputting the at least two geographic entity descriptions and their corresponding geographic context relationship graphs into the graph attention network; aggregating the surrounding geographic environment information of each of the at least two geographic entity descriptions to their respective described target geographic entities based on their respective geographic context relationship graphs to obtain the geographic context features of the at least two geographic entity descriptions; and inputting the geographic context features of the at least two geographic entity descriptions into the second similarity network for similarity calculation to obtain the geographic context similarity between the at least two geographic entity descriptions. By constructing a geographic context relationship graph based on geographic entity descriptions and their surrounding geographic environment information, and using the graph attention network to aggregate the surrounding geographic environment information represented by neighboring nodes to the central node, the geographic context features of the geographic entities can be effectively extracted. The similarity calculation results between these geographic context features enable the model to identify whether the description information of at least two geographic entities points to the same geographic entity based on the surrounding environment information, which can significantly improve the robustness of matching.

[0078] In this embodiment, address information and / or spatial location information contained in the description information of at least two geographic entities are obtained. Based on the surrounding geographic environment information corresponding to each address information and / or spatial location information, a corresponding geographic context graph is constructed for each geographic entity. The geographic context graph includes at least nodes and edges. Nodes include the central node of the target geographic entity and the neighboring nodes of the central node. The central node represents the subject to which the geographic entity description information points. Neighboring nodes represent the surrounding geographic environment information of the target geographic entity. Edges represent the relationships between nodes, mainly reflecting the geographic connections between the target geographic entity and its surrounding geographic environment.

[0079] In this embodiment, the description information of two geographic entities and their respective constructed geographic context relationship graphs are input into a graph attention network. The graph attention network dynamically evaluates the importance of each surrounding environmental node through an attention weight mechanism; and aggregates the geographic environmental information with higher importance weights to the corresponding target geographic entity node, thereby obtaining the geographic context features of the geographic entity description information.

[0080] Furthermore, the geographic context features corresponding to the at least two geographic entity descriptions obtained above are input into a second similarity network for similarity calculation. The second similarity network outputs the geographic context similarity between the at least two geographic entity descriptions. The second similarity network can be a classification network or a regression network, without limitation. The second similarity network can be implemented with reference to the aforementioned first similarity network, with the main difference being in the input and output: the second similarity network calculates similarity based on geographic context features and outputs the geographic context similarity.

[0081] In an optional embodiment, when constructing a geographic context graph corresponding to each of the at least two geographic entity description information based on the geographic environment information surrounding each of the at least two geographic entity description information, the method includes: for any geographic entity description information, marking the target geographic entity described by the geographic entity description information on a map according to the address information and / or spatial location information in the geographic entity description information; determining a target area on the map with the target geographic entity as the center, and selecting points of interest that are compatible with the target geographic entity from the target area; adding edges between the center node and the neighbor nodes, with the target geographic entity as the center node and the points of interest as neighbor nodes, to obtain the geographic context graph corresponding to the geographic entity description information; wherein, the nodes in the geographic context graph represent the encoding features of the names of the target geographic entity or points of interest, and the edges represent the spatial relationship features between the center node and the neighbor nodes.

[0082] In this embodiment, for any geographic entity description information, based on the address information and / or spatial location information (e.g., latitude and longitude coordinates) contained in the geographic entity description information, the target geographic entity pointed to by the geographic entity description information is marked on the map. Then, a target area is determined on the map centered on the target geographic entity, for example, a preset range (e.g., 200 meters, 500 meters, or other ranges) is delineated centered on the target geographic entity. The shape of the target area is not fixed; for example, it can be a circular area; or it can be a square area; or it can be an irregular area.

[0083] Furthermore, geographic environmental information surrounding the target geographic entity is retrieved from the target area, such as points of interest (POIs), including but not limited to residential areas, office buildings, commercial centers, bus stops, subway stations, schools, etc. From the candidate POIs, several POIs with high relevance to the target geographic entity are selected, for example, by sorting them by distance to obtain the Top-K POIs, where K is a positive integer and its value can be flexibly set, such as 10, 15, 30, etc. Further, the target geographic entity is used as the central node, and the POIs are used as neighboring nodes, and a graph structure is constructed based on the spatial relationships between the nodes. In this embodiment, edges are added between the central node and each neighboring node to form a geographic context graph corresponding to the geographic entity description information.

[0084] In one example, Figure 2b This is a schematic diagram illustrating the process of constructing a geographic context graph based on a graph attention network. The graph on the left (hereinafter referred to as the left graph) is constructed based on the geographic entity description information A1 mentioned above, and the graph on the right (hereinafter referred to as the right graph) is constructed based on the geographic entity description information A2. In this embodiment, the left graph is used as an example for explanation. For information on the right graph, please refer to the relevant description of the left graph.

[0085] like Figure 2b As shown, the central node is located in the central area of ​​the left image, and it is the node corresponding to the target geographic entity. Correspondingly, the circles labeled 1, 2, 3, 4, and 5 represent the neighboring nodes of the central node; in other words, the points of interest surrounding the target geographic entity. β1, β2, β3, β4, and β5 are the edges between the neighboring nodes and the central node, forming a network as shown. Figure 2b The diagram on the left shows the geographic context.

[0086] Optionally, when aggregating the geographic environment information surrounding each of the at least two geographic entity description information to the target geographic entity they describe, based on the geographic context relationship graphs corresponding to the at least two geographic entity description information, to obtain the geographic context features of the at least two geographic entity description information, the method includes: for any geographic entity description information corresponding to the geographic context relationship graph, linearly mapping the encoded features represented by each node in the geographic context relationship graph based on a shared weight matrix to obtain the mapping features represented by each node; calculating the attention coefficient between the central node and neighboring nodes based on the mapping features represented by each node and the spatial relationship features represented by the edges in the geographic context relationship graph; performing a weighted summation of the mapping features represented by the neighboring nodes according to the attention coefficient to obtain the aggregated features; and generating the geographic context features of the geographic entity description information based on the aggregated features and the mapping features represented by the central node.

[0087] In this embodiment, for any geographic entity description information, node feature processing is performed on its corresponding geographic context graph. The geographic context graph includes a central node and multiple neighboring nodes. Each node contains its corresponding encoded features. In this embodiment, a shared weight matrix is ​​set in the graph attention network. By linearly mapping the encoded features of each node, the mapped features represented by each node are obtained. The shared weight matrix refers to a set of weight parameters used when linearly transforming the features of all nodes in the graph attention network. By using the shared weight matrix to linearly map the encoded features of each node, it can be ensured that the central node and neighboring nodes are mapped to the same feature space, facilitating subsequent calculation of attention coefficients.

[0088] Furthermore, based on the mapping features of each node in the geographic context graph and the spatial relationship features carried by the edges, the attention coefficient between the central node and each neighboring node is calculated. The attention coefficient measures the strength of the relationship between each neighboring node and the central node, thereby dynamically adjusting the weights of different neighboring nodes during subsequent feature aggregation, allowing them to contribute different levels of feature information according to their importance. For example, the attention coefficient is larger when the semantic features of a neighboring node are closer to those of the central node, or when they are closer in distance. Conversely, the attention coefficient is smaller when the semantic features of a neighboring node differ significantly from those of the central node, or when they are farther apart.

[0089] Furthermore, based on the aforementioned attention coefficients and the mapping features represented by neighboring nodes and the central node, geographic context features for describing any geographic entity can be generated. In this application embodiment, the specific implementation method for generating geographic context features for describing any geographic entity based on the aforementioned attention coefficients and the mapping features represented by neighboring nodes and the central node is not limited. Examples are provided below: In an optional embodiment, the method for generating geographic context features of any geographic entity description information based on the aforementioned attention coefficients and the mapping features represented by neighboring nodes and the central node includes: weighted summation of the mapping features represented by neighboring nodes based on the attention coefficients to obtain aggregated features, which refer to the context information aggregated by the central node from neighboring nodes. During the aggregation process, closer neighboring nodes may receive higher attention coefficients, thus allocating more weight to them. Further, geographic context features of the geographic entity description information are generated based on the aggregated features and the mapping features of the central node. These geographic context features simultaneously reflect the semantic features of the target geographic entity itself and the aggregated features of its surrounding geographic environment, providing a more comprehensive feature representation and thereby improving the accuracy of geographic entity matching.

[0090] In one example, For the first Geographic context features of a geographic entity; This represents the number of neighboring nodes connected to the central node, and is a positive integer. Indicates the first Attention coefficients between each neighboring node and the central node; For the first Mapping characteristics of neighboring nodes; The mapping feature of the central node itself; ReLU is an example of a nonlinear activation function. In the above formula, ReLU means that the nonlinear activation function is used to perform nonlinear processing on the calculation result in parentheses.

[0091] In another alternative embodiment, an attention mechanism is introduced into the graph neural network. In the graph neural network, the attention mechanism, based on the aforementioned attention coefficients and combined with the mapping features represented by neighboring nodes and the central node, generates geographic context features for any geographic entity description information. Furthermore, self-attention and cross-attention mechanisms are simultaneously introduced into the graph neural network to fully extract the association relationships between the central node and various types of neighboring nodes. Specifically, neighboring nodes can be divided into multiple categories based on their attributes; self-attention is calculated between the central node and neighboring nodes of the same category based on the attention coefficients to obtain the first context features between the central node and neighboring nodes of the same category; cross-attention is calculated between the central node and neighboring nodes of different categories based on the attention coefficients to obtain the second context features between the central node and neighboring nodes of different categories; and geographic context features for geographic entity description information are generated based on the first and second context features.

[0092] In the above optional embodiments, the specific implementation method of classifying neighboring nodes is not limited. For example, neighboring nodes can be divided into different types based on their distance from the central node, such as neighboring nodes within 100 meters, neighboring nodes within 100-500 meters, and neighboring nodes within 500-1000 meters. Alternatively, neighboring nodes can be divided into different types based on their function, such as small shops and supermarkets as one type, educational institutions such as schools and kindergartens as another, transportation stations such as buses and subway stations as yet another, and office buildings and commercial buildings as yet another.

[0093] By classifying neighboring nodes, and considering that neighboring nodes of the same type are similar or identical in function, distance, or other classification dimensions, a self-attention mechanism is used to calculate the contextual features between this type of neighboring node and the central node. This enables semantic homogeneity modeling and strengthens the contextual contribution of neighboring nodes of the same type (or the same semantic dimension) to the central node. For example, it can enhance the transportation convenience and convenience services of a geographic entity, and reduce the contextual interference caused by neighboring nodes of different types (or other semantic dimensions). For neighboring nodes of different types, considering that neighboring nodes are different or dissimilar in function, distance, or other classification dimensions, a cross-attention mechanism is used to calculate the contextual features between this type of neighboring node and the central node. This enables semantic heterogeneity modeling, dynamically capturing the differentiated contributions of different types of neighboring nodes (or different semantic dimensions) to the central node, and accurately characterizing the surrounding environmental information of geographic entities. For example, the surrounding environmental information of geographic entities can be comprehensively considered from multiple aspects such as transportation convenience, public services, and educational institutions. Finally, the geographic context features generated from these two types of context features can highlight both single-type and comprehensive types of surrounding environment information, and more accurately describe the geographic context features of geographic entity description information, providing a precise and comprehensive data foundation for subsequent identification of whether they belong to the same geographic entity based on geographic context features.

[0094] Furthermore, the embodiments of this application are not limited to specific implementation methods for generating geographic context features of geographic entity information based on the first context feature and the second context feature. For example, the first context feature and the second context feature can be concatenated to obtain the geographic context feature, or the first context feature and the second context feature can be weighted and summed to obtain the geographic context feature, and so on.

[0095] In an optional embodiment, before linearly mapping the encoded features represented by each node in the geographic context graph based on the shared weight matrix, the method further includes: selecting a benchmark neighbor node from the neighbor nodes in the geographic context graph; the benchmark neighbor node is an interest point with a popularity greater than a set popularity threshold, an interest point of a specific type, or an interest point that is closest to the center node; enhancing the features represented by the benchmark neighbor node in conjunction with a geographic knowledge graph; and / or increasing the number of benchmark neighbor nodes based on the number of neighbor nodes in the geographic context graph. By selecting interest points with higher popularity, interest points of a specific type, or interest points that are closer to the center node as benchmark neighbor nodes, the model can prioritize contextual information that is more representative or relevant to the central geographic entity, thereby reducing interference from noisy nodes. Furthermore, introducing a geographic knowledge graph to enhance the features represented by the benchmark neighbor node can supplement more feature information of the benchmark neighbor node, making the neighbor features more complete and accurate.

[0096] In this embodiment, to avoid semantic dilution caused by all points of interest participating in the encoding on an average basis, for any geographic context graph corresponding to the description information of a geographic entity, neighboring nodes with higher representativeness are selected as benchmark neighboring nodes according to preset rules. Benchmark neighboring nodes include, for example, points of interest with a popularity greater than a set popularity threshold, such as shopping malls, hospitals, schools and other frequently visited locations; for example, they also include points of interest of specific types, such as educational, medical or transportation hub points of interest; and for example, they also include points of interest that are closest to the target geographic entity.

[0097] In this embodiment, the features represented by the baseline neighbor nodes are enhanced by combining a geographic knowledge graph. In some embodiments, the geographic knowledge graph may include information such as the category hierarchy of points of interest, functional attributes, spatial relationships with landmarks or roads, and industry relationships or semantic associations between points of interest.

[0098] In some embodiments, a geographic knowledge graph is constructed based on structured geographic data. This involves extracting the categories, attributes, and spatial proximity relationships of points of interest through a map data interface, structuring the above information, and forming a knowledge graph that includes nodes (points of interest) and their attributes, as well as the relationships between nodes.

[0099] In other embodiments, the knowledge graph is constructed based on spatial and semantic analysis. That is, by using the geographical location, relative distance and name semantics of points of interest, the association between entities is identified through spatial clustering or text relation extraction and then added to the knowledge graph.

[0100] For example, a geographic knowledge graph contains semantic information such as the category and industry relationship of points of interest. By associating it with the knowledge graph, the semantic expression of the points of interest nodes can be completed or enhanced, making it easier for the model to understand the basic attributes and importance of neighboring nodes.

[0101] In this embodiment, when enhancing the number of baseline neighbor nodes based on the number of neighbor nodes in the geographic context graph, the number of neighbor nodes associated with the central node in the geographic context graph can be judged. If the number of neighbor nodes is lower than a preset threshold, points of interest with high relevance can be selected from the geographic knowledge graph to increase the number of neighbor nodes and make the structure of the geographic context graph more complete. If the number of neighbor nodes is not lower than the preset threshold, there is no need to supplement the number of baseline neighbor nodes. Through this number enhancement strategy, even when neighbor nodes are sparse or lack sufficient information, the model can still obtain sufficient and representative contextual information, thereby improving the stability and robustness of node feature representation.

[0102] In one example, in this embodiment, points of interest around a geographic entity are modeled to characterize the surrounding geographic environment information. For the first... A geographic entity has several points of interest surrounding it. The key information of a point of interest is represented as a tuple. This tuple includes: and . Used to indicate the first The geographical entity surrounding the first Name information for each point of interest; Used to indicate that the point of interest is related to the first Spatial distances between geographic entities. By introducing the above tuples, both the name and distance information of points of interest can be captured simultaneously, thus providing a richer environmental description for the geographic context feature encoder and helping to improve the model's discriminative ability.

[0103] In some embodiments, the name information and / or distance information represented by the baseline neighbor nodes are augmented, normalized, or semantically enhanced to improve the multidimensional feature representation of the points of interest. For example, by appropriately augmenting, normalizing, or semantically enhancing the name and distance information of the baseline neighbor nodes, interference caused by missing information, incomplete representation, or numerical anomalies can be reduced. At the same time, by enhancing the expressive power of key dimensions such as name and distance, the distinguishability between different points of interest can be improved, thereby enhancing the accuracy of subsequent encoding and similarity calculation.

[0104] Furthermore, in this embodiment, a classification head network is used to classify and predict geographic semantic similarity, spatial distance features, and geographic context similarity to determine whether at least two geographic entity descriptions point to the same geographic entity. This embodiment provides several different classification strategies, which are discussed below. Figures 3a-3b Let me introduce it.

[0105] Strategy 1: such as Figure 3a As shown, geographic semantic similarity, spatial distance features, and geographic context similarity are respectively input into the three-way sub-classification head network for classification prediction to obtain three initial classification results; the three initial classification results are then input into the overall classification head network for fusion prediction to obtain classification results on whether at least two geographic entity descriptions point to the same geographic entity. In Strategy 1 above, geographic semantic similarity is input into the first sub-classification head network to obtain the first initial classification result; spatial distance features are input into the second sub-classification head network to obtain the second initial classification result; and geographic context similarity is input into the third sub-classification head network to obtain the third initial classification result. The sub-classification head networks can employ a multilayer perceptron structure. Furthermore, the three initial classification results are input into the overall classification head network for fusion and judgment to generate the final classification result. By inputting geographic semantic similarity, spatial distance features, and geographic context similarity into the three sub-classification head networks respectively, the model can learn independently for different types of features. Furthermore, fusing the three initial classification results through the overall classification head network can integrate complementary information between the three types of features at a higher level, improving discrimination stability and thus enhancing overall matching accuracy.

[0106] Strategy Two: Figure 3b As shown, geographic semantic similarity, spatial distance features, and geographic context similarity are concatenated and then input into a classification head network for classification prediction to obtain classification results on whether at least two geographic entity descriptions point to the same geographic entity.

[0107] In Strategy 2 above, geographic semantic similarity, spatial distance features, and geographic context similarity are concatenated according to feature dimensions to form a fused feature. This fused feature is then input into the classification head network, which outputs a classification result indicating whether at least two geographic entity descriptions point to the same geographic entity. By concatenating these three types of features and inputting them into the classification head network, the model can simultaneously utilize geographic semantics, spatial distance, and geographic context similarity within the same classification path, achieving joint learning and enhancing the interaction capabilities between features. This approach is simple in structure, highly efficient, and can improve overall matching performance in scenarios with inconsistent descriptive information.

[0108] Strategy 3: Geographic semantic similarity, spatial distance features, and geographic context similarity are input into three sub-classification head networks for classification prediction to obtain three initial classification results. Then, geographic semantic similarity, spatial distance features, geographic context similarity, and the three initial classification results are input into a final classification head network for fusion prediction to output a classification result indicating whether at least two geographic entity descriptions point to the same geographic entity. Each sub-classification head network performs independent classification prediction based on a single feature; the final classification head network jointly fuses and learns the above three types of features and the three initial classification results to output a classification result indicating whether at least two geographic entity descriptions point to the same geographic entity.

[0109] By inputting the three types of features into three sub-classification head networks to obtain initial classification results, and then inputting these initial classification results along with the three types of features into the overall classification head network, a two-layer semantic fusion mechanism of "original features + high-level prediction" can be constructed. On the one hand, the three types of features provide low-level information for the overall classification head network; on the other hand, the three initial classification results can serve as a high-level abstract expression of the three types of features, directly using the classification intent of each feature dimension as auxiliary input. This two-layer structure can significantly enhance the model's ability to analyze complex information, thereby further improving the robustness of matching.

[0110] In the embodiments of this application, the aforementioned geographic entity information processing method can be executed based on pre-trained models. During training, each feature encoder can be trained separately, or a joint training approach can be used to uniformly optimize the geographic semantic feature encoder, distance feature encoder, geographic context feature encoder, and classification head network.

[0111] In the following embodiments, a joint training approach is adopted, enabling various features to learn collaboratively under a unified optimization objective, thereby further enhancing the overall model's discriminative ability. Furthermore, in this embodiment, when using the joint training approach, both a classification loss function and a ranking loss function are used for model training to obtain a geographic entity information processing model for implementing the geographic entity information processing process described in the foregoing method embodiments. This geographic entity information processing model includes the geographic semantic feature encoder, distance feature encoder, geographic context feature encoder, and classification head network mentioned in the foregoing embodiments. The following describes one joint training method.

[0112] Figure 4 This is a flowchart illustrating a model training method provided for an exemplary embodiment of this application. Figure 4As shown, when training the model based on data from the ontology library and third-party libraries, multiple first sample geographic entity description information is obtained from the third-party library; for each first sample geographic entity description information, at least one similar second sample geographic entity description information is recalled from the ontology library to form a recall context.

[0113] In this embodiment, as Figure 4 As shown, the first sample geographic entity description information and each second sample geographic entity description information in the recall context are combined to form sample pairs, and they are labeled according to whether they point to the same geographic entity. The labeling result is between a first value and a second value, with the first value being less than the second value. Specifically, a value of the first value indicates that they point to different geographic entities (negative samples), and a value of the second value indicates that they point to the same geographic entity (positive samples).

[0114] Furthermore, through methods such as Figure 4 The "Net" shown refers to the geosemantic feature encoder, distance feature encoder, geocontext feature encoder, and classification head network. For each sample pair, the Net computation includes: calculating sample geosemantic similarity, sample spatial distance feature, and sample geocontext similarity using the geosemantic feature encoder, distance feature encoder, and geocontext feature encoder, respectively; and the Net computation also includes: inputting the sample geosemantic similarity, sample spatial distance feature, and sample geocontext similarity into the classification head network, which outputs a first probability score and a second probability score through Softmax; the first probability score represents the probability of belonging to the same geographic entity; and the second probability score represents the probability of belonging to different geographic entities.

[0115] For each sample pair, the classification loss is calculated using the cross-entropy function based on the labeling results of the sample pair; the first probability scores are weighted and summed to obtain the corresponding positive ranking loss; the second probability scores are weighted and summed to obtain the corresponding negative ranking loss; a model loss function is constructed based on the classification loss, positive ranking loss, and negative ranking loss; the geographic semantic feature encoder, distance feature encoder, geographic context feature encoder, and classification head network are jointly trained according to the model loss function until the model loss meets the set conditions.

[0116] In this embodiment, based on the positive and negative sample training mechanism, the model is forced to pay attention not only to similar descriptive information (positive samples) during the learning process, but also to the subtle differences between itself and multiple candidates (i.e., multiple negative samples), thereby enhancing its ability to distinguish between "similar but different" geographical entities. In this way, the model's discriminative ability is significantly improved, further increasing the accuracy of the final mounting results.

[0117] The following describes an example of a joint training process, but it does not constitute a limitation on this embodiment.

[0118] like Figure 4 As shown, when training the model based on data from the ontology library and third-party libraries, the description information of the i-th first sample geographic entity from the third-party library is... It can retrieve descriptions of multiple similar second-sample geographic entities from the ontology database. This creates a recall context. Furthermore, any of the above sample pairs... Input the data into the corresponding encoder for encoding, and obtain the geographic semantic similarity scores respectively. Spatial distance characteristics and geographical context similarity .

[0119] Furthermore, the three types of features are concatenated according to feature dimensions and then input into the overall classification head network. The output is: Where, vector The probability score includes two components: idx=0 corresponding to "pointing to the same geographic entity"; idx=1 corresponds to... The probability score for "should point to different geographical entities". Since Softmax regression and Logistic regression are equivalent in binary classification scenarios when there are 2 categories, the dissimilarity probability output by Softmax when idx=1 can be written as: in, In response to The number of geographic entity descriptions retrieved in the second sample. (Targeting...) With each in the recall context For each sample pair formed, the model outputs a set of two-dimensional probability scores: Furthermore, the probability that the sample pairs are relevant given the current recall context and the input sample pairs is calculated as follows: in Given the current recall context Under the conditions, the first The second sample's geographical description information and The probability of the most relevant information can be defined as: Based on cross-entropy loss, the forward ranking loss is calculated as follows: Based on the above-mentioned forward ranking loss Building upon this foundation, to prevent negative samples from being ignored during training, a negative ranking loss can be further constructed. .

[0120] Specifically, for any sample pair Its corresponding tags This indicates whether the sample pair can be mounted. When At this point, the sample pair is considered irrelevant. In this case, the forward ranking loss... The absence of components from the sample pair may weaken its effectiveness during training.

[0121] Therefore, a negative ranking loss is constructed based on negative samples, and its definition is as follows: Based on the above embodiments, this embodiment further constructs a unified model training loss function to achieve joint training. Specifically, the final loss of the model... The definition is as follows: in, Primary classification loss; For semantic supervision loss; Loss is provided for monitoring geographic context. For positive sorting loss; This represents the negative sorting loss. Parameters These are adjustable weighting coefficients used to balance the contribution ratios of multiple tasks during joint training.

[0122] By employing the aforementioned joint training, the model can simultaneously acquire supervisory signals from semantic information, spatial information, geographic context information, and recall-level ranking relationships, achieving collaborative optimization and thus significantly improving matching accuracy and robustness.

[0123] In this embodiment, the loss function used is no longer a direct binary classification loss, but a ranking loss optimized based on the recall context. Compared to a simple classification loss, this forces the model to not only focus on similar descriptive information (positive samples) during the learning process, but also to pay attention to subtle differences between itself and multiple candidates (i.e., multiple negative samples), thereby enhancing its ability to distinguish between "similar but different" geographical entities. In this way, the model's discriminative ability is significantly improved, further increasing the accuracy of the final mounting results.

[0124] It should be noted that the implementation structure and working principle of each encoder or classifier head network during model training can be found in the descriptions in the foregoing embodiments. The only difference is that the descriptions in the foregoing embodiments are from the perspective of model inference. For those skilled in the art, the model training process can be easily deduced from the model structure and working principle during model inference, so it will not be described in detail here.

[0125] The detailed implementation methods and beneficial effects of each step in this embodiment have been described in detail in the foregoing embodiments, and will not be elaborated here.

[0126] It should be noted that the execution subject of each step of the method provided in the above embodiments can be the same device, or the method can be executed by different devices. For example, the execution subject of steps 101 to 103 can be device A; or the execution subject of steps 101 and 102 can be device A, and the execution subject of step 103 can be device B; and so on.

[0127] Furthermore, some processes described in the above embodiments and accompanying drawings include multiple operations appearing in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or they may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not represent any execution order. Additionally, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.

[0128] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 5 As shown, the electronic device includes a memory 54 and a processor 55.

[0129] Memory 54 is used to store computer programs and can be configured to store various other data to support operation on the electronic device. Examples of this data include instructions for any application or method used to operate on the electronic device, data structures, contact data, phone book data, messages, pictures, videos, etc.

[0130] Processor 55, coupled to memory 54, is configured to execute a computer program in memory 54 for: acquiring description information of at least two geographic entities, the description information including name information, address information, and spatial location information of the geographic entities; calculating the geographic semantic similarity between the at least two geographic entity description information based on the name information and address information in the at least two geographic entity description information using a geographic semantic feature encoder; calculating the spatial distance feature between the at least two geographic entity description information based on the spatial location information in the at least two geographic entity description information using a distance feature encoder; calculating the geographic context similarity between the at least two geographic entity description information based on the at least two geographic entity description information and their respective surrounding geographic environment information using a geographic context feature encoder; inputting the geographic semantic similarity, spatial distance feature, and geographic context similarity into a classification head network, and outputting a classification result indicating whether the at least two geographic entity description information points to the same geographic entity.

[0131] In an optional embodiment, when the processor 55 calculates the geographic semantic similarity between the at least two geographic entity description information based on the name information and address information in the at least two geographic entity description information using a geographic semantic feature encoder, it specifically performs the following: performs regularization enhancement, word segmentation, and embedding processing on the name information and address information in the at least two geographic entity description information to construct the input sequence of the geographic semantic feature encoder; inputs the input sequence into the encoding network in the geographic semantic feature encoder to perform multi-layer bidirectional contextual semantic encoding based on an attention mechanism to obtain the geographic semantic features of the at least two geographic entity description information; and inputs the geographic semantic features of the at least two geographic entity description information into the first similarity network in the geographic semantic feature encoder to calculate similarity to obtain the geographic semantic similarity between the at least two geographic entity description information.

[0132] In an optional embodiment, when the processor 55 performs regular expression enhancement, word segmentation, and embedding processing on the name information and address information in the at least two geographic entity description information to construct the input sequence of the geographic semantic feature encoder, it specifically performs the following steps: for the name information in any geographic entity description information, it identifies the discriminative target word in the name information based on regular expressions, adds enhancement tags to the target word to obtain enhanced name information; according to a preset structure, it concatenates the enhanced name information and address information in the at least two geographic entity description information into a sentence pair; based on the enhancement tags, with the goal of segmenting the target word into a single word, it segments the sentence pair to obtain a word segmentation sequence; in the word segmentation sequence, the target word is segmented into a single word; and it performs word embedding, sentence embedding, and position embedding on the word segmentation sequence to obtain the input sequence of the geographic semantic feature encoder.

[0133] In an optional embodiment, when the processor 55 inputs the input sequence into the encoding network of the geographic semantic feature encoder to perform multi-layer bidirectional contextual semantic encoding based on an attention mechanism to obtain the geographic semantic features of the at least two geographic entity description information, it is specifically configured to: input the input sequence into the encoding network of the geographic semantic feature encoder, and perform the following operations in the encoding network: Based on the segmentation markers, multi-layer bidirectional self-attention calculations are performed on the name and address information of the same geographic entity description information to obtain the fused semantic features of the same geographic entity description information itself. Based on the segmentation markers, multi-layer bidirectional cross-attention calculations are performed on the name information of different geographic entity description information to obtain cross-semantic features in the name dimension. Based on the segmentation markers, multi-layer bidirectional cross-attention calculations are performed on the address information of different geographic entity description information to obtain cross-semantic features in the address dimension. Global attention is calculated on the fused semantic features of each geographic entity description information in the input sequence, as well as the cross semantic features of different geographic entity description information in the name dimension and the address dimension, to obtain the geographic semantic features of each geographic entity description information.

[0134] In an optional embodiment, the processor 55 is further configured to: based on the information dimensions and encoding objectives included in the input sequence, invoke a target language model to generate a target thought chain, wherein the encoding objective is a geographic semantic feature for generating the description information of the at least two geographic entities, and the target thought chain includes a structural parsing step, a spatial reasoning step, and a fusion encoding step required to guide the geographic semantic feature encoder to achieve the encoding objective. Accordingly, when the processor 55 inputs the input sequence into the encoding network of the geographic semantic feature encoder to perform multi-layer bidirectional contextual semantic encoding based on an attention mechanism to obtain the geographic semantic features of the at least two geographic entity description information, it specifically performs the following operations: inputting the target thought chain and the input sequence into the encoding network of the geographic semantic feature encoder, and performing the following operations in the encoding network: under the guidance of the structure parsing step, performing structure parsing on the name information in the input sequence according to a preset information hierarchy and adding corresponding hierarchical information to the name information to obtain a first intermediate sequence; under the guidance of the spatial reasoning step, combining the address knowledge base to perform rationality verification and correction on the address information in the first intermediate sequence to obtain a second intermediate sequence; under the guidance of the fusion encoding step, performing multi-layer bidirectional contextual semantic encoding based on an attention mechanism on the name information with hierarchical information and the corrected address information in the second intermediate sequence to obtain the geographic semantic features of the at least two geographic entity description information.

[0135] In an optional embodiment, when the processor 55 calculates the spatial distance features between the at least two geographic entity descriptions based on the spatial location information in the at least two geographic entity descriptions using a distance feature encoder, it specifically performs the following operations: inputting the spatial location information in the at least two geographic entity descriptions into the distance feature encoder, and performing the following operations in the distance feature encoder: calculating the distance information between the at least two geographic entity descriptions based on the spatial location information; mapping the distance information to a target distance interval according to a set distance variation rule, wherein the target distance interval has an interval number, and the interval number represents the position of the target distance interval in a plurality of sequentially arranged distance intervals; and using an embedding layer to map the interval number of the target distance interval from the numerical space to the vector space to obtain the spatial distance features between the at least two geographic entity descriptions.

[0136] In an optional embodiment, the geographic context feature encoder includes a graph attention network and a second similarity network. When the processor 55 calculates the geographic context similarity between the at least two geographic entity descriptions based on the geographic environment information surrounding each of the at least two geographic entities using the geographic context feature encoder, it specifically performs the following steps: constructing a geographic context relationship graph corresponding to each of the at least two geographic entity descriptions based on the geographic environment information surrounding each of the at least two geographic entity descriptions; inputting the at least two geographic entity descriptions and their corresponding geographic context relationship graphs into the graph attention network; aggregating the geographic environment information surrounding each of the at least two geographic entity descriptions towards the target geographic entity they describe based on the geographic context relationship graphs, to obtain the geographic context features of the at least two geographic entity descriptions; and inputting the geographic context features of the at least two geographic entity descriptions into the second similarity network for similarity calculation, to obtain the geographic context similarity between the at least two geographic entity descriptions.

[0137] In an optional embodiment, when the processor 55 constructs a geographic context graph corresponding to each of the at least two geographic entity description information based on the geographic environment information surrounding each of the at least two geographic entity description information, it is specifically configured to: for any geographic entity description information, mark the target geographic entity described by the geographic entity description information on the map according to the address information and / or spatial location information in the geographic entity description information; determine a target area on the map with the target geographic entity as the center, and select a point of interest that matches the target geographic entity from the target area; take the target geographic entity as the center node, take the point of interest as the neighbor node, and add an edge between the center node and the neighbor node to obtain the geographic context graph corresponding to the geographic entity description information; wherein, the nodes in the geographic context graph represent the encoding features of the name of the target geographic entity or point of interest, and the edges represent the spatial relationship features between the center node and the neighbor node.

[0138] In an optional embodiment, when the processor 55 aggregates the geographic environment information surrounding each of the at least two geographic entity description information to the target geographic entity described by each of the at least two geographic entity description information based on the geographic context relationship graph corresponding to each of the at least two geographic entity description information, to obtain the geographic context features of the at least two geographic entity description information, the processor 55 specifically performs the following: for any geographic entity description information, based on a shared weight matrix, linearly maps the encoded features represented by each node in the geographic context relationship graph to obtain the mapping features represented by each node; calculates the attention coefficient between the central node and its neighboring nodes based on the mapping features represented by each node and the spatial relationship features represented by the edges in the geographic context relationship graph; and generates the geographic context features of the geographic entity description information based on the attention coefficient and the mapping features represented by the neighboring nodes and the central node.

[0139] In an optional embodiment, when the processor 55 generates the geographic context features of the geographic entity description information based on the attention coefficient and the mapping features represented by the neighboring nodes and the central node, it specifically performs the following steps: weighted summation of the mapping features represented by the neighboring nodes based on the attention coefficient to obtain aggregated features; generating the geographic context features of the geographic entity description information based on the aggregated features and the mapping features represented by the central node; or, classifying the neighboring nodes into multiple categories based on their attributes; performing self-attention calculation on the central node and neighboring nodes of the same category based on the attention coefficient to obtain a first context feature between the central node and neighboring nodes of the same category; performing cross-attention calculation on the central node and neighboring nodes of different categories based on the attention coefficient to obtain a second context feature between the central node and neighboring nodes of different categories; and generating the geographic context features of the geographic entity description information based on the first context feature and the second context feature.

[0140] In an optional embodiment, before linearly mapping the encoded features represented by each node in the geographic context graph based on the shared weight matrix, the processor 55 is further configured to: select a benchmark neighbor node from the neighbor nodes in the geographic context graph, wherein the benchmark neighbor node is an interest point with a heat value greater than a set heat value threshold, an interest point of a specific type, or an interest point that is the most recently located from the center node; enhance the features represented by the benchmark neighbor node in conjunction with the geographic knowledge graph; and / or enhance the number of the benchmark neighbor nodes based on the number of neighbor nodes in the geographic context graph.

[0141] In an optional embodiment, when the processor 55 inputs the geographic semantic similarity, spatial distance features, and geographic context similarity into the classification head network and outputs a classification result indicating whether the description information of at least two geographic entities points to the same geographic entity, it specifically performs the following steps: inputting the geographic semantic similarity, spatial distance features, and geographic context similarity into three sub-classification head networks for classification prediction to obtain three initial classification results; inputting the three initial classification results into the overall classification head network for fusion prediction to obtain a classification result indicating whether the description information of at least two geographic entities points to the same geographic entity; or, concatenating the geographic semantic similarity, spatial distance features, and geographic context similarity and inputting them into the classification head network for classification prediction to obtain a classification result indicating whether the description information of at least two geographic entities points to the same geographic entity.

[0142] In an optional embodiment, when the processor 55 acquires at least two geographic entity description information, it specifically performs the following: combines the geographic entity description information in the ontology library with that in the third-party library in pairs, and uses the geographic entity description information in each combination as the at least two geographic entity description information; or, for each geographic entity description information in the third-party library, based on a rule matching strategy and / or a text similarity matching strategy, recalls at least one similar geographic entity description information in the ontology library to form the at least two geographic entity description information.

[0143] Furthermore, such as Figure 5 As shown, the electronic device also includes other components such as a communication component 56, a display 57, a power supply component 58, and an audio component 59. Figure 5 The diagram only shows some components and does not mean that the electronic device includes only these components. Figure 5 The components shown. Additionally... Figure 5 The components within the dashed box are optional, not mandatory, and their specific requirements depend on the product form of the electronic device. The electronic device in this embodiment can be a terminal device such as a desktop computer, laptop computer, smartphone, or IoT device, or a server-side device such as a conventional server, cloud server, or server array. If the electronic device in this embodiment is a terminal device such as a desktop computer, laptop computer, or smartphone, it may include... Figure 5 The components within the dashed box; if the electronic device in this embodiment is implemented as a conventional server, cloud server, or server array, etc., it may be omitted. Figure 5 The component within the dashed box.

[0144] The aforementioned memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0145] The aforementioned communication component is configured to facilitate wired or wireless communication between the device containing the communication component and other devices. The device containing the communication component can access wireless networks based on communication standards, such as 2G, 3G, 4G / LTE, 5G, or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel.

[0146] The aforementioned display includes a screen, which may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a Touch Panel, the screen can be implemented as a touchscreen to receive input signals from the user. The Touch Panel includes one or more touch sensors to sense touches, swipes, and gestures on the Touch Panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation.

[0147] The aforementioned power supply components provide power to various components within the device in which they reside. These power supply components may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device in which they reside.

[0148] The aforementioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.

[0149] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above-described method embodiments. The computer-readable storage medium includes volatile or non-volatile components, or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, CD-ROM, Digital Video Disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium. Accordingly, this application also provides a computer program product, which includes a computer program or instructions that, when executed by a processor, cause the processor to implement the steps in the above method embodiments. It should be understood that each step or combination of steps in the above method flow can be implemented by the computer program or instructions. Furthermore, these computer programs or instructions can be applied to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device, enabling the processor of the general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to function as an apparatus for implementing the corresponding functions in the above method embodiments.

[0150] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0151] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for processing geographic entity information, characterized in that, include: Obtain description information for at least two geographic entities, wherein the description information includes the name, address, and spatial location information of the geographic entities; Using a geographic semantic feature encoder, the geographic semantic similarity between the at least two geographic entity description information is calculated based on the name information and address information in the description information of the at least two geographic entities. Using a distance feature encoder, spatial distance features between the descriptions of the at least two geographic entities are calculated based on spatial location information in the descriptions of the at least two geographic entities. Using a geographic context feature encoder, the geographic context similarity between the description information of the at least two geographic entities is calculated based on their respective surrounding geographic environment information. The geographic semantic similarity, spatial distance features, and geographic context similarity are input into the classification head network, which outputs the classification result of whether the description information of the at least two geographic entities points to the same geographic entity.

2. The method according to claim 1, characterized in that, Using a geographic semantic feature encoder, based on the name and address information in the description information of the at least two geographic entities, the geographic semantic similarity between the description information of the at least two geographic entities is calculated, including: The name and address information in the description information of the at least two geographic entities are subjected to regular expression enhancement, word segmentation and embedding processing to construct the input sequence of the geographic semantic feature encoder; The input sequence is input into the encoding network of the geographic semantic feature encoder to perform multi-layer bidirectional contextual semantic encoding based on the attention mechanism, so as to obtain the geographic semantic features of the description information of the at least two geographic entities. The geographic semantic features of the description information of the at least two geographic entities are input into the first similarity network in the geographic semantic feature encoder to calculate the similarity, so as to obtain the geographic semantic similarity between the description information of the at least two geographic entities.

3. The method according to claim 2, characterized in that, The name and address information in the description information of the at least two geographic entities are subjected to regular expression enhancement, word segmentation, and embedding processing to construct the input sequence of the geographic semantic feature encoder, including: For the name information in the description information of any geographic entity, target words with distinctiveness in the name information are identified based on regular expressions, and enhanced tags are added to the target words to obtain enhanced name information; According to the preset structure, the enhanced name information and address information in the description information of the at least two geographic entities are concatenated into sentence pairs; Based on the enhanced markers, with the goal of segmenting the target word into a single word, the sentence pair is segmented to obtain a segmentation sequence; in the segmentation sequence, the target word is segmented into a single word. Word embedding, sentence embedding, and location embedding are performed on the word segments in the word segmentation sequence to obtain the input sequence of the geographic semantic feature encoder.

4. The method according to claim 2, characterized in that, The input sequence includes segmentation tags for segmenting the description information of the at least two geographic entities. The input sequence is then input into the encoding network of the geographic semantic feature encoder for multi-layer bidirectional contextual semantic encoding based on an attention mechanism, to obtain the geographic semantic features of the description information of the at least two geographic entities, including: The input sequence is fed into the encoding network of the geographic semantic feature encoder, and the following operations are performed in the encoding network: Based on the segmentation markers, multi-level bidirectional self-attention calculations are performed on the name and address information of the same geographic entity description information to obtain the fused semantic features of the same geographic entity description information itself. Based on the segmentation markers, multi-layer bidirectional cross-attention calculations are performed on the name information of different geographic entity description information to obtain cross-semantic features in the name dimension. Based on the segmentation markers, multi-layer bidirectional cross-attention calculations are performed on the address information of different geographic entity description information to obtain cross-semantic features in the address dimension. Global attention is calculated on the fused semantic features of each geographic entity description information in the input sequence, as well as the cross semantic features of different geographic entity description information in the name dimension and the address dimension, to obtain the geographic semantic features of each geographic entity description information.

5. The method according to claim 2, characterized in that, The method further includes: based on the information dimensions and encoding objectives included in the input sequence, calling the target language model to generate a target thought chain, wherein the encoding objective is a geographic semantic feature for generating the geographic entity description information of the at least two geographic entities, and the target thought chain includes a structural parsing step, a spatial reasoning step, and a fusion encoding step required to guide the geographic semantic feature encoder to achieve the encoding objective; The input sequence is fed into the encoding network of the geographic semantic feature encoder, and multi-layer bidirectional contextual semantic encoding based on an attention mechanism is performed to obtain the geographic semantic features of the at least two geographic entity description information, including: The target thought chain and the input sequence are input into the encoding network of the geographic semantic feature encoder, and the following operations are performed in the encoding network: Guided by the structure parsing step, the name information in the input sequence is parsed according to the preset information hierarchy, and the corresponding hierarchy information is added to the name information to obtain the first intermediate sequence; Guided by the spatial reasoning step, the address information in the first intermediate sequence is validated and corrected using the address knowledge base to obtain the second intermediate sequence. Guided by the fusion encoding step, the name information and the corrected address information with hierarchical information in the second intermediate sequence are subjected to multi-layer bidirectional contextual semantic encoding based on an attention mechanism to obtain the geographic semantic features of the at least two geographic entity description information.

6. The method according to any one of claims 1-5, characterized in that, Using a distance feature encoder, based on the spatial location information in the description information of the at least two geographic entities, the spatial distance feature between the description information of the at least two geographic entities is calculated, including: The spatial location information from the description information of the at least two geographic entities is input into the distance feature encoder, and the following operations are performed in the distance feature encoder: Based on the spatial location information, calculate the distance information between the description information of the at least two geographic entities; According to the set distance change rule, the distance information is mapped to a target distance interval, and the target distance interval has an interval number, which represents the position of the target distance interval in a plurality of sequentially arranged distance intervals; An embedding layer is used to map the interval numbers of the target distance interval from the numerical space to the vector space to obtain the spatial distance features between the description information of the at least two geographic entities.

7. The method according to any one of claims 1-5, characterized in that, The geographic context feature encoder includes a graph attention network and a second similarity network; Using a geographic context feature encoder, based on the description information of the at least two geographic entities and their respective surrounding geographic environment information, the geographic context similarity between the description information of the at least two geographic entities is calculated, including: Based on the geographic environment information surrounding each of the at least two geographic entity description information, construct a geographic context relationship graph corresponding to each of the at least two geographic entity description information. The description information of the at least two geographic entities and their corresponding geographic context relationship graphs are input into the graph attention network. Based on the geographic context relationship graphs corresponding to the description information of the at least two geographic entities, the geographic environment information surrounding the description information of the at least two geographic entities is aggregated to the target geographic entity described by each entity to obtain the geographic context features of the description information of the at least two geographic entities. The geographic context features of the description information of the at least two geographic entities are input into the second similarity network for similarity calculation to obtain the geographic context similarity between the description information of the at least two geographic entities.

8. The method according to claim 7, characterized in that, Based on the geographic environment information surrounding each of the at least two geographic entity descriptions, a geographic context relationship graph corresponding to each of the at least two geographic entity descriptions is constructed, including: For any geographic entity description information, mark the target geographic entity described by the geographic entity description information on the map according to the address information and / or spatial location information in the geographic entity description information; Centered on the target geographic entity, a target area is determined on the map, and points of interest that are compatible with the target geographic entity are selected from the target area. The target geographic entity is taken as the central node, the point of interest is taken as the neighbor node, and an edge is added between the central node and the neighbor node to obtain the geographic context relationship graph corresponding to the geographic entity description information. In the geographic context graph, the nodes represent the encoded features of the names of the target geographic entities or points of interest, and the edges represent the spatial relationship features between the central node and its neighboring nodes.

9. The method according to claim 8, characterized in that, Based on the geographic context relationship graphs corresponding to the description information of each of the at least two geographic entities, the geographic environment information surrounding each of the at least two geographic entity description information is aggregated to the target geographic entity described by each entity to obtain the geographic context features of the description information of the at least two geographic entities, including: For any geographic entity description information corresponding to the geographic context relationship graph, based on the shared weight matrix, the encoded features represented by each node in the geographic context relationship graph are linearly mapped to obtain the mapped features represented by each node. Based on the mapping features represented by each node and the spatial relationship features represented by the edges in the geographic context graph, the attention coefficient between the central node and its neighboring nodes is calculated. Based on the attention coefficient, and combined with the mapping features represented by the neighboring nodes and the central node, the geographic context features of the geographic entity description information are generated.

10. The method according to claim 9, characterized in that, Based on the attention coefficient, and combined with the mapping features represented by the neighboring nodes and the central node, the geographic context features of the geographic entity description information are generated, including: Based on the attention coefficient, the mapping features represented by the neighboring nodes are weighted and summed to obtain aggregated features; based on the aggregated features and the mapping features represented by the center node, the geographic context features of the geographic entity description information are generated. or Based on the attributes of the neighboring nodes, the neighboring nodes are divided into multiple categories; based on the attention coefficient, self-attention is calculated between the central node and neighboring nodes of the same category to obtain a first contextual feature between the central node and neighboring nodes of the same category; based on the attention coefficient, cross-attention is calculated between the central node and neighboring nodes of different categories to obtain a second contextual feature between the central node and neighboring nodes of different categories; based on the first contextual feature and the second contextual feature, the geographic contextual feature of the geographic entity description information is generated.

11. The method according to claim 9, characterized in that, Before linearly mapping the encoded features represented by each node in the geographic context graph based on the shared weight matrix, the method further includes: From the neighbor nodes in the geographic context graph, a reference neighbor node is selected. The reference neighbor node is an interest point with a heat value greater than a set heat value threshold, an interest point of a specific type, or an interest point that is the closest to the center node. By combining the geographic knowledge graph, the features represented by the benchmark neighbor nodes are enhanced, and / or, based on the number of neighbor nodes in the geographic context graph, the number of the benchmark neighbor nodes is enhanced.

12. The method according to any one of claims 1-5 or 8-11, characterized in that, The geographic semantic similarity, spatial distance features, and geographic context similarity are input into the classification head network, which outputs a classification result indicating whether the description information of at least two geographic entities points to the same geographic entity, including: The geographic semantic similarity, spatial distance features, and geographic context similarity are respectively input into a three-way sub-classification head network for classification prediction to obtain three initial classification results; the three initial classification results are input into a total classification head network for fusion prediction to obtain a classification result of whether the description information of at least two geographic entities points to the same geographic entity; or The geographic semantic similarity, spatial distance features, and geographic context similarity are concatenated and then input into the classification head network for classification prediction to obtain the classification result of whether the description information of the at least two geographic entities points to the same geographic entity.

13. The method according to any one of claims 1-5 or 8-11, characterized in that, Obtain description information for at least two geographic entities, including: The geographic entity description information in the ontology library and the third-party library are combined in pairs, and the geographic entity description information in each combination is used as the description information of the at least two geographic entities. or For each geographic entity description information in the third-party library, based on rule matching strategy and / or text similarity matching strategy, at least one similar geographic entity description information is recalled in the ontology library to form the at least two geographic entity description information.

14. An electronic device, characterized in that, include: Memory and processor; The memory is used to store a computer program; the processor is coupled to the memory and is used to execute the computer program to implement the steps of the method according to any one of claims 1-13.

15. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it causes the processor to perform the steps of the method according to any one of claims 1-13.

16. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the processor is caused to perform the steps of the method according to any one of claims 1-13.