Method, apparatus, device, and storage medium for removing duplicate points of interest based on graph neural network

By constructing a geographic location-based interest point graph and using a graph convolution neural network model, combining the text and spatial information of the interest points, the problem of mismatch and insufficient feature utilization of identifying duplicate data in the existing interest point deduplication method is solved, and a more efficient interest point deduplication effect is achieved.

CN114491201BActive Publication Date: 2025-07-11SHENZHEN YISHIHUOLALA TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210079770.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-24
Publication Date
2025-07-11
Estimated Expiration
2042-01-24

AI Technical Summary

Technical Problem

The existing deduplication methods for point-of-interest deduplication have problems such as mismatch, difficulty in setting thresholds, cumbersome feature engineering, limited model expression ability and poor non-text feature compatibility when identifying duplicate data, and are especially unable to effectively utilize the spatial positional relationship between points of interest.

Method used

The graph neural network-based method is used to construct a geographic location-based interest point graph. Through the graph convolution neural network model, the text information and spatial position relationship of the points of interest are combined to deduplicate the point of interest, and the clustered neighbor information of the graph neural network is used to identify similar neighbor structures, and the repetition of the point of interest is judged by the embedding distance.

Benefits of technology

It improves the accuracy and recall rate of point-of-interest deduplication, achieves more complete model training, and improves the efficiency and applicability of deduplication effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114491201B_ABST
    Figure CN114491201B_ABST
Patent Text Reader

Abstract

This application relates to a method, apparatus, computer device, and storage medium for deduplicating points of interest based on a graph neural network. The method includes: obtaining all points of interest within a target geographical area to be deduplicated, and constructing a point-of-interest graph based on geographical locations according to all the points of interest; screening out multiple pairs of duplicate points of interest from all the points of interest, and annotating the multiple pairs of duplicate points of interest to obtain multiple pairs of seed duplicate points of interest; iteratively training a graph neural network model according to the point-of-interest graph and the multiple pairs of seed duplicate points of interest to obtain a trained graph neural network model; processing the point-of-interest graph through the trained graph neural network model, and determining all duplicate points of interest among all the points of interest according to the processing result; deleting any one of the points of interest in each pair of duplicate points of interest. The embodiments of this application can improve the effect of deduplicating points of interest, making it efficient, reasonable, and widely applicable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of electronic maps, and particularly to a method, device, computer device, and storage medium for deduplicating points of interest based on a graph neural network. Background Art

[0002] A point of interest (POI), generally containing information such as name, address, longitude and latitude, category, etc., is the most important content of a network electronic map and the foundation of Internet location services. Since there are a large number of duplicate and redundant data in the points of interest on the Internet, that is, two points of interest may have different text descriptions, but may correspond to the same point of interest in the real world. When users use map services, it will seriously affect their experience. Therefore, how to remove duplicate data on the basis of ensuring data richness and display more concise and pure map point-of-interest data to users has become a current research hotspot. Currently, the commonly used method is to identify duplicate points of interest through algorithms, and then delete the duplicate points of interest from the database, which can not only improve the user experience but also save storage space and data maintenance costs.

[0003] Currently, the main deduplication solutions for point-of-interest data are as follows:

[0004] 1. Unsupervised similarity calculation-based solution

[0005] From the target point-of-interest data, calculate the text similarity of the names and addresses of two points of interest in turn. The similarity algorithms include edit distance, TF-IDF (term frequency–inverse document frequency), etc. Calculate the overall similarity by setting a weight for the calculated text similarity of the names and addresses as the similarity score between two points of interest. When the score is higher than a certain threshold, the two points of interest can be considered duplicates.

[0006] 2. Traditional machine learning-based duplicate judgment model solution

[0007] Extract pairs of points of interest with duplicate relationships from the target point-of-interest data as training data, construct features by calculating the text similarity of the names and category similarities of the pairs of points of interest, and use traditional machine learning (GBDT, Xgboost, etc.) methods to train a duplicate judgment model to determine whether two points of interest are in a duplicate relationship.

[0008] 3. Pre-trained deep learning-based duplicate judgment model solution

[0009] Extract pairs of points of interest with duplicate relationships from the target point-of-interest data and use them as training data. Fine-tune on commonly used pre-trained deep models (such as bert, albert, roberta, etc.) to train a duplicate judgment model to determine whether two points of interest have a duplicate relationship.

[0010] The inventors found that the above-mentioned solutions all have some drawbacks in practical applications.

[0011] For example, the drawbacks of the above-mentioned Solution 1:

[0012] (1) Based on the unsupervised similarity score method, for scenarios where two points of interest are truly duplicate data but have large text differences, the matching effect is poor.

[0013] (2) For two points of interest that are very close in text but are not actually duplicate data, it will cause mis-matching.

[0014] (3) The threshold of the similarity score is not easy to set.

[0015] The drawbacks of the above-mentioned Solution 2:

[0016] (1) A large amount of feature engineering work is required to construct features, and the process is relatively cumbersome.

[0017] (2) The model is relatively shallow, with limited expressive ability, and the duplicate judgment effect is average.

[0018] (3) This method assumes that points of interest are independent of each other. However, there is a certain spatial position relationship between actual points of interest. Therefore, the relationship information between points of interest is not used for duplicate judgment, and less information is utilized, resulting in poor effects.

[0019] The drawbacks of the above-mentioned Solution 3:

[0020] (1) Pre-trained deep models generally input pure text information and have poor compatibility with non-text features.

[0021] (2) This method assumes that points of interest are independent of each other. However, there is a certain spatial position relationship between actual points of interest. Therefore, the relationship information between points of interest is not used for duplicate judgment, and less information is utilized, resulting in poor effects. Summary of the Invention

[0022] The present application provides a method, device, computer device, and storage medium for deduplicating points of interest based on a graph neural network in view of the above deficiencies or drawbacks. The embodiments of the present application can improve the effect of deduplicating points of interest, making it efficient, reasonable, and widely applicable.

[0023] According to a first aspect of the present application, a method for deduplicating points of interest based on a graph neural network is provided. In one embodiment, the method includes:

[0024] Obtain all the points of interest within the target geographical area to be de-duplicated, and construct a point-of-interest graph based on geographical location according to all the points of interest;

[0025] Screen out multiple pairs of duplicate points of interest from all the points of interest, label the multiple pairs of duplicate points of interest, and obtain multiple pairs of seed duplicate points of interest;

[0026] Iteratively train the graph neural network model according to the point-of-interest graph and the multiple pairs of seed duplicate points of interest to obtain a trained graph neural network model;

[0027] Process the point-of-interest graph through the trained graph neural network model, and determine all the duplicate points of interest among all the points of interest according to the processing result;

[0028] Delete any one of the points of interest in each pair of duplicate points of interest.

[0029] In one embodiment, any training process of the graph neural network model includes:

[0030] Input the adjacency matrix and node attribute feature matrix of the point-of-interest graph into the graph neural network model to obtain the output data of the graph neural network model. The output data includes the embeddings of each point of interest among all the points of interest;

[0031] Determine the target seed duplicate points of interest for this training from the multiple pairs of seed duplicate points of interest, use the target seed duplicate points of interest as the positive duplicate pairs for this training, and construct the negative duplicate pairs for this training according to the target seed duplicate points of interest;

[0032] Obtain the embeddings of the positive duplicate pairs and the embeddings of the negative duplicate pairs from the output data, and calculate the loss of this training according to the embeddings of the positive duplicate pairs and the embeddings of the negative duplicate pairs;

[0033] Judge whether the stop training condition is satisfied according to the loss;

[0034] If it is satisfied, stop training and use the graph neural network model as the trained graph neural network model;

[0035] If it is not satisfied, update the network parameters of the graph neural network model according to the loss, and perform the next training on the updated graph neural network model.

[0036] In one embodiment, processing the point-of-interest graph through the trained graph neural network model and determining all the duplicate points of interest among all the points of interest according to the processing result includes:

[0037] Process the point-of-interest graph through the trained graph neural network model to obtain the embeddings of each point of interest among all the points of interest output by the trained graph neural network model;

[0038] Pair all the points of interest in pairs to obtain a plurality of point - of - interest pairs, and calculate the embedding distance for each point - of - interest pair; the embedding distance for each point - of - interest pair refers to the distance between the embeddings of the two points of interest included in each point - of - interest pair.

[0039] Determine the point - of - interest duplicate pairs for those point - of - interest pairs whose embedding distance is less than a preset threshold.

[0040] In one embodiment, constructing a location - based point - of - interest graph according to all the points of interest includes:

[0041] Obtain the location information of each point of interest among all the points of interest;

[0042] Calculate the distance between every two points of interest according to the location information of every two points of interest;

[0043] Determine the weight of the edge between every two points of interest according to the distance between every two points of interest, and obtain the location - based point - of - interest graph.

[0044] In one embodiment, determining the weight of the edge between every two points of interest according to the distance between every two points of interest includes:

[0045] When the distance between any two points of interest is less than the preset threshold, determine that an edge relationship is formed between the two points of interest, and set the weight of the edge between the two points of interest to 1;

[0046] When the distance between any two points of interest is greater than or equal to the preset threshold, determine that no edge relationship is formed between the two points of interest, and set the weight of the edge between the two points of interest to 0;

[0047] In one embodiment, the number of point - of - interest graphs is multiple; constructing a location - based point - of - interest graph according to all the points of interest includes:

[0048] Divide the target geographical area into multiple spatial grids;

[0049] Traverse the longitude and latitude attributes of each point of interest among all the points of interest to determine the set of points of interest for each spatial grid;

[0050] Construct a location - based point - of - interest graph for each spatial grid according to the set of points of interest for each spatial grid.

[0051] In one embodiment, processing the point - of - interest graph through a trained graph neural network model, and determining all the point - of - interest duplicate pairs among all the points of interest according to the processing result includes:

[0052] Input the point-of-interest (POI) map of each spatial grid into the trained graph neural network model to obtain the embedding of each POI in the POI set of each spatial grid;

[0053] Pair up the POI sets of each spatial grid in pairs to obtain the POI pair set of each spatial grid;

[0054] Calculate the embedding distance of each POI pair in the POI pair set of each spatial grid; the embedding distance of each POI pair refers to the distance between the embeddings of the two POIs included in each POI pair;

[0055] Determine each POI pair with an embedding distance less than a preset threshold in the POI pair set of each spatial grid as a duplicate POI pair.

[0056] According to a second aspect, the present application provides a duplicate POI removal device based on a graph neural network. In one embodiment, the device includes:

[0057] A POI map construction module, configured to obtain all POIs within a target area to be de-duplicated, and construct a location-based POI map according to all POIs;

[0058] A seed duplicate pair construction module, configured to screen out multiple pairs of duplicate POI pairs from all POIs, and label the multiple pairs of duplicate POI pairs to obtain multiple pairs of seed duplicate POI pairs;

[0059] A model training module, configured to iteratively train a graph neural network model according to the POI map and the multiple pairs of seed duplicate POI pairs to obtain a trained graph neural network model;

[0060] A duplicate pair determination module, configured to process the POI map through the trained graph neural network model, and determine all duplicate POI pairs among all POIs according to the processing result;

[0061] A de-duplication module, configured to delete any one of the POIs in each duplicate POI pair.

[0062] According to a third aspect, the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the embodiments of any of the above methods are implemented.

[0063] According to a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the embodiments of any of the above methods are implemented.

[0064] The embodiments of the present application can bring the following beneficial effects compared with the prior art:

[0065] Traditional machine learning and pre-trained deep models make the assumption that two points of interest are unrelated in dealing with the problem of deduplicating points of interest. In contrast, the embodiments of the present application assume that there is a correlation between two points of interest. Based on this, the attributes of the points of interest themselves and the spatial position relationship between the points of interest are combined to perform deduplication of points of interest.

[0066] Specifically, the inventors found that there are usually similar neighbors around duplicate pairs of points of interest. Based on this discovery, a graph convolutional neural network model based on a graph structure is selected to perform deduplication of points of interest. The text information of the names of the points of interest and the information on the spatial position relationship between the points of interest are used as the model input. The graph neural network can identify isomorphic graphs by transmitting and aggregating neighbor node information, and can better identify similar neighbor structures. If two points of interest have similar neighbors, through the convolutional aggregation operation, the embedding representations (which can be simply referred to as embeddings) of these two points of interest will also be very close, and the embedding distance between the two points of interest will be very close. By calculating the embedding distance between two points of interest, it is determined whether the two points of interest are duplicates. Through the method of graph neural network, the problem of duplicate points of interest can be solved more effectively.

[0067] On the other hand, traditional machine learning and pre-trained deep models generally can only perform model training based on seed pairs of points of interest, and the seed pairs of points of interest need to be manually labeled, with a relatively high acquisition cost. Therefore, the quantity is generally small, which will result in insufficient training. The embodiments of the present application can enable all data to participate in training through the graph convolutional neural network model, which can train the model more fully and make the duplicate judgment effect of the model better.

[0068] In summary, the embodiments of the present application can fundamentally improve the effect of deduplicating points of interest, making it efficient, reasonable, and widely applicable. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 is a schematic flowchart of a method for deduplicating points of interest based on a graph neural network in an embodiment;

[0070] Figure 2 is a flowchart of model training in an embodiment;

[0071] Figure 3 is a flowchart of model prediction in an embodiment;

[0072] Figure 4 is a structural block diagram of a device for deduplicating points of interest based on a graph neural network in an embodiment;

[0073] Figure 5 is an internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0074] To make the objectives, technical solutions and advantages of this application clearer and more understandable, the following further elaborates on this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0075] This application provides a method for deduplicating points of interest based on a graph neural network. In one embodiment, the method includes the steps as Figure 1 shown below, and the method will be described below.

[0076] S110: Obtain all the points of interest within the target area to be deduplicated, and construct a point-of-interest graph based on geographical locations according to all the points of interest.

[0077] S120: Screen out multiple pairs of duplicate point-of-interest pairs from all the points of interest, and label these multiple pairs of duplicate point-of-interest pairs to obtain multiple pairs of seed duplicate point-of-interest pairs.

[0078] S130: Iteratively train a graph neural network model according to the point-of-interest graph and these multiple pairs of seed duplicate point-of-interest pairs to obtain a trained graph neural network model.

[0079] S140: Process the point-of-interest graph through the trained graph neural network model, and determine all the duplicate point-of-interest pairs among all the points of interest according to the processing results.

[0080] S150: Delete any one of the points of interest in each duplicate point-of-interest pair.

[0081] In this embodiment, by obtaining all the points of interest within the target area to be deduplicated, constructing a point-of-interest graph based on geographical locations according to all the points of interest, screening out multiple pairs of duplicate point-of-interest pairs from all the points of interest, and labeling these multiple pairs of duplicate point-of-interest pairs to obtain multiple pairs of seed duplicate point-of-interest pairs; iteratively training a graph neural network model according to the point-of-interest graph and these multiple pairs of seed duplicate point-of-interest pairs to obtain a trained graph neural network model; processing the point-of-interest graph through the trained graph neural network model, and determining all the duplicate point-of-interest pairs among all the points of interest according to the processing results; deleting any one of the points of interest in each duplicate point-of-interest pair.

[0082] The embodiments of this application can bring the following beneficial effects compared with the prior art:

[0083] Traditional machine learning and pre-trained deep models make the assumption that two points of interest are independent in the problem of deduplicating points of interest. The embodiments of this application assume that there is a relationship between two points of interest. On this basis, the attributes of the points of interest themselves and the spatial position relationship between the points of interest are combined to perform deduplication of points of interest.

[0084] Specifically, the inventors found that there are usually similar neighbors around duplicate pairs of points of interest. Based on this discovery, a graph convolutional neural network model based on a graph structure is selected to perform deduplication of points of interest. The text information of the names of the points of interest and the spatial position relationship information between the points of interest are used as model inputs. The graph neural network has the ability to identify isomorphic graphs by transmitting and aggregating neighbor node information, and can better identify similar neighbor structures. If two points of interest have similar neighbors, through the convolutional aggregation operation, the embedding representations (embeddings, which can be simply referred to as embeddings) of these two points of interest will also be very close, and the embedding distance between the two points of interest will be very close. By calculating the embedding distance between two points of interest, it is determined whether the two points of interest are duplicates. Through the graph neural network method, the problem of duplicate points of interest can be solved more effectively.

[0085] On the other hand, traditional machine learning and pre-trained deep models generally can only perform model training based on seed pairs of points of interest, and the seed pairs of points of interest need to be manually labeled, and the acquisition cost is relatively high, so the quantity is generally small, which will lead to insufficient training. The embodiments of the present application can enable all data to participate in training through the graph convolutional neural network model, so that the model can be trained more fully and the deduplication effect of the model is better.

[0086] In summary, the embodiments of the present application can fundamentally improve the deduplication effect of points of interest, making it efficient, reasonable, and widely applicable.

[0087] In one embodiment, constructing a location-based point-of-interest graph according to all points of interest includes: obtaining the location information of each point of interest in all points of interest; calculating the distance between each two points of interest according to the location information of each two points of interest; determining the weight of the edge between each two points of interest according to the distance between each two points of interest, and obtaining a point-of-interest graph based on location information. The distance between two points of interest refers to the spatial distance, such as 100 meters, 50 meters, etc.

[0088] Among them, determining the weight of the edge between each two points of interest according to the distance between each two points of interest includes: when the distance between any two points of interest is less than a preset threshold, determining that an edge relationship is formed between the any two points of interest, and setting the weight of the edge between the any two points of interest to 1; when the distance between any two points of interest is greater than or equal to the preset threshold, determining that no edge relationship is formed between the any two points of interest, and setting the weight of the edge between the any two points of interest to 0. Among them, the preset threshold can be set according to actual needs, such as set to 50 meters, etc., and this embodiment does not make specific limitations on this.

[0089] In one embodiment, any training process of the graph neural network model includes:

[0090] Input the adjacency matrix of the point-of-interest (POI) graph and the node attribute feature matrix into the graph neural network model to obtain the output data of the graph neural network model. The output data includes the embeddings of each POI among all POIs. Determine the target seed POI repeat pairs for the current training from the multiple pairs of seed POIs. Use the target seed POI repeat pairs as the positive repeat pairs for the current training, and construct the negative repeat pairs for the current training based on the target seed POI repeat pairs. Obtain the embeddings of the positive repeat pairs and the embeddings of the negative repeat pairs from the output data, and calculate the loss for the current training based on the embeddings of the positive repeat pairs and the embeddings of the negative repeat pairs. Determine whether the stop training condition is satisfied based on the loss. If satisfied, stop training and use the graph neural network model as the trained graph neural network model. If not satisfied, update the network parameters of the graph neural network model according to the loss, and perform the next training on the updated graph neural network model.

[0091] In one embodiment, process the POI graph through the trained graph neural network model, and determine all POI repeat pairs among all POIs according to the processing result, including: process the POI graph through the trained graph neural network model to obtain the embeddings of each POI among all POIs output by the trained graph neural network model; pair all POIs in pairs to obtain multiple POI pairs, and calculate the embedding distance of each POI pair; the embedding distance of each POI pair refers to the distance between the embeddings of the two POIs included in each POI pair; determine the POI pairs with each embedding distance less than the preset threshold as the POI repeat pairs.

[0092] In another embodiment, the number of the above-mentioned POI graphs is multiple. Correspondingly, the construction of the location-based POI graph according to all POIs includes: divide the target geographical area into multiple spatial grids; traverse the longitude and latitude attributes of each POI among all POIs to determine the POI set of each spatial grid; construct a location-based POI graph for each spatial grid according to the POI set of each spatial grid.

[0093] This embodiment takes into account that in some scenarios, when the range of the target geographical area is large, for example, the target geographical area refers to the whole of China, then the data that needs to be deduplicated (referring to POI deduplication) is the full amount of data across the whole of China. At this time, the data volume is very large. If all the data is directly used to construct a graph (referring to the POI graph), the nodes and adjacency matrix of this POI graph will be very large, which will require very high computing resources. Therefore, in the case of limited computing resources, a graph is divided into many small graphs according to spatial grids, and then deduplication is performed on each small graph. In this way, the nodes and adjacency matrix of each small graph become smaller, and data deduplication can be performed with less computing resources. The specific method can be as follows:

[0094] According to the geographical space coordinates of China, from the westernmost end to the easternmost end, and from the northernmost end to the southernmost end, it is divided into 1-kilometer * 1-kilometer square spatial grids (the grid size can be flexibly adjusted according to actual needs). Each of the four vertices of the grid has corresponding longitude and latitude coordinates, and each point of interest has longitude and latitude attributes. By traversing all points of interest according to their longitude and latitude, the points of interest can fall into the corresponding grids. In this way, each grid will contain points of interest that are relatively close. Subsequently, only the duplicate removal of points of interest needs to be carried out in each spatial grid.

[0095] Correspondingly, processing the point-of-interest graph through the trained graph neural network model, and determining all point-of-interest duplicate pairs among all points of interest according to the processing results, including:

[0096] Inputting the point-of-interest graph of each spatial grid into the trained graph neural network model respectively to obtain the embedding of each point of interest in the point-of-interest set of each spatial grid; pairing the point-of-interest sets of each spatial grid pairwise to obtain the point-of-interest pair set of each spatial grid; calculating the embedding distance of each point-of-interest pair in the point-of-interest pair set of each spatial grid; the embedding distance of each point-of-interest pair refers to the distance between the embeddings of the two points of interest included in each point-of-interest pair; determining each point-of-interest pair with an embedding distance less than a preset threshold in the point-of-interest pair set of each spatial grid as a point-of-interest duplicate pair.

[0097] The above embodiments are illustrated below through a specific application example.

[0098] This application example is specifically divided into seven parts: division of spatial grids, acquisition of seed point-of-interest duplicate pairs, data preprocessing, feature engineering, model training, prediction of duplicate data by the model, and duplicate removal processing.

[0099] This application example takes China as the target geographical area. First, according to the geographical space coordinates of China, the geographical space of China is divided into square grids of 1 kilometer * 1 kilometer. Then, the points of interest at the corresponding spatial positions are dropped into the corresponding grids to form sub-graphs. This application example uses supervised learning to train the graph neural network model. Since it is supervised learning, it is necessary to manually label the repeated pairs of seed points of interest for model training. After obtaining the repeated pairs of seed points of interest, it is necessary to preprocess the data, including operations such as converting full-width characters to half-width characters, removing special symbols, converting English uppercase to lowercase, and converting traditional Chinese to simplified Chinese, to clean the data. Before training the model, it is necessary to first perform feature engineering to obtain the attribute features of the nodes and the adjacency matrix. The attribute features of the nodes and the adjacency matrix are input into the graph convolutional neural network GCN (Graph Convolutional Network), and the model is trained by minimizing the loss function through backpropagation to obtain the weight matrix W, which is the parameter that the model needs to learn. After obtaining the parameter W, the prediction of the repeated pairs of points of interest at the graph level can be performed through forward propagation. Finally, in the full amount of data, one of the points of interest data in each repeated pair of points of interest is deleted to obtain the deduplicated points of interest data.

[0100] The following will explain each of the above parts.

[0101] 1. Division of spatial grids

[0102] Since the data to be deduplicated is the full amount of data across China and the total amount of data is very large, if all the data is directly used to construct a large graph, both the nodes and the adjacency matrix will be extremely large, requiring too much computing resources. Therefore, in the case of limited computing resources, the large graph is divided into many small graphs according to spatial grids, and data deduplication is performed on each small graph. In this way, both the nodes and the adjacency matrix become smaller, and data deduplication can be carried out with less computing resources.

[0103] Specific method: According to the geographical space coordinates of China, from the westernmost end to the easternmost end, and from the northernmost end to the southernmost end, it is divided into square spatial grids of 1 kilometer * 1 kilometer. The four vertices of the grid have corresponding longitude and latitude coordinates, and the points of interest all have longitude and latitude attributes. By traversing all the points of interest according to the longitude and latitude, the points of interest can fall into the corresponding grids. In this way, each grid will have points of interest that are relatively close, and subsequent deduplication of points of interest only needs to be carried out in each spatial grid.

[0104] 2. Obtaining repeated pairs of seed points of interest

[0105] This application example requires seed interest point duplicate pairs to train the model, so it is necessary to label some data as seed interest point duplicate pairs. All the above spatial grids can be traversed, and by using simple text similarity of interest point names (such as edit distance, etc.), suspected duplicate interest point pairs can be roughly found in each spatial grid, and then the labeling personnel are asked to find the truly duplicate interest point pairs. In this way, seed interest point duplicate pairs are constructed.

[0106] 3. Data Preprocessing

[0107] Special symbols and traditional Chinese characters may be included in the text of the interest point name. It is necessary to perform preprocessing first and then construct the feature input model. At the same time, in order to ensure the consistency of the distribution of labeled data and unlabeled data, the same preprocessing operations need to be performed on the labeled data and unlabeled data. The data preprocessing process includes the following four steps:

[0108] (1) Full-width to half-width conversion of characters

[0109] (2) Special symbols

[0110] (3) Capital to lowercase conversion of English

[0111] (4) Traditional Chinese to simplified Chinese conversion

[0112] 4. Feature Engineering

[0113] (1) Generate attribute features of graph nodes

[0114] The input of the graph convolutional neural network GCN includes the topological structure of the graph, that is, the adjacency matrix, and the attribute features of all nodes in the graph. Each node attribute feature is a multi-dimensional feature vector. In this application example, a specified algorithm is used to process each interest point into a 512-dimensional Embedding vector, and this Embedding vector is used as the attribute feature of the node. Among them, the specified algorithm can be any existing algorithm that can map the interest point to the Embedding vector, so it will not be elaborated here.

[0115] (2) Generate the edges and adjacency matrix of all subgraph structures

[0116] In this application example, interest points within 50 meters are regarded as having edge relationships. Specifically, all the interest points in a spatial grid are taken out to form a set. An interest point is taken out from the set, and the distance is calculated with all other interest points in the set except itself. Among them, the interest points with a distance less than 50 meters form an edge relationship with the taken-out interest point, and the weight of the edge is 1; then the interest points are taken out from the set in turn, and the above operations are performed in the same way until all the interest points in the set are taken out, and the edge relationships of all the interest points in this spatial grid subgraph are formed. According to the definition of the graph structure, the adjacency matrix of the subgraph is obtained. The representation of the adjacency matrix is:

[0117]

[0118] Finally, traverse all the spatial grids according to the above method, and the adjacency matrix of all the subgraphs of the spatial grids is generated.

[0119] 5. Model Training

[0120] As Figure 2 shown is the model training flowchart of this application example.

[0121] Specifically, select a subgraph of the spatial grid KG, and the pre-annotated duplicate interest point seed pairs S = {(e i1 , e i2 )} m i=1 .

[0122] The method of this application example finds new duplicate interest point pairs within this subgraph based on the node embedding of GCN. The basic idea is to use GCN (which can also be called the GCN model) to embed the interest points into a unified vector space, hoping that the distance between duplicate interest points is closer and the distance between non-duplicate interest points is farther.

[0123] (1) Input of GCN:

[0124] GCN is a type of neural network that operates directly on graphs. Its input is the node attribute features and the adjacency matrix of the graph, and the purpose is to output the embedding of each interest point, which is then used for subsequent interest point deduplication. For the node attribute features and the adjacency matrix input to the model, both are from the feature engineering in step 4. After inputting the node attribute features and the adjacency matrix into the GCN model, subsequent GCN operations are performed.

[0125] (2) Operations of GCN:

[0126] A GCN model contains multiple GCN layers. In this application example, two layers are selected. The input H (l) ∈R n×d(l) of the l-th layer is a node attribute feature matrix (i.e., all node attribute features), where n is the number of nodes, and d (l) is the number of features of the l-th layer. The output of the l-th layer is a new feature matrix:

[0127]

[0128] where σ is the relu activation function (used for linear activation transformation), A is the n*n adjacency matrix, where I is the identity matrix. is 's diagonal node degree matrix, and W (l) ∈R d(l)×d(l+1)is the weight matrix between two layers and is used for convolution operation, d (l+1) is the dimension of the new layer.

[0129] (3) Output of GCN

[0130] After two layers of GCN operations, the model outputs a 512-dimensional node embedding representation, which can be used for subsequent operations.

[0131] (4) Loss function of GCN

[0132] This application example hopes that the distance between repeated points of interest is small, and the distance between non-repeated points of interest is large. Based on this, a loss function is constructed. The distance between points of interest is the embedding distance between points of interest. For the outputs e1 and e2 of two points of interest, the calculation method of the distance between them is as follows:

[0133] D(e1, e2) = ||h(e1) - h(e2)||1

[0134] The model is trained by minimizing the margin-based loss function:

[0135]

[0136] where, [x] + = max{0, x}, S' (e1,e2) is a negative repeated pair obtained by randomly replacing one point of interest from (e1, e2), and γ is the margin that distinguishes positive and negative repeated pairs. The model is trained by minimizing the loss function through backpropagation to update the weight matrix W in each layer. After several rounds of training, the final model can learn the weight matrix W to predict repeated pairs of points of interest.

[0137] 6. Model repeated data prediction

[0138] As Figure 3 shown is the model prediction flowchart of this application example.

[0139] This application example is suitable for offline graph-level prediction of repeated pairs of points of interest. The prediction is to find more new repeated pairs of points of interest in the constructed graph. During the training process, the weight matrix W is learned. By inputting node attribute features and the adjacency matrix, after the operations of GCN, each node will output an embedding representation.

[0140] For the embedding of a specific output, calculate its embedding distance from all other points of interest in the subgraph, and select the one with the smallest embedding distance among all points of interest. If this embedding distance is less than a certain threshold, the two points of interest are considered duplicates; if it is not less than this threshold, they are considered not duplicates. According to the above method, traverse all points of interest in the subgraph that are not in the seed pairs, and the duplicate points of interest in the subgraph can be obtained. Then, traverse all subgraphs using the above method to obtain the duplicate point pairs in the full data set.

[0141] 7. Deduplication processing

[0142] Extract all the duplicate point pairs obtained in step 6 from the full data set, traverse these duplicate point pairs, delete one of the points of interest, and add the other point of interest to the full data set. Through this scheme, the deduplicated point of interest data for the full data set can be obtained.

[0143] This application example was statistically analyzed based on the other point of interest deduplication methods mentioned above, and the specific benefits are as follows:

[0144] The accuracy rate of the point of interest deduplication method based on the unsupervised similarity measurement scheme is 92.5%, and the recall rate is 85.4%;

[0145] The accuracy rate of the point of interest deduplication method based on the traditional machine learning model is 96.3%, and the recall rate is 90.4%;

[0146] The accuracy rate of the point of interest deduplication method based on the deep learning model is 97.1%, and the recall rate is 92.6%;

[0147] The accuracy rate of the point of interest deduplication method based on the graph neural network in this application example is 98.4%, and the recall rate is 95.3%.

[0148] It can be intuitively seen from the above data that compared with the previous methods, the point of interest deduplication method based on the graph neural network provided in this application example has a significant improvement in effect.

[0149] Figure 1 It is a schematic flow chart of the point of interest deduplication method based on the graph neural network in an embodiment. It should be understood that although Figure 1 the steps in the flow chart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figure 1At least a part of the steps may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0150] Based on the same inventive concept, the present application also provides a point of interest deduplication device based on a graph neural network. In this embodiment, as Figure 4 shown, the point of interest deduplication device based on the graph neural network includes the following modules:

[0151] A point of interest graph construction module 110, configured to obtain all points of interest within a target geographical area to be deduplicated, and construct a point of interest graph based on geographical locations according to all points of interest;

[0152] A seed duplicate pair construction module 120, configured to screen out multiple pairs of point of interest duplicate pairs from all points of interest, and label the multiple pairs of point of interest duplicate pairs to obtain multiple pairs of seed point of interest duplicate pairs;

[0153] A model training module 130, configured to iteratively train a graph neural network model according to the point of interest graph and the multiple pairs of seed point of interest duplicate pairs to obtain a trained graph neural network model;

[0154] A duplicate pair determination module 140, configured to process the point of interest graph through the trained graph neural network model, and determine all point of interest duplicate pairs among all points of interest according to the processing result;

[0155] A deduplication module 150, configured to delete any one of the points of interest in each point of interest duplicate pair.

[0156] In one embodiment, any training process of the model training module for training the graph neural network model includes:

[0157] Inputting the adjacency matrix and node attribute feature matrix of the point of interest graph into the graph neural network model to obtain output data of the graph neural network model, where the output data includes the embeddings of each point of interest among all points of interest;

[0158] Determining target seed point of interest duplicate pairs for this training from the multiple pairs of seed point of interest duplicate pairs, using the target seed point of interest duplicate pairs as positive duplicate pairs for this training, and constructing negative duplicate pairs for this training according to the target seed point of interest duplicate pairs;

[0159] Obtaining the embeddings of the positive duplicate pairs and the embeddings of the negative duplicate pairs from the output data, and calculating the loss of this training according to the embeddings of the positive duplicate pairs and the embeddings of the negative duplicate pairs;

[0160] Judge whether the stop training condition is satisfied according to the loss;

[0161] If satisfied, stop training and use the graph neural network model as the trained graph neural network model;

[0162] If not satisfied, update the network parameters of the graph neural network model according to the loss, and perform the next training on the updated graph neural network model.

[0163] In one embodiment, the repetition determination module is specifically used for:

[0164] Process the point-of-interest graph through the trained graph neural network model to obtain the embedding of each point of interest in all the points of interest output by the trained graph neural network model;

[0165] Pair all the points of interest in pairs to obtain multiple point-of-interest pairs, and calculate the embedding distance of each point-of-interest pair; the embedding distance of each point-of-interest pair refers to the distance between the embeddings of the two points of interest included in each point-of-interest pair;

[0166] Determine the point-of-interest repeat pairs for the point-of-interest pairs with each embedding distance less than the preset threshold.

[0167] In one embodiment, when constructing the point-of-interest graph, the point-of-interest graph construction module is specifically used for:

[0168] Obtain the geographical location information of each point of interest in all the points of interest;

[0169] Calculate the distance between every two points of interest respectively according to the geographical location information of every two points of interest;

[0170] Determine the weight of the edge between every two points of interest according to the distance between every two points of interest, and obtain the point-of-interest graph based on the geographical location information.

[0171] In one embodiment, when the point-of-interest graph construction module is used to determine the weight of the edge between every two points of interest according to the distance between every two points of interest, it is specifically used for:

[0172] When the distance between any two points of interest is less than the preset threshold, determine that an edge relationship is formed between the any two points of interest, and set the weight of the edge between the any two points of interest to 1;

[0173] When the distance between any two points of interest is greater than or equal to the preset threshold, determine that no edge relationship is formed between the any two points of interest, and set the weight of the edge between the any two points of interest to 0;

[0174] In another embodiment, the number of point-of-interest graphs is multiple; when the point-of-interest graph construction module is used to construct the point-of-interest graph based on the geographical location according to all the points of interest, it is specifically used for:

[0175] Divide the target geographical area into multiple spatial grids;

[0176] Traverse the longitude and latitude attributes of each point of interest in all points of interest to determine the set of points of interest for each spatial grid;

[0177] Construct a point-of-interest map based on geographical location for each spatial grid according to the set of points of interest for each spatial grid.

[0178] Correspondingly, in one embodiment, the repeating determination module is specifically configured to:

[0179] Input the point-of-interest map of each spatial grid into the trained graph neural network model respectively to obtain the embedding of each point of interest in the set of points of interest for each spatial grid;

[0180] Pair the sets of points of interest for each spatial grid in pairs to obtain a set of point pairs for each spatial grid;

[0181] Calculate the embedding distance of each point pair in the set of point pairs for each spatial grid; the embedding distance of each point pair refers to the distance between the embeddings of the two points of interest included in each point pair;

[0182] Determine the point pairs with repeated points of interest for each spatial grid, where each embedding distance in the set of point pairs for each spatial grid is less than a preset threshold.

[0183] For the specific limitations of the point-of-interest deduplication device based on the graph neural network, reference can be made to the limitations on the point-of-interest deduplication method based on the graph neural network in the above text, which will not be elaborated here. Each module in the above-mentioned point-of-interest deduplication device based on the graph neural network can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0184] In one embodiment, a computer device is provided, and its internal structure diagram can be as Figure 5 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a point-of-interest deduplication method based on the graph neural network.

[0185] Those skilled in the art can understand that Figure 5 the structure shown in [the figure] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0186] This embodiment also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps in the method provided in any of the above method embodiments.

[0187] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps in the method provided in any of the above method embodiments.

[0188] Those of ordinary skill in the art can understand that to implement all or part of the processes in the above method embodiments, it can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application may include non-volatile and / or volatile memories. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0189] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0190] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A method for removing duplicate points of interest based on graph neural network, characterized in that The method includes: Obtain all the points of interest within the target geographical area to be de-duplicated, and construct a location-based point-of-interest graph based on the all points of interest; Screen out multiple pairs of duplicate point-of-interest pairs from the all points of interest, and label the multiple pairs of duplicate point-of-interest pairs to obtain multiple pairs of seed duplicate point-of-interest pairs; Iteratively train a graph neural network model based on the point-of-interest graph and the multiple pairs of seed duplicate point-of-interest pairs to obtain a trained graph neural network model; Process the point-of-interest graph through the trained graph neural network model, and determine all the duplicate point-of-interest pairs among the all points of interest according to the processing result; Delete any one of the points of interest in each duplicate point-of-interest pair; Any training process of the graph neural network model includes: Input the adjacency matrix and the node attribute feature matrix of the point-of-interest graph into the graph neural network model to obtain the output data of the graph neural network model, where the output data includes the embedding of each point of interest among the all points of interest; Determine the target seed duplicate point-of-interest pairs for the current training from the multiple pairs of seed duplicate point-of-interest pairs, use the target seed duplicate point-of-interest pairs as the positive duplicate pairs for the current training, and construct the negative duplicate pairs for the current training according to the target seed duplicate point-of-interest pairs; Obtain the embeddings of the positive duplicate pairs and the embeddings of the negative duplicate pairs from the output data, and calculate the loss of the current training according to the embeddings of the positive duplicate pairs and the embeddings of the negative duplicate pairs; Judge whether the stop training condition is satisfied according to the loss; If it is satisfied, stop training and use the graph neural network model as the trained graph neural network model; If it is not satisfied, update the network parameters of the graph neural network model according to the loss, and perform the next training on the updated graph neural network model; The processing of the point-of-interest graph through the trained graph neural network model to determine all the duplicate point-of-interest pairs among the all points of interest according to the processing result includes: Process the point-of-interest graph through the trained graph neural network model to obtain the embedding of each point of interest among the all points of interest output by the trained graph neural network model; Pair the all points of interest in pairs to obtain multiple point-of-interest pairs, and calculate the embedding distance of each point-of-interest pair; the embedding distance of each point-of-interest pair refers to the distance between the embeddings of the two points of interest included in each point-of-interest pair; Determine the point-of-interest pairs with each embedding distance less than a preset threshold as duplicate point-of-interest pairs.

2. The method according to claim 1, wherein The construction of the location-based point-of-interest graph based on the all points of interest includes: Obtain the geographical location information of each point of interest among the all points of interest; Calculate the distance between each two points of interest respectively according to the geographical location information of each two points of interest; Determine the weight of the edge between each two points of interest according to the distance between each two points of interest to obtain the point-of-interest graph based on the geographical location information.

3. The method according to claim 2, wherein Determining the weight of the edge between each two points of interest according to the distance between each two points of interest includes: When the distance between any two points of interest is less than a preset threshold, determine that an edge relationship is formed between the any two points of interest, and set the weight of the edge between the any two points of interest to 1; When the distance between any two points of interest is greater than or equal to a preset threshold, it is determined that no edge relationship is formed between the two points of interest, and the weight of the edge between the two points of interest is set to 0.

4. The method according to claim 1, wherein The number of the point-of-interest graphs is multiple; Said constructing a location-based point-of-interest graph according to all the points of interest includes: Dividing the target geographical range into multiple spatial grids; Traversing the longitude and latitude attributes of each point of interest in all the points of interest to determine the set of points of interest in each spatial grid; Constructing a location-based point-of-interest graph for each spatial grid according to the set of points of interest in each spatial grid.

5. The method according to claim 4, characterized in that Said processing the point-of-interest graph through the trained graph neural network model and determining all pairs of duplicate points of interest among all the points of interest according to the processing result includes: Respectively inputting the point-of-interest graph of each spatial grid into the trained graph neural network model to obtain the embedding of each point of interest in the set of points of interest of each spatial grid; Pairing the points of interest in the set of points of interest of each spatial grid pairwise to obtain a set of pairs of points of interest for each spatial grid; Calculating the embedding distance of each pair of points of interest in the set of pairs of points of interest of each spatial grid; the embedding distance of each pair of points of interest refers to the distance between the embeddings of the two points of interest included in each pair of points of interest; Determining, among the set of pairs of points of interest of each spatial grid, each pair of points of interest with an embedding distance less than the preset threshold as a pair of duplicate points of interest.

6. An interest point deduplication device based on a graph neural network, characterized in that The device includes: A point-of-interest graph construction module, configured to obtain all the points of interest within the target geographical range to be deduplicated, and construct a location-based point-of-interest graph according to all the points of interest; A seed duplicate pair construction module, configured to screen out multiple pairs of duplicate points of interest from all the points of interest, and label the multiple pairs of duplicate points of interest to obtain multiple pairs of seed duplicate points of interest; A model training module, configured to perform iterative training on a graph neural network model according to the point-of-interest graph and the multiple pairs of seed duplicate points of interest to obtain a trained graph neural network model; A duplicate pair determination module, configured to process the point-of-interest graph through the trained graph neural network model and determine all pairs of duplicate points of interest among all the points of interest according to the processing result; A deduplication module, configured to delete any one of the points of interest in each pair of duplicate points of interest; Any training process of the graph neural network model includes: Inputting the adjacency matrix and the node attribute feature matrix of the point-of-interest graph into the graph neural network model to obtain the output data of the graph neural network model, where the output data includes the embedding of each point of interest among all the points of interest; Determining target seed duplicate points of interest for the current training from the multiple pairs of seed duplicate points of interest, using the target seed duplicate points of interest as the positive duplicate pairs for the current training, and constructing negative duplicate pairs for the current training according to the target seed duplicate points of interest; Obtaining the embeddings of the positive duplicate pairs and the embeddings of the negative duplicate pairs from the output data, and calculating the loss of the current training according to the embeddings of the positive duplicate pairs and the embeddings of the negative duplicate pairs; Judging whether the stop training condition is satisfied according to the loss; If satisfied, stop training and use the graph neural network model as the trained graph neural network model; If not satisfied, update the network parameters of the graph neural network model according to the loss, and perform the next training on the updated graph neural network model; Processing the point of interest graph through the trained graph neural network model, and determining all point of interest repeat pairs among all the points of interest according to the processing result, including: Processing the point of interest graph through the trained graph neural network model to obtain the embedding of each point of interest in all the points of interest output by the trained graph neural network model; Pairing all the points of interest in pairs to obtain a plurality of point of interest pairs, and calculating the embedding distance of each point of interest pair; the embedding distance of each point of interest pair refers to the distance between the embeddings of the two points of interest included in each point of interest pair; Determine the point of interest repeat pairs for those point of interest pairs with an embedding distance less than a preset threshold.

7. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Geographic position data processing method and device

    CN111460044A

  • Data deduplication method and device and storage medium

    CN112463774A

  • Infectious disease trend prediction method and system based on regional similarity

    CN113838582A