Heterogeneous Gateway Address Matching Method, Device, Computer Equipment and Storage Medium
The use of graph convolutional neural networks to integrate spatial and textual information in door address data structures enhances matching accuracy and efficiency, addressing the limitations of existing methods by leveraging spatial relationships and full dataset participation.
Patent Information
- Application Number
- CN202210460538.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-28
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-04-28
AI Technical Summary
The existing heterologous gate address matching methods have low matching accuracy and efficiency when text differences are large or spatial position relationships are not used. The traditional method assumes that gate addresses are independent and fails to effectively utilize the correlation information between gate addresses.
A gate address matching model based on graph neural network is constructed. By obtaining the gate address set of target regions, constructing the graph data structure, combining geographical location and text information, iteratively trains the graph neural network, identifying gate address matching pairs, and matching using the spatial position relationship and text features between gate addresses.
It improves the accuracy and efficiency of heterologous gate address matching, can more accurately identify gate address matching pairs, and improves the training adequacy of the matching model and graph-level matching performance.
Smart Images

Figure CN114741621B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of electronic maps, and in particular to a method and device for matching heterogeneous addresses, a computer device, and a storage medium. Background Art
[0002] An address is a type of map data, usually including information such as street name, house number, and longitude and latitude. When a user inputs an address, the map search engine can query the corresponding longitude and latitude coordinates based on the input address and mark them on the electronic map. Address data is an important content of network electronic maps and also the core of Internet location services. However, the sources of address data on the Internet are diverse, and the collection and processing processes are also different, resulting in certain differences in aspects such as spatial location, attribute information, and richness of address data. Therefore, how to effectively eliminate the inconsistencies between address data and organize them into a set of accurate and user-readable data has become a current research hotspot. Currently, the commonly used method is to fuse the information of different sources of address data through a matching method to enrich the information of address data and eliminate the inconsistencies between data.
[0003] Currently, the main heterogeneous address matching solutions are as follows:
[0004] 1. Unsupervised similarity calculation-based solution:
[0005] Extract addresses from two heterogeneous address data respectively, calculate the text similarity of the names of these two addresses and the text similarity of the addresses. The similarity algorithms include edit distance, TF-IDF (term frequency–inverse document frequency), etc. Calculate the overall similarity by setting a weight for the calculated text similarity of the name and the text similarity of the address as the similarity score between the two addresses. When the score is higher than a certain threshold, it can be considered that the two addresses have a matching relationship, thereby matching the heterogeneous address data.
[0006] 2. Traditional machine learning model-based text matching solution:
[0007] Extract address pairs with matching relationships from two heterogeneous address data as training data, construct features by calculating the text similarity of the names of the address pairs, physical distance, category similarity, etc., and use traditional machine learning methods such as gradient boosting decision tree GBDT, Xgboost, etc. to train a text matching model to determine whether two addresses have a matching relationship, thereby matching the heterogeneous address data.
[0008] 3. Pre-trained deep learning model-based text matching solution:
[0009] Taking the pairs of addresses with matching relationships as training data, fine-tuning on commonly used pre-trained deep models such as BERT (Bidirectional Encoder Representation from Transformers), ALBERT (A Lite BERT), etc., to train a text matching model to determine whether two addresses have a matching relationship.
[0010] The inventors found that there are some drawbacks in the above-mentioned solutions when applied in practice.
[0011] For example, the drawbacks of the above Solution 1:
[0012] (1) Based on the unsupervised similarity score method, for scenarios where two addresses truly have a matching relationship but are quite different textually, the matching effect is poor.
[0013] (2) Data where two addresses are close textually but actually do not have a matching relationship will cause mis-matching.
[0014] (3) It is not easy to set the threshold of the similarity score.
[0015] The drawbacks of the above Solution 2:
[0016] (1) A large amount of feature engineering work is required to construct features, and the process is rather cumbersome.
[0017] (2) The model is relatively shallow, with limited expressive ability, and the ceiling of the text matching effect is low.
[0018] (3) The matching process is to match addresses one by one, and then traverse all the data for overall matching. The matching efficiency is low and it cannot directly perform matching at the overall data level.
[0019] (4) This method assumes that addresses are independent of each other. However, there is a certain spatial position relationship between actual addresses, so the relationship information between addresses is not used for matching, resulting in less information utilization and poor effects.
[0020] The drawbacks of the above Solution 3:
[0021] (1) Pre-trained deep models generally input pure text information and have poor compatibility with non-text features.
[0022] (2) Similar to traditional machine learning models, the matching process of pre-trained deep models is to match addresses one by one, and then traverse all the data for overall matching. The matching efficiency is low and it cannot directly perform matching at the overall data level.
[0023] (3) This method assumes that the door addresses are independent of each other. However, there is a certain spatial position relationship between the actual door addresses. Therefore, the relationship information between the door addresses is not used for matching, resulting in less information utilization and poor performance. Summary of the Invention
[0024] In view of the above deficiencies or drawbacks, the present application provides a method, device, computer device, and storage medium for matching heterogeneous door addresses. Embodiments of the present application can improve the matching accuracy and speed of heterogeneous door addresses.
[0025] According to a first aspect of the present application, there is provided a method for matching heterogeneous door addresses. In one embodiment, the method includes:
[0026] Obtain a first set of door addresses within a target geographical area, and construct a first door address graph data structure according to the first set of door addresses;
[0027] Obtain a second set of door addresses within the target geographical area, and construct a second door address graph data structure according to the second set of door addresses; Any door address in the first set of door addresses is heterogeneous from any door address in the second set of door addresses.
[0028] Select multiple door address matching pairs from the first set of door addresses and the second set of door addresses, and label each door address matching pair to obtain corresponding sample door address matching pairs;
[0029] Use the first door address graph data structure, the second door address graph data structure, and the multiple sample door address matching pairs to iteratively train a door address matching model to obtain a trained door address matching model;
[0030] Process the first door address graph data structure and the second door address graph data structure through the trained door address matching model, and identify all door address matching pairs in the first set of door addresses and the second set of door addresses according to the processing results.
[0031] In one embodiment, constructing a first door address graph data structure according to the first set of door addresses includes:
[0032] Pair the door addresses in the first set of door addresses pairwise to obtain multiple first door address pairs;
[0033] Calculate the distance of each first door address pair according to the geographical location information of each first door address pair. The distance of the first door address pair refers to the distance between the two door addresses included in the first door address pair;
[0034] Determine the weight of the edge between the two door addresses included in each first door address pair according to the distance of each first door address pair to obtain a first door address graph data structure based on geographical location information;
[0035] Determining the weight of the edge between the two door addresses included in each first door address pair according to the distance of each first door address pair includes:
[0036] When the distance between any two gate addresses in the first set of gate addresses is less than the first preset threshold, it is determined that an edge relationship is formed between the any two gate addresses, and the weight of the edge between the any two gate addresses is set to 1;
[0037] When the distance between any two gate addresses in the first set of gate addresses is greater than or equal to the first preset threshold, it is determined that no edge relationship is formed between the any two gate addresses, and the weight of the edge between the any two gate addresses is set to 0.
[0038] In one embodiment, constructing a second gate address graph data structure according to the second set of gate addresses, including:
[0039] Pairing the gate addresses in the second set of gate addresses pairwise to obtain a plurality of second gate address pairs;
[0040] Calculating the distance of each second gate address pair according to the geographical location information of each second gate address pair, where the distance of the second gate address pair refers to the distance between the two gate addresses included in the second gate address pair;
[0041] Determining the weight of the edge between the two gate addresses included in each second gate address pair according to the distance of each second gate address pair, to obtain a second gate address graph data structure based on geographical location information;
[0042] Determining the weight of the edge between the two gate addresses included in each second gate address pair according to the distance of each second gate address pair, including:
[0043] When the distance between any two gate addresses in the second set of gate addresses is less than the second preset threshold, it is determined that an edge relationship is formed between the any two gate addresses, and the weight of the edge between the any two gate addresses is set to 1;
[0044] When the distance between any two gate addresses in the second set of gate addresses is greater than or equal to the second preset threshold, it is determined that no edge relationship is formed between the any two gate addresses, and the weight of the edge between the any two gate addresses is set to 0.
[0045] In one embodiment, iteratively training a gate address matching model using the first gate address graph data structure, the second gate address graph data structure, and the plurality of sample gate address matching pairs to obtain a trained gate address matching model, including:
[0046] Constructing a gate address matching model, where the gate address matching model is a graph neural network model;
[0047] The first gate address graph data structure, the second gate address graph data structure, and the plurality of sample gate address matching pairs iteratively train the gate address matching model, and when the preset training end condition is met, a trained gate address matching model is obtained;
[0048] Wherein, each training process of the gate address matching model includes:
[0049] Obtain the adjacency matrix and the node attribute feature matrix of the first site map data structure and the second site map data structure, and input them into the site matching model to obtain the output data of the site matching model. The output data includes the embeddings of each site in the first site map data structure and the second site map data structure.
[0050] Determine the target sample site matching pairs for this training from the multiple sample site matching pairs, use the target sample site matching pairs as positive matching pairs, and construct corresponding negative matching pairs according to the target sample site matching pairs.
[0051] Obtain the embeddings of the positive matching pairs and the embeddings of the negative matching pairs from the output data, and calculate the loss of this training according to the embeddings of the positive matching pairs and the embeddings of the negative matching pairs.
[0052] Judge whether the stop training condition is satisfied according to the loss.
[0053] When it is determined that the condition is satisfied, end the training, and use the site matching model trained this time as the trained site matching model.
[0054] When it is determined that the condition is not satisfied, update the network parameters of the site matching model according to the loss, and perform the next training on the site matching model with the updated network parameters.
[0055] In one embodiment, process the first site map data structure and the second site map data structure through the trained site matching model, and identify all site matching pairs in the first site set and the second site set according to the processing results, including:
[0056] Input the first site map data structure into the trained site matching model, and obtain the embeddings of each site in the first site set according to the output of the trained site matching model.
[0057] Input the second site map data structure into the trained site matching model, and obtain the embeddings of each site in the second site set according to the output of the trained site matching model.
[0058] Pair each site in the first site set with each site in the second site set pairwise to obtain a plurality of third site pairs.
[0059] Calculate the embedding distance of each third site pair; the embedding distance of each third site pair refers to the distance between the embeddings of the two sites included in each third site pair.
[0060] Determine the third site pairs with each embedding distance less than the preset threshold as site matching pairs.
[0061] In one embodiment, the number of the first address map data structures and the second address map data structures is the same, and both are multiple; correspondingly, constructing the first address map data structure according to the first address set includes:
[0062] Dividing the target geographical area into a plurality of spatial grids;
[0063] Traversing the longitude and latitude attributes of each address in the first address set to determine the first address subset corresponding to each spatial grid;
[0064] Constructing the corresponding first address map data structure for each spatial grid according to the first address subset corresponding to each spatial grid;
[0065] Constructing the second address map data structure according to the second address set includes:
[0066] Traversing the longitude and latitude attributes of each address in the second address set to determine the second address subset corresponding to each spatial grid;
[0067] Constructing the corresponding second address map data structure for each spatial grid according to the second address subset corresponding to each spatial grid.
[0068] In one embodiment, processing the first address map data structure and the second address map data structure through a trained address matching model, and identifying all address matching pairs in the first address set and the second address set according to the processing result, including:
[0069] Inputting the first address map data structure and the second address map data structure corresponding to each spatial grid into the trained address matching model respectively, and obtaining the embedding of each address in the first address subset corresponding to each spatial grid and the embedding of each address in the second address subset corresponding to each spatial grid according to the output of the trained address matching model;
[0070] Pairing the first address subset and the second address subset corresponding to each spatial grid pairwise to obtain the address pair set of each spatial grid;
[0071] Calculating the embedding distance of each fourth address pair in the address pair set of each spatial grid; wherein, the embedding distance of each fourth address pair refers to the distance between the embeddings of the two addresses included in each fourth address pair;
[0072] Determining each fourth address pair with an embedding distance less than a preset threshold in the address pair set of each spatial grid as an address matching pair.
[0073] The present application provides a heterogeneous address matching device according to the second aspect. In one embodiment, the device includes:
[0074] The first graph construction module is used to obtain the first address set of the target geographical area and construct the first address graph data structure according to the first address set;
[0075] The second graph construction module is used to obtain the second address set of the target geographical area and construct the second address graph data structure according to the second address set; Any address in the first address set and any address in the second address set are not from the same source;
[0076] The sample construction module is used to screen out multiple address matching pairs from the first address set and the second address set, label each address matching pair, and obtain the corresponding sample address matching pairs;
[0077] The training module is used to iteratively train the address matching model using the first address graph data structure, the second address graph data structure, and the multiple sample address matching pairs to obtain a trained address matching model;
[0078] The matching module is used to process the first address graph data structure and the second address graph data structure through the trained address matching model, and identify all address matching pairs in the first address set and the second address set according to the processing results.
[0079] According to a third aspect, the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the embodiments of any of the above methods are implemented.
[0080] According to a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the embodiments of any of the above methods are implemented.
[0081] In the above embodiments of the present application, by obtaining the first address set of the target geographical area, constructing the first address graph data structure according to the first address set, and obtaining the second address set of the target geographical area, constructing the second address graph data structure according to the second address set; wherein, any address in the first address set and any address in the second address set are not from the same source; then, screening out multiple address matching pairs from the first address set and the second address set, labeling each address matching pair, and obtaining the corresponding sample address matching pairs; after that, iteratively training the address matching model using the first address graph data structure, the second address graph data structure, and the multiple sample address matching pairs to obtain a trained address matching model, and finally processing the first address graph data structure and the second address graph data structure through the trained address matching model, and identifying all address matching pairs in the first address set and the second address set according to the processing results, it is possible to more accurately and quickly identify non-homogeneous addresses. Brief Description of the Drawings
[0082] Figure 1 It is a schematic flowchart of a method for heterogeneous address matching in an embodiment;
[0083] Figure 2 It is a structural block diagram of a device for heterogeneous address matching in an embodiment;
[0084] Figure 3 It is an internal structure diagram of a computer device in an embodiment. Specific implementation manners
[0085] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0086] The present application provides a method for heterogeneous address matching. In one embodiment, the method includes steps as Figure 1 shown, and the method can be applied to a cloud server. The method will be described below.
[0087] S110: Obtain a first address set of a target geographical area, and construct a first address graph data structure according to the first address set;
[0088] S120: Obtain a second address set of the target geographical area, and construct a second address graph data structure according to the second address set; any address in the first address set is heterogeneous from any address in the second address set;
[0089] S130: Screen out multiple address matching pairs from the first address set and the second address set, and label each address matching pair to obtain corresponding sample address matching pairs;
[0090] S140: Use the first address graph data structure, the second address graph data structure and the multiple sample address matching pairs to iteratively train an address matching model to obtain a trained address matching model;
[0091] S150: Process the first address graph data structure and the second address graph data structure through the trained address matching model, and identify all address matching pairs in the first address set and the second address set according to the processing results.
[0092] This embodiment can bring the following beneficial effects compared with the prior art:
[0093] Traditional machine learning and pre-trained deep models make the assumption that two addresses are irrelevant in the address deduplication problem. The embodiment of the present application assumes that there is an association between two addresses. On this basis, the heterogeneous address matching is performed by combining the attribute characteristics of the addresses themselves and the spatial position relationship between the addresses.
[0094] Specifically, the inventors found that there are usually similar neighbors around the mutually matching address pairs. Based on this discovery, a graph convolutional neural network model based on a graph structure is selected for heterogeneous address matching. The text information of the address names and the spatial position relationship information between the addresses are used as the model inputs. The graph neural network has the ability to recognize isomorphic graphs by transmitting and aggregating neighbor node information, and can better recognize similar neighbor structures. In the address graph data structure, the address pairs that can usually be matched often have similar neighbors, that is, among the adjacent addresses of two matched addresses, there often contain other equivalent address pairs. And the embedded representation of the node is generated by aggregating neighbor information, so that other equivalent address pairs in the neighbor nodes are also more likely to be matched. In this way, the matching problem between two heterogeneous address graph data structures can be solved more effectively.
[0095] On the other hand, traditional machine learning and pre-trained deep models generally can only perform model training based on sample address pairs, and the sample address pairs need to be manually labeled, with a relatively high acquisition cost, so the quantity is generally small, which will lead to insufficient training. The embodiment of the present application can enable all data to participate in training through the graph convolutional neural network model, so that the model can be trained more fully and the model effect is better.
[0096] In addition, traditional address matching methods all match addresses one by one, and then traverse the entire graph data structure for graph data structure level matching, with relatively low matching efficiency and unable to directly perform matching at the graph data structure level. However, the method using the graph convolutional neural network model in this case can achieve layer-level matching, improving the performance and efficiency of matching.
[0097] In one embodiment, constructing the first address graph data structure according to the first address set as described above includes: pairing the addresses in the first address set pairwise to obtain a plurality of first address pairs; calculating the distance of each first address pair according to the geographical location information of each first address pair, where the distance of the first address pair refers to the distance between the two addresses included in the first address pair; determining the weight of the edge between the two addresses included in each first address pair according to the distance of each first address pair to obtain the first address graph data structure based on geographical location information. The distance between two addresses refers to the spatial distance (or physical distance), such as 200 meters, 100 meters, 50 meters, etc.
[0098] Among them, determining the weight of the edge between the two addresses included in each first address pair according to the distance of each first address pair includes: when the distance between any two addresses in the first address set is less than the first preset threshold, determining that an edge relationship is formed between the any two addresses, and setting the weight of the edge between the any two addresses to 1; when the distance between any two addresses in the first address set is greater than or equal to the first preset threshold, determining that no edge relationship is formed between the any two addresses, and setting the weight of the edge between the any two addresses to 0. Among them, the first preset threshold can be set according to actual needs, such as set to 50 meters, etc., and this embodiment does not specifically limit this.
[0099] In one embodiment, constructing the second address graph data structure according to the second address set includes: pairing the addresses in the second address set pairwise to obtain a plurality of second address pairs; calculating the distance of each second address pair according to the geographical location information of each second address pair, and the distance of the second address pair refers to the distance between the two addresses included in the second address pair; determining the weight of the edge between the two addresses included in each second address pair according to the distance of each second address pair to obtain the second address graph data structure based on geographical location information. The above-mentioned second preset threshold is the same as the first preset threshold in the above embodiment, and for details, please refer to the description of the above embodiment.
[0100] Among them, determining the weight of the edge between the two addresses included in each second address pair according to the distance of each second address pair includes: when the distance between any two addresses in the second address set is less than the second preset threshold, determining that an edge relationship is formed between the any two addresses, and setting the weight of the edge between the any two addresses to 1; when the distance between any two addresses in the second address set is greater than or equal to the second preset threshold, determining that no edge relationship is formed between the any two addresses, and setting the weight of the edge between the any two addresses to 0.
[0101] In one embodiment, iteratively training the address matching model using the first address graph data structure, the second address graph data structure, and the plurality of sample address matching pairs to obtain the trained address matching model includes: constructing the address matching model, and the address matching model is a graph neural network model; iteratively training the address matching model with the first address graph data structure, the second address graph data structure, and the plurality of sample address matching pairs, and when the preset end training condition is satisfied, obtaining the trained address matching model.
[0102] Among them, each training process of the address matching model includes:
[0103] Obtain the adjacency matrix and the node attribute feature matrix of the first site map data structure and the second site map data structure, and input them into the site matching model to obtain the output data of the site matching model. The output data includes the embeddings of each site in the first site map data structure and the second site map data structure;
[0104] Determine the target sample site matching pairs for this training from the multiple sample site matching pairs, use the target sample site matching pairs as positive matching pairs, and construct corresponding negative matching pairs according to the target sample site matching pairs;
[0105] Obtain the embeddings of the positive matching pairs and the embeddings of the negative matching pairs from the output data, and calculate the loss for this training based on the embeddings of the positive matching pairs and the embeddings of the negative matching pairs;
[0106] Judge whether the stop training condition is satisfied according to the loss;
[0107] When it is determined that the condition is satisfied, end the training, and use the site matching model trained this time as the trained site matching model;
[0108] When it is determined that the condition is not satisfied, update the network parameters of the site matching model according to the loss, and perform the next training on the site matching model with the updated network parameters.
[0109] Correspondingly, in one embodiment, processing the first site map data structure and the second site map data structure through the trained site matching model, and identifying all site matching pairs in the first site set and the second site set according to the processing results, includes:
[0110] Input the first site map data structure into the trained site matching model, and obtain the embeddings of each site in the first site set according to the output of the trained site matching model; input the second site map data structure into the trained site matching model, and obtain the embeddings of each site in the second site set according to the output of the trained site matching model; pair each site in the first site set with each site in the second site set pairwise to obtain a plurality of third site pairs; calculate the embedding distance of each third site pair; the embedding distance of each third site pair refers to the distance between the embeddings of the two sites included in each third site pair; determine the third site pairs with each embedding distance less than the preset threshold as site matching pairs.
[0111] In another embodiment, the number of the first site map data structures and the second site map data structures is the same, and both are multiple. Correspondingly, constructing the first site map data structure according to the first site set includes: dividing the target geographical area into multiple spatial grids; traversing the longitude and latitude attributes of each site in the first site set to determine the first site subset corresponding to each spatial grid; and constructing the corresponding first site map data structure for each spatial grid according to the first site subset corresponding to each spatial grid. Constructing the second site map data structure according to the second site set includes: traversing the longitude and latitude attributes of each site in the second site set to determine the second site subset corresponding to each spatial grid; and constructing the corresponding second site map data structure for each spatial grid according to the second site subset corresponding to each spatial grid.
[0112] In this embodiment, it is considered that in some scenarios, when the range of the target geographical area is large. For example, if the target geographical area refers to the whole of China, then the data to be matched (referring to site matching) at this time is the full amount of data across the whole country. At this time, the total amount of data is very large. If all the data is directly used to construct two maps (referring to site map data structures), then the node attribute feature matrix and the adjacency matrix of each site map data structure will be very large, which will require a very high demand for computing resources. Therefore, in the case of limited computing resources, the two maps are divided into many small maps according to spatial grids, and then deduplication is performed on each small map. In this way, the nodes and adjacency matrices of each small map become smaller, and data matching can be performed with less computing resources. The specific method can be as follows:
[0113] According to the geographical spatial coordinates of China, from the westernmost end to the easternmost end, and from the northernmost end to the southernmost end, it is divided into square spatial grids of 1 kilometer * 1 kilometer (the grid size can be flexibly adjusted according to actual needs). The four vertices of the grid have corresponding longitude and latitude coordinates, and the sites all have longitude and latitude attributes. By traversing all the sites according to the longitude and latitude, the sites can fall into the corresponding grids. In this way, each grid will have sites that are relatively close, and subsequent site matching only needs to be performed in each spatial grid.
[0114] Correspondingly, in one embodiment, the above-mentioned processing of the first address map data structure and the second address map data structure by the trained address matching model to identify all address matching pairs in the first address set and the second address set according to the processing result includes: inputting the first address map data structure and the second address map data structure corresponding to each spatial grid into the trained address matching model respectively, and obtaining the embedding of each address in the first address subset corresponding to each spatial grid and the embedding of each address in the second address subset corresponding to each spatial grid according to the output of the trained address matching model; pairing the first address subsets and the second address subsets corresponding to each spatial grid pairwise to obtain the address pair set of each spatial grid; calculating the embedding distance of each fourth address pair in the address pair set of each spatial grid; wherein, the embedding distance of each fourth address pair refers to the distance between the embeddings of the two addresses included in each fourth address pair; and determining each fourth address pair with an embedding distance less than the preset threshold in the address pair set of each spatial grid as an address matching pair.
[0115] The above embodiment is described below through a specific application example.
[0116] This application example is specifically divided into six parts: division of spatial grids, acquisition of sample address matching pairs, data preprocessing, feature engineering, model training, and model prediction.
[0117] This application example takes China as the target geographical area. First, according to the geographical spatial coordinates of China, the geographical space of China is divided into square grids of 1 km * 1 km, and then the corresponding addresses are dropped into the corresponding grids to form subgraphs. This application example uses supervised learning to train the graph neural network model. Since it is a supervised learning task, artificial annotation of sample address matching pairs is required for model training. After obtaining the sample address matching pairs, data preprocessing is required, including operations such as converting full-width characters to half-width characters, removing special symbols, converting English uppercase to lowercase, and converting traditional Chinese to simplified Chinese, to clean the data. Before training the model, feature engineering needs to be done first to obtain the attribute features of each node in each address map data structure and the adjacency matrix of each address map data structure. The node attribute feature matrix and the adjacency matrix of each address map data structure are input into the graph convolutional neural network GCN (Graph Convolutional Network), and the model is trained by minimizing the loss function through backpropagation to obtain the weight matrix W, which is the parameter that the model needs to learn. After obtaining the parameter W, graph-level address matching prediction can be performed through forward propagation.
[0118] The above-mentioned parts are described below.
[0119] 1. Division of spatial grids
[0120] Since the two address map data structures to be matched (which can be simply referred to as map data structures) are full-scale data for the whole of China, and the total amount of data is very large. If all the data is directly used to construct a large graph, both the nodes and the adjacency matrix will be extremely large, requiring too much computing resources. Therefore, in the case of limited computing resources, the large graph is divided into many small graphs according to spatial grids, and the two map data structures are matched on their respective small graphs. In this way, both the nodes and the adjacency matrix become smaller, and the matching can be carried out with less computing resources.
[0121] Specific method: According to the geographical spatial coordinates of China, from the westernmost end to the easternmost end, and from the northernmost end to the southernmost end, it is divided into square spatial grids of 1 km * 1 km. The four vertices of the grid all have corresponding longitude and latitude coordinates, and the addresses all have longitude and latitude attributes. By traversing all the addresses according to the longitude and latitude, the addresses can fall into the corresponding grids. In this way, each grid will have addresses that are relatively close.
[0122] Perform the above operations on both of these address map data structures, and then only need to perform address matching in their respective corresponding spatial grids.
[0123] 2. Obtaining sample address matching pairs
[0124] This application example requires sample address pairs to train the model, so some data needs to be labeled as sample address pairs. Traverse all the above spatial grids, and use simple text similarity of address names (such as edit distance, etc.) to roughly find suspected matching address pairs in each spatial grid, and then hand them over to the annotators to find the real address pairs. In this way, sample address matching pairs are constructed.
[0125] 3. Data preprocessing
[0126] The text of the address name may contain special symbols and traditional Chinese characters, etc., and needs to be preprocessed first before constructing features for input into the model. At the same time, in order to ensure the consistency of the distribution of labeled data and unlabeled data, the same preprocessing operations need to be performed on the labeled data and unlabeled data. The data preprocessing process includes the following four steps:
[0127] (1) Converting full-width characters to half-width characters
[0128] (2) Removing special symbols
[0129] (3) Converting English uppercase to lowercase
[0130] (4) Converting traditional Chinese to simplified Chinese
[0131] 4. Feature engineering
[0132] (1) Generating attribute features of graph nodes
[0133] The input of the graph convolutional neural network GCN includes the topological structure of the graph, i.e., the adjacency matrix, and the attribute features of all nodes in the graph. Each node attribute feature is a multi-dimensional feature vector. In this application example, a specified algorithm is used to process each door address into a 512-dimensional Embedding vector, and this Embedding vector is used as the attribute feature of the node. Among them, the specified algorithm can be any existing algorithm that can map the door address to the Embedding vector, so it will not be elaborated here.
[0134] (2) Generate the edges and adjacency matrix of all sub-graph structures
[0135] In this application example, the door addresses within 50 meters are regarded as having edge relationships. Take out all the door addresses in a spatial grid to form a set. Take out a door address from the set and calculate the distance with all other door addresses in the set except itself. Among them, the door addresses with a distance less than 50 meters form an edge relationship with the taken-out door address, and the weight of the edge is 1; then take out the door addresses from the set in turn and perform the above operations until all the door addresses in the set are taken out, and the edge relationships of all the door addresses in this spatial grid sub-graph are formed. According to the definition of the graph structure, the adjacency matrix of this sub-graph is obtained. The representation of the adjacency matrix is:
[0136]
[0137] Finally, traverse all the spatial grids according to the above method to generate the adjacency matrices of all spatial grid sub-graphs.
[0138] 5. Model training
[0139] Given two corresponding address sub-graphs KG1 and KG2 of spatial grids, and a set of pre-matched address sample pairs S = {(e i1 , e i2 )} m i=1 .
[0140] The method of this application example finds new address matching pairs based on the node embedding of GCN. The basic idea of the method is to use GCN to embed the addresses from different graph data structures into a unified vector space, and at the same time hope that the distance between the matched addresses is closer, and the distance between the unmatched addresses is farther.
[0141] (1) Input of GCN:
[0142] GCN is a type of neural network that operates directly on graphs. Its inputs are the node attribute features and the adjacency matrix of the graph, with the aim of outputting node-level location embeddings, which are then used for subsequent location matching. The model uses two two-layer GCNs, with each GCN processing one KG. Let GCN1 and GCN2 process KG1 and KG2 respectively. For the node attribute features input to the model, they are all derived from the feature engineering in step 4, and the dimensionality of the node attribute features input to GCN1 and GCN2 is 512 dimensions; for the adjacency matrix input to the model, it is also obtained through 4-step feature engineering. After inputting the node attribute features and the adjacency matrix into the GCN model, subsequent GCN operations are performed.
[0143] (2) Operations of GCN:
[0144] A GCN model contains multiple GCN layers. In this application example, two layers are selected. The input H (l) ∈R n×d(l) of the l-th layer is a node attribute feature matrix (i.e., all node attribute features), where n is the number of nodes, and d (l) is the number of features in the l-th layer. The output of the l-th layer is a new feature matrix:
[0145]
[0146] where σ is the relu activation function (used for linear activation transformation), A is an n*n adjacency matrix, where I is the identity matrix. is 's diagonal node degree matrix, and W (l) ∈R d(l)×d(l+1) is the weight matrix between two layers, used for convolution operations, and d (l+1) is the dimension of the new layer.
[0147] (3) Output of GCN
[0148] After passing through two two-layer GCNs, the dimensionalities of the node feature vectors output by GCN1 and GCN2 are the same, both being 512-dimensional embedding representations, and this embedding representation can be used for subsequent location matching.
[0149] (4) Loss function of GCN
[0150] In this application example, it is desired that the distance between the locations to be matched is small, and the distance between locations that cannot be matched is large. Based on this, a loss function is constructed. The distance between locations is the embedding distance between locations. For location pairs e1 and e2, where e1 ∈ KG1 and e2 ∈ KG2, e1 and e2 are the node embeddings output by GCN1 and GCN2 respectively, and the method for calculating the distance between them is as follows:
[0151] D(e1,e2) = ||h(e1) - h(e2)||1
[0152] The model is trained by minimizing the following margin-based loss function:
[0153]
[0154] where [x] + = max{0, x}, S' (e1,e2) is a negative matching pair obtained by randomly replacing one address from (e1, e2), and γ is the margin that distinguishes positive and negative matching pairs. The model is trained by minimizing the loss function through backpropagation to update the weight matrix W in each layer. After several rounds of training, the final model can learn the weight matrix W to predict the address pair matching.
[0155] 6. Model Prediction
[0156] This application example is suitable for offline graph-level matching prediction. The prediction is to find more new address matching pairs in the constructed graph. During the training process, the weight matrix W is learned. By inputting the node attribute features and the adjacency matrix, after the operations of the GCN, each node will output an embedding representation.
[0157] For a specific output embedding e1 ∈ KG1, calculate its embedding distance from all addresses in KG2, and select the one with the smallest embedding distance among all addresses. If this embedding distance is less than a certain threshold, it is considered that these two address pairs match; if it is not less than this threshold, it is considered that they do not match. According to the above method, by traversing all addresses in KG1 that are not in the sample pairs, the corresponding matching addresses in KG2 can be obtained, and thus the matching prediction result can be directly obtained at the layer level.
[0158] This application example has statistically counted the application situations of the heterogeneous address matching method provided in this application and other address matching methods mentioned in the background technology. The specific benefits are as follows: The accuracy rate of the address matching method based on the unsupervised similarity calculation scheme is 95.4%, and the recall rate is 84.1%; the accuracy rate of the address matching method based on the traditional machine learning model is 96.7%, and the recall rate is 91.8%; the accuracy rate of the address matching method based on the deep learning model is 97.8%, and the recall rate is 94.6%; while the accuracy rate of the method provided in this application is 98.5%, and the recall rate is 95.7%. Obviously, compared with the traditional methods, the method provided in this application has a greater improvement in effect.
[0159] Figure 1 is a schematic flowchart of the heterogeneous address matching method in an embodiment. It should be understood that although Figure 1The steps in the flowchart are shown in sequence according to the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 at least a part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in rotation with at least a part of other steps or sub-steps or stages of other steps.
[0160] Based on the same inventive concept, the present application also provides a heterogeneous address matching device. In this embodiment, as Figure 2 shown, the heterogeneous address matching device includes the following modules:
[0161] The first graph construction module 110 is configured to obtain a first address set of a target geographical area and construct a first address graph data structure according to the first address set;
[0162] The second graph construction module 120 is configured to obtain a second address set of the target geographical area and construct a second address graph data structure according to the second address set; any address in the first address set is different in source from any address in the second address set;
[0163] The sample construction module 130 is configured to screen out multiple address matching pairs from the first address set and the second address set, label each address matching pair, and obtain corresponding sample address matching pairs;
[0164] The training module 140 is configured to iteratively train an address matching model using the first address graph data structure, the second address graph data structure, and the multiple sample address matching pairs to obtain a trained address matching model;
[0165] The matching module 150 is configured to process the first address graph data structure and the second address graph data structure through the trained address matching model, and identify all address matching pairs in the first address set and the second address set according to the processing result.
[0166] In one embodiment, the first graph construction module 110 is configured to:
[0167] Pair the addresses in the first address set pairwise to obtain multiple first address pairs;
[0168] Calculate the distance of each first address pair according to the geographical location information of each first address pair, where the distance of the first address pair refers to the distance between the two addresses included in the first address pair;
[0169] Determine the weights of the edges between the two addresses included in each first address pair according to the distances of each first address pair, and obtain the first address graph data structure based on the geographical location information.
[0170] Among them, when the first graph construction module 110 determines the weights of the edges between the two addresses included in each first address pair according to the distances of each first address pair, it is specifically used for:
[0171] When the distance between any two addresses in the first address set is less than the first preset threshold, it is determined that an edge relationship is formed between the two addresses, and the weight of the edge between the two addresses is set to 1;
[0172] When the distance between any two addresses in the first address set is greater than or equal to the first preset threshold, it is determined that no edge relationship is formed between the two addresses, and the weight of the edge between the two addresses is set to 0.
[0173] In one embodiment, the second graph construction module 120 is used for:
[0174] Pair the addresses in the second address set pairwise to obtain a plurality of second address pairs;
[0175] Calculate the distance of each second address pair according to the geographical location information of each second address pair, and the distance of the second address pair refers to the distance between the two addresses included in the second address pair;
[0176] Determine the weights of the edges between the two addresses included in each second address pair according to the distances of each second address pair, and obtain the second address graph data structure based on the geographical location information.
[0177] Among them, when the second graph construction module 120 determines the weights of the edges between the two addresses included in each second address pair according to the distances of each second address pair, it is specifically used for:
[0178] When the distance between any two addresses in the second address set is less than the second preset threshold, it is determined that an edge relationship is formed between the two addresses, and the weight of the edge between the two addresses is set to 1;
[0179] When the distance between any two addresses in the second address set is greater than or equal to the second preset threshold, it is determined that no edge relationship is formed between the two addresses, and the weight of the edge between the two addresses is set to 0.
[0180] In one embodiment, the training module 140 is used for:
[0181] Construct an address matching model, and the address matching model is a graph neural network model;
[0182] Iteratively train the door address matching model with the first door address map data structure, the second door address map data structure, and the multiple sample door address matching pairs, and obtain a trained door address matching model when the preset training end condition is satisfied;
[0183] Among them, each training process of the door address matching model includes:
[0184] Obtain the adjacency matrix and node attribute feature matrix of the first door address map data structure and the second door address map data structure, and input them into the door address matching model to obtain the output data of the door address matching model. The output data includes the embeddings of each door address in the first door address map data structure and the second door address map data structure;
[0185] Determine the target sample door address matching pair for this training from the multiple sample door address matching pairs, use the target sample door address matching pair as the positive matching pair, and construct the corresponding negative matching pair according to the target sample door address matching pair;
[0186] Obtain the embeddings of the positive matching pair and the embeddings of the negative matching pair from the output data, and calculate the loss of this training according to the embeddings of the positive matching pair and the embeddings of the negative matching pair;
[0187] Judge whether the stop training condition is satisfied according to the loss;
[0188] When it is determined to be satisfied, end the training, and use the door address matching model trained this time as the trained door address matching model;
[0189] When it is determined not to be satisfied, update the network parameters of the door address matching model according to the loss, and perform the next training on the door address matching model with the updated network parameters.
[0190] In one embodiment, the matching module 150 is used for:
[0191] Input the first door address map data structure into the trained door address matching model, and obtain the embeddings of each door address in the first door address set according to the output of the trained door address matching model;
[0192] Input the second door address map data structure into the trained door address matching model, and obtain the embeddings of each door address in the second door address set according to the output of the trained door address matching model;
[0193] Pair each door address in the first door address set with each door address in the second door address set to obtain a plurality of third door address pairs;
[0194] Calculate the embedding distance of each third door address pair; the embedding distance of each third door address pair refers to the distance between the embeddings of the two door addresses included in each third door address pair;
[0195] Determine each third gateway pair with an embedding distance less than a preset threshold as a gateway matching pair.
[0196] In another embodiment, the number of the first gateway map data structures and the second gateway map data structures is the same, both being multiple; correspondingly, the first map construction module 110 is further configured to:
[0197] Divide the target geographical area into multiple spatial grids;
[0198] Traverse the longitude and latitude attributes of each gateway in the first gateway set to determine the first gateway subset corresponding to each spatial grid;
[0199] Construct the corresponding first gateway map data structure for each spatial grid according to the first gateway subset corresponding to each spatial grid.
[0200] The first map construction module 110 is further configured to:
[0201] Traverse the longitude and latitude attributes of each gateway in the second gateway set to determine the second gateway subset corresponding to each spatial grid;
[0202] Construct the corresponding second gateway map data structure for each spatial grid according to the second gateway subset corresponding to each spatial grid.
[0203] Correspondingly, in one embodiment, the matching module 150 is further configured to:
[0204] Input the first gateway map data structure and the second gateway map data structure corresponding to each spatial grid into the trained gateway matching model respectively, and obtain the embedding of each gateway in the first gateway subset corresponding to each spatial grid and the embedding of each gateway in the second gateway subset corresponding to each spatial grid according to the output of the trained gateway matching model;
[0205] Pair the first gateway subset and the second gateway subset corresponding to each spatial grid pairwise to obtain the gateway pair set of each spatial grid;
[0206] Calculate the embedding distance of each fourth gateway pair in the gateway pair set of each spatial grid; wherein, the embedding distance of each fourth gateway pair refers to the distance between the embeddings of the two gateways included in each fourth gateway pair;
[0207] Determine each fourth gateway pair with an embedding distance less than a preset threshold in the gateway pair set of each spatial grid as a gateway matching pair.
[0208] For the specific limitations of the heterogeneous gateway matching device, reference may be made to the limitations of the heterogeneous gateway matching method in the foregoing text, which will not be elaborated herein. Each module in the above heterogeneous gateway matching device can be implemented in whole or in part by software, hardware, and their combinations. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the above modules.
[0209] In one embodiment, a computer device is provided, and its internal structure diagram can be as Figure 3 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a heterogeneous gateway matching method.
[0210] Those skilled in the art can understand that Figure 3 the structure shown in
[0211] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0212] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it implements the steps in the method provided in any of the above method embodiments.
[0213] Those of ordinary skill in the art can understand that implementing all or part of the processes in the above method embodiments can be accomplished by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0214] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0215] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of the patent of this application should be subject to the appended claims.
Claims
1. A heterologous gateway matching method, characterized in that The method includes: Obtaining a first set of door addresses within a target geographical range, and constructing a first door address graph data structure according to the first set of door addresses; Obtaining a second set of door addresses within the target geographical range, and constructing a second door address graph data structure according to the second set of door addresses; any door address in the first set of door addresses is not from the same source as any door address in the second set of door addresses; Screening out multiple door address matching pairs from the first set of door addresses and the second set of door addresses, and labeling each door address matching pair to obtain corresponding sample door address matching pairs; Iteratively training a door address matching model using the first door address graph data structure, the second door address graph data structure, and multiple sample door address matching pairs to obtain a trained door address matching model; Processing the first door address graph data structure and the second door address graph data structure through the trained door address matching model, and identifying all door address matching pairs in the first set of door addresses and the second set of door addresses according to the processing results; The door address matching model finds new door address matching pairs based on the node embeddings of two two-layer graph convolutional neural networks (GCNs). Each GCN processes a door address sub-graph KG. Let the first GCN1 and the second GCN2 process the first door address sub-graph KG1 and the second door address sub-graph KG2 respectively. For the door address pair e1 and e2, where e1 ∈ KG1 and e2 ∈ KG2, e1 and e2 are the node embeddings output by the first GCN1 and the second GCN2. For a specific output node embedding e1 ∈ the first door address sub-graph KG1, calculate the embedding distance between the node embedding e1 and all the door addresses in the second door address sub-graph KG2, and select the door address with the smallest embedding distance among all the door addresses. If the embedding distance is less than the threshold, it is considered that these two door address pairs match; if it is not less than the threshold, it is considered that they do not match.
2. The method according to claim 1, characterized in that, The constructing the first door address graph data structure according to the first set of door addresses includes: Pairing the door addresses in the first set of door addresses pairwise to obtain multiple first door address pairs; Calculating the distance of each first door address pair according to the geographical location information of each first door address pair, where the distance of the first door address pair refers to the distance between the two door addresses included in the first door address pair; Determining the weight of the edge between the two door addresses included in each first door address pair according to the distance of each first door address pair to obtain a first door address graph data structure based on geographical location information; The determining the weight of the edge between the two door addresses included in each first door address pair according to the distance of each first door address pair includes: When the distance between any two door addresses in the first set of door addresses is less than a first preset threshold, determining that an edge relationship is formed between the any two door addresses, and setting the weight of the edge between the any two door addresses to 1; When the distance between any two door addresses in the first set of door addresses is greater than or equal to the first preset threshold, determining that no edge relationship is formed between the any two door addresses, and setting the weight of the edge between the any two door addresses to 0.
3. The method according to claim 1, characterized in that The constructing the second door address graph data structure according to the second set of door addresses includes: Pair the addresses in the second set of addresses pairwise to obtain a plurality of second address pairs; Calculate the distance of each second address pair according to the geographical location information of each second address pair, where the distance of the second address pair refers to the distance between the two addresses included in the second address pair; Determine the weight of the edge between the two addresses included in each second address pair according to the distance of each second address pair, and obtain a second address graph data structure based on geographical location information; The determining the weight of the edge between the two addresses included in each second address pair according to the distance of each second address pair includes: When the distance between any two addresses in the second set of addresses is less than a second preset threshold, determine that an edge relationship is formed between the any two addresses, and set the weight of the edge between the any two addresses to 1; When the distance between any two addresses in the second set of addresses is greater than or equal to the second preset threshold, determine that no edge relationship is formed between the any two addresses, and set the weight of the edge between the any two addresses to 0.
4. The method according to claim 1, wherein The using the first address graph data structure, the second address graph data structure and a plurality of sample address matching pairs to iteratively train an address matching model to obtain a trained address matching model includes: Construct an address matching model, where the address matching model is a graph neural network model; The first address graph data structure, the second address graph data structure and a plurality of sample address matching pairs iteratively train the address matching model, and when a preset end training condition is satisfied, obtain a trained address matching model; Among them, each training process of the address matching model includes: Obtain the adjacency matrix and the node attribute feature matrix of the first address graph data structure and the second address graph data structure, and input them into the address matching model to obtain the output data of the address matching model, where the output data includes the embeddings of each address in the first address graph data structure and the second address graph data structure; Determine a target sample address matching pair for the current training from the plurality of sample address matching pairs, use the target sample address matching pair as a positive matching pair, and construct a corresponding negative matching pair according to the target sample address matching pair; Obtain the embedding of the positive matching pair and the embedding of the negative matching pair from the output data, and calculate the loss of the current training according to the embedding of the positive matching pair and the embedding of the negative matching pair; Judge whether the stop training condition is satisfied according to the loss; When it is determined that the condition is satisfied, end the training, and use the address matching model trained in this time as the trained address matching model; When it is determined that the condition is not satisfied, update the network parameters of the address matching model according to the loss, and perform the next training on the address matching model with updated network parameters.
5. The method according to claim 1, characterized in that, The identifying all address matching pairs in the first address set and the second address set according to the processing result by processing the first address graph data structure and the second address graph data structure through the trained address matching model includes: Input the first address map data structure into the trained address matching model, and obtain the embedding of each address in the first address set according to the output of the trained address matching model; Input the second address map data structure into the trained address matching model, and obtain the embedding of each address in the second address set according to the output of the trained address matching model; Pair each address in the first address set with each address in the second address set pairwise to obtain a plurality of third address pairs; Calculate the embedding distance of each third address pair; the embedding distance of each third address pair refers to the distance between the embeddings of the two addresses included in each third address pair; Determine the third address pairs with each embedding distance less than the preset threshold as address matching pairs.
6. The method according to claim 1, wherein The number of the first address map data structure and the second address map data structure is the same, both being a plurality; The constructing the first address map data structure according to the first address set includes: Divide the target geographical range into a plurality of spatial grids; Traverse the longitude and latitude attributes of each address in the first address set, and determine the first address subset corresponding to each spatial grid; Construct the corresponding first address map data structure for each spatial grid according to the first address subset corresponding to each spatial grid; The constructing the second address map data structure according to the second address set includes: Traverse the longitude and latitude attributes of each address in the second address set, and determine the second address subset corresponding to each of the spatial grids; Construct the corresponding second address map data structure for each spatial grid according to the second address subset corresponding to each of the spatial grids.
7. The method according to claim 6, wherein Processing the first address map data structure and the second address map data structure through the trained address matching model, and identifying all address matching pairs in the first address set and the second address set according to the processing result includes: Input the first address map data structure and the second address map data structure corresponding to each spatial grid into the trained address matching model respectively, and obtain the embedding of each address in the first address subset corresponding to each spatial grid and the embedding of each address in the second address subset corresponding to each spatial grid according to the output of the trained address matching model; Pair the first address subset and the second address subset corresponding to each spatial grid pairwise to obtain the address pair set of each spatial grid; Calculate the embedding distance of each fourth address pair in the address pair set of each spatial grid; wherein, the embedding distance of each fourth address pair refers to the distance between the embeddings of the two addresses included in each fourth address pair; Determine the fourth address pairs with each embedding distance less than the preset threshold in the address pair set of each spatial grid as address matching pairs.
8. A heterologous gateway matching device, characterized in that, The device includes: A first graph construction module, configured to obtain a first address set of a target geographical range and construct a first address map data structure according to the first address set; The second graph construction module is used to obtain the second set of portal addresses within the target geographical range and construct a second portal address graph data structure according to the second set of portal addresses; any portal address in the first set of portal addresses is not from the same source as any portal address in the second set of portal addresses; The sample construction module is used to screen out multiple portal address matching pairs from the first set of portal addresses and the second set of portal addresses, label each portal address matching pair, and obtain the corresponding sample portal address matching pairs; The training module is used to iteratively train the portal address matching model using the first portal address graph data structure, the second portal address graph data structure, and multiple sample portal address matching pairs to obtain a trained portal address matching model; The matching module is used to process the first portal address graph data structure and the second portal address graph data structure through the trained portal address matching model, and identify all portal address matching pairs in the first set of portal addresses and the second set of portal addresses according to the processing results; The portal address matching model finds new portal address matching pairs based on the node embeddings of two two-layer graph convolutional neural networks (GCNs). Each GCN processes a portal address subgraph KG. Let the first GCN1 and the second GCN2 process the first portal address subgraph KG1 and the second portal address subgraph KG2 respectively. For the portal address pair e1 and e2, where e1 ∈ KG1 and e2 ∈ KG2, e1 and e2 are the node embeddings output by the first GCN1 and the second GCN2. For a specific output node embedding e1 ∈ the first portal address subgraph KG1, calculate the embedding distance between the node embedding e1 and all the portal addresses in the second portal address subgraph KG2, and select the portal address with the smallest embedding distance among all the portal addresses. If the embedding distance is less than the threshold, it is considered that these two portal address pairs match; if it is not less than the threshold, it is considered that they do not match.
9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method described in any one of claims 1 to 7.
Citation Information
Patent Citations
Address matching method and device, computer equipment and storage medium
CN110442603A
Intelligent contract Pincer cheating detection method and system based on graph matching network
CN113127933A