Door address deduplication method, device, equipment and storage medium
By constructing the gate address map data structure and using the graph neural network model, combining the spatial position relationship and attribute characteristics of the gate address, the problem of low gate address deduplication efficiency in the existing technology is solved, and a more efficient deduplication effect is achieved.
Patent Information
- Application Number
- CN202210452234.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-27
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-04-27
AI Technical Summary
The prior art is not efficient when deduplication is removed from the door address, especially when the text descriptions of the two door addresses are large, the deduplication effect is not good.
Using a graph neural network method, by constructing a gate address map data structure, using the spatial position relationship and attribute characteristics between gate addresses, the gate address deduplication model is trained, thereby improving the deduplication efficiency.
Through the training and application of graph neural network models, duplicate gate address data can be more effectively identified and removed, improving the accuracy and efficiency of deduplication.
Smart Images

Figure CN114780660B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electronic maps, and particularly to a method, apparatus, computer device, and storage medium for address deduplication. Background Art
[0002] An address is a type of map data, usually including information such as street name, house number, and longitude and latitude. Users can input an address, and the map search engine can query the corresponding longitude and latitude coordinates based on the input address and mark them on the electronic map.
[0003] Due to a large amount of duplicate and redundant data in the addresses on the Internet, when users use map services, it will seriously affect their experience. Therefore, how to remove duplicate data on the basis of ensuring comprehensive and rich data has become an urgent problem to be solved.
[0004] Address deduplication is to find duplicate and redundant data. Although two addresses may have different text descriptions, they may correspond to the same address data in the real world. In the past, when performing address deduplication, it was often a feature-based duplicate detection method based on a similarity function. For example, using text similarity, the edit distance of the address pair names was calculated to determine whether they were the same address. However, this deduplication method is not efficient. For example, if two addresses are duplicate data, but the text descriptions of the two addresses are quite different, in this case, the deduplication effect of the above method is not good. Summary of the Invention
[0005] In view of the above deficiencies or drawbacks, this application provides a method, apparatus, computer device, and storage medium for address deduplication. Embodiments of this application can improve the efficiency of address deduplication.
[0006] According to a first aspect of this application, a method for address deduplication is provided. In one embodiment, the method includes:
[0007] Obtain each address within the target geographical area, and construct an address graph data structure based on the geographical location according to the obtained addresses;
[0008] Select all address duplicate pairs from the above-obtained addresses, mark all address duplicate pairs, and obtain multiple sample address duplicate pairs for training;
[0009] Use the address graph data structure and the above multiple sample address duplicate pairs to train an address deduplication model;
[0010] Use the trained address deduplication model to process the address graph data structure to obtain a processing result;
[0011] Determine each address duplicate pair in the above-obtained addresses according to the processing result, and delete any one address in each address duplicate pair.
[0012] In one embodiment, the training of the address deduplication model is repeated using the address graph data structure and the above-mentioned multiple sample address duplicates, including:
[0013] Construct an address deduplication model, which is a graph neural network model;
[0014] Iteratively train the address deduplication model repeatedly using the address graph data structure and the above-mentioned multiple sample address duplicates, and when the preset end training condition is satisfied, obtain the trained address deduplication model;
[0015] Among them, each training process of the address deduplication model includes:
[0016] Obtain the adjacency matrix and node attribute feature matrix of the address graph data structure, and input them into the address deduplication model to obtain the output data of the address deduplication model. The output data includes the embeddings of each address in the above-obtained addresses;
[0017] Obtain the target sample address duplicate pair for this training from the above-mentioned multiple sample address duplicate pairs, use the target sample address duplicate pair as the positive duplicate pair, and construct the corresponding negative duplicate pair according to the target sample address duplicate pair;
[0018] Obtain the embeddings of the positive duplicate pair and the negative duplicate pair from the output data, and calculate the loss of this training according to the embeddings of the positive duplicate pair and the negative duplicate pair;
[0019] Judge whether the preset end training condition is satisfied according to the loss;
[0020] When it is determined that the condition is satisfied, end the training, and use the address deduplication model trained this time as the trained address deduplication model;
[0021] When it is determined that the condition is not satisfied, update the network parameters of the address deduplication model according to the loss, and perform the next training on the address deduplication model with the updated network parameters.
[0022] In one embodiment, use the trained address deduplication model to process the address graph data structure to obtain a processing result, and determine each address duplicate pair in the above-obtained addresses according to the processing result, including:
[0023] Use the trained address deduplication model to process the address graph data structure to obtain the embeddings of each address in the above-obtained addresses;
[0024] Pair the above-obtained addresses in pairs to obtain multiple address pairs;
[0025] Determine the embedding distance of each address pair, and determine the address pair with an embedding distance less than the preset threshold as the address duplicate pair, where the embedding distance of each address pair refers to the distance between the embeddings of the two addresses included in each address pair.
[0026] In one embodiment, a location-based address map data structure is constructed according to the obtained addresses, including:
[0027] Pair the obtained addresses in pairs to obtain a plurality of address pairs;
[0028] Calculate the distance of each address pair according to the location information of each address pair. The distance of the address pair refers to the distance between the two addresses included in the address pair;
[0029] Determine the weight of the edge between the two addresses included in each address pair according to the distance of each address pair, and obtain the address map data structure based on the location information.
[0030] In one embodiment, determining the weight of the edge between the two addresses included in any address pair according to the distance of the address pair includes:
[0031] When the distance between the two addresses included in any address pair is less than a preset threshold, it is determined that an edge relationship is formed between the two addresses, and the weight of the edge between the two addresses is set to 1;
[0032] When the distance between the two addresses included in any address pair is greater than or equal to the preset threshold, it is determined that no edge relationship is formed between the two addresses, and the weight of the edge between the two addresses is set to 0.
[0033] In one embodiment, the number of address map data structures is multiple; constructing a location-based address map data structure according to the obtained addresses includes:
[0034] Divide the target geographical area into multiple spatial grids;
[0035] Traverse the longitude and latitude attributes of each address in the obtained addresses, and determine the address set of each spatial grid;
[0036] Construct a location-based address map data structure for each spatial grid according to the address set of each spatial grid.
[0037] In one embodiment, use the trained address deduplication model to process the address map data structure to obtain a processing result, and determine each address duplicate pair in the obtained addresses according to the processing result, including:
[0038] Input the address map data structure of each spatial grid into the trained address deduplication model respectively to obtain the embedding of each address in the address set of each spatial grid;
[0039] Pair the addresses in the address set of each spatial grid in pairs to obtain the address pair set of each spatial grid;
[0040] Calculate the embedding distance of each address pair in the set of address pairs for each spatial grid; the embedding distance of each address pair refers to the distance between the embeddings of the two addresses included in each address pair.
[0041] Determine, for each spatial grid, as address duplicate pairs those address pairs in the set of address pairs for which the embedding distance is less than a preset threshold.
[0042] According to a second aspect, the present application provides an address deduplication device. In one embodiment, the device includes:
[0043] A construction module, configured to obtain each address within a target geographical area and construct a location-based address map data structure based on the obtained addresses.
[0044] A sample construction module, configured to select all address duplicate pairs from the obtained addresses, label all address duplicate pairs, and obtain multiple sample address duplicate pairs for training.
[0045] A training module, configured to train an address deduplication model using the address map data structure and the multiple sample address duplicate pairs.
[0046] A processing module, configured to process the address map data structure using the trained address deduplication model to obtain a processing result.
[0047] A deletion module, configured to determine each address duplicate pair among the obtained addresses according to the processing result and delete any one address from each address duplicate pair.
[0048] According to a third aspect, the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the embodiments of any of the above methods are implemented.
[0049] According to a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the embodiments of any of the above methods are implemented.
[0050] In the above embodiments of the present application, by obtaining each address within the target geographical area and constructing a location-based address map data structure based on the obtained addresses; then selecting all address duplicate pairs from the obtained addresses, marking the all address duplicate pairs to obtain a plurality of sample address duplicate pairs for training; further using the address map data structure and the plurality of sample address duplicate pairs to train an address deduplication model; then using the trained address deduplication model to process the address map data structure to obtain a processing result; finally, determining each address duplicate pair in the obtained addresses according to the processing result and deleting any one address in each address duplicate pair, the efficiency of address deduplication can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is a flowchart of a method for address deduplication in one or more embodiments;
[0052] Figure 2 is a training flowchart of an address deduplication model in one or more embodiments;
[0053] Figure 3 is a prediction flowchart of an address deduplication model in one or more embodiments;
[0054] Figure 4 is a structural block diagram of an address deduplication device in one or more embodiments;
[0055] Figure 5 is an internal structure diagram of a computer device in one or more embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0057] The present application provides a method for address deduplication. In one embodiment, the method includes steps as Figure 1 shown below, and the method will be described below.
[0058] S110: Obtain each address within the target geographical area and construct a location-based address map data structure (which can also be simply referred to as an address map) based on the obtained addresses.
[0059] S120: Select all address duplicate pairs from the obtained addresses, mark the all address duplicate pairs to obtain a plurality of sample address duplicate pairs for training.
[0060] S130: Use the address map data structure and the plurality of sample address duplicate pairs to train an address deduplication model.
[0061] S140: Process the address map data structure using the trained address deduplication model to obtain a processing result.
[0062] S150: Determine each address duplicate pair among the obtained addresses according to the processing result, and delete any one address in each address duplicate pair.
[0063] Exemplarily, the address duplicate data can be as shown in Table 1 below:
[0064] Table 1:
[0065] Serial number Door address text data 1 No. 188, Middle Huaihai Road, Huangpu District, Shanghai 2 188 Middle Huaihai Road, Shanghai 3 No. 188, Middle Huaihai Road, Huangpu District, Shanghai 4 Lane 188, Middle Huaihai Road, Huangpu District, Shanghai 5 188 Huangpu Huaihai Road, Shanghai 6 No. 188, Middle Huaihai Road, Shanghai 7 No. 188, Middle Huaihai Road, Huangpu District 8 No. 188, Middle Huaihai Road, Shanghai
[0066] After applying the address deduplication method provided in this embodiment, only one address out of the 8 addresses included in the above address duplicate data will be finally retained, and the other addresses will be deleted.
[0067] In this embodiment, by obtaining each address within the target geographical area, constructing an address map data structure based on geographical location according to the obtained addresses; then selecting all address duplicate pairs from the obtained addresses, marking the all address duplicate pairs to obtain multiple sample address duplicate pairs for training; further training an address deduplication model using the address map data structure and the multiple sample address duplicate pairs; then processing the address map data structure using the trained address deduplication model to obtain a processing result; and finally determining each address duplicate pair among the obtained addresses according to the processing result, and deleting any one address in each address duplicate pair, the following technical effects can be achieved:
[0068] Traditional machine learning and pre-trained deep models make the assumption that two addresses are independent in the address deduplication problem. In the embodiments of this application, it is assumed that there is an association between two addresses. On this basis, the address deduplication is carried out by combining the attribute characteristics of the addresses themselves and the spatial position relationship between the addresses.
[0069] The inventors of this application noticed that there are usually similar neighbors around duplicate address pairs. Based on this, a graph convolutional neural network model based on a graph structure can be selected for address deduplication, using the text information of the address names and the spatial position relationship information between the addresses as model inputs. Specifically, the graph neural network has the ability to identify isomorphic graphs by passing and aggregating neighbor node information, and can better identify similar neighbor structures. If two addresses have similar neighbors, through the convolution aggregation operation, the embedding representations (embedding, which can be simply referred to as embedding) of these two addresses will also be very close, and the embedding distance between the two addresses will be very close. By calculating the embedding distance between two addresses to determine whether the two addresses are duplicates. Therefore, address deduplication through the graph neural network can more effectively solve the problem of address duplication.
[0070] On the other hand, traditional machine learning and pre-trained deep models generally can only perform model training based on sample address pairs, and sample address pairs require manual annotation, with a relatively high acquisition cost, so the quantity is generally small, which will lead to insufficient training. Through the graph convolutional neural network model in the embodiments of the present application, all data can participate in the training, so that the model can be trained more fully and the duplicate detection effect of the model is better.
[0071] In summary, the embodiments of the present application can fundamentally improve the effect of address duplicate removal, making it efficient, reasonable, and widely applicable.
[0072] In one embodiment, constructing the address graph data structure based on geographical location according to the obtained addresses includes: pairing the obtained addresses in pairs to obtain a plurality of address pairs; calculating the distance of each address pair according to the geographical location information of each address pair, where the distance of the address pair refers to the distance between the two addresses included in the address pair; determining the weight of the edge between the two addresses included in each address pair according to the distance of each address pair to obtain the address graph data structure based on geographical location information. The distance between two addresses refers to the spatial distance, such as 100 meters, 50 meters, etc.
[0073] Among them, determining the weight of the edge between the two addresses included in any address pair according to the distance of any address pair includes: when the distance between the two addresses included in any address pair is less than the preset threshold, determining that an edge relationship is formed between the two addresses, and setting the weight of the edge between the two addresses to 1; when the distance between the two addresses included in any address pair is greater than or equal to the preset threshold, determining that no edge relationship is formed between the two addresses, and setting the weight of the edge between the two addresses to 0. Among them, the preset threshold can be set according to actual needs, such as set to 50 meters, etc., and this embodiment does not make specific limitations on this.
[0074] In one embodiment, using the address graph data structure and the above-mentioned multiple sample addresses to repeatedly train the address duplicate removal model includes: constructing an address duplicate removal model, where the address duplicate removal model is a graph neural network model; using the address graph data structure and the above-mentioned multiple sample addresses to repeatedly perform iterative training on the address duplicate removal model, and obtaining a trained address duplicate removal model when the preset end training condition is met.
[0075] Among them, each training process of the address deduplication model includes: obtaining the adjacency matrix and the node attribute feature matrix of the address graph data structure, and inputting them into the address deduplication model to obtain the output data of the address deduplication model. The output data includes the embeddings of each address in the obtained addresses; obtaining the target sample address duplicate pairs for this training from the above-mentioned multiple sample address duplicate pairs, taking the target sample address duplicate pairs as positive duplicate pairs, and constructing corresponding negative duplicate pairs according to the target sample address duplicate pairs; obtaining the embeddings of the positive duplicate pairs and the embeddings of the negative duplicate pairs from the output data, and calculating the loss of this training according to the embeddings of the positive duplicate pairs and the embeddings of the negative duplicate pairs; judging whether the stop preset end training condition is satisfied according to the loss; when it is determined that the condition is satisfied, end the training, and take the address deduplication model trained this time as the trained address deduplication model; when it is determined that the condition is not satisfied, update the network parameters of the address deduplication model according to the loss, and perform the next training on the address deduplication model with the updated network parameters.
[0076] In one embodiment, using the trained address deduplication model to process the address graph data structure to obtain a processing result, and determining each address duplicate pair in the obtained addresses according to the processing result, includes: using the trained address deduplication model to process the address graph data structure to obtain the embeddings of each address in the obtained addresses; pairing the obtained addresses in pairs to obtain a plurality of address pairs; determining the embedding distance of each address pair, and determining the address pairs with the embedding distance less than the preset threshold as address duplicate pairs, where the embedding distance of each address pair refers to the distance between the embeddings of the two addresses included in each address pair.
[0077] In another embodiment, the number of the above-mentioned address graphs is multiple. Correspondingly, the construction of the address graph data structure based on the geographical location according to the obtained addresses includes: dividing the target geographical area into a plurality of spatial grids; traversing the longitude and latitude attributes of each address in the obtained addresses to determine the address set of each spatial grid; constructing the address graph data structure based on the geographical location for each spatial grid according to the address set of each spatial grid.
[0078] In this embodiment, it is considered that in some scenarios, when the range of the target geographical area is large, for example, the target geographical area refers to the whole of China, then the data that needs to be deduplicated (referring to address deduplication) at this time is the full amount of data in the whole of China. At this time, the data volume is very large. If all the data is directly used to construct a graph (referring to the address graph), the nodes and adjacency matrix of this address graph will be very large, and this will require very high computing resources. Therefore, in the case of limited computing resources, a graph is divided into many small graphs according to spatial grids, and then deduplication is performed on each small graph. In this way, the nodes and adjacency matrix of each small graph become smaller, and data deduplication can be performed with less computing resources. The specific method can be as follows:
[0079] According to the geographical spatial coordinates of China, from the westernmost end to the easternmost end, and from the northernmost end to the southernmost end, it is divided into square spatial grids of 10 kilometers * 10 kilometers (the grid size can be flexibly adjusted according to actual needs). Each of the four vertices of the grid has corresponding longitude and latitude coordinates, and each door address has longitude and latitude attributes. By traversing all the door addresses according to the longitude and latitude, the door addresses can fall into the corresponding grids. In this way, each grid will have door addresses that are relatively close. Subsequently, only the duplication of door addresses needs to be removed in each spatial grid.
[0080] Correspondingly, for the above-mentioned processing of the door address map data structure using the trained door address duplication removal model to obtain a processing result, and determining each door address duplication pair in the obtained door addresses according to the processing result, it includes: respectively inputting the door address map data structure of each spatial grid into the trained door address duplication removal model to obtain the embedding of each door address in the door address set of each spatial grid; pairing the door addresses in the door address set of each spatial grid pairwise to obtain the door address pair set of each spatial grid; calculating the embedding distance of each door address pair in the door address pair set of each spatial grid; the embedding distance of each door address pair refers to the distance between the embeddings of the two door addresses included in each door address pair; determining each door address pair with an embedding distance less than a preset threshold in the door address pair set of each spatial grid as a door address duplication pair.
[0081] The above embodiments will be described below through a specific application example.
[0082] This application example is specifically divided into seven parts: division of spatial grids, acquisition of seed door address duplication pairs, data preprocessing, feature engineering, model training, model prediction of duplicate data, and duplication removal processing.
[0083] This application example takes China as the target geographical area. First, according to the geographical spatial coordinates of China, the geographical space of China is divided into square grids of 10 kilometers * 10 kilometers, and then the door addresses at the corresponding spatial positions are dropped into the corresponding grids to form sub-graphs. This application example uses supervised learning to train the graph neural network model. Since it is supervised learning, it is necessary to manually label sample door address duplicate pairs for model training. After obtaining the sample door address duplicate pairs, it is necessary to preprocess the data, including operations such as converting full-width characters to half-width characters, removing special symbols, converting English uppercase to lowercase, and converting traditional Chinese to simplified Chinese, to clean the data. Before training the model, it is necessary to perform feature engineering to obtain the attribute features of the nodes and the adjacency matrix, input the attribute features of the nodes and the adjacency matrix into the graph convolutional neural network GCN (Graph Convolutional Network), and train the model by minimizing the loss function through backpropagation to obtain the weight matrix W, which is the parameter that the model needs to learn. After obtaining the parameter W, it is possible to predict the door address duplicate pairs at the graph level through forward propagation. Finally, in the full amount of data, one of the door address data in each door address duplicate pair is deleted to obtain the deduplicated door address data.
[0084] The following will explain each of the above parts.
[0085] 1. Division of spatial grids
[0086] Since the data to be deduplicated is the full amount of data across the whole of China and the total amount of data is very large, if all the data is directly used to construct a large graph, both the nodes and the adjacency matrix will be extremely large, requiring too much computing resources. Therefore, in the case of limited computing resources, the large graph is divided into many small graphs according to spatial grids, and data deduplication is carried out on each small graph. In this way, both the nodes and the adjacency matrix become smaller, and data deduplication can be carried out with relatively limited computing resources.
[0087] Specific method: According to the geographical spatial coordinates of China, from the westernmost end to the easternmost end, and from the northernmost end to the southernmost end, it is divided into square spatial grids of 10 kilometers * 10 kilometers. The four vertices of the grid all have corresponding longitude and latitude coordinates, and the door addresses all have longitude and latitude attributes. By traversing all the door addresses according to the longitude and latitude, the door addresses can fall into the corresponding grids. In this way, each grid will have the door addresses that are relatively close, and subsequent deduplication of the door addresses only needs to be carried out in each spatial grid.
[0088] 2. Obtaining sample door address duplicate pairs
[0089] This application example requires sample address pairs to be repeated to train the model, so it is necessary to label some data as sample address pairs. All the above spatial grids can be traversed, and using simple text similarity of address names (such as edit distance, etc.), suspected duplicate address pairs can be roughly found in each spatial grid, and then the labeling personnel are asked to find the truly duplicate address pairs. In this way, sample address pairs are constructed.
[0090] 3. Data Preprocessing
[0091] Special symbols and traditional Chinese characters may be included in the text of the address name. It is necessary to perform preprocessing first and then construct features for input into the model. At the same time, to ensure the consistency of the distribution of labeled data and unlabeled data, the same preprocessing operations need to be performed on the labeled data and unlabeled data. The data preprocessing process includes the following four steps:
[0092] (1) Full-width to half-width conversion of characters
[0093] (2) Special symbols
[0094] (3) Capital English to lowercase
[0095] (4) Traditional Chinese to simplified Chinese
[0096] 4. Feature Engineering
[0097] (1) Generate attribute features of graph nodes
[0098] The input of the graph convolutional neural network GCN includes the topological structure of the graph, that is, the adjacency matrix, and the attribute features of all nodes in the graph. Each node attribute feature is a multi-dimensional feature vector. In this application example, a specified algorithm is used to process each address into a 512-dimensional Embedding vector, and this Embedding vector is used as the attribute feature of the node. Among them, the specified algorithm can be any existing algorithm that can map an address to an Embedding vector, so it will not be elaborated here.
[0099] (2) Generate the edges and adjacency matrix of all subgraph structures
[0100] In this application example, addresses within 50 meters are considered to have edge relationships. Specifically, all addresses in a spatial grid are taken out to form a set. An address is taken out from the set, and the distance is calculated with all other addresses in the set except itself. Among them, addresses with a distance less than 50 meters form an edge relationship with the taken-out address, and the weight of the edge is 1; then addresses are taken out from the set in turn, and the above operations are also performed until all addresses in the set are taken out, and the edge relationships of all addresses in this spatial grid subgraph are formed. According to the definition of the graph structure, the adjacency matrix of the subgraph is obtained. The representation of the adjacency matrix is:
[0101]
[0102] Finally, traverse all the spatial grids according to the above method, and the adjacency matrix of all the subgraphs of the spatial grids is generated.
[0103] 5. Model Training
[0104] As Figure 2 shown in the model training flowchart of this application example.
[0105] Specifically, select a subgraph KG of the spatial grid and the pre-annotated duplicate gate address sample pairs S = {(e i1 , e i2 )} m i=1 .
[0106] The method of this application example finds new duplicate gate address pairs within the subgraph based on the node embedding of GCN. The basic idea is to use GCN (which can also be called the GCN model) to embed the gate addresses into a unified vector space, hoping that the distance between duplicate gate addresses is closer and the distance between non-duplicate gate addresses is farther.
[0107] (1) Input of GCN:
[0108] GCN is a type of neural network that operates directly on graphs. Its input is the node attribute features and the adjacency matrix of the graph, and the purpose is to output the embedding of each gate address, which is then used for subsequent gate address deduplication. For the node attribute features and the adjacency matrix input to the model, both are from the feature engineering in step 4. After inputting the node attribute features and the adjacency matrix into the GCN model, the subsequent GCN operations are performed.
[0109] (2) Operations of GCN:
[0110] A GCN model contains multiple GCN layers. In this application example, two layers are selected. The input H (l) ∈R n×d(l) of the l-th layer is a node attribute feature matrix (i.e., all node attribute features), where n is the number of nodes, and d (l) is the number of features in the l-th layer. The output of the l-th layer is a new feature matrix:
[0111]
[0112] where σ is the relu activation function (used for linear activation transformation), A is the n*n adjacency matrix, where I is the identity matrix. is the diagonal node degree matrix of, and W (l) ∈R d(l)×d(l+1) is the weight matrix between two layers, used for convolution operations, and d(l+1) It is a new layer dimension.
[0113] (3) Output of GCN
[0114] After two layers of GCN operations, the model outputs a 512-dimensional node embedding representation, which can be used for subsequent operations.
[0115] (4) Loss function of GCN
[0116] This application example hopes that the distance between repeated addresses is small, and the distance between non-repeated addresses is large, and constructs a loss function based on this. The distance between addresses is the embedding distance between addresses. For the outputs e 1 and e 2 of two addresses, the distance calculation method between them is as follows:
[0117] D(e 1 , e 2 ) = ||h(e 1 ) - h(e 2 )|| 1
[0118] The model is trained by minimizing the margin-based loss function:
[0119]
[0120] where [x] + = max{0, x}, S' (e1,e2) is a negative repeated pair obtained by randomly replacing one address from (e 1 , e 2 ), and γ is the margin that distinguishes positive repeated pairs and negative repeated pairs. The model is trained by minimizing the loss function through backpropagation, updating the weight matrix W in each layer. After several rounds of training, the final model can learn the weight matrix W to predict address repeated pairs.
[0121] 6. Model repeated data prediction
[0122] As Figure 3 shown is the model prediction flowchart of this application example.
[0123] This application example is suitable for offline graph-level address repeated pair prediction. The prediction is to find more new address repeated pairs in the constructed graph. During the training process, the weight matrix W is learned. By inputting node attribute features and the adjacency matrix, after the operations of GCN, each node will output an embedding representation.
[0124] For the embedding of a specific output, calculate its embedding distance from all other gate addresses in the subgraph, and select the one with the smallest embedding distance among all gate addresses. If this embedding distance is less than a certain threshold, these two gate addresses are considered duplicates; if it is not less than this threshold, they are considered non-duplicates. According to the above method, traverse all the gate addresses in the subgraph that are not in the sample pairs, and the duplicate gate addresses in the subgraph can be obtained. Then, traverse all subgraphs using the above method to obtain the duplicate gate address pairs in the full data.
[0125] 7. Deduplication processing
[0126] Extract all the duplicate gate address pairs obtained in step 6 from the full data, traverse these duplicate gate address pairs, delete one of the gate addresses, and add the other gate address to the full data. Through this scheme, the deduplicated gate address data of the full data can be obtained.
[0127] This application example is statistically based on other gate address deduplication methods mentioned above, and the specific benefits are as follows:
[0128] The accuracy rate of the gate address deduplication method based on the unsupervised similarity measurement scheme is 91.3%, and the recall rate is 83.6%;
[0129] The accuracy rate of the gate address deduplication method based on the traditional machine learning model is 94.8%, and the recall rate is 89.9%;
[0130] The accuracy rate of the gate address deduplication method based on the deep learning model is 96.2%, and the recall rate is 91.8%;
[0131] The accuracy rate of the gate address deduplication method in this application example is 97.9%, and the recall rate is 94.6%.
[0132] It can be intuitively seen from the above data that compared with the previous methods, the gate address deduplication method provided in this application example has a significant improvement in effect.
[0133] Figure 1 It is a flowchart showing the process of the gate address deduplication method in an embodiment. It should be understood that although Figure 1 the steps in the flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figure 1 at least a part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0134] Based on the same inventive concept, this application also provides an address deduplication device. In this embodiment, as Figure 4 shown, the address deduplication device includes the following modules:
[0135] A construction module 110, configured to obtain each address within the target geographical area and construct a geographical location-based address graph data structure according to the obtained addresses;
[0136] A sample construction module 120, configured to select all address duplicate pairs from the obtained addresses, label all address duplicate pairs, and obtain multiple sample address duplicate pairs for training;
[0137] A training module 130, configured to train an address deduplication model using the address graph data structure and the multiple sample address duplicate pairs;
[0138] A processing module 140, configured to process the address graph data structure using the trained address deduplication model to obtain a processing result;
[0139] A deletion module 150, configured to determine each address duplicate pair in the obtained addresses according to the processing result, and delete any one address in each address duplicate pair.
[0140] In one embodiment, the training module 130 is configured to:
[0141] Construct an address deduplication model, where the address deduplication model is a graph neural network model;
[0142] Iteratively train the address deduplication model using the address graph data structure and the multiple sample address duplicate pairs, and obtain a trained address deduplication model when a preset end training condition is satisfied;
[0143] Wherein, each training process of the address deduplication model includes:
[0144] Obtain the adjacency matrix and the node attribute feature matrix of the address graph data structure, input them into the address deduplication model, and obtain the output data of the address deduplication model. The output data includes the embedding of each address in the obtained addresses;
[0145] Obtain the target sample address duplicate pair for this training from the multiple sample address duplicate pairs, use the target sample address duplicate pair as a positive duplicate pair, and construct a corresponding negative duplicate pair according to the target sample address duplicate pair;
[0146] Obtain the embedding of the positive duplicate pair and the embedding of the negative duplicate pair from the output data, and calculate the loss of this training according to the embedding of the positive duplicate pair and the embedding of the negative duplicate pair;
[0147] Judge whether the preset end training condition is satisfied according to the loss;
[0148] When the determination is satisfied, end the training, and use the trained address deduplication model obtained in this training as the trained address deduplication model;
[0149] When the determination is not satisfied, update the network parameters of the address deduplication model according to the loss, and perform the next training on the address deduplication model with the updated network parameters.
[0150] In one embodiment, the processing module 140 and the deletion module 150 are used for:
[0151] Process the address map data structure using the trained address deduplication model to obtain the embedding of each address in the addresses obtained above;
[0152] Pair the addresses obtained above in pairs to obtain a plurality of address pairs;
[0153] Determine the embedding distance of each address pair, and determine the address pairs with an embedding distance less than a preset threshold as duplicate address pairs, where the embedding distance of each address pair refers to the distance between the embeddings of the two addresses included in each address pair.
[0154] In one embodiment, the construction module 110 is used for:
[0155] Pair the addresses obtained above in pairs to obtain a plurality of address pairs;
[0156] Calculate the distance of each address pair according to the geographical location information of each address pair, where the distance of the address pair refers to the distance between the two addresses included in the address pair;
[0157] Determine the weight of the edge between the two addresses included in each address pair according to the distance of each address pair to obtain the address map data structure based on geographical location information.
[0158] In one embodiment, when the construction module 110 determines the weight of the edge between the two addresses included in any address pair according to the distance of the address pair, it is specifically used for:
[0159] When the distance between the two addresses included in any address pair is less than the preset threshold, determine that an edge relationship is formed between the two addresses, and set the weight of the edge between the two addresses to 1;
[0160] When the distance between the two addresses included in any address pair is greater than or equal to the preset threshold, determine that no edge relationship is formed between the two addresses, and set the weight of the edge between the two addresses to 0.
[0161] In one embodiment, the number of address map data structures is multiple; the construction module 110 is further used for:
[0162] Divide the target geographical area into multiple spatial grids;
[0163] Traverse the longitude and latitude attributes of each door address obtained above to determine the set of door addresses for each spatial grid;
[0164] Construct a door address map data structure based on geographical location for each spatial grid according to the set of door addresses for each spatial grid.
[0165] Correspondingly, in one embodiment, the processing module 140 and the deletion module 150 are further configured to:
[0166] Input the door address map data structure of each spatial grid into the trained door address deduplication model respectively to obtain the embedding of each door address in the set of door addresses for each spatial grid;
[0167] Pair up the door addresses in the set of door addresses for each spatial grid to obtain a set of door address pairs for each spatial grid;
[0168] Calculate the embedding distance of each door address pair in the set of door address pairs for each spatial grid; the embedding distance of each door address pair refers to the distance between the embeddings of the two door addresses included in each door address pair;
[0169] Determine the door address duplicate pairs for each spatial grid from the door address pairs in which each embedding distance is less than a preset threshold.
[0170] For the specific definition of the door address deduplication device, reference can be made to the definition of the door address deduplication method in the above text, which will not be elaborated here. Each module in the above door address deduplication device can be implemented in whole or in part by software, hardware and their combination. The above modules can be embedded in the processor in the computer device in the form of hardware or independent of it, or stored in the memory in the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the above modules.
[0171] In one embodiment, a computer device is provided, and its internal structure diagram can be as Figure 5 shown. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements a door address deduplication method.
[0172] Those skilled in the art can understand, Figure 5The structure shown is only a block diagram of some of the structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component arrangement.
[0173] This embodiment also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps in the method provided in any of the above method embodiments are implemented.
[0174] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the method provided in any of the above method embodiments are implemented.
[0175] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the above method embodiments. Among them, any reference to the memory, storage, database, or other media used in the various embodiments provided in this application may include non-volatile and / or volatile memories. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0176] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0177] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A method for deduplicating door addresses, characterized in that, the method includes: Obtain each door address within the target geographical area, and construct a door address graph data structure based on geographical location according to the obtained door addresses; Select all door address duplicate pairs from the obtained door addresses, label the all door address duplicate pairs, and obtain multiple sample door address duplicate pairs for training; Use the door address graph data structure and the multiple sample door address duplicate pairs to train a door address deduplication model; Use the trained door address deduplication model to process the door address graph data structure to obtain a processing result; Determine each door address duplicate pair in the obtained door addresses according to the processing result, and delete any one door address in each door address duplicate pair; The using the door address graph data structure and the multiple sample door address duplicate pairs to train a door address deduplication model includes: Construct a door address deduplication model, and the door address deduplication model is a graph neural network model; Use the door address graph data structure and the multiple sample door address duplicate pairs to iteratively train the door address deduplication model, and obtain a trained door address deduplication model when the preset end training condition is satisfied; Wherein, each training process of the door address deduplication model includes: Obtain the adjacency matrix and node attribute feature matrix of the door address graph data structure, and input them into the door address deduplication model to obtain the output data of the door address deduplication model. The output data includes the embedding of each door address in the obtained door addresses; Obtain the target sample door address duplicate pair for this training from the multiple sample door address duplicate pairs, use the target sample door address duplicate pair as a positive duplicate pair, and construct a corresponding negative duplicate pair according to the target sample door address duplicate pair; Obtain the embedding of the positive duplicate pair and the embedding of the negative duplicate pair from the output data, and calculate the loss of this training according to the embedding of the positive duplicate pair and the embedding of the negative duplicate pair; Judge whether the preset end training condition is satisfied according to the loss; When it is determined that the condition is satisfied, end the training, and use the door address deduplication model trained this time as the trained door address deduplication model; When it is determined that the condition is not satisfied, update the network parameters of the door address deduplication model according to the loss, and perform the next training on the door address deduplication model with updated network parameters.
2. The method according to claim 1, characterized in that, the using the trained door address deduplication model to process the door address graph data structure to obtain a processing result, and determining each door address duplicate pair in the obtained door addresses according to the processing result includes: Use the trained door address deduplication model to process the door address graph data structure to obtain the embedding of each door address in the obtained door addresses; Pair the obtained door addresses in pairs to obtain multiple door address pairs; Determine the embedding distance of each door address pair, and determine the door address pair with an embedding distance less than a preset threshold as a door address duplicate pair, where the embedding distance of each door address pair refers to the distance between the embeddings of the two door addresses included in each door address pair.
3. The method according to claim 1, characterized in that, the constructing a door address graph data structure based on geographical location according to the obtained door addresses includes: Pair the obtained door addresses in pairs to obtain multiple door address pairs; Calculate the distance between each pair of address locations based on the geographical location information of each pair of address locations, where the distance between the pair of address locations refers to the distance between the two address locations included in the pair of address locations; Determine the weight of the edge between the two address locations included in each pair of address locations according to the distance between each pair of address locations, and obtain the address graph data structure based on the geographical location information.
4. The method according to claim 3, characterized in that, Determining the weight of the edge between the two address locations included in any pair of address locations according to the distance of any pair of address locations includes: When the distance between the two address locations included in any pair of address locations is less than the preset threshold, it is determined that an edge relationship is formed between the two address locations, and the weight of the edge between the two address locations is set to 1; When the distance between the two address locations included in any pair of address locations is greater than or equal to the preset threshold, it is determined that no edge relationship is formed between the two address locations, and the weight of the edge between the two address locations is set to 0.
5. The method according to claim 1, characterized in that, The number of the address graph data structures is multiple; The constructing the address graph data structure based on the geographical location according to the obtained address locations includes: Dividing the target geographical area into multiple spatial grids; Traversing the longitude and latitude attributes of each address location in the obtained address locations, and determining the address set of each spatial grid; Constructing an address graph data structure based on the geographical location for each spatial grid according to the address set of each spatial grid.
6. The method according to claim 5, characterized in that, Using the trained address deduplication model to process the address graph data structure to obtain a processing result, and determining each address duplicate pair in the obtained address locations according to the processing result includes: Inputting the address graph data structure of each spatial grid into the trained address deduplication model respectively to obtain the embedding of each address in the address set of each spatial grid; Pairing the addresses in the address set of each spatial grid in pairs to obtain the address pair set of each spatial grid; Calculating the embedding distance of each address pair in the address pair set of each spatial grid; the embedding distance of each address pair refers to the distance between the embeddings of the two addresses included in each address pair; Determining each address pair with an embedding distance less than the preset threshold in the address pair set of each spatial grid as an address duplicate pair.
7. An address deduplication device, characterized in that, The device includes: A construction module, configured to obtain each address location within the target geographical area, and construct an address graph data structure based on the geographical location according to the obtained address locations; A sample construction module, configured to select all address duplicate pairs from the obtained address locations, label the all address duplicate pairs, and obtain multiple sample address duplicate pairs for training; A training module, configured to train an address deduplication model using the address graph data structure and the multiple sample address duplicate pairs; A processing module, configured to use the trained address deduplication model to process the address graph data structure to obtain a processing result; A deletion module, configured to determine each address duplicate pair in the obtained address locations according to the processing result, and delete any one address in each address duplicate pair; Repeatedly training the training address deduplication model using the address graph data structure and the multiple sample address duplicates includes: Constructing an address deduplication model, where the address deduplication model is a graph neural network model; Repeatedly using the address graph data structure and the multiple sample address duplicates to iteratively train the address deduplication model, and when a preset end training condition is satisfied, obtaining a trained address deduplication model; Among them, each training process of the address deduplication model includes: Obtaining the adjacency matrix and the node attribute feature matrix of the address graph data structure, and inputting them into the address deduplication model to obtain the output data of the address deduplication model, where the output data includes the embeddings of each address in the obtained addresses; Obtaining the target sample address duplicate pair for this training from the multiple sample address duplicates, using the target sample address duplicate pair as a positive duplicate pair, and constructing a corresponding negative duplicate pair according to the target sample address duplicate pair; Obtaining the embeddings of the positive duplicate pair and the embeddings of the negative duplicate pair from the output data, and calculating the loss of this training according to the embeddings of the positive duplicate pair and the embeddings of the negative duplicate pair; Judging whether the preset end training condition is satisfied according to the loss; When it is determined to be satisfied, end the training, and use the address deduplication model trained this time as the trained address deduplication model; When it is determined not to be satisfied, update the network parameters of the address deduplication model according to the loss, and perform the next training on the address deduplication model with updated network parameters.
8. A computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, on which a computer program is stored, wherein, when the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.