Address consistency identification method and device, related equipment and computer program product

The method improves address consistency identification by recognizing geographic entities and applying weight-based encoding with BERT and OPTICS clustering, addressing format and cultural variations while reducing manual library maintenance, enhancing accuracy and efficiency.

CN120316319APending Publication Date: 2025-07-15ANHUI IFLYHEALTH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510453232.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-15

Smart Images

  • Figure CN120316319A_ABST
    Figure CN120316319A_ABST
Patent Text Reader

Abstract

The invention discloses an address consistency identification method and device, related equipment and a computer program product. Geographic entities and administrative division levels to which the geographic entities belong are identified from address data. A corresponding hierarchy weight can be preset for each administrative division hierarchy, and the hierarchy weight indicates the importance of the geographic information of the corresponding administrative division hierarchy. And according to the hierarchical weight of the administrative division to which the geographic entities belong, carrying out weighted fusion on the coding features of the geographic entities in the address data to obtain the coding features of the geographic data, and enhancing the accuracy of feature representation. The similarity between every two address data is calculated according to the coding characteristics of the address data, the address data are clustered according to the similarity, and the address data in the same cluster have address consistency. By adopting the method provided by the invention, the accuracy of an address consistency identification result can be improved. In addition, the method does not depend on an address standardization library and a rule library, and the labor cost can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing technology. More specifically, it relates to an address consistency recognition method, device, related equipment, and computer program product. Background Art

[0002] In fields such as e-commerce, logistics distribution, and financial services, address information plays a crucial role. Identifying the consistency of a large number of address data not only helps improve the data management efficiency of enterprises but also greatly reduces operating costs. In actual application scenarios, address information in text often exists in various forms of expression, such as different formats, language styles, punctuation marks, and ellipsis forms, which makes it an important technical challenge to accurately identify different expressions of the same location.

[0003] Currently, the solutions for address consistency recognition mainly include two types:

[0004] 1. The method based on string matching and edit distance

[0005] This method mainly measures the similarity of two address texts by calculating the edit distance (such as Levenshtein distance) between address strings or using a string matching algorithm. This method has low accuracy when dealing with addresses with complex geographical hierarchical structures.

[0006] 2. The method based on a rule library and address standardization

[0007] This method uses a pre-constructed address standardization library and rule library to parse and standardize address texts. Through word segmentation technology and regular expressions, key elements in the address (such as province, city, district, street, house number) are identified, and then they are mapped to the canonical expressions in the standard address library. After that, the consistency of the standardized addresses is judged. This method requires maintaining and updating the address standardization library, which requires a large investment of human resources and cannot process addresses not included in the library. Summary of the Invention

[0008] In view of the above problems, this application is proposed to provide an address consistency recognition method, device, related equipment, and computer program product to improve the accuracy of address consistency recognition while reducing labor costs. The specific solutions are as follows:

[0009] In a first aspect, an address consistency recognition method is provided, including:

[0010] Obtain a set of address data to be processed, where the set includes several address data;

[0011] Identify the geographical entities in the address data and the administrative division levels to which the geographical entities belong;

[0012] Obtain the coding features of the geographical entities in the address data, and perform weighted fusion on the coding features of the geographical entities in the address data according to the hierarchical weights of the geographical entities, to obtain the coding features of the address data, where the hierarchical weights of the geographical entities correspond to the administrative division levels to which the geographical entities belong;

[0013] Based on the coding features of the address data, calculate the similarity between pairwise address data in the set, and cluster the address data in the set according to the similarity. Each address data belonging to the same cluster in the clustering result has address consistency.

[0014] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, the administrative division levels include at least a first level and a second level, and the first level is higher than the second level;

[0015] The hierarchical weight of the geographical entity belonging to the first level is less than the hierarchical weight of the geographical entity belonging to the second level.

[0016] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, before identifying the geographical entities in the address data, it further includes:

[0017] Preprocess the address data in the set, and the preprocessing includes removing the numbers therein;

[0018] Then, the clustering result of clustering the address data in the set according to the similarity is used as a coarse-grained address consistency recognition result, and the coarse-grained address consistency recognition result ignores the difference in house numbers.

[0019] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, it further includes:

[0020] For each cluster in the clustering result, restore the address data therein to address data including the house number in digital form;

[0021] Perform verification processing on each cluster to obtain each processed cluster as a refined address consistency recognition result, and the refined address consistency recognition result takes into account the difference in house numbers;

[0022] Among them, the process of verifying any clustering cluster includes: if the clustering cluster contains address data with more than two different house numbers, then the clustering cluster is split into several different clustering clusters, and the house numbers of the address data in each of the split clustering clusters are the same.

[0023] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, it further includes:

[0024] For each outlier in the clustering result, count the number of occurrences of the address data corresponding to the outlier in the set;

[0025] If the number of occurrences of the address data corresponding to the outlier in the set exceeds a set threshold, then redefine the address data corresponding to the outlier as an independent clustering cluster.

[0026] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, the process of weighted fusion of the coding features of the geographical entities in the address data according to the hierarchical weights of the geographical entities includes:

[0027] Weight the coding features of the geographical entities in the address data according to the hierarchical weights of the geographical entities to obtain the weighted coding features of the geographical entities;

[0028] Combine the weighted coding features of the geographical entities in the address data through vector concatenation or pooling operations.

[0029] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, before clustering the address data in the set according to the similarity, it further includes:

[0030] For the similarity between any two address data in the set, if there is a target similarity less than the set similarity threshold, then set the target similarity to the set minimum value.

[0031] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, the process of clustering the address data in the set according to the similarity includes:

[0032] Based on the similarity between any two address data in the set, determine the distance between the two address data, and the distance is negatively correlated with the similarity;

[0033] Cluster the address data in the set by using a density-based clustering algorithm according to the distance between the two address data.

[0034] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, the process of obtaining the coding features of the geographical entity in the address data includes:

[0035] Using a pre-trained language model to extract the character features of each character in the geographical entity;

[0036] Summing or averaging the character features of each character in the geographical entity, and using the result as the coding features of the geographical entity.

[0037] In a second aspect, there is provided an address consistency recognition device, including:

[0038] An address set acquisition unit, configured to acquire a set of address data to be processed, where the set includes several pieces of address data;

[0039] An entity recognition unit, configured to recognize the geographical entity in the address data and the administrative division level to which the geographical entity belongs;

[0040] An address data coding unit, configured to obtain the coding features of the geographical entity in the address data, and perform weighted fusion on the coding features of the geographical entity in the address data according to the hierarchical weights of the geographical entity, to obtain the coding features of the address data, where the hierarchical weights of the geographical entity correspond to the administrative division level to which the geographical entity belongs;

[0041] An address data clustering unit, configured to calculate the similarity between any two address data in the set based on the coding features of the address data, and cluster the address data in the set according to the similarity, and the address data belonging to the same cluster in the clustering result has address consistency.

[0042] In a third aspect, there is provided an electronic device, including: a memory and a processor;

[0043] The memory is configured to store a program;

[0044] The processor is configured to execute the program to implement each step of the address consistency recognition method described in any one of the foregoing first aspects of the present application.

[0045] In a fourth aspect, there is provided a readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, each step of the address consistency recognition method described in any one of the foregoing first aspects of the present application is implemented.

[0046] In a fifth aspect, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, each step of the address consistency recognition method described in any one of the foregoing first aspects of the present application is implemented.

[0047] With the above technical solution, the present application identifies geographical entities from address data, as well as the administrative division levels to which the geographical entities belong (i.e., the geographical hierarchical structure, such as provinces, cities, districts, streets, etc.). For each administrative division level, a corresponding level weight can be preset in advance. The level weight indicates the importance of the geographical information of the corresponding administrative division level, that is, the importance of the geographical information of different administrative division levels can be distinguished by the level weight. According to the level weights of the administrative divisions to which the geographical entities belong, the coding features of each geographical entity in the address data are weighted and fused, and the coding features of the geographical data can be obtained. The coding features can cover the administrative division level information in the geographical data and enhance the accuracy of feature representation. By calculating the similarity between two address data according to the coding features of the address data, the accuracy of the similarity calculation result can be improved. By clustering the address data according to the similarity, the address data in the same clustering cluster have address consistency. Using the method of the present application can improve the accuracy of the similarity calculation result of the address data, thereby improving the accuracy of the final clustering result and the accuracy of the address consistency recognition result, especially for address data with a complex geographical hierarchical structure.

[0048] In addition, the method of the present application does not rely on an address standardization library and a rule library, and can greatly reduce labor costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0050] Figure 1 is a schematic diagram of an implementation system architecture of the address consistency recognition method provided by an embodiment of the present application;

[0051] Figure 2 is a schematic diagram of a flow of an address consistency recognition method provided by an embodiment of the present application;

[0052] Figure 3 is a schematic diagram of the structure of an address consistency recognition device disclosed by an embodiment of the present application;

[0053] Figure 4 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0055] It can be understood that before using the technical solutions disclosed in the embodiments of the present application, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present application should be informed to users and the authorization of users should be obtained through appropriate means in accordance with relevant laws and regulations.

[0056] The data involved in the technical solutions disclosed in the embodiments of the present application (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related regulations.

[0057] In the current address consistency recognition scheme, the following challenges are often faced:

[0058] Diverse address expressions: The expression methods of addresses can vary due to reasons such as regions, cultures, and language styles. For example, "College Road, Haidian District, Beijing" and "Haidian District, College Road, Beijing" are different expressions of the same address. Traditional string matching algorithms are difficult to accurately identify these differences, and existing standardization tools cannot cover all possible expression forms. In addition, address information often contains redundant characters, symbols, or due to human factors such as typing errors during actual input, the matching difficulty is further increased.

[0059] Non-standardized colloquial addresses: In many natural language application scenarios, users often express address information in an informal and colloquial way. For example, expressions such as "I live near Zhongguancun" or "My home is over there near Lujiazui Financial Center" are very common. Compared with the matching of traditional standardized address libraries, this unstructured and semantically ambiguous expression method is more difficult to process, and the research and application of existing technologies in this field are still relatively limited.

[0060] Data scale and processing complexity: In large-scale application scenarios, address data is usually in the millions or even hundreds of millions, which poses higher requirements for computing resources and processing efficiency. How to ensure high accuracy while ensuring the scalability of the algorithm in large-scale data is an important topic in current technical research. Especially in fields such as logistics and e-commerce, the amount of address data that needs to be processed every day is huge, and how to effectively deduplicate, standardize, and perform consistency matching has become an urgent problem to be solved.

[0061] Lack of effective similarity measurement criteria: Address matching is different from general text similarity calculation. Address information usually includes specific geographical names, building names, streets, numbers, etc. Simply relying on general similarity measurement methods such as traditional string edit distance and n-gram often makes it difficult to obtain the semantic similarity between addresses. And due to the different importance of different parts in address data, directly applying existing deep learning or machine learning models cannot accurately capture the key semantic features between addresses, resulting in poor matching effects.

[0062] Currently, the solutions for address consistency recognition mainly include two types:

[0063] 1. Methods based on string matching and edit distance

[0064] This method mainly measures the similarity of two address texts by calculating the edit distance between address strings or using string matching algorithms, and then decides whether two addresses are consistent.

[0065] This method measures address similarity by comparing the character differences between two address strings, and has the advantages of simplicity and efficiency. However, since this method can only compare at the character level and cannot recognize the geographical hierarchical structure in the address, it cannot distinguish the different importance of geographical information such as provinces, cities, districts, and streets, is overly sensitive to minor differences caused by character addition and deletion, is prone to introducing misjudgments, and shows low accuracy when dealing with addresses with complex structures.

[0066] 2. Methods based on rule libraries and address standardization

[0067] This method uses a pre-constructed address standardization library and rule library to parse and standardize address texts. Through word segmentation technology and regular expressions, key elements in the address (such as provinces, cities, districts, streets, and house numbers) are identified, and then mapped to the canonical expressions in the standard address library. After that, consistency discrimination is performed on the standardized addresses.

[0068] This method has the following defects:

[0069] High maintenance cost: It is necessary to continuously maintain and update the address standard library to ensure that new addresses and special addresses can be included. Building and maintaining a complete address standard library requires a large amount of manual intervention and resource investment.

[0070] Lack of flexibility: For addresses not included in the standard library (such as newly developed communities or streets), this method cannot perform matching, making it difficult to adapt to dynamic scenarios.

[0071] Over-reliance on the standardization library: Once the standard library is incomplete or not updated in a timely manner, the recognition effect will be greatly reduced, and it cannot accurately recognize multiple expressions of the same location.

[0072] Therefore, how to accurately and quickly identify various forms of the same address based on limited address information and effectively cluster these addresses without investing too much manpower in maintaining the address standardization library is an important challenge in the current fields of natural language processing and geographic information systems.

[0073] This application aims to provide an efficient and low-cost solution to the above problems by combining text similarity calculation and clustering algorithms and introducing the administrative division levels of different geographical entities (such as provinces, cities, districts, streets, etc.). The solution of this application can not only handle the diversity and non-standardization problems of address expressions, but also cover the administrative division level information in the encoding feature representation of address data, thereby improving the accuracy of address data similarity calculation and clustering. In addition, the solution of this application does not need to rely on the maintenance of a large address standardization library, can flexibly meet the address consistency recognition requirements in different scenarios, and at the same time maintain high computing efficiency and scalability.

[0074] This application provides an address consistency recognition method, which can be applied to a system architecture as shown in Figure 1 . The system may include a terminal 100 and a server 200. The server 200 may include one or more servers ( Figure 1 illustrated by taking one server as an example).

[0075] Either the terminal 100 or the server 200 can be used alone to execute the address consistency recognition method provided in the embodiments of this application. In addition, the terminal 100 and the server 200 can also be used in cooperation to execute the address consistency recognition method provided in the embodiments of this application.

[0076] Next, describe Figure 1 the product form of the terminal 100 in

[0077] The terminal 100 in the embodiments of this application can be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc., and this application does not make any restrictions on this.

[0078] The embodiments of this application provide an address consistency recognition method. Taking the application of this method to a computer device as an example, the computer device can specifically be Figure 1 the terminal 100 inFigure 2 , the address consistency recognition method specifically includes the following steps:

[0079] Step S100, obtain the set of address data to be processed.

[0080] The set of address data to be processed (abbreviated as the set) includes several pieces of address data. This application needs to perform consistency recognition on each piece of address data in the set, such as clustering each piece of address data representing the same address into one category. The set of address data to be processed can be data provided by a third party or data collected by this application itself, and this application does not limit this.

[0081] Optionally, for the obtained set of address data, preprocessing operations can be further performed, and the preprocessed set of address data is used as the object to be processed in the subsequent steps. The preprocessing operations include but are not limited to: data cleaning, standardization processing, etc.

[0082] For the data cleaning process, it can include removing numbers, punctuation marks, special characters, etc. from the address data, which helps to reduce noise and unify the text format.

[0083] For the standardization processing process, it can include converting all address data into text format, further unifying the case form, removing redundant spaces, ensuring the consistency of address data, and eliminating the influence caused by format differences.

[0084] Step S110, identify the geographical entities in the address data and the administrative division levels to which the geographical entities belong.

[0085] Among them, the geographical entity is an entity word representing geographical information. The administrative division level to which the geographical entity belongs can be understood as the geographical hierarchical structure to which the geographical entity belongs. The administrative division level can divide geographical information into several different levels of administrative divisions according to regulations. Exemplarily, in the order from high to low, the administrative division level can include: province → city → district → street → community, etc.

[0086] In this embodiment, a rule-based matching method can be used to identify the geographical entities in the address data and the administrative division levels to which the geographical entities belong. In another possible implementation, named entity recognition (NER) technology can be used to identify the geographical entities in each piece of address data in the set and the administrative division levels to which the geographical entities belong.

[0087] An optional example is that a geographical entity recognition model can be pre-trained. The geographical entity recognition model can adopt a sequence labeling model or a generative model. In this embodiment, the sequence labeling model is taken as an example for illustration. In the training stage, sample address data with geographical entity type labels can be collected in advance for training the geographical entity recognition model. Among them, the geographical entity type represents the administrative division level to which the geographical entity belongs, and examples are "province", "city", "district", "street", "landmark", "house number", etc.

[0088] The annotation method of the geographical entity type label in the sample address data can adopt BIO or other annotation methods. Taking the BIO annotation method as an example, where B (Begin): represents the starting word of the entity; I (Inside): represents the middle word of the entity; O (Outside): represents a non-entity word.

[0089] For example, for the address "Lujiazui Financial Center, Pudong New Area, Shanghai", we annotate it character by character:

[0090] "Shang", "Hai", "Shi" are respectively annotated as B-city, I-city, I-city.

[0091] "Pu", "Dong", "Xin", "Qu" are respectively annotated as B-district, I-district, I-district, I-district.

[0092] "Lu", "Jia", "Zui", "Jin", "Rong", "Zhong", "Xin" are respectively annotated as B-landmark, I-landmark, I-landmark, I-landmark, I-landmark, I-landmark, I-landmark.

[0093] Through this annotation method of geographical entity type labels, the model can learn the role of each character in the address and its corresponding entity type, so that it can accurately identify the head and tail and type of geographical entities during prediction.

[0094] The geographical entity recognition model can include a pre-trained language model and a cascaded annotation module. Taking the BERT model as an example of the pre-trained language model for illustration.

[0095] Based on the trained geographical entity recognition model, geographical entities in the address data and the administrative division levels to which the geographical entities belong can be recognized.

[0096] Specifically, the address data can be input into the geographical entity recognition model. By leveraging the powerful context understanding ability of the BERT model, the input address data is encoded to obtain the encoding features of each character. Based on the encoding features of each character, the annotation module predicts the entity type labels of each character (such as B - City, I - City, B - Street, I - Street, O, etc.). By identifying consecutive B - labels and I - labels, the character range of the geographical entity and the type of the geographical entity can be determined.

[0097] According to the label sequence predicted by the geographical entity recognition model, the geographical entities in the address data can be extracted, and the start position, end position, and entity type of each geographical entity can be determined. For example, it is recognized that "Shanghai" is a geographical entity of the city type, and "Pudong New Area" is a geographical entity of the district and county type.

[0098] In this embodiment, through address word segmentation and named entity recognition based on BERT, the semantic and structural information of the address data can be deeply understood, and key geographical entities can be accurately extracted. This lays a solid foundation for subsequent address data encoding, similarity calculation, and clustering analysis, and significantly improves the overall effect of address consistency recognition.

[0099] Step S120: Obtain the encoding features of the geographical entities in the address data, and perform weighted fusion on the encoding features of the geographical entities in the address data according to the hierarchical weights of the geographical entities to obtain the encoding features of the address data.

[0100] Among them, the hierarchical weight of the geographical entity corresponds to the administrative division level to which the geographical entity belongs. In this application, the corresponding hierarchical weights can be set in advance for different administrative division levels according to the importance of geographical information at different administrative division levels. The size of the hierarchical weight is positively correlated with the importance of the geographical information at the corresponding administrative division level for measuring the similarity of address data. Further, according to the administrative division level to which the geographical entity belongs, the corresponding hierarchical weight of the geographical entity can be determined, that is, the hierarchical weight of the administrative division level to which the geographical entity belongs is determined as the corresponding hierarchical weight of the geographical entity.

[0101] Exemplarily, the administrative division level includes at least a first level and a second level, and the first level is higher than the second level. Then, the hierarchical weight of the geographical entity belonging to the first level is smaller than the hierarchical weight of the geographical entity belonging to the second level.

[0102] Taking the administrative division levels including provinces, cities, communities, and streets as an example. For higher levels (such as provinces and cities), a lower level weight can be set, and for lower levels (such as communities and streets), a higher level weight can be set, so that the encoding features of address data pay more attention to low-level geographical information, improving the influence of low-level geographical information in the subsequent address data similarity measurement process, thereby enhancing the accuracy of identifying different expressions of the same location.

[0103] In one possible implementation, in this step of obtaining the encoding features of geographical entities in the address data, multiple encoding strategies can be adopted. Exemplarily, the n-gran features of geographical entities can be statistically calculated as the encoding features. In another alternative example, a pre-trained language model can also be used to extract the character features of each character in the geographical entity, and then the character features of each character in the geographical entity are summed or averaged, and the result is used as the encoding feature of the geographical entity.

[0104] Among them, examples of pre-trained language models are the BERT model or other models.

[0105] When the process of identifying geographical entities and types in the address data in step S110 is implemented through a geographical entity recognition model, if the geographical entity recognition model includes the BERT model, then the character features of each character in the input address data can be obtained during the processing of the geographical entity recognition model. In this step, the character features obtained by the BERT model encoding each character belonging to the geographical entity can be directly obtained, and then the character features of each character in the geographical entity are summed or averaged, and the result is used as the encoding feature of the geographical entity.

[0106] The bidirectional encoding ability of the BERT model enables it to understand complex address structures, process address texts of different lengths, and nested or irregular entity expressions. By using the character features encoded by the BERT model to determine the encoding features of geographical entities, the high-dimensional representation of the encoding features of geographical entities is ensured, making it have rich semantic information.

[0107] After obtaining the encoding features of geographical entities, the encoding features of geographical entities in the address data can be weighted and fused according to the hierarchical weights of geographical entities to obtain the encoding features of the address data.

[0108] In one possible implementation, the encoding features of geographical entities in the address data can be weighted according to the hierarchical weights of geographical entities to obtain the weighted encoding features of geographical entities.

[0109] Furthermore, the weighted encoding features of each geographical entity in the address data are combined through vector concatenation or pooling operations to obtain the encoding features of the address data.

[0110] Taking the administrative division levels including "Province V, City, District, Street, Community, Others" as an example, the coding feature V of the address data can be expressed as:

[0111]

[0112] Among them, Concat represents vector concatenation, represents the coding feature of the geographical entity belonging to the administrative division level of "Province" in the address data, represents the coding feature of the geographical entity belonging to the administrative division level of "City" in the address data, represents the coding feature of the geographical entity belonging to the administrative division level of "District / County" in the address data, represents the coding feature of the geographical entity belonging to the administrative division level of "Street" in the address data, represents the coding feature of the geographical entity belonging to the administrative division level of "Community" in the address data, represents the coding feature of the geographical entity belonging to the administrative division level of "Others" in the address data.

[0113] Taking the address data of "Guanshan Avenue, Wuchang District, Wuhan City, Hubei Province" as an example, its coding feature V is expressed as:

[0114]

[0115]

[0116]

[0117]

[0118]

[0119] Among them, a1 - a4 respectively represent the layer weights of the four administrative division levels of "Province", "City", "District / County", and "Street". Exemplarily, a1 < a2 < a3 < a4.

[0120] In this embodiment, by weighted fusion of the coding features of the geographical entities in the address data according to the layer weights of the geographical entities, the coding feature of the address data is obtained, which can make the coding feature of the address data retain the hierarchical information of the administrative division level and the details of each level of geographical elements are retained.

[0121] Step S130: Calculate the similarity between pairwise address data in the set based on the coding feature of the address data, and cluster the address data in the set according to the similarity. Each address data belonging to the same cluster in the clustering result has address consistency.

[0122] After obtaining the coding features of each address data in the set in the above steps, the similarity between two address data can be calculated based on the coding features of the address data.

[0123] In the similarity calculation process, the vector similarity between the coding features of two address data can be calculated. Exemplarily, the Jaccard similarity can be used to measure the similarity between two address data, and a similarity value between 0 and 1 can be obtained. Of course, other similarity calculation methods can also be adopted, which will not be listed one by one in this embodiment.

[0124] By calculating the similarity between two address data in the set, an n×n similarity matrix can be obtained, where each element represents the similarity between a pair of address data. n represents the number of address data in the set.

[0125] The similarity represents the degree of similarity between two address data and measures the possibility that two address data belong to the same address. Therefore, the address data in the set can be clustered according to the similarity. Each address data belonging to the same cluster in the clustering result has address consistency, that is, each address data in the same cluster can be understood as different expressions of the same address.

[0126] In a possible implementation, based on the similarity between two address data in the set (the similarity matrix corresponding to the set), the distance between two address data can be determined, and this distance is negatively correlated with the similarity. According to the distance between two address data, a density-based clustering algorithm is used to cluster the address data in the set to obtain a clustering result.

[0127] When determining the distance between two address data based on the similarity between two address data, the distance can be set as the difference result of 1 minus the similarity, ensuring that the higher the similarity between two address data, the smaller the distance between them, thus facilitating the clustering process.

[0128] In a possible implementation, the similarity matrix corresponding to the set calculated above can be converted into a distance matrix. Each element in the distance matrix represents the distance between two address data, and the distance can be equal to the difference result of 1 minus the similarity. Based on the distance matrix, the address data in the set is clustered.

[0129] In a possible implementation, in order to reduce noise interference, after calculating the similarity between two address data in the set, a low similarity filtering operation can be performed, that is, for the similarity between two address data in the set, if there is a target similarity less than the set similarity threshold, the target similarity is set to the set minimum value. For example, the minimum value is 0, so as to only retain the address data with higher similarity for the subsequent clustering process.

[0130] Among them, the set similarity threshold can be flexibly adjusted by the user to be applicable to a wider range of data scenarios, reducing the instability caused by improper parameter setting. It can be understood that the higher the similarity threshold, the fewer address data are retained, and the fewer address data are divided into the same cluster in subsequent clustering processing, that is, the more stringent the requirements for address consistency judgment. Therefore, when the usage scenario has strict requirements for address consistency, the user can set the similarity threshold higher; conversely, when the usage scenario has relatively loose requirements for address consistency, the user can set the similarity threshold lower.

[0131] Generally, the similarity threshold can be set to 0.54 or other values.

[0132] In this embodiment, a density-based clustering algorithm can be used to cluster the address data in the set. Exemplarily, the clustering algorithm can adopt the OPTICS algorithm, which can process address data with irregular distributions and provides strong adaptability for diverse address representations.

[0133] When using a density-based clustering algorithm to cluster the address data in the set, the minimum sample number can be set, that is, the minimum number of address data included in each cluster. Exemplarily, the minimum sample number can be 3 or other values greater than 1 to ensure clustering density and reduce unnecessary outliers. Among them, the minimum sample number can be flexibly set by the user.

[0134] The address consistency recognition method provided in the embodiment of the present application identifies geographical entities and the administrative division levels to which the geographical entities belong (that is, the geographical hierarchical structure, such as provinces, cities, districts, streets, etc.) from the address data. For each administrative division level, a corresponding level weight can be preset. The level weight indicates the importance of the geographical information of the corresponding administrative division level, that is, the importance of the geographical information of different administrative division levels can be distinguished by the level weight. According to the level weights of the administrative divisions to which the geographical entities belong, the coding features of each geographical entity in the address data are weighted and fused to obtain the coding features of the geographical data. These coding features can cover the administrative division level information in the geographical data and enhance the accuracy of feature representation. By calculating the similarity between two address data according to the coding features of the address data, the accuracy of the similarity calculation result can be improved. Clustering the address data according to the similarity, and the address data in the same cluster have address consistency. Using the method of the present application can improve the accuracy of the address data similarity calculation result, and further improve the accuracy of the final clustering result and the accuracy of the address consistency recognition result, especially for address data with complex geographical hierarchical structures.

[0135] In addition, the method of this embodiment does not rely on an address standardization library and a rule library, which can greatly reduce labor costs.

[0136] In some embodiments of the present application, another alternative implementation of the address consistency recognition method is introduced. Based on the foregoing embodiments, the following post-processing steps can be further added to the method of this embodiment:

[0137] For each outlier in the clustering result, count the number of occurrences of the address data corresponding to the outlier in the set.

[0138] If the number of occurrences of the address data corresponding to an outlier in the set exceeds a set threshold, redefine the address data corresponding to the outlier as an independent clustering cluster. The system can assign a new clustering label (the number of the clustering cluster) to the redefined clustering cluster. The assignment method of the new clustering label can be incremented from the maximum value of the existing clustering labels, or assigned in other ways.

[0139] Furthermore, if the number of occurrences of the address data corresponding to an outlier in the set does not exceed the set threshold, it means that the outlier belongs to noise and can be removed.

[0140] Among them, the set threshold can be set by the user, or can be set based on the minimum number of samples. For example, the set threshold can be equal to or greater than M times the minimum number of samples, where M is a value greater than or equal to 2. For example, M = 3.

[0141] In the clustering analysis of address data, density-based clustering algorithms can identify address data points with low density and mark them as outliers. These outliers usually refer to independent addresses that do not form obvious clusters with other addresses. However, in practical applications, some frequently occurring addresses may be mislabeled as outliers due to data characteristics. To solve this problem, this embodiment provides an outlier processing strategy based on the number of occurrences, that is, if the address data represented by the outlier appears in the set more than the set threshold, it means that the address data is relatively important. Therefore, the address data corresponding to the outlier can be redefined as an independent clustering cluster to ensure that important address information is not missed. By counting the occurrence frequency of outliers in the set, this embodiment can better distinguish random noise and repeatedly occurring addresses, and finally output each clustering cluster as the consistency recognition result.

[0142] In some embodiments of the present application, another alternative implementation solution of the address consistency recognition method is introduced. Before identifying geographical entities in the address data in the foregoing step S110, the address data in the obtained address data set can be preprocessed, and this preprocessing process can include removing the numbers therein to avoid the influence of numbers on the calculation of address similarity.

[0143] According to the preprocessed address data, subsequent processing steps are performed for clustering to obtain a clustering result, and this clustering result is used as the coarse-grained address consistency recognition result. Here, the coarse-grained address consistency recognition result ignores the difference in house numbers.

[0144] In some business scenarios, it may only be necessary to make a coarse-grained judgment on whether two addresses are consistent. For example, if two addresses belong to the same street or the same community, etc., it can be determined that the two addresses are consistent without distinguishing whether the house numbers in the addresses are the same. In this case, using the coarse-grained address consistency recognition result can meet the business requirements.

[0145] In some other business scenarios, it may be necessary to accurately compare whether two addresses are consistent. At this time, it is necessary to be accurate to the house number, and the coarse-grained address consistency recognition result cannot meet the business requirements. Therefore, based on the above embodiments, the method of the present application can further include the following processing steps:

[0146] For each clustering cluster in the clustering result, the address data therein is restored to address data including the house number in digital form.

[0147] Check each clustering cluster to obtain each processed clustering cluster as the refined address consistency recognition result, and the refined address consistency recognition result takes into account the difference in house numbers.

[0148] Among them, the process of checking any clustering cluster includes: if the clustering cluster contains address data with more than two different house numbers, the clustering cluster is split into several different clustering clusters, and the house numbers of the address data in each split clustering cluster are the same.

[0149] It can be understood that in the coarse-grained address consistency recognition result, the address data in each clustering cluster is the address after removing the numbers, and it does not include the house number in digital form. Therefore, in this embodiment, the address data in the clustering cluster is restored to address data including the house number in digital form. On this basis, the house numbers of the address data in the same clustering cluster are compared, and the address data with different house numbers in the same cluster are split into two different clustering clusters to ensure that in each split clustering cluster, the house numbers of the address data are the same.

[0150] Each clustering cluster obtained after the above processing can be used as the refined address consistency recognition result, which takes into account the difference in the house numbers of the address data and is applicable to business scenarios with higher requirements for address consistency.

[0151] The address consistency recognition method provided in the foregoing embodiments of the present application can be applied to a variety of scenarios, including but not limited to the following example scenarios:

[0152] 1. Address consistency verification:

[0153] The clustering results obtained by the method of this application can be used to verify whether multiple expressions belong to the same actual location. For example, in some business scenarios, there may be multiple expressions such as "Lujiazui Financial Center, Pudong New Area, Shanghai" and "Lujiazui Financial Center". Through the clustering analysis of this method, the consistency of these address expressions can be confirmed.

[0154] 2. Business data analysis and management:

[0155] The clustering results obtained by the method of this application help to improve the quality of business data, such as the cleaning of customer data and the unification of location data. By aggregating addresses representing the same location, data redundancy can be reduced and the accuracy of business analysis can be improved.

[0156] 3. Precise push and service:

[0157] For services related to geographical locations (such as food delivery, express delivery, etc.), by identifying different expressions of the same location, the location of customers can be more accurately located and service efficiency can be optimized.

[0158] 4. Log recording and audit tracking:

[0159] In actual applications, this application can also generate log files by recording the process of clustering analysis (such as similarity adjustment, outlier processing, etc.). These log files contribute to the transparency of the system and audit tracking, ensuring the traceability of the entire processing process.

[0160] Next, the address consistency recognition device provided by the embodiments of this application will be described. The address consistency recognition device described below can be mutually corresponding and referred to the address consistency recognition method described above.

[0161] See Figure 3 , Figure 3 which is a schematic structural diagram of an address consistency recognition device disclosed in the embodiments of this application.

[0162] As Figure 3 shown, the device may include:

[0163] An address set acquisition unit 11, configured to acquire a set of address data to be processed, where the set includes several pieces of address data;

[0164] An entity recognition unit 12, configured to recognize geographical entities in the address data and the administrative division level to which the geographical entities belong;

[0165] An address data encoding unit 13, configured to obtain the encoding features of the geographical entity in the address data, and perform weighted fusion on the encoding features of the geographical entity in the address data according to the hierarchical weights of the geographical entity, so as to obtain the encoding features of the address data, where the hierarchical weights of the geographical entity correspond to the administrative division levels to which the geographical entity belongs;

[0166] An address data clustering unit 14, configured to calculate the similarity between pairwise address data in the set based on the encoding features of the address data, and cluster the address data in the set according to the similarity, and the address data belonging to the same cluster in the clustering result has address consistency.

[0167] In a possible implementation, the administrative division levels at least include a first level and a second level, and the first level is higher than the second level;

[0168] The hierarchical weight of the geographical entity belonging to the first level is less than the hierarchical weight of the geographical entity belonging to the second level.

[0169] In a possible implementation, the device of the present application may further include:

[0170] A preprocessing unit, configured to preprocess the address data in the set before the entity recognition unit processes it, and the preprocessing includes removing the numbers therein;

[0171] Then, the clustering result of clustering the address data in the set by the address data clustering unit according to the similarity is used as a coarse-grained address consistency recognition result, and the coarse-grained address consistency recognition result ignores the difference in house numbers.

[0172] In a possible implementation, the device of the present application may further include:

[0173] An address data restoration unit, configured to restore the address data in each cluster in the clustering result to address data including a numerical house number;

[0174] A cluster verification unit, configured to perform verification processing on each cluster to obtain each processed cluster as a refined address consistency recognition result, and the refined address consistency recognition result takes into account the difference in house numbers;

[0175] Wherein, the process of performing verification processing on any cluster includes: if the cluster contains address data with more than two different house numbers, the cluster is split into several different clusters, and the house numbers of the address data in each split cluster are the same.

[0176] In a possible implementation, the device of the present application may further include:

[0177] An outlier analysis unit is configured to, for each outlier in the clustering result, count the number of occurrences of the address data corresponding to the outlier in the set; if the number of occurrences of the address data corresponding to the outlier in the set exceeds a set threshold, redefine the address data corresponding to the outlier as an independent clustering cluster.

[0178] In a possible implementation, the process of the address data encoding unit performing weighted fusion on the encoding features of the geographical entities in the address data according to the hierarchical weights of the geographical entities includes:

[0179] Weight the encoding features of the geographical entities in the address data according to the hierarchical weights of the geographical entities to obtain the weighted encoding features of the geographical entities;

[0180] Combine the weighted encoding features of the geographical entities in the address data through vector concatenation or pooling operations.

[0181] In a possible implementation, before clustering the address data in the set according to the similarity, the address data clustering unit is further configured to: for the similarity between any two address data in the set, if there is a target similarity less than the set similarity threshold, set the target similarity to a set minimum value.

[0182] In a possible implementation, the process of the address data clustering unit clustering the address data in the set according to the similarity includes:

[0183] Based on the similarity between any two address data in the set, determine the distance between the two address data, and the distance is negatively correlated with the similarity;

[0184] Cluster the address data in the set by using a density-based clustering algorithm according to the distance between the two address data.

[0185] In a possible implementation, the process of the address data encoding unit obtaining the encoding features of the geographical entities in the address data includes:

[0186] Use a pre-trained language model to extract the character features of each character in the geographical entity;

[0187] Sum or average the character features of each character in the geographical entity, and the result is used as the encoding feature of the geographical entity.

[0188] An electronic device is further provided in an embodiment of the present application. Refer to Figure 4As shown, it shows a schematic structural diagram of an electronic device suitable for implementing the electronic device in the embodiments of the present application. The electronic device in the embodiments of the present application may include, but is not limited to, fixed terminals such as mobile phones, tablet computers, teaching large screens, wearable devices, and the like. Figure 4 The electronic device shown is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0189] As Figure 4 shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage device 608 into the random access memory (RAM) 603, so as to implement the address consistency recognition method in the foregoing embodiments of the present application. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.

[0190] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a memory card, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 4 it shows an electronic device having various devices, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0191] In the embodiments of the present application, there is also provided a computer program product including computer-readable instructions, which, when running on an electronic device, cause the electronic device to implement any one of the address consistency recognition methods provided in the embodiments of the present application.

[0192] In the embodiments of the present application, there is also provided a computer-readable storage medium carrying one or more computer programs, which, when executed by an electronic device, can cause the electronic device to implement any one of the address consistency recognition methods provided in the embodiments of the present application.

[0193] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in this application, the connection relationships between the modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines.

[0194] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits or dedicated circuits. However, for this application, in more cases, software program implementation is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disc of a computer, and includes several instructions to enable a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of this application.

[0195] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0196] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from a website, a computer, a training device, or a data center to another website, a computer, a training device, or a data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, or a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0197] The embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The embodiments can be combined as needed, and the same or similar parts can be referred to each other.

Claims

1. An address consistency recognition method, characterized in that, Including: Obtain a set of address data to be processed, where the set includes several pieces of address data; Identify the geographical entities in the address data and the administrative division levels to which the geographical entities belong; Obtain the coding features of the geographical entities in the address data, and perform weighted fusion on the coding features of the geographical entities in the address data according to the hierarchical weights of the geographical entities to obtain the coding features of the address data. The hierarchical weights of the geographical entities correspond to the administrative division levels to which the geographical entities belong; Based on the coding features of the address data, calculate the similarity between every two address data in the set, and cluster the address data in the set according to the similarity. Each address data belonging to the same cluster in the clustering result has address consistency.

2. The method according to claim 1, characterized in that, The administrative division levels include at least a first level and a second level, and the first level is higher than the second level; The hierarchical weight of the geographical entity belonging to the first level is less than the hierarchical weight of the geographical entity belonging to the second level.

3. The method according to claim 1, wherein Before identifying the geographical entities in the address data, it further includes: Preprocess the address data in the set, and the preprocessing includes removing the numbers therein; Then, the clustering result after clustering the address data in the set according to the similarity is used as a coarse-grained address consistency recognition result, and the coarse-grained address consistency recognition result ignores the difference in house numbers.

4. The method according to claim 3, characterized in that, It further includes: For each cluster in the clustering result, restore the address data therein to address data including the house number in digital form; Perform verification processing on each cluster to obtain each processed cluster as a refined address consistency recognition result, and the refined address consistency recognition result takes into account the difference in house numbers; Among them, the process of performing verification processing on any cluster includes: if the cluster contains more than two address data with different house numbers, then split the cluster into several different clusters, and the house numbers of the address data in each split cluster are the same.

5. The method according to claim 1, characterized in that, It further includes: For each outlier in the clustering result, count the number of times the address data corresponding to the outlier appears in the set; If the number of times the address data corresponding to the outlier appears in the set exceeds a set threshold, then redefine the address data corresponding to the outlier as an independent cluster.

6. The method according to claim 1, wherein The process of performing weighted fusion on the coding features of the geographical entities in the address data according to the hierarchical weights of the geographical entities includes: Weight the coding features of the geographical entities in the address data according to the hierarchical weights of the geographical entities to obtain the weighted coding features of the geographical entities; Combine the weighted coding features of each geographical entity in the address data through vector splicing or pooling operations.

7. The method according to claim 1, characterized in that, Before clustering the address data in the set according to the similarity, it further includes: For the similarity between every two address data in the set, if there is a target similarity less than the set similarity threshold, then set the target similarity to the set minimum value.

8. The method according to any one of claims 1 to 7, characterized in that, The process of clustering the address data in the set according to the similarity includes: Based on the similarity between pairwise address data in the set, determine the distance between the pairwise address data, where the distance is negatively correlated with the similarity; According to the distance between the pairwise address data, use a density-based clustering algorithm to cluster the address data in the set.

9. The method according to any one of claims 1 to 7, characterized in that The process of obtaining the coding features of the geographical entity in the address data includes: Use a pre-trained language model to extract the character features of each character in the geographical entity; Sum or average the character features of each character in the geographical entity, and the result is used as the coding feature of the geographical entity.

10. An address consistency recognition device, characterized in that, It includes: An address set acquisition unit for acquiring a set of address data to be processed, where the set includes several pieces of address data; An entity recognition unit for recognizing the geographical entity in the address data and the administrative division level to which the geographical entity belongs; An address data coding unit for obtaining the coding features of the geographical entity in the address data, and performing weighted fusion on the coding features of the geographical entity in the address data according to the hierarchical weight of the geographical entity to obtain the coding features of the address data, where the hierarchical weight of the geographical entity corresponds to the administrative division level to which the geographical entity belongs; An address data clustering unit for calculating the similarity between pairwise address data in the set based on the coding features of the address data, and clustering the address data in the set according to the similarity. Each address data belonging to the same clustering cluster has address consistency.

11. An electronic device, characterized in that, It includes: A memory and a processor; The memory is used to store programs; The processor is used to execute the program to implement each step of the address consistency recognition method described in any one of claims 1 to 9.

12. A readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by the processor, it implements each step of the address consistency recognition method described in any one of claims 1 to 9.

13. A computer program product, comprising a computer program, characterized in that, When this computer program is executed by the processor, it implements each step of the address consistency recognition method described in any one of claims 1 to 9.