Intelligent address standardization and matching processing engine and method based on big data
Through the intelligent address standardization and matching processing engine based on big data, the problem of inconsistent address format and incomplete information is solved, the accurate analysis and automated processing of unstructured addresses are realized, the efficiency and accuracy of address matching are improved, and it is suitable for logistics, government affairs and map services.
Patent Information
- Application Number
- CN202510431437.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-29
AI Technical Summary
The existing address processing and matching technologies have problems such as inconsistent address format, incomplete information, inconsistent and inaccurate information. Traditional methods are difficult to process massive data and cannot dynamically adapt to complex address scenarios.
It adopts an intelligent address standardization and matching processing engine based on big data, including address segmentation module, CRF machine learning word segmentation module, POI address library completion error correction module, and Trie tree and inverted index matching module. It realizes accurate analysis of unstructured addresses through five-segment twenty-level segmentation standards and CRF machine learning models, combining POI library completion error correction and hybrid index technology.
It significantly improves address matching efficiency and accuracy, can automatically process large amounts of address data, reduce manual intervention and errors, improve data consistency and comparability, reduce enterprise costs, and is suitable for logistics, government affairs and map services and other fields.
Smart Images

Figure CN120386826A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of data processing and geographic information technology, and particularly to an intelligent address standardization and matching processing engine and method based on big data. Background Art
[0002] At present, significant progress has been made in address processing and matching engine technology driven by technologies such as big data. Big data technology can effectively identify and parse key information in structured address data, such as province, city, street, etc., and standardize it. Big data technology supports the efficient processing and analysis of large-scale address data.
[0003] However, there are problems of inconsistent address formats in address processing and matching, and at the same time, there are problems of incomplete, inconsistent, and inaccurate address information. Traditional rule matching or simple word segmentation algorithms are difficult to process massive data and cannot dynamically adapt to complex address scenarios.
[0004] Therefore, we propose an intelligent address standardization and matching processing engine and method based on big data. Summary of the Invention
[0005] The present invention mainly solves the technical problems existing in the above-mentioned prior art, and provides an intelligent address standardization and matching processing engine and method based on big data.
[0006] To achieve the above object, the present invention adopts the following technical solutions. An intelligent address standardization and matching processing engine based on big data includes an address segmentation module, a CRF machine learning word segmentation module, a POI address library completion and error correction module, and a Trie tree and inverted index matching module. The address segmentation module parses unstructured addresses into structured fields based on the five-segment and twenty-level standard.
[0007] Preferably, the CRF machine learning word segmentation module uses a conditional random field (CRF) model combined with an active learning mechanism to perform word segmentation and noise filtering on address texts, with an accuracy rate ≥ 94%.
[0008] Preferably, the POI address library completion and error correction module constructs an inverted index using the crawled POI data, and completes the missing district and county fields and corrects error information through multi-field matching and voting mechanisms.
[0009] Preferably, the Trie tree and inverted index matching module combines the Trie tree to correct the address hierarchy relationship, and realizes fast retrieval and optimal matching of a massive address library based on the inverted index.
[0010] An intelligent address standardization and matching processing method based on big data includes the above-mentioned intelligent address standardization and matching processing engine based on big data, and specifically includes the following steps:
[0011] Step 1: Multi-level segmentation and CRF word segmentation: Input the original address text, perform word segmentation and structured parsing through the CRF model, and output the standardized segmentation result;
[0012] Step 2: Complete missing fields and correct errors: If there are missing or incorrect parsing results (such as missing districts or counties), call the POI address library to complete and correct them, and generate a complete standardized address;
[0013] Step 3: Correct the address hierarchy relationship: Based on the Trie tree, perform hierarchical verification on the standardized address and correct the errors in the spatial constraint relationship;
[0014] Step 4: Retrieve the candidate set and output the optimal longitude and latitude: Use the inverted index technology to screen the candidate address set from the massive address library, determine the optimal matching result through similarity calculation (such as cosine similarity), and output the longitude and latitude coordinates.
[0015] The present invention provides an intelligent address standardization and matching processing engine and method based on big data. It has the following beneficial effects:
[0016] 1. For the intelligent address standardization and matching processing engine and method based on big data, by setting up an address segmentation module, a CRF machine learning word segmentation module, a POI address library completion and error correction module, and a Trie tree and inverted index matching module, through the five-segment and twenty-level segmentation standard and the CRF machine learning model, the accurate parsing of unstructured addresses is realized. Combining the POI library completion and error correction and the hybrid index technology, the address matching efficiency and accuracy are significantly improved, and it can be widely applied in fields such as logistics, government affairs, and map services, solving the problem of low business efficiency caused by chaotic address data.
[0017] 2. For the intelligent address standardization and matching processing engine and method based on big data, by setting up a Trie tree and inverted index matching module, in business scenarios that need to process a large amount of address data, such as logistics distribution and customer information management, the address processing and matching engine can automatically complete tasks such as address recognition, standardization, and error correction, significantly improving the efficiency of the business process, reducing manual intervention and errors, and reducing the enterprise's dependence on manpower by automatically processing address data, and reducing the repetitive work and cost waste caused by address errors.
[0018] 3. The intelligent address standardization and matching processing engine and method based on big data can improve the consistency and comparability of data, facilitate data storage, management and analysis, automatically identify and correct errors in addresses, such as spelling mistakes and format errors, and complete missing address information by setting up a POI address library complement and error correction module, thereby improving the accuracy and integrity of the data. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is the system architecture diagram of the present invention;
[0020] Figure 2 is the method flow diagram of the present invention;
[0021] Figure 3 is the address processing logic diagram of the present invention;
[0022] Figure 4 is the schematic diagram of AddressTree of the present invention;
[0023] Figure 5 is the address to longitude and latitude flow chart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and for those of ordinary skill in the art, without creative efforts, other implementation drawings can be obtained according to the provided drawings.
[0025] The structures, ratios, sizes, etc. shown in this specification are only used to cooperate with the content disclosed in the specification for those who are familiar with this technology to understand and read, and are not used to limit the limited conditions for the implementation of the present invention. Therefore, they do not have technical essence. Any modification of the structure, change of the proportional relationship or adjustment of the size should still fall within the scope that can be covered by the technical content disclosed by the present invention without affecting the effects that the present invention can produce and the purposes that can be achieved.
[0026] It should be noted that similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0027] In the description of the embodiments of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "inner", "outer", "side", etc. is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the invention product is usually placed during use. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. In addition, the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0028] In the description of the embodiments of the present invention, it should also be noted that unless otherwise clearly specified and limited, the terms "set", "install", "connect", and "couple" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the embodiments of the present invention can be understood according to specific situations.
[0029] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.
[0030] Embodiment 1: An intelligent address standardization and matching processing engine based on big data, as Figures 1 - 3 shown, includes an address segmentation module, a CRF machine learning word segmentation module, a POI address library completion and error correction module, and a Trie tree and inverted index matching module. Based on the "Public Security Industry Standard of the People's Republic of China - House Number Address Data Standard" (ICS 35.240.70) issued by the Ministry of Public Security and the actual needs of the business, a "five-segment and twenty-level address segmentation and grading standard" is formulated, and the specific standard is as follows.
[0031] (1) Segment A, the official administrative division segment
[0032] Referring to the administrative division grading standard of the Ministry of Public Security, this segment is divided into six levels:
[0033] The first level: Country
[0034] The second level: Province, municipality directly under the Central Government, autonomous region, special administrative region;
[0035] The third level: Region, prefecture-level city, provincial capital city, autonomous prefecture, league, etc.;
[0036] Fourth level: counties, county-level cities, city districts, autonomous counties, banners, forest areas, industrial and agricultural areas, special zones, etc.;
[0037] Fifth level: townships, towns, ethnic townships, sub-district offices, etc.;
[0038] Sixth level: natural villages.
[0039] (2) Section B, street (road) section
[0040] Refer to the public security standard, that is, the road address. This section is divided into two levels:
[0041] First level: roads;
[0042] Second level: house numbers.
[0043] (3) Section C, area section
[0044] It includes areas, communities, neighborhood committees, business districts, landmark buildings, industrial parks, science and technology parks, economic development zones, etc. It is found in the analysis process that since there are many included items, the classification method of a single level for Section C is likely to affect the segmentation and processing effect. Therefore, Section C is further divided into four levels:
[0045] First level: industrial parks, science and technology parks, economic development zones, business districts, communities, areas, etc.;
[0046] Second level: landmark buildings (buildings, office buildings, squares, markets, large shopping centers, etc.);
[0047] Third level: institutions (companies, enterprises, schools, hospitals, hotels, village committees, sub-district offices, neighborhood committees, police stations, departments, supermarkets, restaurants, merchants, etc.);
[0048] Fourth level: transportation places (bus stops, airports, bridges, subway stations, toll stations, etc.).
[0049] (4) Section D, residential area section
[0050] It includes the residential areas defined in the public security standard and has only one level.
[0051] (5) Section E, detailed address section
[0052] It includes detailed address information such as building numbers, units, blocks, floors, room numbers, etc. and is divided into four levels:
[0053] First level: building numbers;
[0054] Second level: units / blocks;
[0055] Third level: floors;
[0056] Fourth level: room numbers.
[0057] The above five paragraphs with seventeen levels of address segmentation are combined with levels for possible squad / group levels that may appear in processing partial data, levels for identifying directional word information in the address, and other levels for classifying all texts that do not belong to address-related information in order to filter and clean address data, jointly forming a twenty-level classification standard for address standardization and determining the label system "Address Standardization Classification Table". In subsequent address standardization processing, this standard is used as the basis for implementation.
[0058] "Address Standardization Classification Table"
[0059]
[0060]
[0061]
[0062] Example 2: On the basis of Example 1, as Figures 1 - 5 shown, under the premise of a given input sequence X, the probability P(Y|X) is modeled based on the assumption of a Markov random field through a conditional random field model, and its modeling formula is:
[0063]
[0064] where i represents the current node position, k represents the ordinal number of the feature function, and each different feature function has a weight λ k, for each token in the sequence, multiple feature functions are constructed, and weighted summation is performed on each feature function during the final modeling. Z(X) is a normalization term used to form probability values. Here, P(Y|X) represents the probability value of a hidden state sequence given the input sequence X. During the actual sequence labeling task, there may be multiple possible hidden state sequences. The Viterbi algorithm is used to solve for the hidden state sequence with the globally optimal probability as the output sequence. The training process of the conditional random field model often uses common optimization methods such as maximum likelihood estimation, gradient descent, and quasi-Newton descent for parameter optimization. For the task of address standardization word segmentation, the above-mentioned conditional random field model algorithm is adopted, and hundreds of thousands of feature functions are extracted as features of the address text string using the feature templates of CRF++. The conditional random field model is trained using some manually labeled data as training data. At the same time, based on the active learning mechanism, the overall probability value of the feedback result of the conditional random field model is used to screen the labeled data, and only the samples with an overall probability value of less than 0.9 in the model prediction result are selected as the data to be labeled in the next iteration. With this strategy, samples with low confidence in the model prediction are accurately screened for labeling in each iteration, significantly reducing the cost of manual labeling and improving the efficiency of iterative improvement. By using the conditional random field model machine learning algorithm, address standardization word segmentation successfully solves problems such as unclear features and difficulty in division of some internal fields of the address. At the same time, it also filters out the noise during the structural analysis of the address. After iteration, the current address standardization word segmentation achieves an overall accuracy of 94.81%, a recall rate of 93.40%, and an F1 value of 94.10% for each field of the address. By setting up an address segmentation module, a CRF machine learning word segmentation module, a POI address library completion and error correction module, and a Trie tree and inverted index matching module, accurate parsing of unstructured addresses is achieved through a five-segment and twenty-level segmentation standard and a CRF machine learning model. Combining POI library completion and error correction and hybrid indexing technology significantly improves the efficiency and accuracy of address matching and can be widely applied in fields such as logistics, government affairs, and map services, solving the problem of low business efficiency caused by chaotic address data.
[0065] Example 3: On the basis of Example 1 and Example 2, as Figures 1 - 5As shown, based on the POI address library, information completion and error correction are performed on incorrect addresses to help supplement the information fields not filled by users and correct the district fields filled incorrectly by users to improve the standardized results. The POI address library crawls the POI point information of Baidu Map, Amap, etc. through web crawlers, and after standardizing and segmenting the results, stores them in the database. Generally, the POI point information crawled is relatively complete in terms of fields and contains the district field information where it is located. Based on this characteristic, the information in the crawled POI address library can be used. After standardizing and segmenting, the subsequent field information corresponding to the Coun field is extracted, and an inverted index is established hierarchically to summarize the other field information under all districts crawled in the POI library. For the data with the problem of missing districts during the parsing process, using the results after its standardization and segmentation, the A segment, B segment, C segment, and D segment address fields of its existing information are used to query the district hierarchical inverted index constructed by the POI address library. If a matching item is found, the corresponding district is selected to complete the district field of this address. If a certain community or road name appears in the results of multiple districts at the same time, other fields are expanded for multiple matches, and finally, based on the voting mechanism, the district with the highest number of matches is selected to complete the address. If the multi-field matches are still the same, it proves that the information currently contained in this address cannot clarify its district, then all the matched districts are filled in this address to generate multiple addresses with districts and returned for further manual determination after display. For the district error correction task, first, it is necessary to identify the addresses with incorrect districts, that is, during the field matching process for each address, if no match can be found in the lower-level hierarchy of the district in the original address, it is determined that there is an incorrect district in this input address. Then, a strategy similar to district completion is adopted to match the other address segments except the district in the global district hierarchical index, find the district that contains it, and use it as the correct district to correct the original district field to complete the district error correction function. By setting up the Trie tree and inverted index matching module, in business scenarios that need to process a large amount of address data, such as logistics distribution and customer information management, the address processing and matching engine can automatically complete tasks such as address recognition, standardization, and error correction, significantly improving the efficiency of the business process, reducing manual intervention and errors, reducing the enterprise's dependence on manpower through automatic processing of address data, and reducing repetitive work and cost waste caused by incorrect addresses.
[0066] Example 4: On the basis of Example 1, Example 2, and Example 3, as Figures 1 - 5As shown, in order to facilitate statistical analysis in business for many problems of different writings of the same address in actual business, the address is normalized. Here, the normalization mainly involves two aspects. One is the normalization of the habitual writing structure, such as whether to add "City" after the city, whether to add "District" after the county or district, whether the building unit adopts the "-" separation form or the form of "Building XX, Unit XX", etc. On the other hand, it is the problem of different names for communities, landmark buildings, etc., such as the full name and abbreviation of each university, various alternative names of buildings or squares, etc. For these two types of problems, two different sets of technologies are adopted to solve them. First, for the normalization of the first type of writing structure, in the form of a rule template, the common writing structures of each field are constructed into a template for recognition, and finally unified into a normalized output result containing the characteristic words of each field, in the form of "XX City, XX District, XX Road, No. XXX, XXXX Community, Building XX, Unit XX, Room XXX". For the second type of problem of abbreviations and alternative names, first, some alias and abbreviation information of universities and communities is crawled through websites such as Baidu Encyclopedia and Lianjia to construct an alias and abbreviation mapping dictionary. At the same time, large-scale poi data is also used. In the poi data, a poi address group with a longitude and latitude distance less than 10 meters is found. The word vectors calculated by the word2vec algorithm for all community fields, regional fields, and landmark building fields in the address group are added and averaged as features to calculate the cosine text similarity. If the similarity is higher than 80%, it is considered that these addresses are aliases or abbreviations of the same actual address and are added to the alias and abbreviation dictionary. Subsequently, when processing address fields that often appear with abbreviations or alternative names, the abbreviations and alternative names will be matched first. If the match is successful, the abbreviations and alternative names will be replaced with the corresponding full names. With the analysis of the large-scale poi address knowledge base and the crawling of open data sources, effective address normalization is achieved, helping to meet the business requirements for address statistical analysis. By setting up a POI address library completion and error correction module, by converting address information in various formats into a unified standard format, the address processing and matching engine can improve the consistency and comparability of data, facilitate data storage, management and analysis, automatically identify and correct errors in addresses, such as spelling mistakes, format errors, etc., and complete missing address information, thereby improving the accuracy and integrity of data.
[0067] Example Five: On the basis of Example One, Example Two, Example Three and Example Four, as Figures 1 - 5As shown in the figure, according to the standard data of country, province, city, district and street, a Trie tree is established to correct or complete the province, city, district and street in the address. The inverted index module completes the task of finding the document set containing certain words from a large number of documents with a time complexity of O(1) or O(logn), where n is the number of documents in the index. That is to say, by using the inverted index technology, a retrieval complexity that is basically independent of the size of the document set can be achieved, which is crucial for the retrieval of massive content. The inverted index stems from the need to find records according to the values of attributes in practical applications. Each item in this index table includes an attribute value and the addresses of each record with this attribute value. Since it is not the record that determines the attribute value, but the attribute value that determines the position of the record, it is called an inverted index (InvertedIndex). A file with an inverted index is called an inverted index file, abbreviated as an inverted file (InvertedFile). In this address-to-latitude-longitude algorithm, the address is used as the inverted file (InvertedFile), and each field of the hierarchical address is used as the key. For the requested address, based on the corresponding inverted file obtained through the key, the specific process of the address-to-latitude-longitude algorithm is as follows: S1: For the addresses in the POI library, divide or complete them according to the addresses of the national administrative region standard through the AddressTree, and use this standard value as the key, that is, PCD to establish a Map. The content in the Map is all the addresses belonging to this key, that is, Address; S2: Establish an inverted index through the hierarchical address fields, that is, Word, and each address field corresponds to its corresponding address information; S3: When requesting the latitude and longitude of a new address, divide or complete the new address according to the addresses of the national administrative region standard through the AddressTree in the same way, obtain its corresponding key, and then obtain the corresponding Address information in the Map; and perform grading based on the address standardization segmentation algorithm to obtain its address segmentation field, that is, Word; find its corresponding Address information as the candidate set through the address segmentation field Word; S4: According to the business requirements, a model can be used to find the address most similar to the requested address from the candidate set; according to the grading field, compare with each field in the candidate set, and adjust the matching granularity of the fields according to different business requirements, calculate the optimal candidate address, and take out its corresponding latitude and longitude as the final latitude and longitude result.
[0068] Working principle of the present invention: First, the original address data is segmented through a preset address standardization hierarchical classification table. The address data is carefully divided into multiple levels, such as country, province, municipality directly under the Central Government, autonomous region, special administrative region, region, prefecture-level city, provincial capital city, autonomous prefecture, league, county, county-level city, city district, autonomous county, banner, forest area, industrial and agricultural area, special zone, township, town, ethnic township, sub-district office, natural village, road, house number, team / group, industrial park, science and technology park, economic development zone, business district, community, area, landmark building, institution, transportation venue, residential area, building, unit / block, floor, room number, as well as orientation words and other text contents that do not belong to the address classification, ensuring the exhaustiveness and accuracy of the address information. During the processing, the engine uses the conditional random field model to segment and label the address, effectively solving the problem that some field features inside the address are not obvious and not easy to divide. At the same time, through the active learning mechanism, the engine can accurately screen out samples with low confidence in model prediction for manual annotation, thus greatly reducing the cost of manual annotation and improving the efficiency of effect iteration. After multiple iterations, the address standardization word segmentation achieves high accuracy, recall rate, and F1 value for each field in the address. In addition, the engine also has the functions of supplementing and correcting incorrect address information. Based on a large-scale poi address library, it supplements and corrects the data lacking districts or with incorrect districts during the parsing process, ensuring the integrity and accuracy of the address information. At the same time, aiming at the problem of different writings for the same address encountered in actual business, the engine also performs address normalization processing, unifying the habitual writing structures and different names of residential areas, landmark buildings, etc., facilitating statistical analysis in business. Finally, the engine realizes the correction or supplementation operation of the standard data of the country, province, city, and district streets, as well as the rapid retrieval of a large amount of content by establishing a Trie tree and an inverted index module. When the user requests the longitude and latitude of a new address, the engine can quickly find the address most similar to the requested address and extract its corresponding longitude and latitude as the final longitude and latitude result, thus meeting the requirements of various application scenarios.
[0069] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. An intelligent address standardization and matching processing engine based on big data, characterized in that, It includes an address segmentation module, a CRF machine learning word segmentation module, a POI address library completion and error correction module, and a Trie tree and inverted index matching module. The address segmentation module parses unstructured addresses into structured fields based on the five-segment and twenty-level standard.
2. The intelligent address standardization and matching processing engine based on big data according to claim 1, characterized in that: The CRF machine learning word segmentation module uses a conditional random field (CRF) model combined with an active learning mechanism to segment and filter noise from address texts, with an accuracy rate of ≥94%.
3. The intelligent address standardization and matching processing engine based on big data according to claim 1, characterized in that: The POI address library completion and error correction module constructs an inverted index using the crawled POI data, and completes the missing district / county fields and corrects error information through multi-field matching and voting mechanisms.
4. The intelligent address standardization and matching processing engine based on big data according to claim 1, characterized in that: The Trie tree and inverted index matching module combines the Trie tree to correct the address hierarchy relationship, and realizes fast retrieval and optimal matching of a massive address library based on the inverted index.
5. An intelligent address standardization and matching processing method based on big data, characterized in that, It includes the intelligent address standardization and matching processing engine based on big data described in any one of claims 1-4, specifically including the following steps: The first step: multi-level segmentation and CRF word segmentation: Input the original address text, perform word segmentation and structured parsing through the CRF model, and output the standardized segmentation result. The second step: complete the missing fields and correct errors: If there are missing or incorrect (such as missing district / county) in the parsing result, call the POI address library for completion and error correction to generate a complete standardized address. The third step: correct the address hierarchy relationship: Based on the Trie tree, perform hierarchical verification on the standardized address and correct the errors in the spatial constraint relationship. The fourth step: retrieve the candidate set and output the optimal longitude and latitude: Use the inverted index technology to screen the candidate address set from the massive address library, determine the optimal matching result through similarity calculation (such as cosine similarity), and output the longitude and latitude coordinates.