A method and system for updating a customer standard address database
The configuration table and Trie tree structure double supplementary processing of the text addresses entered by users in different places, solving the problem that cannot be effectively standardized in the prior art and achieving efficient standard address database updates.
Patent Information
- Application Number
- CN202211259838.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-14
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-10-14
AI Technical Summary
The prior art cannot effectively standardize the original address in the form of text input by the user, resulting in inefficient update of the standard address database.
By configuring the table to split text information into the region address array and the detailed address array, combining the third-party address standardization API and Trie tree structure for double supplementation, calculating the address hierarchy weight sum, selecting the optimal address for standardization, and updating the standard address database.
It realizes high-accuracy standardization of remote input text-form addresses, improving the update efficiency and quality of standard address databases.
Smart Images

Figure CN115438061B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly relates to a method and system for updating a customer standard address database. Background Art
[0002] In various service industries involving address usage such as logistics distribution and car navigation, the data sufficiency in the standard address database and the standardization of each address are closely related to service efficiency and service quality. Therefore, it is necessary to continuously store new addresses into the standard address database, and perform standardization processing on the corresponding new addresses before storing them into the standard address database.
[0003] In the prior art, for the standardization processing of new addresses, it is mostly carried out through the following steps: First, obtain corresponding address parameters based on the user's current location request; then, screen several address nodes similar to the address parameters in the local ES library based on a third-party address coding API; finally, compare the address parameters with each address node respectively and select the one with the smallest offset as the standard address corresponding to the user's current location request, and store it into the standard address database.
[0004] However, this method is only applicable to the address standardization situation where the new address comes from the user's location request. In actual use, new addresses mostly appear in various text information forms input by users. At the same time, affected by the user input process, these text information have more form defects compared with the address parameters obtained based on the location request, resulting in the inability to effectively apply the standardization method based on the location request in this type of situation. In particular, when the text information corresponding to the new address is input by the user from a different place. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and system for updating a customer standard address database, which is used to solve the technical problem that it is impossible to standardize text-based original addresses, especially text-based original addresses input from other places, and to update the standard address database based on them.
[0006] To achieve the above object, the present invention proposes the following technical solutions:
[0007] A method for updating a customer standard address database includes:
[0008] Obtain the text information corresponding to the original address, and split the text information based on the configuration table to obtain the corresponding regional address array and detailed address array; wherein, the regional address array includes, from high to low: provincial address, municipal address, county address, town address, and community address; the detailed address array includes, from high to low: building address and doorplate address;
[0009] Concatenate all address levels in the area address array and the highest address level in the detail address array to obtain a first concatenated address, and supplement the first concatenated address based on a third-party address standardization API to obtain a first completed address, and the corresponding latitude and longitude data of the first completed address;
[0010] Match the area address array to the word segmentation matching table based on the Trie tree structure, and supplement the area address array with the associated addresses in the corresponding child nodes to obtain a supplemented area address array;
[0011] Concatenate all address levels in the supplemented area address array and the highest address level in the detail address array to obtain a second concatenated address, and supplement the second concatenated address based on a third-party address standardization API to obtain a second completed address, and the corresponding latitude and longitude data of the second completed address;
[0012] Through Calculate the weight sum of each address level in the first completed address and the second completed address respectively, and take the first completed address or the second completed address corresponding to the larger weight sum as the pre-standard address; where, k is the total number of address levels, y i represents i whether the y i = 0 indicates a null value, y i = 1 indicates a filled value, x i represents the hit rate of fuzzy matching between the i th address level in the first completed address or the second completed address and the i th address level in the original address, x j represents the hit rate of fuzzy matching between the j th address level in the first completed address or the second completed address and the j th address level in the original address, f ij represents the influence coefficient of the j th address level on the i th address level after the th address level is hit;
[0013] Supplement the pre-standard address based on the detail address array to be the standard address, and store the standard address and the corresponding latitude and longitude data in the standard address database to update it.
[0014] Further, before splitting the text information based on the configuration table to obtain the corresponding area address array and detailed address array, it includes:
[0015] Processing the text information based on the fuzzy semantics algorithm to correct the incorrect expression information or defective expression information therein.
[0016] Further, after storing the standard address and the corresponding latitude and longitude data into the standard address database, it includes:
[0017] Performing string matching between the standard address and the word segmentation matching table based on the Trie tree structure and multi-pattern matching algorithm;
[0018] If the matching fails, constructing a new address node in the word segmentation matching table based on the standard address.
[0019] Further, after storing the standard address and the corresponding latitude and longitude data into the standard address database, it includes:
[0020] Comparing the standard address with the original addresses in the standard address database to supplement the missing address levels in the original addresses, or modifying the incorrect address levels in the original addresses.
[0021] An update system for a customer standard address database, including:
[0022] An acquisition module, configured to acquire text information corresponding to the original address, and split the text information based on the configuration table to obtain the corresponding area address array and detailed address array; wherein, the area address array includes, from high to low: provincial address, municipal address, county address, town address, and community address in sequence; the detailed address array includes, from high to low: building address and house number address in sequence;
[0023] A first standardization module, configured to splice all address levels in the area address array and the highest address level in the detailed address array to obtain a first spliced address, and supplement the first spliced address based on a third-party address standardization API to obtain a first complemented address, and the corresponding latitude and longitude data of the first complemented address;
[0024] A first preprocessing module, configured to match the area address array to the word segmentation matching table based on the Trie tree structure, and supplement the area address array with the associated addresses in the corresponding child nodes to obtain a supplemented area address array;
[0025] A second standardization module, configured to splice all address levels in the supplementary area address array and the highest address level in the detail address array to obtain a second spliced address, supplement the second spliced address based on a third-party address standardization API to obtain a second supplemented address, and the corresponding latitude and longitude data of the second supplemented address;
[0026] A comparison module, configured to calculate the weight sums of each address level in the first supplemented address and the second supplemented address respectively, and take the first supplemented address or the second supplemented address corresponding to the larger weight sum as a pre-standard address; where k is the total number of address levels, y i represents i whether the y i th address level is a null value, y i = 0 indicates a null value, x i represents the hit rate of fuzzy matching between the i th address level in the first supplemented address or the second supplemented address and the i th address level in the original address, x j represents the hit rate of fuzzy matching between the j th address level in the first supplemented address or the second supplemented address and the j th address level in the original address, f ij represents the influence coefficient of the j th address level on the i th address level after the i th address level is hit;
[0027] A first update module, configured to supplement the pre-standard address based on the detail address array as a standard address, and store the standard address and the corresponding latitude and longitude data in a standard address database for updating it.
[0028] Further, it includes:
[0029] A second preprocessing module, configured to process the text information based on a fuzzy semantics algorithm to correct the incorrect expression information or defective expression information therein.
[0030] Further, it includes:
[0031] A matching module, configured to perform string matching between the standard address and the word segmentation matching table based on a Trie tree structure and a multi-pattern matching algorithm;
[0032] A new module is used to construct a new address node in the word segmentation matching table based on the standard address if the matching fails.
[0033] Furthermore, it includes:
[0034] A second update module is used to compare the standard address with the original address in the standard address database to supplement the missing address levels in the original address or modify the incorrect address levels in the original address.
[0035] Beneficial effects:
[0036] As can be seen from the above technical solutions, the technical solution of the present invention provides a method for updating a customer standard address database to improve the technical defects that in the existing update process of the standard address database, the standardization processing of the original address in text form, especially the text form of off-site input, cannot be performed, and the corresponding standard address database cannot be updated.
[0037] The method first splits the text information corresponding to the original address (i.e., the new address to be standardized) through a configuration table to obtain a regional address array and a detailed address array reflecting specific address level information. Secondly, a third-party address standardization API is used to effectively supplement the split regional address array and detailed address array to perform standardization processing on the original address; and at the same time, the difference in the completion result caused by the influence of the search value when supplementing based on the third-party address standardization API is considered, and further the difference in the accuracy of the standard address is caused. In the specific standardization processing, it is carried out in two stages. In the first stage, the third-party address standardization API is directly used to supplement the spliced regional address array and the detailed address array including the highest address level to obtain a first completed address and the corresponding longitude and latitude data. In the second stage, the associated addresses in the word segmentation matching table are used to pre-supplement the regional address array based on the Trie tree structure to obtain a supplementary regional area array, and then the third-party address standardization API is used to supplement the spliced supplementary regional address array and the detailed address array including the highest address level to obtain a second completed address and the corresponding longitude and latitude data. Finally, based on the matching hit rates between the first completed address and the second completed address and the original address, the weight sums of each address level in the first completed address and the second completed address are respectively calculated, and the completed address with the higher weight sum is selected as the pre-standard address, and after adding the remaining detailed address array to it, the required standard address is obtained. Storing it and the corresponding longitude and latitude data in the standard address database completes one update of the standard address database.
[0038] It can be seen that the technical solution of the present invention can achieve a relatively high-accuracy original address in text form without the participation of a positioning system, especially for the off-site standardization processing of the original address in text form input from other places, thereby realizing the update of the standard address database in such cases.
[0039] It should be understood that all combinations of the foregoing concepts and additional concepts described in more detail below can be regarded as part of the inventive subject matter of the present disclosure as long as such concepts do not conflict with each other.
[0040] The foregoing and other aspects, embodiments, and features of the teachings of the present invention can be more fully understood from the following description in conjunction with the accompanying drawings. Other additional aspects of the present invention, such as the features and / or beneficial effects of exemplary embodiments, will be apparent in the following description or will be learned through the practice of specific embodiments in accordance with the teachings of the present invention. Description of the Drawings
[0041] The drawings are not intended to be drawn to scale. In the drawings, each identical or nearly identical component shown in each figure may be represented by the same reference numeral. For clarity, not every component is labeled in each figure. Now, embodiments of various aspects of the present invention will be described by way of example and with reference to the drawings, wherein:
[0042] Figure 1 is a flowchart of the method for updating the standard address database according to this embodiment;
[0043] Figure 2 is for Figure 1 the flowchart of preprocessing the text information;
[0044] Figure 3 is for Figure 1 the flowchart of updating the word segmentation matching table;
[0045] Figure 4 is the flowchart of updating the old standard address in the standard address database. Detailed Embodiments
[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meaning as understood by those of ordinary skill in the art to which the present invention pertains.
[0047] In the description of this invention patent application and the claims, the terms "first", "second" and similar terms do not denote any order, quantity or importance, but are only used to distinguish different components. Similarly, unless the context clearly indicates otherwise, singular terms such as "a", "an" or "the" do not denote a limitation of quantity, but mean that there is at least one. Terms such as "comprising" or "including" mean that the elements or objects appearing before "comprising" or "including" cover the features, wholes, steps, operations, elements and / or components listed after "comprising" or "including", and do not exclude the existence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations. Terms such as "upper", "lower", "left" and "right" are only used to represent relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0048] In the prior art, in the update of the standard address database, since it is necessary to simultaneously standardize the original address by means of address parameters based on positioning requests and a third-party address standardization API, this type of method is limited to the case where the original address corresponds to the current address parameters obtained by requesting the positioning system. In actual applications, however, the original address generally comes from the user's text input, so the above address standardization method cannot be effectively applied, especially when the original address comes from the user's text input from other places. This brings inconvenience to the efficient development of the corresponding service industry. Therefore, this embodiment aims to provide a method for updating the customer standard address database to improve the above-mentioned defects existing in the update of the existing standard address database.
[0049] The following will specifically introduce the method for updating the customer standard address database disclosed in this embodiment with reference to the accompanying drawings.
[0050] As Figure 1 shown, the update method includes:
[0051] Step S102: Obtain the text information corresponding to the original address, and split the text information based on the configuration table to obtain the corresponding regional address array and detailed address array.
[0052] In this step, the original address can be manually entered by the user or obtained by image recognition through photographing; the input process can be local input, and more particularly, it can be input from other places.
[0053] Specifically, the regional address array includes, from high to low: provincial address, municipal address, county address, town address and community address; the detailed address array includes, from high to low: building address and doorplate address.
[0054] When splitting the text information based on the configuration table to obtain the regional address array, the two-way maximum matching principle is followed. That is, first, the text information is roughly segmented according to punctuation marks and decomposed into several sentences, and then these sentences are scanned and segmented using the forward maximum matching method and the reverse maximum matching method. If the matching results obtained from the two word segmentation processes are the same, the word segmentation is considered correct; otherwise, it is processed according to the minimum set.
[0055] The specific structure of the obtained regional address array at this time is as follows:
[0056] {"province": a1, "city": a2, "district": a3, "street": a4, "community": a5}.
[0057] After obtaining the regional address array, the text information is further split based on the configuration table to obtain the detailed address array. The specific process is as follows: take the lowest-level address hierarchy in the regional address array that is not a null value, and then sequentially match keywords from the highest to the lowest address hierarchy according to the regular expression to obtain the detailed address array.
[0058] The specific structure of the obtained detailed address array at this time is as follows:
[0059] {"number": d1, "detail": d2}.
[0060] As a specific implementation manner, since the text information is input by the user, there are often cases where there are typos in the address name, duplicate input of the address name, or incomplete address name. Therefore, in order to obtain effective regional address arrays and detailed address arrays, as Figure 2 shown, step S102 also includes:
[0061] Step S102.2: Process the text information based on the fuzzy semantic algorithm to correct the incorrect expression information or defective expression information therein.
[0062] Step S104: Concatenate all address hierarchies in the regional address array and the highest address hierarchy in the detailed address array to obtain the first concatenated address, and supplement the first concatenated address based on the third-party address standardization API to obtain the first completed address, and the corresponding latitude and longitude data of the first completed address.
[0063] The highest address hierarchy in the detailed address array in this step is: {"number": d1}.
[0064] In this embodiment, a third-party address standardization API is used to complete the first spliced address to implement the standardization processing of the original address. At the same time, the inventor found in actual application that affected by the search value, there will be significant differences in the addresses obtained by the standardized processing. Therefore, the following steps are continued:
[0065] Step S106: Match the area address array to the word segmentation matching table based on the Trie tree structure, and take the associated address in the corresponding sub-node to supplement the area address array to obtain a supplemented area address array.
[0066] Step S108: Concatenate all address levels in the supplemented area address array and the highest address level in the detail address array to obtain a second spliced address, and supplement the second spliced address based on the third-party address standardization API to obtain a second completed address and the longitude and latitude data corresponding to the second completed address.
[0067] The third-party address standardization API used in this step and step S104 can be any publicly available address engine interface or an address engine interface created by the customer himself.
[0068] By supplementing the area address array based on the Trie tree structure and the word segmentation matching table in step S106, a search value different from that in step S104 is formed, and then through concatenation in step S108 and supplementation by the third-party address standardization API, the re-standardization processing of the original address is realized.
[0069] To determine whether to use the first completed address or the second completed address as the final pre-standard address, considering their matching hit rates with the original address, the following steps are continued:
[0070] Step S110: By Calculate the weight sum of each address level in the first completed address and the second completed address respectively, and take the first completed address or the second completed address corresponding to the larger weight sum as the pre-standard address.
[0071] Where, k is the total number of address levels, y i represents i whether the y i th address level is a null value, y i = 0 indicates a null value, x i represents the i th address level in the first completed address or the second completed address and the iThe hit rate after fuzzy matching at each address level, x j indicating the hit rate after fuzzy matching between the j nth address level in the first or second complemented address and the j nth address level in the original address, f ij indicating the influence coefficient on the j mth address level after the nth address level in the first or second complemented address hits. i For ease of calculation, in the specific weights and calculations, each of the hit rates and influence coefficients uses the value obtained by expanding the actual hit rate and influence coefficient by 10 times.
[0072] Step S112: Supplement the pre-standard address based on the detailed address array to obtain the standard address, and store the standard address and its corresponding latitude and longitude data in the standard address database for updating.
[0073] In order to achieve further iterative optimization during the update of the standard address database, on the one hand, as
[0074] shown, after each update, the corresponding word segmentation matching table is also updated and improved. The specific steps include: Figure 3
[0075] Step S114.2: Perform string matching between the standard address and the word segmentation matching table based on the Trie tree structure and multi-pattern matching algorithm.
[0076] The multi-pattern string matching algorithm is also known as the AC automaton algorithm. Based on the Trie tree, it adds a next array similar to the KMP algorithm to search for multiple pattern strings in a main string. When the text information corresponding to the original address is input, the text information is used as the main string and starts matching from the first character in the Trie tree. When it matches to the leaf node of the Trie tree or encounters a character that does not match during the process, the starting position of the main string for matching is shifted one position backward, and the matching starts from the next character until the entire match is completed.
[0077] Step S114.4: If the matching fails, construct a new address node in the word segmentation matching table based on the standard address.
[0078] If the matching fails, it indicates that there is no corresponding address node in the word segmentation matching table, so a new one is added to update the word segmentation matching table.
[0079] On the other hand, affected by the actual area division and the address maintenance frequency, even the addresses existing in the standard address database (i.e., the old standard addresses) may no longer conform to the current address rules over time, thus bringing inconvenience to the corresponding service industries. However, the existing technologies rarely pay attention to this objective defect existing in the addresses of the standard address database. Therefore, as Figure 4 shown, after step S112, it further includes:
[0080] Step S114.2': Compare the standard address with the original addresses in the standard address database to supplement the missing address levels in the original addresses or modify the incorrect address levels in the original addresses.
[0081] If all the address levels in the area address array of an old standard address in the standard address database are the same as those in the area address array of the newly stored standard address except for one address level (such as the county-level address level) in the middle position, it indicates that the information of the aforementioned exceptional address level in the old standard address is incorrect or missing. Therefore, it is supplemented based on the newly input standard address. Furthermore, the maintenance and update of the old standard addresses in the standard address database are realized.
[0082] As can be seen from the above, this embodiment provides an update method for a customer standard address database, which performs a dual-process standardization process on the original address in text form supplemented based on a third-party address standardization API, and calculates the weight sum based on the matching hit rate to obtain the final required standard address. Furthermore, the standardization process of the original address in text form, especially the original address in text form input from other places, and the update of the corresponding standard address database are realized. It meets the service requirements of high efficiency and high quality in the corresponding service industries.
[0083] The above program can run in a processor or can also be stored in a memory (or referred to as a computer-readable storage medium). A computer-readable medium includes permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0084] These computer programs can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate computer-implemented processing. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or Figure 1 one block or multiple blocks. Corresponding to different steps, different modules can be used to implement them.
[0085] This embodiment also provides an update system for a customer standard address database. The system includes:
[0086] An acquisition module, configured to acquire text information corresponding to an original address, and split the text information based on a configuration table to obtain a corresponding area address array and a detailed address array; wherein, the area address array includes, from high to low in sequence: provincial address, municipal address, county address, town address, and community address; the detailed address array includes, from high to low in sequence: building address and house number address.
[0087] A first standardization module, configured to splice all address levels in the area address array and the highest address level in the detailed address array to obtain a first spliced address, supplement the first spliced address based on a third-party address standardization API to obtain a first complemented address, and longitude and latitude data corresponding to the first complemented address.
[0088] The first preprocessing module is used to match the area address array to the word segmentation matching table based on the Trie tree structure, and supplement the area address array with the associated addresses in the corresponding child nodes to obtain a supplemented area address array.
[0089] The second normalization module is used to splice all address levels in the supplemented area address array and the highest address level in the detailed address array to obtain a second spliced address, and supplement the second spliced address based on a third-party address normalization API to obtain a second completed address and the corresponding latitude and longitude data.
[0090] The comparison module is used to calculate the weight sums of each address level in the first completed address and the second completed address respectively, and take the first completed address or the second completed address corresponding to the larger weight sum as the pre-standard address; where k is the total number of address levels, y i represents i whether the y i th address level is a null value, y i =0 indicates a null value, x i represents the hit rate of fuzzy matching between the i th address level in the first completed address or the second completed address and the i th address level in the original address, x j represents the hit rate of fuzzy matching between the j th address level in the first completed address or the second completed address and the j th address level in the original address, f ij represents the influence coefficient of the j th address level on the i th address level after the
[0091] The first update module is used to supplement the pre-standard address based on the detailed address array to obtain a standard address, and store the standard address and the corresponding latitude and longitude data in the standard address database for updating.
[0092] This system is used to implement the steps of the above method, so those that have been described will not be elaborated here.
[0093] For example, the system further includes:
[0094] A second preprocessing module, configured to process the text information based on a fuzzy semantics algorithm to correct incorrect expression information or defective expression information therein.
[0095] For example, the system further includes:
[0096] A matching module, configured to perform string matching between the standard address and the word segmentation matching table based on a Trie tree structure and a multi-pattern matching algorithm.
[0097] An adding module, configured to construct a new address node in the word segmentation matching table based on the standard address if the matching fails.
[0098] For example, the system further includes:
[0099] A second updating module, configured to compare the standard address with the original addresses in the standard address database to supplement missing address levels in the original addresses, or modify incorrect address levels in the original addresses.
[0100] Since the system is built based on the method, the system can also implement the standardization processing of the original address in text form, especially the original address in text form input from other places, and the update of the corresponding standard address database. Thereby, the service efficiency and service quality of the corresponding service industry are greatly improved.
[0101] Although the present invention has been disclosed above with preferred embodiments, it is not intended to limit the present invention. Those with ordinary knowledge in the technical field to which the present invention pertains can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention shall be subject to that defined by the claims.
Claims
1. A method for updating a customer standard address database, characterized in that Including: Obtain the text information corresponding to the original address, and split the text information based on the configuration table to obtain the corresponding area address array and detailed address array; wherein, the area address array includes, from high to low: provincial address, municipal address, county address, town address and community address; the detailed address array includes, from high to low: building address and house number address; Concatenate all address levels in the area address array and the highest address level in the detailed address array to obtain the first concatenated address, and supplement the first concatenated address based on the third-party address standardization API to obtain the first complemented address, and the latitude and longitude data corresponding to the first complemented address; Match the area address array to the word segmentation matching table based on the Trie tree structure, and supplement the area address array with the associated addresses in the corresponding child nodes to obtain a supplemented area address array; Concatenate all address levels in the supplemented area address array and the highest address level in the detailed address array to obtain the second concatenated address, and supplement the second concatenated address based on the third-party address standardization API to obtain the second complemented address, and the latitude and longitude data corresponding to the second complemented address; By calculate the weight sums of each address level in the first completed address and the second completed address respectively, and take the first completed address or the second completed address corresponding to the larger weight sum as the pre-standard address; wherein, k is the total number of address levels, y i represents whether the i th address level is a null value, y i = 0 indicates a null value, y i = 1 indicates a filled value, x i represents the hit rate after fuzzy matching between the i th address level in the first completed address or the second completed address and the i th address level in the original address, x j represents the hit rate after fuzzy matching between the j th address level in the first completed address or the second completed address and the j th address level in the original address, f ij represents the influence coefficient of the j th address level on the i th address level after the th address level is hit; Supplement the pre-standard address based on the detailed address array to be the standard address, and store the standard address and the corresponding latitude and longitude data in the standard address database to update it.
2. The method for updating the customer standard address database according to claim 1, wherein Before splitting the text information based on the configuration table to obtain the corresponding area address array and detailed address array, including: Process the text information based on the fuzzy semantic algorithm to correct the incorrect expression information or defective expression information therein.
3. The method for updating the customer standard address database according to claim 1, wherein After storing the standard address and the corresponding latitude and longitude data in the standard address database, including: Perform string matching between the standard address and the word segmentation matching table based on the Trie tree structure and multi-pattern matching algorithm; If the matching fails, construct a new address node in the word segmentation matching table based on the standard address.
4. The method for updating the customer standard address database according to claim 1, characterized in that After storing the standard address and the corresponding latitude and longitude data in the standard address database, including: Compare the standard address with the original address in the standard address database to supplement the missing address levels in the original address, or modify the incorrect address levels in the original address.
5. An update system for a customer standard address database, characterized in that, Including: An acquisition module for obtaining the text information corresponding to the original address, and splitting the text information based on the configuration table to obtain the corresponding area address array and detailed address array; wherein, the area address array includes, from high to low: provincial address, municipal address, county address, town address and community address; the detailed address array includes, from high to low: building address and house number address; A first standardization module for concatenating all address levels in the area address array and the highest address level in the detailed address array to obtain the first concatenated address, and supplementing the first concatenated address based on the third-party address standardization API to obtain the first complemented address, and the latitude and longitude data corresponding to the first complemented address; The first preprocessing module is used to match the area address array into the word segmentation matching table based on the Trie tree structure, and supplement the area address array with the associated addresses in the corresponding child nodes to obtain a supplemented area address array; The second normalization module is used to splice all address levels in the supplemented area address array and the highest address level in the detailed address array to obtain a second spliced address, and supplement the second spliced address based on a third-party address normalization API to obtain a second supplemented address and the corresponding latitude and longitude data; A comparison module, for passing through calculating the weight sums of each address level in the first completed address and the second completed address respectively, and taking the first completed address or the second completed address corresponding to the larger weight sum as a pre-standard address; wherein, k is the total number of address levels, y i represents whether the i th address level is a null value, y i =0 indicates a null value, y i =1 indicates a filled value, x i represents the hit rate after fuzzy matching between the i th address level in the first completed address or the second completed address and the i th address level in the original address, x j represents the hit rate after fuzzy matching between the j th address level in the first completed address or the second completed address and the j th address level in the original address, f ij represents the influence coefficient of the j th address level on the i th address level after the th address level is hit; The first update module is used to supplement the pre-standard address based on the detailed address array as the standard address, and store the standard address and the corresponding latitude and longitude data in the standard address database for updating.
6. The update system for the customer standard address database according to claim 5, characterized in that, including: The second preprocessing module is used to process the text information based on the fuzzy semantics algorithm to correct the incorrect expression information or defective expression information therein.
7. The update system of the customer standard address database according to claim 5, characterized in that, including: The matching module is used to perform string matching between the standard address and the word segmentation matching table based on the Trie tree structure and the multi-pattern matching algorithm; The new addition module is used to construct a new address node in the word segmentation matching table based on the standard address if the matching fails.
8. The update system for the customer standard address database according to claim 5, characterized in that, including: The second update module is used to compare the standard address with the original address in the standard address database to supplement the missing address levels in the original address or modify the incorrect address levels in the original address.
Citation Information
Patent Citations
Intelligent address searching system and method based on deep learning
CN115017420A
Communication address resolution service updating method and device, equipment, medium and product
CN115168385A