An address standardization method and device, electronic equipment and storage medium

By predicting and validating address text, a preset format address file is generated, which solves the problem of low standardization accuracy of non-standard Chinese addresses and achieves complete and accurate matching of administrative regions.

CN116167367BActive Publication Date: 2026-05-19QIZHI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
QIZHI TECH CO LTD
Filing Date
2022-12-30
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing technologies have low accuracy in standardizing non-standard Chinese addresses, especially in corporate registration addresses, where incomplete or mismatched address text is difficult to resolve.

Method used

By predicting address text, the system determines administrative regions at all levels, obtains the corresponding administrative region codes, and generates address files in a preset format, including identifying, labeling, verifying, and supplementing administrative regions to improve accuracy.

Benefits of technology

It improves the standardization accuracy of non-standard Chinese addresses, ensures that administrative regions are complete and mutually matched, and reduces the interference of duplicate region names.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116167367B_ABST
    Figure CN116167367B_ABST
Patent Text Reader

Abstract

The application relates to the field of data processing, in particular to an address standardization method and device, electronic equipment and a storage medium, which comprises the following steps: acquiring an address text; predicting based on the address text to determine the predicted administrative regions of each level corresponding to the address text; acquiring the administrative region codes corresponding to the predicted administrative regions of each level respectively, and generating an address file in a preset format based on the address text, the predicted administrative regions of each level and the administrative region codes corresponding to the predicted administrative regions of each level respectively. The application has the effect of improving the accuracy of standardizing non-standard Chinese addresses.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and in particular to an address standardization method, apparatus, electronic device, and storage medium. Background Technology

[0002] In related technologies, the standardization of Chinese addresses largely relies on Named Entity Recognition (NER) technology, which extracts administrative division information elements at each level of the address, cleans and splices them to obtain a standardized address.

[0003] However, when Chinese address text is written in a non-standard or incomplete manner, the standardization results obtained through this method are poor. Especially with business registration addresses, there may be instances where local industrial and commercial departments have irregular management practices, resulting in poor-quality address text entered into the industrial and commercial system. Many address texts are difficult to identify and distinguish regarding the province, city, district (county), and street (township) information they represent. Therefore, improving the accuracy of standardizing non-standard Chinese addresses is a problem that urgently needs to be solved. Summary of the Invention

[0004] To improve the accuracy of standardizing non-standard Chinese addresses, this application provides an address standardization method, apparatus, electronic device, and storage medium.

[0005] Firstly, this application provides an address standardization method, which adopts the following technical solution:

[0006] An address standardization method, comprising:

[0007] Get the address text;

[0008] Based on the address text, a prediction is made to determine the predicted administrative regions at each level corresponding to the address text;

[0009] Obtain the administrative region codes corresponding to the predicted administrative regions at each level;

[0010] An address file in a preset format is generated based on the address text, the predicted administrative regions at each level, and the administrative region codes corresponding to each predicted administrative region.

[0011] By employing the aforementioned technical solution, the predicted administrative regions at each level within the address text can be obtained, and then the administrative region codes for each predicted administrative region can be acquired, reducing interference from duplicate region names. Simultaneously, the predicted administrative regions at each level can be verified, corrected, and supplemented to ensure that the address text contains complete and mutually matching administrative regions. Subsequently, based on the address text, the predicted administrative regions at each level, and their corresponding administrative region codes, a pre-formatted address text is generated, thereby improving the accuracy of standardizing address texts with incomplete administrative regions or mismatched administrative regions.

[0012] In one possible implementation, the step of predicting based on the address text to determine the predicted administrative regions at various levels corresponding to the address text includes:

[0013] Input the address text into the first prediction model to determine the predicted city corresponding to the address text;

[0014] Obtain the second prediction model corresponding to the predicted city;

[0015] Input the address text into the second prediction model to determine the predicted street corresponding to the address text;

[0016] Based on the predicted city and the predicted street, the predicted administrative regions at various levels corresponding to the address text are determined.

[0017] In one possible implementation, before making a prediction based on the address text and determining the predicted administrative regions at various levels corresponding to the address text, the method further includes:

[0018] The address text is identified to determine the names of administrative regions at various levels within the address text;

[0019] The names of administrative regions at each level in the address text are labeled respectively.

[0020] In one possible implementation, inputting the address text into a first prediction model to determine the predicted city corresponding to the address text includes:

[0021] The address text is input into the first prediction model to obtain the confidence scores of each city corresponding to the address text.

[0022] The predicted city corresponding to the address text is determined based on the confidence level of each city corresponding to the address text.

[0023] The step of inputting the address text into any prediction model to determine the predicted street corresponding to the address text includes:

[0024] The address text is input into the second prediction model to obtain the confidence scores of each street corresponding to the address text;

[0025] The predicted street corresponding to the address text is determined based on the confidence level of each street corresponding to the address text.

[0026] In one possible implementation, determining the predicted administrative regions at various levels corresponding to the address text based on the predicted city and the predicted street includes:

[0027] Determine the minimum administrative level in the address text;

[0028] Based on the predicted city and the predicted street, determine the minimum administrative level corresponding to the address text and the predicted administrative regions at each level preceding the minimum administrative level.

[0029] By adopting the above technical solution, when determining an address, the lower administrative level region can be used to determine the larger administrative level region to which it belongs, but the larger administrative level region cannot be used to determine the smaller administrative level region. Therefore, the minimum administrative level in the address text is determined, and then the various larger administrative level regions before the minimum administrative level are determined, which helps to improve accuracy.

[0030] In one possible implementation, generating an address file in a preset format based on the address text, the predicted administrative regions at each level, and the administrative region codes corresponding to each predicted administrative region at each level includes:

[0031] Determine whether the administrative regions at each level in the address text match each other;

[0032] If not, the address text is corrected based on the predicted administrative regions at each level to obtain the corrected address text;

[0033] An address file in a preset format is generated based on the corrected address text and the administrative region codes corresponding to each level of administrative region.

[0034] In one possible implementation, generating an address file in a preset format based on the address text, the predicted administrative regions at each level, and the administrative region codes corresponding to each predicted administrative region at each level includes:

[0035] Determine whether the administrative regions at all levels in the address text are complete;

[0036] If incomplete, the address text is supplemented based on the predicted administrative regions at each level to obtain the supplemented address text;

[0037] An address file in a preset format is generated based on the supplemented address text and the administrative region codes corresponding to each level of administrative region.

[0038] Secondly, this application provides an address standardization device, which adopts the following technical solution:

[0039] An address standardization device, comprising:

[0040] The address text retrieval module is used to retrieve address text.

[0041] The administrative region prediction module is used to predict the corresponding administrative regions at all levels based on the address text.

[0042] The administrative region code acquisition module is used to acquire the administrative region codes corresponding to the predicted administrative regions at each level;

[0043] The generation module is used to generate an address file in a preset format based on the address text, the predicted administrative regions at each level, and the administrative region codes corresponding to the predicted administrative regions at each level.

[0044] By employing the above technical solution, this device predicts the administrative regions at each level in the address text, thereby obtaining the predicted administrative regions at each level. It then acquires the administrative region codes for each predicted administrative region, reducing interference from duplicate region names. Simultaneously, the predicted administrative regions at each level can verify, correct, and supplement the administrative regions contained in the address text, ensuring a complete and mutually matching address text. Subsequently, based on the address text, the predicted administrative regions at each level, and their corresponding administrative region codes, a preset format address text is generated, thus improving the accuracy when standardizing address texts with incomplete administrative regions or mismatched administrative regions.

[0045] In one possible implementation, when the administrative region prediction module makes a prediction based on the address text and determines the predicted administrative regions at various levels corresponding to the address text, it is specifically used for:

[0046] Input the address text into the first prediction model to determine the predicted city corresponding to the address text;

[0047] Obtain the second prediction model corresponding to the predicted city;

[0048] Input the address text into the second prediction model to determine the predicted street corresponding to the address text;

[0049] Based on the predicted city and the predicted street, the predicted administrative regions at various levels corresponding to the address text are determined.

[0050] In one possible implementation, the device further includes:

[0051] The identification module is used to identify the address text and determine the names of administrative regions at various levels in the address text;

[0052] The annotation module is used to annotate the names of administrative regions at each level in the address text.

[0053] In one possible implementation, when the administrative region prediction module inputs the address text into the first prediction model to determine the predicted city corresponding to the address text, it is specifically used for:

[0054] The address text is input into the first prediction model to obtain the confidence scores of each city corresponding to the address text.

[0055] The predicted city corresponding to the address text is determined based on the confidence level of each city corresponding to the address text.

[0056] Specifically, when the administrative region prediction module inputs the address text into any prediction model to determine the predicted street corresponding to the address text, it is used for:

[0057] The address text is input into the second prediction model to obtain the confidence scores of each street corresponding to the address text;

[0058] The predicted street corresponding to the address text is determined based on the confidence level of each street corresponding to the address text.

[0059] In one possible implementation, when the administrative region prediction module determines the predicted administrative regions at various levels corresponding to the address text based on the predicted city and the predicted street, it is specifically used for:

[0060] Determine the minimum administrative level in the address text;

[0061] Based on the predicted city and the predicted street, determine the minimum administrative level corresponding to the address text and the predicted administrative regions at each level preceding the minimum administrative level.

[0062] In one possible implementation, when the generation module generates an address file in a preset format based on the address text, the predicted administrative regions at each level, and the administrative region codes corresponding to each predicted administrative region, it is specifically used for:

[0063] Determine whether the administrative regions at each level in the address text match each other;

[0064] If not, the address text is corrected based on the predicted administrative regions at each level to obtain the corrected address text;

[0065] An address file in a preset format is generated based on the corrected address text and the administrative region codes corresponding to each level of administrative region.

[0066] In one possible implementation, when the generation module generates an address file in a preset format based on the address text, the predicted administrative regions at each level, and the administrative region codes corresponding to each predicted administrative region, it is specifically used for:

[0067] Determine whether the administrative regions at all levels in the address text are complete;

[0068] If incomplete, the address text is supplemented based on the predicted administrative regions at each level to obtain the supplemented address text;

[0069] An address file in a preset format is generated based on the supplemented address text and the administrative region codes corresponding to each level of administrative region.

[0070] Thirdly, this application provides an electronic device that adopts the following technical solution:

[0071] An electronic device comprising:

[0072] At least one processor;

[0073] Memory;

[0074] At least one application, wherein the at least one application is stored in memory and configured to be executed by at least one processor, the at least one application being configured to: perform the address normalization method described above.

[0075] Fourthly, this application provides a computer-readable storage medium, which adopts the following technical solution:

[0076] A computer-readable storage medium includes: a computer program stored thereon that can be loaded by a processor and execute the address normalization method described above.

[0077] In summary, this application includes at least one of the following beneficial technical effects:

[0078] 1. By predicting the administrative regions at each level in the address text, the predicted administrative regions at each level can be obtained, and then the administrative region codes of each predicted administrative region can be obtained to reduce interference from duplicate region names. Simultaneously, the predicted administrative regions at each level can be verified, corrected, and supplemented to ensure that the address text contains complete and mutually matching administrative regions. Then, based on the address text, the predicted administrative regions at each level, and the corresponding administrative region codes, a preset format address text is generated, thereby improving the accuracy when standardizing address texts with incomplete administrative regions or mismatched administrative regions at each level.

[0079] 2. When determining an address, a lower administrative level region can be used to determine the larger administrative level region to which it belongs, but a larger administrative level region cannot be used to determine a smaller administrative level region. Therefore, the lowest administrative level in the address text should be determined, and then the various larger administrative level regions before that lowest administrative level should be determined, which helps to improve accuracy. Attached Figure Description

[0080] Figure 1 This is a flowchart illustrating the address standardization method in an embodiment of this application;

[0081] Figure 2 This is a schematic diagram of the address standardization device in the embodiments of this application;

[0082] Figure 3 This is a schematic diagram of the structure of the electronic device in the embodiments of this application. Detailed Implementation

[0083] The following is in conjunction with the appendix Figure 1 -Appendix Figure 3 This application will be described in further detail.

[0084] After reading this specification, those skilled in the art may make modifications to this embodiment without contributing any inventive step, but such modifications are protected by patent law as long as they fall within the scope of the claims of this application.

[0085] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0086] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article, unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship.

[0087] This application provides an address standardization method, executed by an electronic device, as shown in the following embodiments. Figure 1 The method includes steps S11-S14, wherein:

[0088] Step S11: Obtain the address text.

[0089] In this embodiment of the application, when obtaining address text, one or multiple addresses can be obtained. The format of the obtained address text differs depending on the request method. For example, in a GET request, multiple address texts are separated by commas, such as "No. 55, Guanhou Lane, Ercun, Tianqian Street, Honghai Bay District, Shenzhen, Guangdong Province, China; Zhuoyue East Blue Coast, Shihua Avenue"; while in a POST request, the obtained address text is encapsulated in JSON format.

[0090] Step S12: Make predictions based on the address text and determine the predicted administrative regions at each level corresponding to the address text.

[0091] In this embodiment, the address text may have several possibilities: complete and mutually matching administrative regions at all levels; incomplete administrative regions at all levels; or mismatched administrative regions at all levels. Based on the address text, prediction is performed to determine the corresponding predicted administrative regions at all levels within the address represented by the address text. This allows for the verification, correction, and supplementation of the administrative regions at all levels contained in the address text based on the predicted administrative regions, ultimately resulting in an address text with complete and mutually matching administrative regions at all levels.

[0092] Furthermore, before predicting the corresponding administrative regions at each level based on the address text, it is necessary to identify the address text, determine the names of the administrative regions at each level within the address text, and label these names to associate the region names in the address text with their corresponding administrative regions at each level. For example, if the obtained address text is: No. 55, Guanhou Lane, Ercun, Tianqian Street, Honghai Bay District, Shenzhen City, Guangdong Province, identifying the address text yields: Guangdong Province (provincial-level administrative region) Shenzhen City (municipal-level administrative region) Honghai Bay District (district / county-level administrative region) No. 55, Guanhou Lane, Ercun, Tianqian Street (street-level administrative region).

[0093] Step S13: Obtain the administrative region codes corresponding to each level of predicted administrative region;

[0094] Furthermore, there may be duplicate regional names, but the duplicate regional names may belong to different administrative regions. However, the administrative region code corresponding to each regional name is unique and will not be repeated. Therefore, after obtaining the predicted administrative regions at all levels, the administrative region codes corresponding to each level of administrative region are obtained to reduce the interference of duplicate regional names.

[0095] Step S14: Generate an address file in a preset format based on the address text, the predicted administrative regions at each level, and the corresponding administrative region codes for each predicted administrative region.

[0096] In this embodiment, the names of administrative regions at all levels in the address text are corrected and / or supplemented based on the predicted administrative regions at each level, so that the administrative regions at each level in the address text are complete and mutually matched, resulting in corrected address text. Then, based on the corrected address text and the administrative region codes corresponding to each predicted administrative region, a preset format address text is generated. The preset format is not specifically limited in this embodiment; it is subject to the requirements of the specific usage environment, and users can modify the preset format.

[0097] In related technologies, address standardization often employs Natural Entity Recognition (NER) to identify and label elements within the address. Entity alignment then uses this technology to concatenate the core entities of the address. However, if the address contains errors such as typos, missing information, cluttered information, or outdated information, entity alignment struggles to handle this "dirty data," resulting in poor generalization ability and low accuracy in address standardization. For example, the address "No. 55, Guanhou Lane, Ercun, Tianqian Street, Honghai Bay District, Shenzhen, Guangdong Province" is typically attributed to Shenzhen in these technologies. However, "Tianqian Street, Honghai Bay District" actually corresponds to the administrative region of "Shanwei City / City District / Tianqian Street."

[0098] In this embodiment, by predicting the administrative regions at each level in the address text, the predicted administrative regions at each level corresponding to the address text can be obtained. Then, the administrative region codes of the predicted administrative regions at each level are obtained to reduce interference from duplicate region names. Simultaneously, the predicted administrative regions at each level can verify, correct, and supplement the administrative regions at each level contained in the address text, so as to obtain an address text with complete and mutually matching administrative regions at each level. Then, based on the address text, the predicted administrative regions at each level, and the corresponding administrative region codes, a preset format address text is generated, thereby improving the accuracy when standardizing address texts with incomplete administrative regions or address texts with mismatched administrative regions at each level.

[0099] Further, based on the address text, prediction is performed to determine the corresponding administrative regions at various levels, including steps S121 (not shown in the figure) - S124 (not shown in the figure), wherein:

[0100] Step S121: Input the address text into the first prediction model to determine the predicted city corresponding to the address text.

[0101] In this embodiment, the first prediction model is obtained by training a TextRCNN neural network model architecture. This first prediction model is trained based on the names of all cities in the target region and the names of all levels of administrative regions under each city. For example, in this embodiment, the training data is the names of all provinces and municipalities in China, as well as the names of all levels of administrative regions under each city. Multiple manually annotated correct address texts are used as training samples to train the initial TextRCNN neural network model, where the cities to which the training samples belong are all cities in the target region.

[0102] Specifically, the address text is input into the first prediction model to obtain the confidence scores of each city corresponding to the address text. Then, based on the confidence scores of each city corresponding to the address text, the predicted city corresponding to the address text is determined. When obtaining the confidence scores of each city corresponding to the address text through the first prediction model, the city with the highest confidence score is selected as the predicted city.

[0103] Step S122: Obtain the second prediction model corresponding to the predicted city;

[0104] Step S123: Input the address text into the second prediction model to determine the predicted street corresponding to the address text.

[0105] Furthermore, each city in the target region independently corresponds to a second prediction model, and each second prediction model is obtained by training a TextRCNN neural network model architecture. Each second prediction model is based on all street names under the corresponding city as training data, wherein multiple manually annotated correct address texts are used as training samples to train the initial TextRCNN neural network model, and the streets corresponding to the training samples are all cities corresponding to any second pre-stored model.

[0106] Specifically, the address text is input into the second prediction model to obtain the confidence score of each street corresponding to the address text. Then, based on the confidence scores of each street corresponding to the address text, the predicted street corresponding to the address text is determined. When obtaining the confidence scores of each street corresponding to the address text through the second prediction model, the street with the highest confidence score is selected as the predicted street.

[0107] Specifically, if the predicted street corresponding to the address text can be obtained through the second pre-stored model, then the smallest administrative region in the address text should include the street-level administrative region or the sub-administrative region of the street-level administrative region.

[0108] Step S124: Determine the predicted administrative regions at all levels corresponding to the address text based on the predicted city and predicted street.

[0109] Specifically, after determining the names of administrative regions at all levels in the address text, the lowest administrative level in the address text is determined. Then, based on the predicted city and predicted street, the lowest administrative level corresponding to the address text, as well as the predicted administrative regions at all levels preceding the lowest administrative level, are determined. In other words, if the administrative regions at all levels in the address text are incomplete, higher-level administrative regions can be predicted based on lower-level administrative regions, but lower-level administrative regions cannot be predicted based on higher-level administrative regions.

[0110] Further, based on the address text, the predicted administrative regions at each level, and the corresponding administrative region codes for each predicted administrative region, an address file in a preset format is generated, including steps S141 (not shown in the figure) - S143 (not shown in the figure), wherein:

[0111] Step S141: Determine whether the administrative regions at each level in the address text match each other.

[0112] Specifically, for the administrative regions in the address text, it is determined whether each administrative region in the address text belongs to its superior administrative region. If each administrative region belongs to its superior administrative region, it means that the administrative regions at all levels in the address text match each other.

[0113] Step S142: If not, then correct the address text based on the predicted administrative regions at each level to obtain the corrected address text;

[0114] Step S143: Generate an address file in a preset format based on the corrected address text and the corresponding administrative region codes for each level of administrative region.

[0115] Specifically, if not, it indicates that there is an incorrect administrative region name in the address text. The address text is then corrected based on the predicted administrative regions at each level, resulting in the corrected address text. Continuing with the example above, the address text is: No. 55, Ercun Guanhou Lane, Tianqian Street, Honghai Bay District, Shenzhen City, Guangdong Province. The predicted administrative regions at each level are: Guangdong Province, Shanwei City, Urban District, Tianqian Street, No. 55, Ercun Guanhou Lane. Therefore, the corrected address text is: No. 55, Ercun Guanhou Lane, Tianqian Street, Urban District, Shanwei City, Guangdong Province.

[0116] Furthermore, when generating a preset-format address file based on the address text, the predicted administrative regions at each level, and the corresponding administrative region codes for each level, it is also necessary to determine whether the administrative regions at each level in the address text are complete. If incomplete, the address text is supplemented based on the predicted administrative regions at each level to obtain the supplemented address text. Then, the preset-format address file is generated based on the supplemented address text and the corresponding administrative region codes for each level.

[0117] Further, we can define the smallest administrative region corresponding to the complete address text. Then, after determining the names of each level of administrative region in the address text, we determine whether the address text is complete based on these names and the smallest administrative region corresponding to the complete address text. For example, if the smallest administrative region corresponding to the complete address text is a street-level administrative region, then the address text "55, Guanhou Lane, Ercun, Tianqian Street" is incomplete, lacking provincial and municipal administrative regions. By supplementing the address text with predicted administrative regions at each level, we obtain a more accurate and complete address text.

[0118] The above embodiments describe an address standardization method from the perspective of process flow. The following embodiments describe an address standardization device from the perspective of virtual module or virtual unit. For details, please refer to the following embodiments.

[0119] This application provides an address standardization device, such as... Figure 2 As shown, the device 200 may specifically include an address text acquisition module 201, an administrative region prediction module 202, an administrative region code acquisition module 203, and a generation module 204, wherein:

[0120] Address text acquisition module 201 is used to acquire address text;

[0121] The administrative region prediction module 202 is used to make predictions based on address text and determine the predicted administrative regions at all levels corresponding to the address text.

[0122] The administrative division code acquisition module 203 is used to acquire the administrative division codes corresponding to the predicted administrative regions at each level;

[0123] The generation module 204 is used to generate an address file in a preset format based on the address text, the predicted administrative regions at each level, and the administrative region codes corresponding to the predicted administrative regions at each level.

[0124] In one possible implementation, when the administrative region prediction module 202 makes predictions based on address text and determines the predicted administrative regions at various levels corresponding to the address text, it is specifically used for:

[0125] Input the address text into the first prediction model to determine the predicted city corresponding to the address text;

[0126] Obtain the second prediction model corresponding to the predicted city;

[0127] Input the address text into the second prediction model to determine the predicted street corresponding to the address text;

[0128] Based on the predicted city and predicted street, the corresponding predicted administrative regions at various levels are determined from the address text.

[0129] In one possible implementation, the device 200 further includes:

[0130] The identification module is used to identify address text and determine the names of administrative regions at various levels within the address text;

[0131] The annotation module is used to annotate the names of administrative regions at each level in the address text.

[0132] In one possible implementation, when the administrative region prediction module 202 inputs the address text into the first prediction model to determine the predicted city corresponding to the address text, it is specifically used for:

[0133] Input the address text into the first prediction model to obtain the confidence scores of each city corresponding to the address text;

[0134] The predicted city corresponding to the address text is determined based on the confidence level of each city corresponding to the address text.

[0135] Specifically, when the administrative region prediction module 202 inputs address text into any prediction model to determine the predicted street corresponding to the address text, it is used for:

[0136] Input the address text into the second prediction model to obtain the confidence scores of each street corresponding to the address text;

[0137] The predicted street corresponding to the address text is determined based on the confidence level of each street corresponding to the address text.

[0138] In one possible implementation, when the administrative region prediction module 202 determines the predicted administrative regions at various levels corresponding to the address text based on the predicted city and predicted street, it is specifically used for:

[0139] Determine the lowest administrative level in the address text;

[0140] Based on the predicted city and predicted street, determine the minimum administrative level corresponding to the address text, as well as the predicted administrative regions at each level preceding the minimum administrative level.

[0141] In one possible implementation, when the generation module 204 generates an address file in a preset format based on the address text, the predicted administrative regions at each level, and the administrative region codes corresponding to each predicted administrative region, it is specifically used for:

[0142] Determine whether the administrative regions at each level in the address text match each other;

[0143] If not, the address text is corrected based on the predicted administrative regions at each level to obtain the corrected address text;

[0144] An address file in a preset format is generated based on the corrected address text and the corresponding administrative region codes for each level of administrative region.

[0145] In one possible implementation, when the generation module 204 generates an address file in a preset format based on the address text, the predicted administrative regions at each level, and the administrative region codes corresponding to each predicted administrative region, it is specifically used for:

[0146] Determine if the administrative regions at all levels in the address text are complete;

[0147] If incomplete, the address text is supplemented based on the predicted administrative regions at each level to obtain the supplemented address text;

[0148] An address file in a preset format is generated based on the supplemented address text and the corresponding administrative region codes for each level of administrative region.

[0149] This application provides an electronic device, such as... Figure 3 As shown, Figure 3 The illustrated electronic device 300 includes a processor 301 and a memory 303. The processor 301 and the memory 303 are connected, for example, via a bus 302. Optionally, the electronic device 300 may also include a transceiver 304. It should be noted that in practical applications, the transceiver 304 is not limited to one type, and the structure of this electronic device 300 does not constitute a limitation on the embodiments of this application.

[0150] Processor 301 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 301 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0151] Bus 302 may include a pathway for transmitting information between the aforementioned components. Bus 302 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 302 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 3 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0152] The memory 303 may be a ROM (Read Only Memory) or other type of static storage device capable of storing static information and instructions, RAM (Random Access Memory) or other type of dynamic storage device capable of storing information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.

[0153] The memory 303 is used to store application code that executes the solution of this application, and its execution is controlled by the processor 301. The processor 301 is used to execute the application code stored in the memory 303 to implement the content shown in the foregoing method embodiments.

[0154] Electronic devices include, but are not limited to: mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and in-vehicle terminals (such as in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Servers can also be included. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0155] This application provides a computer-readable storage medium storing a computer program that, when run on a computer, enables the computer to execute the corresponding content in the aforementioned method embodiments.

[0156] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0157] The above description is only a partial embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. An address standardization method, characterized in that, include: S11. Obtain the address text; S12. Based on the address text, perform prediction to determine the predicted administrative regions at various levels corresponding to the address text, including: S121. Inputting the address text into a first prediction model to determine the predicted city corresponding to the address text includes: inputting the address text into the first prediction model to obtain the confidence level of each city corresponding to the address text; determining the predicted city corresponding to the address text based on the confidence level of each city corresponding to the address text; wherein, the first prediction model is trained based on the names of all cities in the target area and the names of the administrative regions under each city, and is used to establish the affiliation relationship between the administrative regions under each city and the city; S122. Obtain the second prediction model corresponding to the predicted city; wherein, the second prediction model independently corresponds to the predicted city and is trained based on all street names under the predicted city; S123. Inputting the address text into the second prediction model to determine the predicted street corresponding to the address text includes: inputting the address text into the second prediction model to obtain the confidence level of each street corresponding to the address text; and determining the predicted street corresponding to the address text based on the confidence level of each street corresponding to the address text. S124. Determining the predicted administrative regions at various levels corresponding to the address text based on the predicted city and the predicted street includes: determining the minimum administrative level in the address text; determining the minimum administrative level corresponding to the address text and the predicted administrative regions at various levels before the minimum administrative level based on the predicted city and the predicted street; wherein, the lower-level administrative regions corresponding to the minimum administrative level are used to predict the higher-level administrative regions; S13. Obtain the administrative region codes corresponding to the predicted administrative regions at each level; S14. Generate an address file in a preset format based on the address text, the predicted administrative regions at each level, and the administrative region codes corresponding to the predicted administrative regions at each level.

2. The address standardization method according to claim 1, characterized in that, Before making predictions based on the address text and determining the corresponding administrative regions at various levels, the process also includes: The address text is identified to determine the names of administrative regions at various levels within the address text; The names of administrative regions at each level in the address text are labeled respectively.

3. The address standardization method according to claim 1, characterized in that, The process of generating an address file in a preset format based on the address text, the predicted administrative regions at each level, and the administrative region codes corresponding to each predicted administrative region includes: S141. Determine whether the administrative regions at all levels in the address text match each other; S142. If not, then the address text is corrected based on the predicted administrative regions at each level to obtain the corrected address text; S143. Generate an address file in a preset format based on the corrected address text and the administrative region codes corresponding to each level of administrative region.

4. The address standardization method according to claim 1, characterized in that, The process of generating an address file in a preset format based on the address text, the predicted administrative regions at each level, and the administrative region codes corresponding to each predicted administrative region includes: Determine whether the administrative regions at all levels in the address text are complete; If incomplete, the address text is supplemented based on the predicted administrative regions at each level to obtain the supplemented address text; An address file in a preset format is generated based on the supplemented address text and the administrative region codes corresponding to each level of administrative region.

5. An address standardization device, characterized in that, The method described by any one of claims 1-4 includes: The address text retrieval module is used to retrieve address text. The administrative region prediction module is used to predict the corresponding administrative regions at all levels based on the address text. The administrative region code acquisition module is used to acquire the administrative region codes corresponding to the predicted administrative regions at each level; The generation module is used to generate an address file in a preset format based on the address text, the predicted administrative regions at each level, and the administrative region codes corresponding to the predicted administrative regions at each level.

6. An electronic device, characterized in that, The electronic device includes: At least one processor; Memory; At least one application, wherein the at least one application is stored in memory and configured to be executed by at least one processor, the at least one application being configured to: perform the address normalization method of any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, include: The computer program is stored that can be loaded by a processor and execute the address normalization method as described in any one of claims 1-4.