Address Information Parsing Method, Device, Electronic Device, and Computer-Readable Medium
The address information parsing method improves the reliability and uniqueness of address data by using tokenization and clustering to identify and verify address words, addressing the limitations of human-dependent data collection.
Patent Information
- Application Number
- CN202210109312.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-28
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-01-28
AI Technical Summary
Logistics distribution personnel are only familiar with the address information of their distribution area, which leads to the co-occurrence words collected that are not reliable and address unique words are not unique, affecting the accuracy of logistics distribution.
By performing word segmentation processing on the target logistics address information, the first and second address word sequences are generated, and the co-occurring address words and unique words are identified using the preset address word information table and category combination table, and the analysis process is performed to ensure the accuracy of the address information.
It improves the reliability of co-occurrence words in the region, ensures the uniqueness of address unique words in the region, and improves the accuracy of logistics distribution.
Smart Images

Figure CN114443920B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technologies, and more particularly, to methods, apparatuses, electronic devices, and computer-readable media for address information parsing. Background Art
[0002] Among the words representing addresses, there are words describing address attributes, such as "hotel", "university", "hospital"; there are words that can uniquely represent a location or area, such as "XX Headquarter Building", "Restaurant (XX Branch)". However, a single "XX" or "restaurant" is a word with a chain attribute and cannot represent a unique location. We usually refer to words like "XX Headquarter Building" that can uniquely identify a location as address unique words, and words like "hotel", "internet cafe", "restaurant" as address co-occurrence words (because they often co-occur with other words). Accurately identifying and obtaining address unique words and co-occurrence words play an important role in the construction of logistics basic data and the precise sorting and distribution of logistics.
[0003] Currently, the common way to identify and obtain address unique words and co-occurrence words is usually to collect address unique words and co-occurrence words in the area by logistics delivery personnel.
[0004] However, the above method usually has the following technical problems: Logistics delivery personnel are only familiar with the address information in their own delivery areas, resulting in the unreliability of the collected co-occurrence words (for example, whether the co-occurrence words collected by logistics delivery personnel in a certain area are still co-occurrence words in the whole region / city), and the non-uniqueness of the collected address unique words (for example, whether the unique words collected by logistics delivery personnel in a certain area are still unique in the whole region / city). Summary of the Invention
[0005] The content part of the present disclosure is used to briefly introduce concepts, which will be described in detail in the following detailed implementation part. The content part of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0006] Some embodiments of the present disclosure propose methods, apparatuses, electronic devices, and computer-readable media for address information parsing to solve one or more of the technical problems mentioned in the above background art part.
[0007] In a first aspect, some embodiments of the present disclosure provide an address information parsing method, which includes: performing word segmentation on target logistics address information to generate a first address word sequence and a second address word sequence; generating at least one co-occurring address word corresponding to the first address word sequence based on a preset address word information table and the first address word sequence; identifying a unique word in the second address word sequence based on a preset address word category combination table; and performing parsing processing on the target logistics address information based on the at least one co-occurring address word and the unique word.
[0008] Optionally, the performing word segmentation on target logistics address information to generate a first address word sequence and a second address word sequence includes: selecting, as candidate address words, address words with a word frequency greater than or equal to an initial word frequency threshold from each address word included in the target logistics address information based on the word frequency of the address words included in the address word information table, to obtain a candidate address word group; and performing first word segmentation on the target logistics address information according to the candidate address word group to generate a first address word sequence.
[0009] Optionally, the performing word segmentation on target logistics address information to generate a first address word sequence and a second address word sequence includes: performing second word segmentation on the target logistics address information according to a preset address word category table to generate a second address word sequence.
[0010] Optionally, the identifying a unique word in the second address word sequence based on a preset address word category combination table includes: for each address word category combination in the address word category combination table, performing association processing on the corresponding second address word in the second address word sequence through the address word category combination to generate an associated address word; for each generated associated address word, identifying, from a preset associated address word information table, the associated address word information corresponding to the associated address word as candidate associated address word information, to obtain a candidate associated address word information group, where the candidate associated address word information includes a word frequency; and determining the associated address word corresponding to the candidate associated address word information with the largest word frequency included in the candidate associated address word information group as the unique word in the second address word sequence.
[0011] Optionally, the method further includes: determining, as candidate regional address information, the regional address information including the unique word in a preset regional address information table, to obtain a candidate regional address information group; and determining the unique word as a regional unique word based on the candidate regional address information group.
[0012] Optionally, determining the unique word as the regional unique word based on the above-mentioned alternative area address information group includes: marking the address points corresponding to each alternative area address information in the alternative area address information group on the map page; performing clustering processing on the marked address points on the map page to obtain an address point clustering map; removing the address points that do not meet the preset conditions in the address point clustering map to update the address point clustering map; determining the circumscribed rectangle of the updated address point clustering map; and in response to the diagonal length of the circumscribed rectangle being less than or equal to the preset length, determining the unique word as the regional unique word.
[0013] Optionally, generating at least one co-occurring address word corresponding to the first address word sequence based on the preset address word information table and the first address word sequence includes: determining the word frequency of each first address word in the first address word sequence through the word frequency of the address words included in the address word information table to obtain a word frequency sequence; determining the sum of every two word frequencies included in the word frequency sequence as the word frequency threshold to obtain a word frequency threshold sequence; and merging the first address words corresponding to each word frequency threshold greater than or equal to the preset threshold in the word frequency threshold sequence to generate co-occurring address words.
[0014] Optionally, the method further includes: storing the at least one co-occurring address word and the regional unique word in a preset logistics address search server.
[0015] In a second aspect, some embodiments of the present disclosure provide an address information parsing device, which includes: a word segmentation unit configured to perform word segmentation processing on the target logistics address information to generate a first address word sequence and a second address word sequence; a generation unit configured to generate at least one co-occurring address word corresponding to the first address word sequence based on the preset address word information table and the first address word sequence; an identification unit configured to identify the unique word in the second address word sequence based on the preset address word category combination table; and an analysis unit configured to perform analysis processing on the target logistics address information based on the at least one co-occurring address word and the unique word.
[0016] Optionally, the word segmentation unit is further configured to: select, based on the word frequency of the address words included in the address word information table, the address words with a word frequency greater than or equal to the initial word frequency threshold from each address word included in the target logistics address information as alternative address words to obtain an alternative address word group; and perform first word segmentation processing on the target logistics address information according to the alternative address word group to generate a first address word sequence.
[0017] Optionally, the word segmentation unit is further configured to: perform second word segmentation processing on the target logistics address information according to the preset address word category table to generate a second address word sequence.
[0018] Optionally, the recognition unit is further configured to: for each address word category combination in the above address word category combination table, perform an association process on the corresponding second address word in the second address word sequence through the above address word category combination to generate an associated address word; for each generated associated address word, identify the associated address word information corresponding to the above associated address word from a preset associated address word information table as the alternative associated address word information, to obtain a group of alternative associated address word information, where the above alternative associated address word information includes word frequency; determine the associated address word corresponding to the alternative associated address word information with the largest word frequency in the above group of alternative associated address word information as the unique word in the second address word sequence.
[0019] Optionally, the apparatus further includes: a first determination unit configured to determine the regional address information including the above unique word in a preset regional address information table as alternative regional address information, to obtain a group of alternative regional address information; a second determination unit configured to, based on the above group of alternative regional address information, determine the above unique word as the regional unique word.
[0020] Optionally, the second determination unit is further configured to: mark the address points corresponding to each alternative regional address information in the above group of alternative regional address information on the map page; perform a clustering process on the marked address points on the above map page to obtain an address point clustering map; remove the address points that do not meet the preset conditions in the above address point clustering map to update the above address point clustering map; determine the circumscribed rectangle of the updated address point clustering map; in response to the diagonal length of the above circumscribed rectangle being less than or equal to a preset length, determine the above unique word as the regional unique word.
[0021] Optionally, the generation unit is further configured to: determine the word frequency of each first address word in the first address word sequence through the word frequency of the address words included in the above address word information table, to obtain a word frequency sequence; determine the sum of every two word frequencies included in the above word frequency sequence as the word frequency threshold, to obtain a word frequency threshold sequence; perform a merging process on the first address words corresponding to each word frequency threshold greater than or equal to a preset threshold in the above word frequency threshold sequence to generate co-occurring address words.
[0022] Optionally, the apparatus further includes: a storage unit configured to store the above at least one co-occurring address word and the above regional unique word in a preset logistics address search server.
[0023] In a third aspect, some embodiments of the present disclosure provide an electronic device, including: one or more processors; a storage device having stored thereon one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the method described in any implementation manner of the first aspect above.
[0024] Fourthly, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation manner of the first aspect above is implemented.
[0025] The above embodiments of the present disclosure have the following beneficial effects: Through the address information parsing method of some embodiments of the present disclosure, the reliability of the co-occurring words collected within the region is improved, and the uniqueness of the address unique words within the region is ensured. Specifically, the reason for the non-uniqueness of the collected address unique words is that: logistics delivery personnel are only familiar with the address information of their own delivery areas, resulting in the unreliability of the co-occurring words collected (for example, whether the co-occurring words in the area collected by the logistics delivery personnel are still co-occurring words in the whole region / city), and resulting in the non-uniqueness of the collected address unique words (for example, whether the unique words in the area collected by the logistics delivery personnel are still unique in the whole region / city). Based on this, the address information parsing method of some embodiments of the present disclosure first performs word segmentation on the target logistics address information to generate a first address word sequence and a second address word sequence. Thus, data support is provided for accurately identifying the co-occurring words and unique words corresponding to the target logistics address information subsequently. Then, based on a preset address word information table and the above first address word sequence, at least one co-occurring address word corresponding to the first address word sequence is generated. Thus, the co-occurring address words in the first address word sequence can be accurately identified according to the preset address word information table. Then, based on a preset address word category combination table, the unique words in the second address word sequence are identified. Thus, the address unique words corresponding to the target logistics address information can be accurately identified. Finally, based on the above at least one co-occurring address word and the above unique words, parsing processing is performed on the above target logistics address information. Thus, the reliability of the co-occurring words collected within the region is improved, and the uniqueness of the address unique words within the region is ensured. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In combination with the accompanying drawings and referring to the following specific embodiments, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the elements and elements are not necessarily drawn to scale.
[0027] Figure 1 is a schematic diagram of an application scenario of the address information parsing method of some embodiments of the present disclosure;
[0028] Figure 2 is a flowchart according to some embodiments of the address information parsing method of the present disclosure;
[0029] Figure 3It is a flowchart of some other embodiments of the address information parsing method according to the present disclosure;
[0030] Figures 4 - 6 It is an application scenario diagram of the address point clustering map in the address information parsing method according to the present disclosure;
[0031] Figure 7 It is a schematic structural diagram of some embodiments of the address information parsing device according to the present disclosure;
[0032] Figure 8 It is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed implementation manners
[0033] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0034] In addition, it should be noted that for the sake of convenience of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.
[0035] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence relationship of the functions performed by these devices, modules or units.
[0036] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0037] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0038] The present disclosure will be described in detail below with reference to the drawings and in combination with the embodiments.
[0039] Figure 1 It is a schematic diagram of an application scenario of the address information parsing method according to some embodiments of the present disclosure.
[0040] In Figure 1In the application scenario, first, the computing device 101 can perform word segmentation on the target logistics address information 102 to generate a first address word sequence 103 and a second address word sequence 104. Then, the computing device 101 can generate at least one co-occurring address word 106 corresponding to the first address word sequence 103 based on the preset address word information table 105 and the first address word sequence 103. Next, the computing device 101 can identify the unique word 108 in the second address word sequence 104 based on the preset address word category combination table 107. Finally, the computing device 101 can perform parsing processing on the target logistics address information based on the at least one co-occurring address word 106 and the unique word 108.
[0041] It should be noted that the computing device 101 described above can be hardware or software. When the computing device is hardware, it can be implemented as a distributed cluster composed of multiple servers or terminal devices, or as a single server or a single terminal device. When the computing device is embodied as software, it can be installed in the above-listed hardware devices. It can be implemented as, for example, multiple software or software modules for providing distributed services, or as a single software or software module. No specific limitation is made here.
[0042] It should be understood that Figure 1 the number of computing devices in
[0043] Continuing to refer to Figure 2 , a process 200 of some embodiments of the address information parsing method according to the present disclosure is shown. The address information parsing method includes the following steps:
[0044] Step 201, perform word segmentation on the target logistics address information to generate a first address word sequence and a second address word sequence.
[0045] In some embodiments, the execution subject of the address information parsing method (such as Figure 1The computing device 101 shown in the figure can firstly perform word segmentation processing on the target logistics address information through the Jieba word segmenter with a preset configuration dictionary to generate a first address word sequence. Then, the above-mentioned execution subject can perform word segmentation processing on the target logistics address information through the preset regular expression {"[一二三四五六七八九十一百千000a-zA-Z0-9]+[-#]";"[甲乙丙丁戊己庚辛壬癸临东西方南南]*[a-zA-Z0-9一二三四五六七八九十佰千00]+号(省|市|区|巷|园|舍|座|楼|院|门|单元|号|室|巷|栋|层|米|楼|院)} to generate a second address word sequence. Address word sequence. Here, the word segmentation processing method corresponding to the second address word sequence may be Jieba word segmentation. Here, the target logistics address information may refer to the delivery address information collected by the logistics delivery personnel. Here, the Jieba word segmenter of the preset configuration dictionary may refer to the Jieba word segmenter that pre-stores high-frequency words. Here, the Jieba word segmenter of the preset configuration dictionary segments high-frequency words. For example, if "Garden" and "Tianbao" are stored in the Jieba word segmenter of the preset configuration dictionary, "Tianbao Garden" is segmented into "Tianbao / Garden".
[0046] As an example, the target logistics address information may be “Building 5, Unit 2, Daxiong Tulip House, Tianbao Garden, Huayuan Road, Shuibo Town, ZZ District, YY City, XX Province”.
[0047] First, the above-mentioned execution entity can perform word segmentation processing on the above-mentioned target logistics address information "Building No. 5, Unit 2, Daxiong Tulip House, Tianbao Garden, Shuibo Town, Huayuan Road, ZZ District, YY City, XX Province" through the Jieba word segmenter with a preset configuration dictionary to generate a first address word sequence "XX / Province / YY / City / ZZ / District / Shuibo / Town / Garden / Road / Tianbao / Garden / Daxiong / Tulip / She / 2 / Unit / No. / Building".
[0048] Then, the above-mentioned execution entity can perform word segmentation processing on the target logistics address information "Unit 2, Building 5, Daxiong Tulip House, Tianbao Garden, Huayuan Road, Shuibo Town, ZZ District, YY City, XX Province" through the preset regular expression {"[一二三四五六七八九十一百千00a-zA-Z0-9]+[-#]"; "[甲乙丙丁戊己庚辛壬癸临东东南南]*[a-zA-Z0-9一二三四五六七八九十佰千00]+号(省|市|区|巷|园|院|座|楼|院|门|Unit|号|室|巷|楼|层|米|楼|院)} to generate the second address word sequence "XX Province / YY City / ZZ District / Shuibo Town / Huayuan Road / Tianbao Garden / Daxiong Tulip House / Unit 2 / Building 5".
[0049] In some optional implementations of some embodiments, the execution subject may generate the first address word sequence through the following steps:
[0050] First step, based on the word frequencies of the address words included in the address word information table, select the address words with word frequencies greater than or equal to the initial word frequency threshold from each address word included in the above target logistics address information as candidate address words, and obtain a candidate address word group. Here, the address word information in the address word information table may include: the address word and the word frequency corresponding to the above address word.
[0051] In practice, for each address word included in the above target logistics address information, first, the above execution entity can find the word frequency corresponding to the above address word from the address word information table to obtain a word frequency group. Then, the address words corresponding to the word frequencies greater than or equal to the initial word frequency threshold in the word frequency group can be used as candidate address words to obtain a candidate address word group.
[0052] Second step, according to the above candidate address word group, perform a first word segmentation process on the above target logistics address information to generate a first address word sequence. In practice, first, each candidate address word in the above candidate address word group and the word frequency corresponding to the above candidate address word can be associated to generate candidate address word information, and a candidate address word information group is obtained. Then, the above candidate address word information group can be added to the Jieba word segmenter to update the Jieba word segmenter. Finally, the updated Jieba word segmenter is used to perform a first word segmentation process on the target logistics address information to generate a first address word sequence.
[0053] In some optional implementation manners of some embodiments, the above execution entity can perform a second word segmentation process on the above target logistics address information according to a preset address word category table to generate a second address word sequence. Here, the address word category table may refer to a table containing multiple address word attribute categories. In practice, first, for each address word category in the above address word category table, the above execution entity can perform a marking process on the address words corresponding to the above address word category in the above target logistics address information. Then, the above execution entity can perform a word segmentation process at each marked address word in the marked target logistics address information to generate a second address word sequence.
[0054] For example, the address word category table may be:
[0055] Address Word Category Address Word Category Address Word Category Province Road Garden Estate City Garden Number District Courtyard Dormitory Town Unit ···
[0056] The target logistics address information may be "Building 5, Unit 2, Tianbaoyuan, Huayuan Road, Shuibo Town, ZZ District, YY City, XX Province". The address word "province" corresponding to the address word category "province" can be marked. According to the description in the above implementation manner, thus, a second address word sequence "XX Province / YY City / ZZ District / Shuibo Town / Huayuan Road / Tianbaoyuan / 2 Unit / 5 Building" can be generated.
[0057] Step 202: Based on the preset address word information table and the above first address word sequence, generate at least one co-occurring address word corresponding to the above first address word sequence.
[0058] In some embodiments, the above execution subject may generate at least one co-occurring address word corresponding to the above first address word sequence based on the preset address word information table and the above first address word sequence. Here, the preset address word information table may refer to an information table of each address word configured in advance. Here, the address word information in the address word information table may include: the address word and the word frequency corresponding to the above address word.
[0059] In practice, based on the preset address word information table and the above first address word sequence, the above execution subject may generate at least one co-occurring address word corresponding to the above first address word sequence through the following steps:
[0060] First step: Through the above address word information table, look up the word frequency of each first address word in the above first address word sequence to obtain a word frequency sequence.
[0061] For example, the address word information table may be:
[0062] Address Word Word Frequency Address Word Word Frequency Address Word Word Frequency Address Word Word Frequency Province 60 Road 60 Tian 45 Yu 51 City 55 Garden 55 Bao 24 Jin 49 District 52 Courtyard 48 Da 18 Xiang 52 Town 80 Unit 102 Xiong 12 Dormitory 55
[0063] For example, the first address word sequence may be "big / male / jasmine / house". It can be found that the word frequency of the first address word "big" is "18". It can be found that the word frequency of the first address word "male" is "12". It can be found that the word frequency of the first address word "jas" is "51". It can be found that the word frequency of the first address word "mine" is "49". It can be found that the word frequency of the first address word "fra" is "52". It can be found that the word frequency of the first address word "house" is "55". Thus, a word frequency sequence "18, 12, 51, 49, 52, 55" is obtained.
[0064] Second step: According to the above word frequency sequence, determine the information entropy value of each first address word in the above first address word sequence to obtain an information entropy value group. Here, the information entropy value may be the left-right information entropy. Here, the information entropy value of each first address word in the above first address word sequence may be determined through the information entropy formula. For example, the information entropy formula may be the left-right information entropy formula.
[0065] Third step: Sort the above information entropy value group in descending order to obtain an information entropy value sequence.
[0066] Fourth step: Determine the information entropy values corresponding to the serial numbers less than or equal to the preset serial number in the above information entropy value sequence as the target information entropy values to obtain a target information entropy value group. For example, the preset serial number may be "2". That is, the first two information entropy values in the above information entropy value sequence may be determined as the target information entropy value group.
[0067] In the fifth step, perform a merging process on the first address words corresponding to each target information entropy value in the above target information entropy value group and the first address words on the left side of the positions of the above first address words in the first address word sequence to generate co-occurring address words. For example, the respective first address words corresponding to the target information entropy value group may be "gold, fragrance, house". The first address word sequence may be "big / male / flourishing / gold / fragrance / house". Thus, "flourishing / gold / fragrance / house" can be merged into "tulip house" as the co-occurring address word.
[0068] Step 203: Based on a preset address word category combination table, identify the unique words in the above second address word sequence.
[0069] In some embodiments, the above execution entity may, based on a preset address word category combination table, identify the unique words in the above second address word sequence. Here, the preset address word category combination table may refer to a table that is preset and contains multiple address word category combinations. Here, an address word category combination may be a category combination in which at least one address word category is combined together. The address word categories in the address word category combination table may include, but are not limited to, at least one of the following: province, city, district, town, road, main identifier (e.g., garden / courtyard / garden complex), number, unit. Each word in the second address word sequence corresponds to an address word category. For example, in the second address word sequence "XX Province / YY City / ZZ District / Shuibo Town / Garden Road / Tianbao Garden / Unit 2 / No. 5 Building", "Garden Road" corresponds to the address word category "road".
[0070] In practice, based on a preset address word category combination table, the above execution entity may perform the following processing steps for each address word category combination in the address word category combination table:
[0071] In the first step, perform an association process on the address words in the above second address word sequence whose corresponding address word categories are the same as the address word categories included in the above address word category combination to generate associated address words.
[0072] As an example, the address word category combination table may be:
[0073]
[0074]
[0075] The second address word sequence may be "XX Province / YY City / ZZ District / Shuibo Town / Garden Road / Tianbao Garden / Unit 2 / No. 5 Building".
[0076] Thus, the address words in the second address word sequence "XX Province / YY City / ZZ District / Shuibo Town / Huayuan Road / Tianbao Garden / Unit 2 / No. 5 Building" whose corresponding address word categories are the same as those in the above address word category combination "main identifier - number - unit" can be associated to generate an associated address word "Unit 2, No. 5, Tianbao Garden".
[0077] Step 2: Locate the address point represented by the above associated address word on the map page.
[0078] Step 3: In response to the found address points all representing the same address, determine the above associated address word as the unique word.
[0079] Step 4: In response to the found address points not representing the same address, use the address word category combination table after removing the above address word category combination as the new address word category combination table, and execute the above processing steps again.
[0080] Step 204: Based on the above at least one co-occurring address word and the above unique word, perform parsing processing on the above target logistics address information.
[0081] In some embodiments, the above execution subject can perform parsing processing on the above target logistics address information based on the above at least one co-occurring address word and the above unique word. Here, the parsing processing may refer to parsing out the simplified logistics address of the unique address represented by the above target logistics address information. For example, first, for each co-occurring address word in the at least one co-occurring address word, extract the unique address in the above target logistics address information that only includes the co-occurring address word to obtain a unique address group. Then, select the unique address with the least number of characters from the above unique address group as the alternative address. Finally, determine the above unique word and the word with the least number of characters in the above alternative address as the simplified logistics address. For example, the target logistics address information may be "XX Province / YY City / ZZ District / Shuibo Town / Huayuan Road / Tianbao Garden / No. 105 / KFC". Among them, "Tianbao Garden" and "KFC" are co-occurring address words. "ZZ District, Shuibo Town, Huayuan Road, No. 105" is the unique word. The unique address that only includes "Tianbao Garden" can be extracted as "ZZ District, Shuibo Town, Tianbao Garden, No. 105". The unique address that only includes "KFC" can be extracted as "ZZ District, Shuibo Town, KFC". Thus, the simplified logistics address with the least number of characters "ZZ District, Shuibo Town, KFC" can be parsed out.
[0082] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: Through the address information parsing method of some embodiments of the present disclosure, the reliability of the co-occurring words collected within the region is improved, and the uniqueness of the address unique words within the region is ensured. Specifically, the reason for the non-uniqueness of the collected address unique words is that: logistics delivery personnel are only familiar with the address information in their own delivery areas, resulting in the unreliability of the co-occurring words collected (for example, the co-occurring words in the area collected by logistics delivery personnel are not co-occurring words throughout the region / city), and the non-uniqueness of the collected address unique words (for example, the unique words in the area collected by logistics delivery personnel are not unique throughout the region / city). Based on this, in the address information parsing method of some embodiments of the present disclosure, first, the target logistics address information is segmented to generate a first address word sequence and a second address word sequence. Thus, it provides data support for accurately identifying the co-occurring words and unique words corresponding to the target logistics address information subsequently. Then, based on the preset address word information table and the above-mentioned first address word sequence, at least one co-occurring address word corresponding to the first address word sequence is generated. Thus, the co-occurring address words in the first address word sequence can be accurately identified according to the preset address word information table. Finally, based on the preset address word category combination table, the unique words in the second address word sequence are identified. Thus, the address unique words corresponding to the target logistics address information can be accurately identified. Finally, based on the above-mentioned at least one co-occurring address word and the above-mentioned unique words, the target logistics address information is parsed. Thus, the reliability of the co-occurring words collected within the region is improved, and the uniqueness of the address unique words within the region is ensured.
[0083] Further referring to Figure 3 , a flowchart showing some other embodiments of the address information parsing method according to the present disclosure is shown. The address information parsing method includes the following steps:
[0084] Step 301, segment the target logistics address information to generate a first address word sequence and a second address word sequence.
[0085] In some embodiments, the specific implementation of step 301 and the technical effects brought can refer to Figure 2 Step 201 in the corresponding embodiments, which will not be elaborated here.
[0086] Step 302, based on the preset address word information table and the above-mentioned first address word sequence, generate at least one co-occurring address word corresponding to the first address word sequence.
[0087] In some embodiments, based on the preset address word information table and the above-mentioned first address word sequence, the execution subject can generate at least one co-occurring address word corresponding to the first address word sequence through the following steps:
[0088] First step: Determine the word frequency of each first address word in the first address word sequence through the word frequencies of the address words included in the above address word information table, obtaining a word frequency sequence. In practice, for each first address word in the first address word sequence, the above-mentioned execution entity can look up the word frequency corresponding to the first address word in the address word information table to obtain the word frequency sequence.
[0089] Second step: Determine the sum of every two word frequencies included in the above word frequency sequence as the word frequency threshold, obtaining a word frequency threshold sequence. In practice, the above-mentioned execution entity can determine the sum of every two word frequencies included in the above word frequency sequence as the word frequency threshold, obtaining the word frequency threshold sequence.
[0090] Third step: Merge the first address words corresponding to each word frequency threshold greater than or equal to the preset threshold in the above word frequency threshold sequence to generate co-occurring address words.
[0091] As an example, the first address word sequence can be: [De / Yun / Street / Yu / Jin / Xiang / She / Building 7 / Floor 5]. The word frequency sequence can be: {"De": 7; "Yun": 7; "Street": 7; "Yu": 4; "Jin": 4; "Xiang": 4; "She": 4. "Building 7": 1; "Floor 5": 5}. Among them, the word frequency threshold of the first address words "De" and "Yun" is "14". The word frequency threshold of the first address words "Yun" and "Street" is "14". The word frequency threshold of the first address words "Street" and "Yu" is "11". The word frequency threshold of the first address words "Yu" and "Jin" is "8". The word frequency threshold of the first address words "Jin" and "Xiang" is "8". The word frequency threshold of the first address words "Xiang" and "She" is "8". The word frequency threshold of the first address words "She" and "Building 7" is "5". The word frequency threshold of the first address words "Building 7" and "Floor 5" is "6". Merge the two first address words whose word frequency thresholds of every two first address words included in the first address word sequence [De / Yun / Street / Yu / Jin / Xiang / She / Building 7 / Floor 5] are greater than or equal to the predetermined threshold "12" to generate the co-occurring address word [DeYunStreet].
[0092] Step 303: Identify the unique words in the second address word sequence based on a preset address word category combination table.
[0093] In some embodiments, based on the preset address word category combination table and the address word categories corresponding to the second address words in the second address word sequence, the above-mentioned execution entity can identify the unique words in the second address word sequence through the following steps:
[0094] First step: For each address word category combination in the above address word category combination table, perform an association process on the corresponding second address words in the second address word sequence through the above address word category combination to generate associated address words.
[0095] As an example, the address word category combination table can be as follows:
[0096] Combination of Address Word Categories Main Identification - Unit - Building Road - Number Road - Number - Main Identification Town - Road - Number - Main Identification ···
[0097] The second address word sequence can be "XX Province / YY City / ZZ District / Shuibo Town / Huayuan Road / No. 5 / Tianbaoyuan / Unit 2 / Building 2".
[0098] The second address words in the second address word sequence "XX Province / YY City / ZZ District / Shuibo Town / Huayuan Road / No. 5 / Tianbaoyuan / Unit 2" whose corresponding address word categories are the same as those in the above address word category combination "main identifier - unit - building" can be associated to generate the associated address word "Building 2, Unit 2 of Tianbaoyuan".
[0099] Thus, an associated address word group "Building 2, Unit 2 of Tianbaoyuan; No. 5, Huayuan Road; Tianbaoyuan, No. 5, Huayuan Road; Tianbaoyuan, No. 5, Huayuan Road, Shuibo Town" can be obtained.
[0100] In the second step, for each generated associated address word, the associated address word information corresponding to the above - mentioned associated address word is identified from a preset associated address word information table as the alternative associated address word information, and an alternative associated address word information group is obtained. Among them, the above - mentioned alternative associated address word information includes word frequency. Here, the associated address word information table can be a table containing multiple associated address words and the word frequencies corresponding to the associated address words.
[0101] As an example, the associated address word information table can be as follows:
[0102] Related Address Word Word Frequency Tianbaoyuan, Unit 2, Building 2 99 No. 5, Huayuan Road 55 Tianbaoyuan, No. 5, Huayuan Road 102 Tianbaoyuan, No. 5, Huayuan Road, Shuibo Town 180 ···
[0103] According to the description of the second step of the above step 303, an alternative associated address word information group "Building 2, Unit 2 of Tianbaoyuan: 99; No. 5, Huayuan Road: 55; Tianbaoyuan, No. 5, Huayuan Road: 102; Tianbaoyuan, No. 5, Huayuan Road, Shuibo Town: 180" can be obtained.
[0104] In the third step, the associated address word corresponding to the alternative associated address word information with the largest word frequency in the above - mentioned alternative associated address word information group is determined as the unique word in the above - mentioned second address word sequence. As an example, the associated address word "Tianbaoyuan, No. 5, Huayuan Road, Shuibo Town" corresponding to the alternative associated address word information "Tianbaoyuan, No. 5, Huayuan Road, Shuibo Town: 180" with the largest word frequency in the alternative associated address word information group "Building 2, Unit 2 of Tianbaoyuan: 99; No. 5, Huayuan Road: 55; Tianbaoyuan, No. 5, Huayuan Road: 102; Tianbaoyuan, No. 5, Huayuan Road, Shuibo Town: 180" can be determined as the unique word in the above - mentioned second address word sequence.
[0105] Step 304: Based on the above at least one co-occurring address word and the above unique word, perform parsing processing on the above target logistics address information.
[0106] In some embodiments, for the specific implementation of step 304 and the resulting technical effects, reference can be made to Figure 2 step 204 in the corresponding embodiments, which will not be elaborated here.
[0107] Step 305: Determine the regional address information containing the above unique word in the preset regional address information table as alternative regional address information, and obtain a group of alternative regional address information.
[0108] In some embodiments, the above execution subject may determine the regional address information containing the above unique word in the preset regional address information table as alternative regional address information, and obtain a group of alternative regional address information. Here, the regional address information table may be a table composed of multiple regional address information representing unique addresses. Here, the regional address information representing unique addresses may have specific different expression forms.
[0109] As an example, the regional address information table may be:
[0110] Regional Address Information Building 2, Unit 2, Tianbaoyuan, AA District, YY City, XX Province No. 5, Huayuan Road, YY City, XX Province Tianbaoyuan, Huayuan Road, CC District, YY City, XX Province Tianbaoyuan, No. 5, Huayuan Road, Shuibo Town, BB District, YY City, XX Province Tianbaoyuan, No. 5, Huayuan Road, Shuibo Town, YY City Tianbaoyuan, No. 5, Huayuan Road, Shuibo Town, BB District Tianbaoyuan, No. 5, Huayuan Road, Shuibo Town, HH City ···
[0111] The unique word may be "No. 5, Tianbao Garden, Huayuan Road, Shuibo Town".
[0112] According to the description of the above embodiments, a group of alternative regional address information "XX Province, YY City, BB District, No. 5, Huayuan Road, Shuibo Town; YY City, No. 5, Huayuan Road, Shuibo Town; BB District, No. 5, Huayuan Road, Shuibo Town; HH City, No. 5, Huayuan Road, Shuibo Town" can be obtained.
[0113] Step 306: Based on the above group of alternative regional address information, determine the above unique word as the regional unique word.
[0114] In some embodiments, based on the above group of alternative regional address information, the above execution subject may determine the above unique word as the regional unique word through the following steps:
[0115] First step: Mark the address points corresponding to each alternative regional address information in the above group of alternative regional address information on the map page. In practice, first, the above execution subject may call the map page. Then, in the map page, mark the address points represented by each alternative regional address information in the above group of alternative regional address information.
[0116] Step 2: Perform clustering on the address points marked on the above map page to obtain an address point clustering map. In practice, the density clustering algorithm DBSCAN (Density-Based Spatial Clustering of Applications with Noise) can be used to perform clustering on the address points marked on the above map page to obtain an address point clustering map (as Figure 4 shown).
[0117] Step 3: Remove the address points that do not meet the preset conditions in the above address point clustering map to update the above address point clustering map. Here, the preset condition can refer to "the address points connected into clusters in the address point clustering map". As Figure 5 shown, remove the address points that are not connected into clusters in the address point clustering map.
[0118] Step 4: Determine the bounding rectangle of the updated address point clustering map. As Figure 6 shown, the address points in the address point clustering map can be enclosed by the smallest matrix.
[0119] Step 5: In response to the diagonal length of the above bounding rectangle being less than or equal to the preset length, determine the above unique word as the regional unique word. In practice, the above execution entity can, in response to the diagonal length of the above bounding rectangle being less than or equal to the preset length, determine the above unique word as the regional unique word.
[0120] Optionally, store the above at least one co-occurring address word and the above regional unique word in a preset logistics address search server.
[0121] In practice, the above execution entity can store the above at least one co-occurring address word and the above regional unique word in a preset logistics address search server. Here, the logistics address search server can refer to a server that stores multiple logistics address information. Thus, it is convenient for subsequent logistics delivery personnel and users to query addresses. Therefore, it is convenient for logistics delivery personnel to determine the logistics delivery address based on the queried co-occurring address words and regional unique words for logistics delivery.
[0122] From Figure 3 it can be seen that compared with the description of some embodiments corresponding to Figure 2 , Figure 3 the process 300 of the address information parsing method in some embodiments corresponding to ensures the reliability and uniqueness of the address unique word through the verification of the identified unique word, so that the unique word can be accurately pushed to the logistics delivery personnel subsequently, facilitating accurate sorting and delivery of logistics.
[0123] For further reference, Figure 7, as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of an address information parsing device, and these device embodiments correspond to Figure 2 the method embodiments shown, and the device can be specifically applied to various electronic devices.
[0124] As Figure 7 shown, the address information parsing device 700 in some embodiments includes: a word segmentation unit 701, a generation unit 702, an identification unit 703, and an analysis unit 704. Among them, the word segmentation unit 701 is configured to perform word segmentation processing on the target logistics address information to generate a first address word sequence and a second address word sequence; the generation unit 702 is configured to generate at least one co-occurring address word corresponding to the first address word sequence based on a preset address word information table and the above first address word sequence; the identification unit 703 is configured to identify the unique word in the second address word sequence based on a preset address word category combination table and the address word category corresponding to the second address word in the above second address word sequence. The analysis unit 704 is configured to perform analysis processing on the target logistics address information based on the at least one co-occurring address word and the unique word.
[0125] Optionally, the word segmentation unit 701 is further configured to: select, from each address word included in the target logistics address information, an address word with a word frequency greater than or equal to an initial word frequency threshold as an alternative address word based on the word frequency of the address words included in the address word information table, to obtain an alternative address word group; perform first word segmentation processing on the target logistics address information according to the alternative address word group to generate a first address word sequence.
[0126] Optionally, the word segmentation unit 701 is further configured to: perform second word segmentation processing on the target logistics address information according to a preset address word category table to generate a second address word sequence.
[0127] Optionally, the identification unit 703 is further configured to: for each address word category combination in the address word category combination table, perform association processing on the corresponding second address word in the second address word sequence through the address word category combination to generate an associated address word; for each generated associated address word, identify the associated address word information corresponding to the associated address word from a preset associated address word information table as an alternative associated address word information to obtain an alternative associated address word information group, where the alternative associated address word information includes a word frequency; determine the associated address word corresponding to the alternative associated address word information with the largest word frequency included in the alternative associated address word information group as the unique word in the second address word sequence.
[0128] Optionally, the device 700 further includes: a first determination unit configured to determine, as alternative area address information, the area address information including the above-mentioned unique word in a preset area address information table, so as to obtain a group of alternative area address information; and a second determination unit configured to determine the above-mentioned unique word as an area unique word based on the above-mentioned group of alternative area address information.
[0129] Optionally, the second determination unit is further configured to: mark the address points corresponding to each of the alternative area address information in the above-mentioned group of alternative area address information in the map page; perform clustering processing on the marked address points in the above-mentioned map page to obtain an address point clustering map; remove the address points that do not meet the preset conditions in the above-mentioned address point clustering map to update the above-mentioned address point clustering map; determine the circumscribed rectangle of the updated address point clustering map; and in response to the diagonal length of the above-mentioned circumscribed rectangle being less than or equal to a preset length, determine the above-mentioned unique word as an area unique word.
[0130] Optionally, the generating unit 702 is further configured to: determine the word frequency of each first address word in the above-mentioned first address word sequence through the word frequency of the address words included in the above-mentioned address word information table to obtain a word frequency sequence; determine the sum of every two word frequencies included in the above-mentioned word frequency sequence as a word frequency threshold to obtain a word frequency threshold sequence; and perform merging processing on the first address words corresponding to each word frequency threshold greater than or equal to a preset threshold in the above-mentioned word frequency threshold sequence to generate co-occurring address words.
[0131] Optionally, the device 700 further includes: a storage unit configured to store the above-mentioned at least one co-occurring address word and the above-mentioned area unique word in a preset logistics address search server.
[0132] It can be understood that the units described in the device 700 correspond to the respective steps in the method described with reference to Figure 2 Therefore, the operations, features, and beneficial effects described above for the method also apply to the device 700 and the units included therein, and will not be elaborated herein.
[0133] Next, with reference to Figure 8 , which shows a schematic structural diagram of an electronic device (e.g., Figure 1 the computing device 101 in ) 800 suitable for implementing some embodiments of the present disclosure. The electronic devices in some embodiments of the present disclosure may include, but are not limited to, mobile terminals such as laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 8 The electronic device shown in is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0134] As Figure 8 shown, the electronic device 800 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 801, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 802 or the program loaded from the storage device 808 into the random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 are also stored. The processing device 801, the ROM 802, and the RAM 803 are connected to each other through the bus 804. The input / output (I / O) interface 805 is also connected to the bus 804.
[0135] Generally, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 8 the electronic device 800 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had. Figure 8 Each block shown in
[0136] particular, according to some embodiments of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the method shown in the flowchart. In such some embodiments, the computer program may be downloaded and installed from the network through the communication device 809, or installed from the storage device 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above functions defined in the method of some embodiments of the present disclosure are executed.
[0137] It should be noted that the computer-readable media described in some embodiments of the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0138] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed network.
[0139] The above computer-readable medium may be included in the above electronic device; or it may exist independently without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to: perform word segmentation on the target logistics address information to generate a first address word sequence and a second address word sequence; generate at least one co-occurring address word corresponding to the first address word sequence based on a preset address word information table and the first address word sequence; identify unique words in the second address word sequence based on a preset address word category combination table; and perform parsing processing on the target logistics address information based on the at least one co-occurring address word and the unique words.
[0140] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, execute as a stand-alone software package, execute partially on the user's computer and partially on a remote computer, or execute entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0141] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0142] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes a word segmentation unit, a generation unit, an identification unit, and an analysis unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases. For example, the identification unit can also be described as "the unit that identifies the unique word in the above-mentioned second address word sequence based on a preset address word category combination table".
[0143] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Product (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.
[0144] The above description is only some preferred embodiments of the present disclosure and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the embodiments of the present disclosure.
Claims
1. An address information parsing method, comprising: Performing word segmentation processing on the target logistics address information to generate a first address word sequence and a second address word sequence; Generating at least one co-occurring address word corresponding to the first address word sequence based on a preset address word information table and the first address word sequence; Identifying the unique words in the second address word sequence based on a preset address word category combination table, wherein the address words in the second address word sequence whose corresponding address word categories are the same as the address word categories included in the address word category combination are associated to generate associated address words, and the address points represented by the associated address words are searched for in the map page; in response to all the searched address points representing the same address, determining the associated address words as unique words; Performing parsing processing on the target logistics address information based on the at least one co-occurring address word and the unique words.
2. The method according to claim 1, wherein The performing word segmentation processing on the target logistics address information to generate a first address word sequence and a second address word sequence includes: Selecting, as candidate address words, the address words with a word frequency greater than or equal to an initial word frequency threshold from each address word included in the target logistics address information based on the word frequency of the address words included in the address word information table, to obtain a candidate address word group; Performing first word segmentation processing on the target logistics address information according to the candidate address word group to generate a first address word sequence.
3. The method according to claim 1, wherein, The performing word segmentation processing on the target logistics address information to generate a first address word sequence and a second address word sequence includes: Performing second word segmentation processing on the target logistics address information according to a preset address word category table to generate a second address word sequence.
4. The method according to claim 1, wherein, The identifying the unique words in the second address word sequence based on a preset address word category combination table includes: For each address word category combination in the address word category combination table, performing association processing on the corresponding second address words in the second address word sequence through the address word category combination to generate associated address words; For each generated associated address word, identifying the associated address word information corresponding to the associated address word from a preset associated address word information table as candidate associated address word information, to obtain a candidate associated address word information group, wherein the candidate associated address word information includes word frequency; Determining the associated address word corresponding to the candidate associated address word information with the maximum word frequency included in the candidate associated address word information group as the unique word in the second address word sequence.
5. The method according to claim 1, wherein The method further includes: Determining the regional address information including the unique word in a preset regional address information table as candidate regional address information, to obtain a candidate regional address information group; Determining the unique word as a regional unique word based on the candidate regional address information group.
6. The method according to claim 5, wherein, The determining the unique word as a regional unique word based on the candidate regional address information group includes: Marking the address points corresponding to each candidate regional address information in the candidate regional address information group in the map page; Performing clustering processing on the marked address points in the map page to obtain an address point clustering map; Remove the address points in the address point clustering map that do not meet the preset conditions to update the address point clustering map; Determine the bounding rectangle of the updated address point clustering map; In response to the diagonal length of the bounding rectangle being less than or equal to the preset length, determine the unique word as the region unique word.
7. The method according to claim 1, wherein The generating, based on the preset address word information table and the first address word sequence, at least one co-occurring address word corresponding to the first address word sequence includes: Determine the word frequency of each first address word in the first address word sequence through the word frequencies of the address words included in the address word information table to obtain a word frequency sequence; Determine the sum of every two word frequencies included in the word frequency sequence as the word frequency threshold to obtain a word frequency threshold sequence; Perform a merging process on the first address words corresponding to each word frequency threshold greater than or equal to the preset threshold in the word frequency threshold sequence to generate co-occurring address words.
8. The method according to claim 5, wherein The method further includes: Store the at least one co-occurring address word and the region unique word into a preset logistics address search server.
9. An address information parsing device, including: A word segmentation unit configured to perform word segmentation processing on target logistics address information to generate a first address word sequence and a second address word sequence; A generating unit configured to generate at least one co-occurring address word corresponding to the first address word sequence based on the preset address word information table and the first address word sequence; An identifying unit configured to identify the unique word in the second address word sequence based on the preset address word category combination table, where the address words in the second address word sequence whose corresponding address word categories are the same as the address word categories included in the address word category combination are associated to generate associated address words, and the address points represented by the associated address words are searched for in the map page; in response to all the found address points representing the same address, determine the associated address word as the unique word; An analysis unit configured to perform analysis processing on the target logistics address information based on the at least one co-occurring address word and the unique word.
10. An electronic device, including: One or more processors; A storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-8.
11. A computer-readable medium having a computer program stored thereon, wherein, The program, when executed by the processor, implements the method according to any one of claims 1-8.
Citation Information
Patent Citations
Address abbreviation generation method, model training method and related equipment
CN112488194A
Address recognition method and device, and storage medium
CN113609290A
Address information processing method and device, electronic equipment and computer readable medium
CN113722580A