Text address processing method and device
By recognizing and splitting text addresses and storing them in segments with preset address levels, the problems of redundant and unclear text addresses in the prior art are solved, efficient storage is achieved and the accuracy of geocoding and search is improved.
Patent Information
- Application Number
- CN202311704781.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-12
- Publication Date
- 2025-06-20
AI Technical Summary
In the prior art, text addresses with similar geographical coordinates have high overlap, resulting in storage redundancy and occupies unnecessary storage resources. At the same time, the geographical level of text addresses is not clear enough, affecting the accuracy of geocoding and search.
By recognizing and splitting the text address word segmentation, segmented text corresponding to the preset address level is obtained and segmented storage is performed to avoid the generation of redundant text, save storage resources, and improve the clarity of the address level and the accuracy of geocoding and search.
It realizes clear and accurate segmented text storage at the structural level, avoids the generation of redundant text, saves storage resources, and improves the accuracy of geocoding and search.
Smart Images

Figure CN120181079A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a method and apparatus for processing text addresses. Background Art
[0002] With the application of digital information technology in the fields of geocoding and search, more and more software tools provide services for converting text addresses into geographic coordinates. Users can perform address searches by entering text addresses, and merchants can also perform pre-sorting operations on traded items through text addresses. Currently, the common method for processing text addresses mainly involves performing simple standardization processing on the text addresses, storing the complete text addresses in disk space, and then using related technologies for geocoding and search.
[0003] In the process of implementing the present invention, the inventors found the following problems in the prior art:
[0004] For text addresses with similar geographic coordinate positions, the text overlap is very high. The existing storage method for complete text addresses will occupy a lot of disk space, there is a large amount of redundant text, and unnecessary storage resources are occupied; in addition, the geographic hierarchy of text addresses is not clear enough, which is not conducive to ensuring the accuracy of geocoding and search. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a method and apparatus for processing text addresses. Based on the segmented text obtained by word segmentation recognition of the text address, the segmented text is split to obtain segmented texts corresponding to a preset address hierarchy, and each segmented text is stored in segments. Embodiments of the present invention combine a preset address hierarchy based on the segmented text to obtain segmented texts with clear and accurate structural hierarchies, and store each segmented text in segments, avoiding the generation of redundant text, saving storage resources, and the segmented texts with clear and accurate structural hierarchies also improve the accuracy of geocoding and search.
[0006] To achieve the above object, according to one aspect of the embodiments of the present invention, a method for processing a text address is provided, including:
[0007] In response to receiving a processing request for a text address, perform word segmentation recognition on the text address, and according to the word segmentation recognition result, perform a first split on the text address to obtain segmented text;
[0008] According to a preset address hierarchy, perform hierarchy recognition on each of the segmented texts, and according to the result of the hierarchy recognition, perform a second split on each of the segmented texts to obtain segmented texts corresponding to the address hierarchy;
[0009] Store each of the segmented texts in segments.
[0010] Optionally, word segmentation recognition is performed on the text address, including: performing normalization processing on the text address to obtain a standard text address; using a word segmentation component to perform word segmentation recognition on the standard text address.
[0011] Optionally, the segmented text is an administrative division text or a keyword text, and the keyword text includes a regular recognition text, a first text to be distinguished, and a second text to be distinguished; according to a preset address level, level recognition is performed on each segmented text, including: in the case where the segmented text is an administrative division text, using a preset standard administrative division name to recognize the administrative counties and development zones in the segmented text; in the case where the segmented text is a keyword text, for the regular recognition text, according to the preset address level, determining a regular expression, and performing level recognition on the regular recognition text through the regular expression; for the first text to be distinguished, using a preset first distinction rule to perform distinction recognition between an area of interest and a point of interest on the first text to be distinguished; for the second text to be distinguished, using a preset second distinction rule to perform distinction recognition between a house number and a building number on the second text to be distinguished.
[0012] Optionally, the texts included in the segmented text are in order; using a preset first distinction rule to perform distinction recognition between an area of interest and a point of interest on the first text to be distinguished, including: matching the first text to be distinguished with a preset area of interest dictionary, and recognizing the text with a successful match as an area of interest; or, determining whether there is a flag text of an area of interest in the first text to be distinguished, and recognizing the text with the flag text as an area of interest; or, obtaining the next text after the first text to be distinguished, and in the case where the next text is recognized as a building number, recognizing the first text to be distinguished as an area of interest; in the case where the first text to be distinguished is not recognized as an area of interest by the above method, regarding the first text to be distinguished as a point of interest.
[0013] Optionally, the texts included in the segmented text are in order; using a preset second distinction rule to perform distinction recognition between a house number and a building number on the second text to be distinguished, including: obtaining the previous text of the second text to be distinguished; in the case where the previous text is a road name, recognizing the second text to be distinguished as a house number; in the case where the previous text is an area of interest, recognizing the second text field as a building number.
[0014] Optionally, the segmented text has a segmented sequence code; storing each segmented text separately, including: storing each segmented text according to the corresponding address level, and storing the association relationship between each segmented text according to the segmented sequence code.
[0015] Optionally, the text address has an address identifier; each segmented text is stored according to the corresponding address level, and encoded according to the segmentation order, and the association relationship between each segmented text is stored, including: for each segmented text, storing the segmented text into a data table at the corresponding address level, and recording the storage location identifier of the segmented text; storing the storage location identifier into an association relationship table according to the address identifier and the segmentation order encoding of the segmented text.
[0016] Optionally, before storing the segmented text into a data table at the corresponding address level, the method further includes: confirming that there is no such segmented text in the data table at the address level.
[0017] According to a second aspect of an embodiment of the present invention, there is provided a processing device for a text address, including:
[0018] A word segmentation text acquisition module, configured to, in response to receiving a processing request for a text address, perform word segmentation recognition on the text address, and perform a first split on the text address according to the word segmentation recognition result to obtain a word segmentation text;
[0019] A segmented text acquisition module, configured to perform level recognition on each of the word segmentation texts according to a preset address level, and perform a second split on each of the word segmentation texts according to the result of the level recognition to obtain a segmented text corresponding to the address level;
[0020] A segmented text storage module, configured to store each of the segmented texts in segments.
[0021] According to a third aspect of an embodiment of the present invention, there is provided a processing electronic device for a text address, including:
[0022] One or more processors;
[0023] A storage device, configured to store one or more programs,
[0024] When the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the first aspect of the embodiment of the present invention.
[0025] According to a fourth aspect of an embodiment of the present invention, there is provided a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, the method provided in the first aspect of the embodiment of the present invention is implemented.
[0026] One embodiment of the described invention has the following advantages or beneficial effects: By responding to a processing request for a text address, performing word segmentation recognition on the text address, and based on the word segmentation recognition result, performing a first split on the text address to obtain segmented text; according to a preset address hierarchy, performing hierarchy recognition on each segmented text, and based on the result of the hierarchy recognition, performing a second split on each segmented text to obtain segmented text corresponding to the address hierarchy; the technical solution of storing each segmented text separately realizes obtaining segmented text with a clear and accurate structural hierarchy based on the segmented text in combination with the preset address hierarchy, storing each segmented text separately, avoiding the generation of redundant text, saving storage resources, and the clear and accurate structural hierarchy of the address hierarchy also improves the accuracy of geocoding and searching. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The drawings are used to better understand the present invention and do not constitute an improper limitation to the present invention. Among them:
[0028] Figure 1 is a schematic diagram of the main process of the method for processing a text address according to an embodiment of the present invention;
[0029] Figure 2 is a schematic diagram of the storage principle of the data table of the segmented text according to an embodiment of the present invention;
[0030] Figure 3 is a schematic diagram of the main modules of the device for processing a text address according to an embodiment of the present invention;
[0031] Figure 4 is an exemplary system architecture diagram to which an embodiment of the present invention can be applied;
[0032] Figure 5 is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] It should be noted that in the technical solution of the present invention, in terms of the collection / collection, update, analysis, use, transmission, storage, etc. of the user's personal information, all comply with the provisions of relevant laws and regulations, are used for legal and reasonable purposes, are not shared, leaked or sold outside these legal uses, etc., and are subject to the supervision and management of the national regulatory authorities. Necessary measures should be taken for the user's personal information to selectively prevent the use or access to personal information data to prevent illegal access to such personal information data, ensure that the personnel with the right to access personal information data comply with the provisions of relevant laws and regulations, and ensure the security of the user's personal information. In addition, once these user personal information data are no longer needed, the risk should be minimized by restricting or even prohibiting data collection and / or deleting the data.
[0034] The following describes exemplary embodiments of the present invention with reference to the accompanying drawings. Various details of the embodiments of the present invention are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.
[0035] Currently, in the method for processing text addresses, for text addresses with similar geographical coordinate positions, the text overlap is very high. The existing storage method for complete text addresses will occupy a lot of disk space, there are a large number of redundant texts, occupying unnecessary storage resources. In addition, the geographical hierarchy of text addresses is not clear enough, which is not conducive to ensuring the accuracy of geocoding and searching, and cannot well meet the actual application.
[0036] To solve the above problems existing in the prior art, the present invention proposes a method for processing text addresses. Based on the segmented text obtained by segmenting and recognizing the text address, the segmented text is split to obtain segmented texts corresponding to the preset address hierarchy, and each segmented text is stored in segments. The embodiments of the present invention are based on the segmented text, combined with the preset address hierarchy, to obtain segmented texts with clear and accurate structural hierarchies, and store each segmented text in segments, avoiding the generation of redundant texts, saving storage resources, and the segmented texts with clear and accurate structural hierarchies also improve the accuracy of geocoding and searching.
[0037] In the introduction of the embodiments of the present invention, the nouns involved and their meanings are as follows:
[0038] Administrative division: It is the abbreviation of administrative region division for the convenience of administrative management;
[0039] CRF: Conditional Random Field, which is a mathematical algorithm based on a probabilistic graphical model following the Markov property. Combining the characteristics of the maximum entropy model and the hidden Markov model, it is an undirected graph model, and has achieved good results in sequence labeling tasks such as word segmentation, part-of-speech tagging, and named entity recognition in recent years;
[0040] AOI: Area of Interest, also called Information Area, referring to the regional geographical entities in map data;
[0041] POI: Point of Interest, which is a record of a location on the map that a person considers useful or interesting.
[0042] Figure 1 It is a schematic diagram of the main process of the method for processing text addresses according to the embodiments of the present invention. As Figure 1 shown, the method for processing text addresses according to the embodiments of the present invention includes the following steps S101 to step S103.
[0043] Step S101: In response to receiving a processing request for a text address, perform word segmentation recognition on the text address, and according to the word segmentation recognition result, perform a first split on the text address to obtain segmented text.
[0044] Specifically, geocoding and searching mainly rely on text addresses, and the standardization degree of text addresses will affect the accuracy of address coding and searching. Currently, text addresses are generally generated by means of text input or voice input. Considering the individual habit differences of the inputters, there may be problems such as duplicates and abbreviations of administrative divisions in text addresses. Therefore, it is necessary to process and process text addresses to make the text addresses hierarchical, so as to improve the accuracy of geocoding and searching.
[0045] According to an embodiment of the present invention, performing word segmentation recognition on the text address includes: performing standardization processing on the text address to obtain a standard text address; using a word segmentation component to perform word segmentation recognition on the standard text address.
[0046] Specifically, considering that text addresses may be non-standard due to individual expression habit differences, it is necessary to standardize text addresses, which mainly involves data completion and special symbol processing. For example, supplementing missing provincial names, removing special symbols such as \n`~!@#$%^& in the text address, converting traditional Chinese characters in the text address to simplified Chinese characters, and converting Chinese numerals to Arabic numerals. For the standardized standard text address, use the CRF word segmentation service component to perform word segmentation recognition, which can recognize four-level administrative divisions (province, city, district / county, town / street), specific keywords (community, orientation, business district, road number, unit building, etc.), and non-keyword postscripts (call by phone when arriving, at the door, etc.). Further, according to the word segmentation recognition result of the text address, perform a first split on the text address. It should be noted that: the non-keywords recognized by word segmentation have little value for geocoding and searching, and this part can be ignored and removed to obtain segmented text valuable for geocoding and searching.
[0047] Step S102: According to the preset address hierarchy, perform hierarchy recognition on each segmented text, and perform a second split on each segmented text according to the result of the hierarchy recognition to obtain segmented text corresponding to the address hierarchy.
[0048] Specifically, the embodiment of the present invention adopts an address hierarchy composed of a combination of numbers and letters according to the actual address characteristics, with a total of 6 major levels and 20 minor levels. The specific address hierarchy is shown in the following table:
[0049]
[0050]
[0051] According to the address hierarchy, the segmented text of the above text address is hierarchically recognized according to the address hierarchy, and more fine-grained address elements with a clear hierarchical structure can be obtained. Then, according to the recognized hierarchy, the second splitting is performed on each segmented text to obtain the segmented text corresponding to the above hierarchy.
[0052] According to an embodiment of the present invention, the segmented text is an administrative division text or a keyword text, and the keyword text includes a regular recognition text, a first text to be distinguished, and a second text to be distinguished; according to a preset address hierarchy, hierarchical recognition is performed on each segmented text, including: when the segmented text is an administrative division text, using a preset standard administrative division name to recognize the administrative districts and development zones in the segmented text; when the segmented text is a keyword text, for the regular recognition text, according to the preset address hierarchy, a regular expression is determined, and the regular recognition text is hierarchically recognized through the regular expression; for the first text to be distinguished, a preset first distinction rule is used to distinguish and recognize the area of interest and the point of interest of the first text to be distinguished; for the second text to be distinguished, a preset second distinction rule is used to distinguish and recognize the house number and building number of the second text to be distinguished.
[0053] Specifically, through the above-mentioned CRF word segmentation recognition, the administrative division and keywords can be recognized. Correspondingly, the first splitting of the text address results in two types of segmented texts: administrative division text and keyword text. When the segmented text is an administrative division text, the names of administrative districts and development zones / industrial areas are prone to confusion, and the administrative districts and development zones / industrial areas are distinguished by matching the administrative division text with the preset standard administrative division name.
[0054] When the segmented text is a keyword text, the keyword text includes text information such as villages, communities, groups, teams, communities / AOIs, road signs / POIs, house numbers, and building numbers in the address hierarchy of the embodiment of the present invention. Among these text information, villages, communities, groups, and teams can be recognized through regular expressions. The names of communities / AOIs and road signs / POIs are similar and need to be distinguished. House numbers and building numbers are both Arabic numerals and need to be distinguished. The text information of villages, communities, groups, and teams that can be recognized through regular expressions is used as the regular recognition text, the text information of communities / AOIs and road signs / POIs is used as the first text to be distinguished, and the text information of house numbers and building numbers is used as the second text to be distinguished. For the regular recognition text, the regular expressions of villages, communities, groups, and teams are determined according to the address hierarchy, and villages, communities, groups, and teams are recognized through the regular expressions. For the first text to be distinguished and the second text to be distinguished, the first distinction rule for distinguishing the area of interest AOI and the point of interest POI and the second distinction rule for distinguishing the house number and the building number are used to distinguish the area of interest AOI and the point of interest POI, and the house number and the building number.
[0055] According to another embodiment of the present invention, the text included in the segmented text is in order; using a preset first discrimination rule to perform discrimination and recognition of an area of interest and a point of interest on the first text to be discriminated, including: matching the first text to be discriminated with a preset area-of-interest dictionary, and identifying the text with a successful match as an area of interest; alternatively, determining whether there is a marker text of an area of interest in the first text to be discriminated, and identifying the text with the marker text as an area of interest; alternatively, obtaining the next text after the first text to be discriminated, and when the next text is identified as a building number, identifying the first text to be discriminated as an area of interest; in the case where the first text to be discriminated is not identified as an area of interest by the above method, using the first text to be discriminated as a point of interest.
[0056] Specifically, for the first text to be discriminated, define an area-of-interest AOI dictionary, which can include known area-of-interest texts, and maintain the newly added area-of-interest texts in the dictionary. Alternatively, define texts ending with shopping centers, shopping malls, squares, medicine, etc. as area-of-interest AOIs. Match the first text to be discriminated with the area-of-interest dictionary, and identify the text with a successful match as an area of interest. If the match is unsuccessful, it can be identified according to the marker text. In the embodiment of the present invention, the first text to be discriminated maintains the order of the text address. Specifically, it can be determined whether the first text to be discriminated contains "letters / data / direction words" and ends with "district / phase / lane", such as XXX Phase III, XXX District V, etc. If it conforms to the marker text of this area of interest, it is identified as an area of interest. Additionally, according to the dependency relationship between the area of interest and the building number, the next text after the first text to be discriminated can be obtained. If the next text is identified as a building number, it also indicates that the first text to be discriminated is an area of interest. In the case where all the above identifications of the area of interest fail, it indicates that the first text to be discriminated is a point of interest, thus realizing the discrimination between the area of interest and the point of interest. It should be noted that the specific identification steps of the above three identification methods for the area of interest are not limited. It can be identified first through the area-of-interest dictionary as described above, then by searching for the marker text, and finally by the building number. It can also be identified in parallel by multiple threads, or first by searching for the marker text, then through the area-of-interest dictionary, and finally by the building number.
[0057] According to still another embodiment of the present invention, the text included in the segmented text is in order; using a preset second discrimination rule to perform discrimination and recognition of a house number and a building number on the second text to be discriminated, including: obtaining the previous text of the second text to be discriminated; when the previous text is a road name, identifying the second text to be discriminated as a house number; when the previous text is an area of interest, identifying the second text field as a building number.
[0058] Specifically, the segmented text maintains the same order as the text address, and the text included in the segmented text is arranged in an orderly manner according to the text address. The distinction between the house number and the building number is mainly determined by the previous text of the second text to be distinguished. According to the obtained previous text, it is determined whether it is a house number or a building number. If the previous text is the road name, correspondingly, the second text to be distinguished should be the house number; if the previous text is the area of interest, as mentioned above, the area of interest and the building number generally coexist, so the second text to be distinguished field is the building number. In addition, for the description of "2-1-1501", the embodiment of the present invention uses regular matching to identify it as the indoor level, specifically the building number-unit number-room number.
[0059] Further, according to the above hierarchical recognition result, the administrative division text and the keyword text are secondarily split into segmented texts corresponding to the address levels, and fine-grained segmented texts with hierarchical relationships are obtained.
[0060] Step S103, store each of the segmented texts in segments.
[0061] Specifically, according to the address levels of the embodiments of the present invention, storage units corresponding to the address levels and having hierarchical association relationships are set, and each segmented text is stored in segments in the corresponding storage units.
[0062] According to an embodiment of the present invention, the segmented text has a segmented sequence code; storing each of the segmented texts in segments includes: storing each of the segmented texts according to the corresponding address levels, and storing the association relationships between each of the segmented texts according to the segmented sequence code.
[0063] Specifically, in the embodiments of the present invention, two types of storage tables are set up. One type is the data table for storing segmented text, and the other type is the association table for storing the association relationships between segmented texts. The data table for storing text can be defined as multiple data tables according to the address hierarchy. For example, provinces, cities, districts / counties can be defined as one data table, and road boundaries, buildings, and interiors can be each defined as one data table, so a total of four data tables are used to store text. Additionally, in the embodiments of the present invention, in order to further improve the accuracy of geocoding and searching, for address hierarchies with ranges such as provinces, cities, districts / counties, and road boundaries, corresponding geographic information databases are also pre-established to store the current boundary geographic information of each province, city, district / county, and road boundary. For example, the specific longitude and latitude information of the geographic range of a county. Considering that the boundary geographic information may sometimes change due to other factors, in order to ensure that the text addresses involving the boundaries of provinces, cities, districts, and road boundaries can be accurate in subsequent geocoding and searching, the id of the latest boundary geographic information in the geographic information database for the segmented text is stored in the storage table of the segmented text. Further, because it is the storage of segmented text, it is necessary to record the association relationships between segmented texts. According to the segmentation order coding of each segmented text, the association relationships between segmented texts are stored in the association table.
[0064] According to another embodiment of the present invention, the text address has an address identifier; storing each of the segmented texts according to the corresponding address hierarchy and storing the association relationships between the segmented texts according to the segmentation order coding includes: for each of the segmented texts, storing the segmented text in the data table of the corresponding address hierarchy and recording the storage location identifier of the segmented text; storing the storage location identifier in the association table according to the address identifier and the segmentation order coding of the segmented text.
[0065] Specifically, for each segmented text, according to the defined data table of the address hierarchy, the segmented text is stored in the data table of the corresponding address hierarchy, and the storage location identifier of the segmented text in the data table is recorded. For the association table, according to the address identifier of the text address, a storage location for the association relationship of the text address is established in the association table, and then according to the segmentation order coding of the segmented text, the specific location where the storage location identifier of the segmented text is stored in the association table is determined. For example, if the segmentation order coding is 2, the storage location identifier is stored in the second position of the storage location of the text address in the association table. This facilitates subsequent reading. According to the address identifier of the text address, according to the association order stored in the association table, each segmented text is read from the data table in sequence to form a complete and standard text address.
[0066] According to another embodiment of the present invention, before storing the segmented text into a data table at the corresponding address level, the method further includes: confirming that the segmented text does not exist in the data table at the address level.
[0067] Specifically, considering that there is overlapping text content between text addresses with similar geographical coordinate positions. For example, the segmented text corresponding to text address one includes Province A, City B, and District C, and the segmented text corresponding to text address two includes Province A, City B, and District D. Both text addresses have Province A and City B. In this case, the segmented storage of the segmented text in the embodiment of the present invention only stores Province A and City B once, which avoids storing duplicate Province A and City B. Therefore, before storing the segmented text into the data table at the corresponding address level, query whether the segmented text has already been stored in the data table. If not, the segmented text can be stored. If it already exists, there is no need to store the segmented text, and only the storage location identifier of the existing segmented text needs to be recorded.
[0068] Figure 2 It is a schematic diagram of the data table storage principle of the segmented text in the embodiment of the present invention. This figure is illustrated by taking the storage of the data table of the segmented text as an example and does not involve the association relation table. The national standard fence, AOI data, road network data, and business district data in the figure are the above-mentioned geographical information databases, which store the current boundary geographical information of each province, city, district / county, and road boundary. Taking the text address of "XX Building, XXX Store Name, XXXX Branch Name, 8th Floor, No. n8001, 39th Wujie Street, Ding Street, District C, City B, Province A" as an example, the text address is split according to the result of hierarchical recognition. According to the above-mentioned preset address levels, the segmented text is obtained: Province A corresponds to level 1, City B corresponds to level 2, District C corresponds to level 3A, Ding Street corresponds to level 4A, Wujie Street corresponds to level 4E-1, No. 39 corresponds to level 5A-1, XX Building corresponds to level 5B, XXX Store Name corresponds to level 5B, XXXX Branch Name corresponds to level 5B, 8th Floor corresponds to level 6B, and No. n8001 corresponds to level 6C. Search for the ids of the boundary geographical information of Province A, City B, District C, and Ding Street in the national standard fence and road network data, and store the ids corresponding to Province A, City B, District C, and Ding Street and the segmented text into the corresponding province scope, city scope, administrative county, and township street in the data table; for the segmented texts of No. 39, XX Building, XXX Store Name, XXXX Branch Name, 8th Floor, and No. n8001, which do not involve boundary geographical information, they can be directly stored into the corresponding house number, landmark / POI, floor, and room number in the data table according to the address levels they belong to.
[0069] Through the processing method of the text address in the embodiment of the present invention, the area of interest and the point of interest, as well as the house number and building number, can be effectively distinguished, providing address information with clear text structure and distinct segmentation for geocoding and search; and through the segmented storage of the segmented text, storage resources are saved and redundant text is avoided.
[0070] Figure 3 It is a schematic diagram of the main modules of a text address processing device according to an embodiment of the present invention. As Figure 3 shown, the text address processing device 300 mainly includes a segmented text acquisition module 301, a segmented text acquisition module 302, and a segmented text storage module 303.
[0071] The segmented text acquisition module 301 is configured to, in response to receiving a text address processing request, perform word segmentation recognition on the text address, and perform a first split on the text address according to the word segmentation recognition result to obtain segmented text;
[0072] The segmented text acquisition module 302 is configured to perform hierarchical recognition on each of the segmented texts according to a preset address hierarchy, and perform a second split on each of the segmented texts according to the result of the hierarchical recognition to obtain segmented text corresponding to the address hierarchy;
[0073] The segmented text storage module 303 is configured to store each of the segmented texts in segments.
[0074] According to an embodiment of the present invention, the segmented text acquisition module 301 is further configured to: perform normalization processing on the text address to obtain a standard text address; use a word segmentation component to perform word segmentation recognition on the standard text address.
[0075] According to another embodiment of the present invention, the segmented text is an administrative division text or a keyword text, and the keyword text includes a regular recognition text, a first text to be distinguished, and a second text to be distinguished; the segmented text acquisition module 302 is further configured to: in the case where the segmented text is an administrative division text, use a preset standard administrative division name to recognize the administrative districts and development zones in the segmented text; in the case where the segmented text is a keyword text, for the regular recognition text, determine a regular expression according to a preset address hierarchy, and perform hierarchical recognition on the regular recognition text through the regular expression; for the first text to be distinguished, use a preset first distinction rule to perform distinction recognition on the interest surface and interest point of the first text to be distinguished; for the second text to be distinguished, use a preset second distinction rule to perform distinction recognition on the house number and building number of the second text to be distinguished.
[0076] According to still another embodiment of the present invention, the text included in the segmented text is in order; the segmented text acquisition module 302 is further configured to: match the first text to be distinguished with a preset dictionary of areas of interest, and identify the text with a successful match as an area of interest; or, determine whether there is a flag text of an area of interest in the first text to be distinguished, and identify the text with the flag text as an area of interest; or, obtain the next text of the first text to be distinguished, and when the next text is identified as a building number, identify the first text to be distinguished as an area of interest; in the case where the first text to be distinguished is not identified as an area of interest by the above method, use the first text to be distinguished as a point of interest.
[0077] According to yet another embodiment of the present invention, the text included in the segmented text is in order; the segmented text acquisition module 302 is further configured to: obtain the previous text of the second text to be distinguished; when the previous text is a road name, identify the second text to be distinguished as a house number; when the previous text is an area of interest, identify the second text field as a building number.
[0078] According to another embodiment of the present invention, the segmented text has a segmented sequence code; the segmented text storage module 303 is further configured to: store each of the segmented texts according to the corresponding address level, and store the association relationship between each of the segmented texts according to the segmented sequence code.
[0079] According to still another embodiment of the present invention, the text address has an address identifier; the segmented text storage module 303 is further configured to: for each of the segmented texts, store the segmented text into a data table of the corresponding address level, and record the storage location identifier of the segmented text; according to the address identifier and the segmented sequence code of the segmented text, store the storage location identifier into an association relationship table.
[0080] According to yet another embodiment of the present invention, the processing device 300 of the text address further includes a segmented text confirmation module (not shown in the figure), configured to: before storing the segmented text into a data table of the corresponding address level, confirm that there is no such segmented text in the data table of the address level.
[0081] Figure 4 is an exemplary system architecture diagram to which embodiments of the present invention can be applied.
[0082] As Figure 4As shown, the system architecture 400 may include terminal devices 401, 402, 403, a network 404, and a server 405. The network 404 serves as a medium for providing communication links between the terminal devices 401, 402, 403 and the server 405. The network 404 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0083] Users can use the terminal devices 401, 402, 403 to interact with the server 405 via the network 404 to receive or send messages, etc. Various communication client applications, such as a processing application for text addresses (only as an example), may be installed on the terminal devices 401, 402, 403.
[0084] The terminal devices 401, 402, 403 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0085] The server 405 can be a server that provides various services, such as a background management server (only as an example) that supports text addresses used by users with the terminal devices 401, 402, 403. The background management server can, in response to receiving a processing request for a text address, perform word segmentation recognition on the text address, perform a first split on the text address according to the word segmentation recognition result to obtain segmented text; perform hierarchical recognition on each of the segmented texts according to a preset address hierarchy, and perform a second split on each of the segmented texts according to the result of the hierarchical recognition to obtain segmented texts corresponding to the address hierarchy; perform processing such as storing each of the segmented texts separately, and feedback the processing result (such as the segmented text, etc. - only as an example) to the terminal device.
[0086] It should be noted that the method for processing text addresses provided in the embodiments of the present invention is generally executed by the server 405. Correspondingly, the device for processing text addresses is generally set in the server 405.
[0087] It should be understood that Figure 4 the numbers of terminal devices, networks, and servers in
[0088] are merely illustrative. According to implementation requirements, there can be any number of terminal devices, networks, and servers.
[0088] Next, referring to Figure 5 which shows a schematic structural diagram of a computer system of a terminal device or a server suitable for implementing the embodiments of the present invention. Figure 5 The shown terminal device or server is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.
[0089] As in Figure 5As shown, computer system 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from storage section 508 into random access memory (RAM) 503. In RAM 503, various programs and data required for the operation of system 500 are also stored. CPU 501, ROM 502, and RAM 503 are connected to each other via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.
[0090] The following components are connected to I / O interface 505: input section 506 including a keyboard, a mouse, etc.; output section 507 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; storage section 508 including a hard disk, etc.; and communication section 509 including a network interface card such as a LAN card, a modem, etc. Communication section 509 performs communication processing via a network such as the Internet. Drive 510 is also connected to I / O interface 505 as required. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on drive 510 as required so that a computer program read from it can be installed into storage section 508 as required.
[0091] Specifically, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via communication section 509, and / or installed from removable medium 511. When the computer program is executed by central processing unit (CPU) 501, the above-mentioned functions defined in the system of the present invention are executed.
[0092] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0093] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0094] The units involved in the embodiments of the present invention can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes a word segmentation text acquisition module, a segmented text acquisition module, and a segmented text storage module.
[0095] Among them, the names of these modules do not constitute a limitation on the modules themselves in some cases. For example, the segmented text storage module can also be described as "a module for storing each of the segmented texts separately".
[0096] On the other hand, the present invention also provides a computer-readable medium, which can be included in the device described in the embodiments; or it can exist separately without being assembled into the device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the device, the device includes: in response to receiving a processing request for a text address, performing word segmentation recognition on the text address, and according to the word segmentation recognition result, performing a first split on the text address to obtain word segmentation text; according to a preset address hierarchy, performing hierarchy recognition on each of the word segmentation texts, and according to the result of the hierarchy recognition, performing a second split on each of the word segmentation texts to obtain segmented texts corresponding to the address hierarchy; storing each of the segmented texts separately.
[0097] According to the technical solution of the embodiments of the present invention, the following advantages or beneficial effects are achieved: by in response to receiving a processing request for a text address, performing word segmentation recognition on the text address, and according to the word segmentation recognition result, performing a first split on the text address to obtain word segmentation text; according to a preset address hierarchy, performing hierarchy recognition on each of the word segmentation texts, and according to the result of the hierarchy recognition, performing a second split on each of the word segmentation texts to obtain segmented texts corresponding to the address hierarchy; storing each of the segmented texts separately, it is realized to obtain segmented texts with clear and accurate structural levels based on the word segmentation text in combination with the preset address hierarchy, and store each of the segmented texts separately, avoiding the generation of redundant texts, saving storage resources, and the clear and accurate address hierarchy of the structural level also improves the accuracy of geocoding and searching.
[0098] The specific implementation manners do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for processing a text address, characterized in that, Including: In response to receiving a processing request for a text address, perform word segmentation recognition on the text address, and according to the word segmentation recognition result, perform a first split on the text address to obtain segmented texts; According to a preset address hierarchy, perform hierarchy recognition on each of the segmented texts, and according to the result of the hierarchy recognition, perform a second split on each of the segmented texts to obtain segmented texts corresponding to the address hierarchy; Store each of the segmented texts in segments.
2. The method according to claim 1, characterized in that, Performing word segmentation recognition on the text address includes: Perform normalization processing on the text address to obtain a standard text address; Use a word segmentation component to perform word segmentation recognition on the standard text address.
3. The method according to claim 1, characterized in that, The segmented text is an administrative division text or a keyword text, and the keyword text includes a regular recognition text, a first text to be distinguished, and a second text to be distinguished; According to a preset address hierarchy, performing hierarchy recognition on each of the segmented texts includes: In the case where the segmented text is an administrative division text, use a preset standard administrative division name to recognize the administrative counties and development zones in the segmented text; In the case where the segmented text is a keyword text, for the regular recognition text, according to a preset address hierarchy, determine a regular expression, and perform hierarchy recognition on the regular recognition text through the regular expression; for the first text to be distinguished, use a preset first distinction rule to perform distinction recognition of an area of interest and a point of interest on the first text to be distinguished; for the second text to be distinguished, use a preset second distinction rule to perform distinction recognition of a house number and a building number on the second text to be distinguished.
4. The method according to claim 3, characterized in that, The texts included in the segmented text are in order; Using a preset first distinction rule to perform distinction recognition of an area of interest and a point of interest on the first text to be distinguished includes: Match the first text to be distinguished with a preset area of interest dictionary, and recognize the text with a successful match as an area of interest; Alternatively, determine whether there is a flag text of an area of interest in the first text to be distinguished, and recognize the text with the flag text as an area of interest; Alternatively, obtain the next text after the first text to be distinguished, and in the case where the next text is recognized as a building number, recognize the first text to be distinguished as an area of interest; In the case where the first text to be distinguished is not recognized as an area of interest by the above method, use the first text to be distinguished as a point of interest.
5. The method according to claim 3, characterized in that, The texts included in the segmented text are in order; Using a preset second distinction rule to perform distinction recognition of a house number and a building number on the second text to be distinguished includes: Obtain the previous text of the second text to be distinguished; In the case where the previous text is a road name, recognize the second text to be distinguished as a house number; In the case where the previous text is an area of interest, recognize the second text field to be distinguished as a building number.
6. The method according to claim 1, characterized in that, The segmented text has a segmented sequence code; Storing each of the segmented texts in segments includes: Store each of the segmented texts according to the corresponding address hierarchy, and store the association relationship between each of the segmented texts according to the segmented sequence code.
7. The method according to claim 6, characterized in that, The text address has an address identifier; Each of the segmented texts is stored according to the corresponding address level, encoded according to the segmentation order, and the association relationship between the segmented texts is stored, including: For each of the segmented texts, the segmented text is stored in a data table at the corresponding address level, and the storage location identifier of the segmented text is recorded; according to the address identifier and the segmentation order encoding of the segmented text, the storage location identifier is stored in the association relationship table.
8. The method according to claim 7, characterized in that, Before storing the segmented text in the data table at the corresponding address level, the method further includes: Confirming that the segmented text is not in the data table at the address level.
9. A device for processing a text address, characterized in that, Including: A segmented text acquisition module, configured to, in response to receiving a processing request for a text address, perform word segmentation recognition on the text address, and perform a first split on the text address according to the word segmentation recognition result to obtain a segmented text; A segmented text acquisition module, configured to perform level recognition on each of the segmented texts according to a preset address level, and perform a second split on each of the segmented texts according to the result of the level recognition to obtain a segmented text corresponding to the address level; A segmented text storage module, configured to store each of the segmented texts in segments.
10. A mobile electronic device terminal, characterized in that, Including: One or more processors; A storage device, configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-8.
11. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the method according to any one of claims 1-8 is implemented.