Information processing device, information processing method, and information processing program

By tokenizing address information and predicting similarity, the problems of tokenizing address information in the prior art are solved, and the geocoding effect with high accuracy is achieved.

JP7678906B1Active Publication Date: 2025-05-16RAKUTEN GROUP INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024013999
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-02-01
Publication Date
2025-05-16
Estimated Expiration
2044-02-01

AI Technical Summary

Technical Problem

The prior art has complexity and inaccuracy in the tokenization process of address information, especially when tokenization based on incorrect address information, the accuracy of geocoding is difficult to guarantee.

Method used

The target address information is converted into a plurality of tokens corresponding to hierarchical address classification by generating units to generate a token sequence; the prediction unit predicts the similarity between the token sequence and the address information in the address information list; the decision unit determines its corresponding latitude and longitude information from the database based on the matching address information with the highest similarity.

Benefits of technology

It realizes high accuracy when geocoding using tokenized address information, can effectively handle the wrong address information, and ensure the accuracy of geocoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007678906000001_ABST
    Figure 0007678906000001_ABST
Patent Text Reader

Abstract

Geocoding is performed with high accuracy using tokens tokenized from address information. [Solution] An information processing device tokenizes target address information into multiple tokens corresponding to hierarchical address categories to generate a token string, predicts the similarity between the token string and address information included in a predetermined address information list, identifies address information from the list that has the highest similarity to the token string as matching address information, and determines the latitude and longitude information corresponding to the matching address information from a database that stores the list and latitude and longitude information corresponding to the address information included in the list as the latitude and longitude information of the target address information.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to techniques for geocoding. [Background technology]

[0002] Conventionally, geocoding processing technology has been developed to convert address information into latitude and longitude information. For example, Patent Document 1 discloses an address / latitude and longitude conversion device that converts address information into latitude and longitude information at high speed in real time by using a tree structure hierarchically constructed by administrative districts.

[0003] Specifically, the address / latitude and longitude conversion device in this document converts address information into a tree structure, generates address tree information, and stores it in memory. Here, the address tree information is composed of the root node at the first level (top level), prefectures at the second level, wards or cities, towns, villages at the third level, cities, towns, villages or wards at the fourth level, and house addresses (block codes and house numbers) or parcel numbers at the fifth level and below. Furthermore, the address tree information is configured to associate latitude and longitude information with the nodes at the lowest level. This makes it possible to efficiently search for latitude and longitude information for given address information. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] JP 2016-091315 A Summary of the Invention [Problem to be solved by the invention]

[0005] In this way, geocoding has conventionally been performed using character strings representing administrative districts such as prefectures and cities, towns, and villages that are separated from address information, that is, character strings (tokens) that are tokenized from address information. Here, when tokenization is performed manually, not only is the work cumbersome, but the accuracy of tokenization can vary depending on knowledge about addresses. In addition, when tokenization is performed based on incorrect address information, the address information does not actually exist, so geocoding may not be performed accurately.

[0006] The present invention has been made in consideration of the above-mentioned problems, and has an object to provide a technique for performing geocoding with high accuracy using tokens tokenized from address information. [Means for solving the problem]

[0007] In order to solve the above problem, one aspect of an information processing device according to the present invention has a generation unit that tokenizes target address information into a plurality of tokens corresponding to hierarchical address categories to generate a token sequence; a prediction unit that predicts the similarity between address information included in a list of predetermined address information and the token sequence; an identification unit that identifies address information from the list that has the highest similarity to the token sequence as matching address information; and a determination unit that determines, from a database that stores the list and latitude and longitude information corresponding to the matching address information, as the latitude and longitude information of the target address information.

[0008] In order to solve the above problem, one aspect of an information processing method according to the present invention includes tokenizing target address information into a plurality of tokens corresponding to hierarchical address categories to generate a token sequence, predicting a similarity between address information included in a list of predetermined address information and the token sequence, identifying address information from the list that has the highest similarity to the token sequence as matching address information, and determining, from a database that stores the list and latitude and longitude information corresponding to the address information included in the list, the latitude and longitude information corresponding to the matching address information as the latitude and longitude information of the target address information.

[0009] In order to solve the above problem, one aspect of an information processing program according to the present invention is an information processing program for causing a computer to execute information processing, the program causing the computer to execute processes including a generation process for tokenizing target address information into a plurality of tokens corresponding to hierarchical address categories to generate a token sequence, a prediction process for predicting the similarity between address information included in a list of predetermined address information and the token sequence, an identification process for identifying address information from the list that has the highest similarity to the token sequence as matching address information, and a determination process for determining, from a database in which the list and latitude and longitude information corresponding to the address information included in the list are stored, the latitude and longitude information corresponding to the matching address information as the latitude and longitude information of the target address information. Effect of the Invention

[0010] According to the present invention, it is possible to perform geocoding with high accuracy using tokens generated from address information. The above-mentioned objects, aspects, and advantages of the present invention, as well as objects, aspects, and advantages of the present invention not described above, will be understood by those skilled in the art from the following detailed description of the invention by referring to the accompanying drawings and the claims. [Brief description of the drawings]

[0011] [Figure 1]FIG. 1 illustrates an example of a functional configuration of an information processing device according to an embodiment. [Diagram 2] FIG. 2 shows an example of the contents stored in the address point database. [Figure 3A] FIG. 3A shows an example of a token string generated from address information. [Figure 3B] FIG. 3B shows a conceptual diagram of a procedure for generating a token string using a tokenization model. [Figure 4] FIG. 4 illustrates an example of a hardware configuration of an information processing device according to an embodiment. [Diagram 5] FIG. 5 is a flowchart of the geocoding process executed by the information processing device. [Figure 6] FIG. 6 shows an example of a token string generated from address information according to a modified example. [Figure 7] FIG. 7 shows an example of the first address information and the second address information according to a modified example. [Figure 8] FIG. 8 shows an example of an address text token string and a code token string according to a modified example. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0012] Hereinafter, with reference to the attached drawings, an embodiment for carrying out the present invention will be described in detail. Among the components disclosed below, those having the same functions are given the same reference numerals, and their description will be omitted. Note that the embodiment disclosed below is an example of a means for realizing the present invention, and should be appropriately modified or changed depending on the configuration of the device to which the present invention is applied and various conditions, and the present invention is not limited to the following embodiment. Furthermore, not all of the combinations of features described in the present embodiment are necessarily essential to the solution of the present invention.

[0013] [Example of functional configuration of information processing device] 1 shows an example of the functional configuration of an information processing device 10 according to this embodiment. As an example of the functional configuration, the information processing device 10 has an address information acquisition unit 101, a token string generation unit 102, a similarity prediction unit 103, a matching address identification unit 104, a geocoding unit 105, and a learning model storage unit 110. The learning model storage unit 110 is configured to be able to store a tokenization model 111 and a similarity prediction model 112.

[0014] The information processing device 10 may be, for example, a device such as a desktop PC (Personal Computer), a notebook PC, a tablet, or a general-purpose machine. The entire information processing device 10 may not be provided in one device, but may be provided in a plurality of devices. For example, at least a part of the information processing device 10 may be realized in an external server or a virtual server built in a cloud server. In this case, the functions shown in this embodiment are realized by the cooperation between the information processing device 10 and the server.

[0015] The information processing device 10 is configured to be able to communicate with an external address point database 120. In Fig. 1, the address point database 120 is arranged separately from the information processing device 10, but the information processing device 10 may be configured to include the address point database 120. The address point database 120 is a database that stores an address information list, which is a list of address information that actually exists, and latitude and longitude information (latitude information and longitude information) corresponding to the address information included in the list.

[0016] An example of the contents stored in the address point database 120 is shown in FIG. 2. As shown in FIG. 2, the address point database 120 stores an address information list 20, which is a list of address information that actually exists, and a latitude and longitude list 21, which is a list of latitude and longitude information corresponding to the address information included in the list. In this embodiment, the address information is text information (character string) such as "1-2-3, D-cho, ABC-ku, Tokyo", and has a hierarchical structure arranged geographically in order from a large division (higher) to a small division (lower). In this embodiment, each division from the large division to the small division is called an address division. The address division in the address information can be, in order from the top (top), any division of a city, a prefecture, a city, a ward, a county, or a village (or a division equivalent to a city, a ward, a county, or a village; the same applies below), a division of a town or a village (or a division equivalent to a town or a village; the same applies below), or a division of a code representing at least one of a block code and a house number. Therefore, the address information "1-2-3, D-cho, ABC-ku, Tokyo" is composed of the address categories, from the top down, of "Tokyo" (any of the following categories: city, ward, county, or village), "D-cho" (town or area category), and "1-2-3" (code category).

[0017] Furthermore, the address information may include other address categories following the code category. For example, the address information may include a building category including a building name following the code category. For example, in the case of address information of "XY Building, 1-2-3, D-cho, ABC-ku, Tokyo", "XY Building" corresponds to the building category. In addition, in the case of address information of "XY Mansion, Room 203, 1-2-3, D-cho, ABC-ku, Tokyo", "XY Mansion" and "Room 203" (or "XY Mansion, Room 203") correspond to the building category.

[0018] The address information included in the address information list 20 is configured so that each address section is distinguished (identified). For example, the address information included in the address information list 20 may have a space inserted between each address section. In FIG. 2, in "1-2-3, D-cho, ABC-ku, Tokyo," a space is inserted between each address section. That is, spaces are inserted between "Tokyo" and "ABC-ku," between "ABC-ku" and "D-cho," and between "D-cho" and "1-2-3." In addition, the actual address information is changed (for example, increased or decreased) according to land readjustment by the government, construction of buildings, etc. Therefore, the address information list 20 and the latitude and longitude list 21 in the address point database 120 can be updated at regular intervals.

[0019] The address information acquisition unit 101 acquires address information for which latitude and longitude information (latitude information and longitude information) is to be acquired (i.e., which is to be geocoded) from input information to the address information acquisition unit 101. The address information acquisition unit 101 acquires address information from input information input by an operator (user), for example. Alternatively, the address information acquisition unit 101 may accept input information set in advance in the information processing device 10 according to a predetermined program, and acquire address information.

[0020] The address information acquired by the address information acquisition unit 101 is configured with hierarchical address divisions, similar to the address information included in the address information list 20. The address divisions in the address information can be, in order from the top (top), any one of the divisions of a city, a district, a county, or a village, a town or village division, and a code division representing at least one of a block code and a house number. Thus, the address information of "1-2-3, ABC Ward, Tokyo" is configured with the address divisions of "Tokyo" (any one of the divisions of a city, a district, a county, or a village), "D Town" (town or village division), and "1-2-3" (code division) from the top. "1-2-3" represents "1-chome, 2-3-go", and can also be other notations representing "1-chome, 2-3-go", such as "1-chome, 2-3". The address information may also include other address divisions, such as a building division including a building name following the code division.

[0021] When the input information to the address information acquisition unit 101 is composed of only address information, the address information acquisition unit 101 acquires the input information itself as address information. When the input information is composed of address information and other information such as image information, the address information acquisition unit 101 may extract and acquire the address information from the input information. When the input information is audio information including address information, the address information acquisition unit 101 converts the audio information into text information (i.e., transcribes it) and acquires the converted information as address information. The conversion from audio information to text information can be performed by using a known voice recognition function or the like.

[0022] The address information acquiring unit 101 may be configured to acquire address information with a division of a city, a prefecture, or a prefecture added thereto when acquiring address information that does not include any division of a city, a prefecture, or a prefecture at the beginning. For example, when the address information acquiring unit 101 acquires address information of "ABC Ward, D Town 1-2-3" from input information, the address information acquiring unit 101 may acquire address information of "Tokyo" by adding "Tokyo" to the beginning. This may be performed on a rule basis using a predetermined lookup table or the like. For example, the address information acquiring unit 101 may inquire into a predetermined lookup table or the address information list 20 about divisions corresponding to city, ward, county, and village and / or divisions of town or block in the address information, acquire any division of a city, a prefecture, or a prefecture, and add it to the address information.

[0023] The token string generating unit 102 functions as a tokenizer, and tokenizes (divides) the address information acquired by the address information acquiring unit 101 into a plurality of tokens corresponding to hierarchical address divisions to generate a token string. In this embodiment, the token string generating unit 102 generates a token string using a tokenization model 111 stored in the learning model storage unit 110. The tokenization model 111 is a machine learning model trained for natural language processing. For example, the tokenization model 111 may be a model based on character-by-character sequence labeling using Bi-LSTM (Bidirectional-Long Short Term Memory) such as nagisa. In addition, the tokenization model 111 may be a model that searches for a context (morpheme path) that maximizes or minimizes the sum of a language model score for a virtual token string (morpheme) candidate based on an RNNLM (Recurrent Neural Network Language Model) such as Juman++ and a feature score including a concatenation cost and an occurrence cost of the token string candidate. Furthermore, the tokenization model 111 may be a model that divides an input character string (corresponding to address information) such as a SentencePiece into subwords, and may be a model that allows reversible division of an input character string.

[0024] The token string generation unit 102 may generate a token string from address information on a rule basis (using a dictionary) using a predetermined lookup table or the like. For example, the token string generation unit 102 may be a tokenizer that searches for a context (morpheme path) that maximizes or minimizes the concatenation cost and occurrence cost of a token string candidate, such as MeCab. The token string generation unit 102 may also be a tokenizer having multiple morpheme units, such as Sudachi. The token string generation unit 102 may also be a tokenizer to which a predetermined dictionary or a predetermined corpus, such as IPAdic or NEologd, has been added or reflected. When a tokenizer such as Sudachi is used, the token string generating unit 102 may, for example, select a short word as a morpheme unit (for example, address division units such as "Tokyo", "ABC ward", "D town", and "1-2-3"), or may select a morpheme unit in which an affix, a compound verb, a compound noun, or an idiomatic phrase is combined with a short word (for example, a single address information unit such as "1-2-3 D town, ABC ward, Tokyo"). The token string generating unit 102 may generate a token string by combining the above-mentioned models and configurations of the tokenizer.

[0025] 3A shows an example of a token string generated from address information according to this embodiment. In this example, the token string generating unit 102 tokenizes address information 30, "abc Corporation, 1-2-3, D Town, ABC Ward, Tokyo," into multiple tokens corresponding to hierarchical address divisions, and generates a token string 31. Specifically, the token string generating unit 102 tokenizes into "Tokyo," "ABC Ward," "D Town," "1-2-3," "Inc.", and "abc," which are multiple tokens corresponding to any of the divisions of city, prefecture, or prefecture (tokens of city, prefecture, or prefecture), division of city, ward, county, or village (corresponding to tokens of city, ward, county, or village), division of town or village (corresponding to tokens of town or village), division of code representing at least one of block code and house number (corresponding to token of code), and division of building (corresponding to token of building name), and generates a token string 31 arranged in this order. "Tokyo" is a token for a city, a district, a prefecture, "ABC Ward" is a token for a city, a ward, a county, a village, "D Town" is a token for a town or a village, "1-2-3" is a code token, and "Inc." and "abc" are tokens for the building name. Depending on the address information, two or more tokens for a city, a ward, a county, or a village and / or a town or a village may be included.

[0026] 3B shows a conceptual diagram of a procedure for generating a token string using the tokenization model 111. The token string generation unit 102 generates a token string 31 by inputting address information 30 to the tokenization model 111. When the address information 30 is input, the tokenization model 111 is configured to output the address information 30 as a plurality of tokens and the position of each of the plurality of tokens in the address information 30. The token string generation unit 102 generates the token string 31 from the plurality of tokens and the position of each of the tokens output from the tokenization model 111. The tokenization model 111 is generated by a learning unit (not shown) of the information processing device 10 or another device learning a model using teacher data, and is stored in the learning model storage unit 110.

[0027] The reference address information 321 is actually existing address information starting from any one of the address divisions of metropolis, prefecture, or city. The reference token string 322 is a token string consisting of multiple reference tokens that are accurately tokenized (divided) from the reference address information 321 in accordance with the hierarchical address divisions. In the example of FIG. 3B, the reference token string 322 is composed of "Tokyo", "CD Ward", "E Town", and "1-2-3" from the beginning. When the reference address information 321 includes the name of a building (building division), the reference token string 322 has one or more tokens according to the building division. The position information 323 is information that indicates the position of each of the multiple reference tokens in the reference token string 322 in the reference address information 321. In the example of FIG. 3B, the position information 323 is set to "0" for the position of the first character (i.e., the leftmost character) of the first token in the reference token string 322, and includes the first index and the last index for each token that are counted for each character. For example, "Tokyo" corresponds to the first to third characters in reference token sequence 322, and location information 323 is {0,2}. Similarly, in reference token sequence 322, "CD-ku" corresponds to the fourth to sixth characters, and location information 323 is {3,5}, "E-machi" corresponds to the seventh to eighth characters, and location information 323 is {6,7}, and "1-2-3" corresponds to the ninth to thirteenth characters, and location information 323 is {8,12}.

[0028] When the token string generation unit 102 generates a token string, the similarity prediction unit 103 predicts the similarity between the token string and address information included in the address information list 20. As described above, the address information included in the address information list 20 is configured so that each address section is distinguished (identified), and the similarity prediction unit 103 predicts the similarity between one token string and one piece of address information from each token in the token string and each address section of the address information included in the address information list 20.

[0029] In this embodiment, the similarity prediction unit 103 predicts (derives) the similarity using the similarity prediction model 112 stored in the learning model storage unit 110. The similarity prediction model 112 is a machine-learned natural language processing model, and is configured to predict and output the similarity (relationship) between the first data and the second data when the first data and the second data are input. In this embodiment, the first data is a token string, and the second data is any address information included in the address information list 20. The similarity output from the similarity prediction model 112 is expressed, for example, by a numerical value (also referred to as a similarity score). The smaller (or higher) the numerical value, the higher the similarity. The similarity prediction model 112 may be a trained model capable of evaluating the similarity between a plurality of token strings, or may be a trained model capable of evaluating the similarity between each of the plurality of token strings and the corresponding embedding vector (embedded expression, feature expression). The similarity prediction model 112 is, for example, a Siamese network / Siamese network and / or a Transformer-based natural language processing model. In addition, the similarity prediction model 112 may be, for example, a Transformer-based natural language processing model in which a pre-trained BERT (Bidirectional Encoder Representations from Transformers) and a network corresponding to a classifier connected to the BERT are fine-tuned.

[0030] When the Siamese network is used for the similarity prediction model 112, the first data and the second data are first coded in parallel and converted into feature representations. Then, the similarity between the first data and the second data is evaluated using the two converted feature representations by a distance function. The similarity is evaluated by, for example, Euclidean distance or cosine similarity.

[0031] The similarity prediction unit 103 may predict the similarity of the token string with all address information included in the address information list 20, but in that case, the calculation process becomes enormous. Therefore, the similarity prediction unit 103 may select multiple candidates (hereinafter also referred to as matching address information candidates) from the address information list 20 based on the token string (or the address information acquired by the address information acquisition unit 101) and derive the similarity with the matching address information candidates. For example, the similarity prediction unit 103 selects, from the address information list 20, multiple address information having the same division as the token string's division of any one of prefectures, cities, wards, counties, and villages, or division of towns or villages, as matching address information candidates. Then, the similarity prediction unit 103 may predict the similarity between the matching address information candidates and the token string. In this manner, the similarity prediction unit 103 predicts similarities between a plurality of pieces of address information included in the address information list 20 and a token string, and obtains a plurality of similarities.

[0032] The matching address identifying unit 104 uses the multiple similarities obtained by the similarity prediction unit 103 to identify, from the address information list 20, address information that is most likely to match (is deemed to match) the token string as the matching address information. That is, the matching address identifying unit 104 identifies the address information that most closely matches the token string as the matching address information. For example, the matching address identifying unit 104 identifies the address information with the highest similarity as the matching address information. When there are multiple pieces of address information with the highest similarity, the matching address identifying unit 104 may identify any one of the address information as the matching address information based on, for example, a predetermined rule.

[0033] The geocoding unit 105 queries the address point database 120 for the matching address information identified by the matching address identification unit 104, thereby identifying latitude and longitude information corresponding to the matching address information. Referring to Fig. 2, the geocoding unit 105 searches the latitude and longitude list 21 for and identifies latitude and longitude information corresponding to the matching address information in the address information list 20. The geocoding unit 105 determines the identified latitude and longitude information as the latitude and longitude information of the address information acquired by the address information acquisition unit 101. The geocoding unit 105 can output the determined latitude and longitude information to the outside.

[0034] [Hardware configuration of information processing device] 4 is a block diagram showing an example of a hardware configuration of the information processing device 10 according to the present embodiment. The information processing device 10 can be implemented on a single or multiple computers, mobile devices, or any other processing platform. 4, the information processing device 10 is illustrated as being implemented in a single computer, but the information processing device 10 according to the present embodiment may be implemented in a computer system including multiple computers. The multiple computers may be connected to each other via a wired or wireless network so as to be able to communicate with each other.

[0035] 4, the information processing device 10 may include a CPU (Central Processing Unit) 41, a ROM (Read Only Memory) 42, a RAM (Random Access Memory) 43, a HDD (Hard Disk Drive) 44, an input unit 45, a display unit 46, a communication I / F 47, and a system bus 48. The information processing device 10 may also include an external memory. The CPU 41 generally controls the operation of the information processing device 10, and controls each of the components (42 to 47) via a system bus 48, which is a data transmission path.

[0036] The ROM 42 is a non-volatile memory that stores control programs and the like necessary for the CPU 41 to execute processing. Note that the programs may be stored in a non-volatile memory such as the HDD 44 or an SSD (Solid State Drive) or an external memory such as a removable storage medium (not shown). The RAM 43 is a volatile memory and functions as the main memory, work area, etc. of the CPU 41. That is, when executing a process, the CPU 41 loads necessary programs, etc. from the ROM 42 into the RAM 43 and executes the programs, etc. to realize various functional operations. The learning model storage unit 110 and the address point database 120 (when the address point database 120 is provided in the information processing device 10) shown in FIG. 1 can be configured in the RAM 43.

[0037] The HDD 44 stores, for example, various data and various information required when the CPU 41 performs processing using a program. The HDD 44 also stores, for example, various data and various information obtained when the CPU 41 performs processing using a program. The input unit 45 is composed of a keyboard and a pointing device such as a mouse. The display unit 46 is configured with a monitor such as a liquid crystal display (LCD), etc. The display unit 46 may be configured in combination with the input unit 45 to function as a GUI (Graphical User Interface).

[0038] The communication I / F 47 is an interface that controls communication between the information processing device 10 and an external device. The communication I / F 47 provides an interface with a network and executes communication with the external device via the network. Various data, various parameters, and the like are transmitted and received between the information processing device 10 and the external device via the communication I / F 47. In this embodiment, the communication I / F 47 may execute communication via a wired LAN (Local Area Network) or a dedicated line that complies with a communication standard such as Ethernet (registered trademark). However, the network that can be used in this embodiment is not limited to this, and may be configured as a wireless network. This wireless network includes wireless PANs (Personal Area Networks) such as Bluetooth (registered trademark), ZigBee (registered trademark), and UWB (Ultra Wide Band). It also includes wireless LANs (Local Area Networks) such as Wi-Fi (Wireless Fidelity) (registered trademark), and wireless MANs (Metropolitan Area Networks) such as WiMAX (registered trademark). It also includes wireless WANs (Wide Area Networks) such as 4G and 5G defined by 3GPP (Third Generation Partnership Project) (registered trademark). It should be noted that the network is sufficient as long as it can connect the devices to each other so that they can communicate with each other, and the communication standard, scale, and configuration are not limited to those described above.

[0039] At least some of the functions of the information processing device 10 shown in Fig. 1 can be realized by the CPU 21 in the information processing device 10 executing a program. However, at least some of the functions of the information processing device 10 shown in Fig. 1 may be operated as dedicated hardware. In this case, the dedicated hardware may operate under the control of the CPU 41.

[0040] [Geocoding process flow] Next, the geocoding process according to this embodiment will be described with reference to Fig. 5. Fig. 5 is a flowchart of the geocoding process executed by the information processing device 10. The process shown in Fig. 5 can be performed by the CPU 41 executing a control program stored in the information processing device 10.

[0041] In S51, the address information acquisition unit 101 acquires address information to be geocoded (hereinafter referred to as target address information). The address information acquisition unit 101 can acquire the target address information from input information by an operator (user) or a predetermined program. As described above, the target address information is composed of a plurality of hierarchical address divisions. The address divisions may include, but are not limited to, any division of prefecture, city, ward, county, or village, town or village division, and a code division representing at least one of a block code and a house number.

[0042] In S52, the token string generating unit 102 generates a token string by tokenizing the target address information into a plurality of tokens corresponding to hierarchical address classifications. As described with reference to FIG. 3A, the token string generating unit 102 generates a token string (e.g., the token string 31) from the target address information (e.g., the address information 30). As described with reference to FIG. 3B, the token string generating unit 102 can generate a token string using the tokenization model 111 stored in the learning model storage unit 110. Specifically, the token string generating unit 102 inputs the target address information to the tokenization model 111 and acquires a plurality of tokens and the position of each of the plurality of tokens in the target address information. Then, the token string generating unit 102 generates a token string from the plurality of tokens and the position of each of the tokens. Alternatively, the token string generating unit 102 can generate a token string on a rule basis.

[0043] In S53, the similarity prediction unit 103 predicts the similarity between the address information included in the address information list 20 in the address point database 120 and the token string generated in S52. As described above, in this embodiment, the similarity prediction unit 103 predicts the similarity between the address information included in the address information list 20 and the token string using the similarity prediction model 112 stored in the learning model storage unit 110, and obtains multiple similarities.

[0044] In S54, the matching address identification unit 104 uses the multiple similarities acquired by the similarity prediction unit 103 to identify, as the matching address information, the address information that is most likely to match (is deemed to match) the token string generated in S52 from the address information included in the address information list 20. For example, the matching address identification unit 104 identifies, as the matching address information, the address information with the highest similarity acquired in S53 from among the address information included in the address information list 20.

[0045] In S55, the geocoding unit 105 determines the latitude and longitude information corresponding to the matching address information identified in S54 from the address point database 120 as the latitude and longitude information of the target address information. The geocoding unit 105 may output the determined latitude and longitude information to the outside. The output may be any output process, and may be output (distribution) to an external device via a communication I / F (communication I / F 27 in FIG. 4) or display output on the display unit 46.

[0046] In this way, the information processing device 10 according to the present embodiment divides the target address information into a plurality of tokens corresponding to hierarchical address divisions to generate a token string, and identifies address information having the highest similarity to the token string from the address information list 20. Then, the information processing device 10 determines the latitude and longitude information of the target address information by inquiring the address point database 120 about the identified address information (matching address information). Here, when the tokenization model 111 is used for generating the token string, it is possible to generate a token string faster and with higher accuracy than when a token string is generated manually. In addition, by predicting the similarity between the token string and the address information included in the address information list 20 using the similarity prediction model 112, it is possible to predict the similarity between the token string and the address information included in the address information list 20 with higher accuracy even for a token string composed of a large number of tokens or a token string including a token with some characters including an error. Then, the matching address information identified from the predicted similarity in this way is address information that actually exists. Even if the target address information is at least partially incorrect, the address information that is most likely to match the target address information (is deemed to match) is identified as the matching address information. The matching address information is converted to latitude and longitude information using the address point database 120. This series of processes enables geocoding with high accuracy.

[0047] In this embodiment, the processes of the similarity prediction unit 103 and the matching address identification unit 104 are performed separately, but the matching address identification unit 104 may be configured to perform the process of the similarity prediction unit 103. For example, the matching address identification unit 104 may predict the similarity between the address information included in the address information list 20 and the target address information, and identify the address information with the highest similarity from the address information list 20 as the matching address information.

[0048] [Variation 1] In the above embodiment, the token string generating unit 102 generates a token string using all of the address information. That is, referring to FIG. 3A, the token string generating unit 102 generates a token string 31 consisting of "Tokyo", "ABC Ward", "D Town", "1-2-3", "Inc.", and "abc" from address information 30 of "abc Corporation, 1-2-3 D Town, ABC Ward, Tokyo". In this modification, the token string generating unit 102 deletes (omits) one or more tokens at the rear of the multiple tokens included in the generated token string to generate a corrected token string (incomplete token string). That is, the address information has a hierarchical structure arranged in order from a major division (upper) to a minor division (lower), and the token string generating unit 102 generates a corrected token string in which a minor division at a later stage in the token string has been deleted.

[0049] FIG. 6 shows an example of a token string generated from address information according to this modification. As in FIG. 3A, the address information 30 is "abc Corporation, 1-2-3, D Town, ABC Ward, Tokyo," and the token string generation unit 102 generates a token string 31 including multiple tokens from the address information 30 according to hierarchical address classification. Furthermore, the token string generation unit 102 deletes one or more tokens at the end of the multiple tokens included in the token string 31 to generate a corrected token string 60. In this modification, the one or more tokens at the end are one or more tokens after the building classification. Therefore, in the example of FIG. 6, the token string generation unit 102 deletes one or more tokens after the building classification in the token string 31 to generate a corrected token string 60. That is, the token string generation unit 102 deletes two tokens, "abc Corporation" and "abc," from the token string 31 to generate a corrected token string 60 (a token string consisting of "Tokyo," "BC Ward," "D Town," and "1-2-3"). In addition, when the modified token string is generated by deleting one or more tokens following the building division, if the token string does not include a token of the building division, the modified token string does not need to be generated.

[0050] After the token string generation unit 102 generates the corrected token string, the information processing device 10 performs the same process as in the above embodiment using the corrected token string. That is, the similarity prediction unit 103 predicts the similarity between the address information included in the address information list 20 and the corrected token string. The procedure for predicting the similarity is the same as in the above embodiment, and is performed using the similarity prediction model 112 stored in the learning model storage unit 110, and multiple similarities are obtained. The matching address identification unit 104 uses the multiple similarities obtained by the similarity prediction unit 103 to identify address information from the address information list 20 that is most likely to match the corrected token string (can be considered to match) as matching address information. The geocoding unit 105 queries the address point database 120 for the matching address information identified by the matching address identification unit 104, and determines the latitude and longitude information of the matching address information as the latitude and longitude information of the target address information.

[0051] In this way, by identifying matching address information using the correction token string and identifying latitude and longitude information from the identified matching address information, it is possible to obtain highly accurate latitude and longitude information for the target address information even if the building name or information after the building name is incorrect in the target address information. Also, if incomplete address information in which the address after the building name has been deleted is stored in the address list 20 in the address point database 120, an address with a higher similarity can be identified as the matching address information by using the correction token string.

[0052] [Variation 2] As Modification 2, an embodiment will be described in which a token string generated from address information and a corrected token string generated from the token string according to Modification 1 are used to identify matching address information from the address information list 20. The token string generation unit 102 generates a token string and a corrected token string from the address information acquired by the address information acquisition unit 101. With reference to FIG. 6, the token string generation unit 102 generates a token string 31 and a corrected token string 60.

[0053] Next, the similarity prediction unit 103 predicts and obtains a plurality of similarities for each of the token string and the corrected token string. The procedure for predicting the similarity is as described in the above embodiment. Next, the matching address identification unit 104 selects, from the address information list 20, address information having the highest similarity to the token string as the first address information. Also, the matching address identification unit 104 selects, from the address information list 20, address information having the highest similarity to the corrected token string as the second address information. FIG. 7 shows an example of the first address information and the second address information based on FIG. 6. In FIG. 7, the first address information 70 is address information having the highest similarity to the token string 31 among the address information included in the address information list 20. The second address information 71 is address information having the highest similarity to the corrected token string 60 among the address information included in the address information list 20.

[0054] Then, the matching address identification unit 104 compares the first address information with the second address information, and if the first address information and the second address information match, identifies one of the address information as the matching address information. On the other hand, if the first address information and the second address information are different, the matching address identification unit 104 identifies the second address information, i.e., the address information having the highest similarity to the corrected token string, as the matching address information. Alternatively, if the first address information and the second address information are different, the matching address identification unit 104 may identify the first address information, i.e., the address information having the highest similarity to the token string, as the matching address information. After that, the geocoding unit 105 determines the latitude and longitude information of the matching address information as the latitude and longitude information of the target address information by inquiring the address point database 120 about the matching address information identified by the matching address identification unit 104.

[0055] Alternatively, when the first address information and the second address information are different, the user may specify either one of the address information as the matching address information. For this purpose, for example, when the first address information and the second address information are different, the matching address specification unit 104 may output the first address information and the second address information to the outside so that the user can specify the matching address information. As one form of output, the matching address specification unit 104 can output the target address information, the first address information, and the second address information to the display unit 46, and specify the first address information or the second address information selected based on the user's operation as the matching address information. Thereafter, the geocoding unit 105 determines the latitude and longitude information of the matching address information as the latitude and longitude information of the target address information by inquiring the address point database 120 about the matching address information specified by the matching address specification unit 104.

[0056] In this manner, in this modification, the first address information and the second address information are identified from the address information list 20 using the similarity degrees obtained for the token string and the corrected token string, respectively, and matching address information is identified from the first address information and the second address information. When the first address information and the second address information are different, either address information may be identified as the matching address information, or it may be identified by the user. By using two types of similarity degrees, matching address information can be identified with high accuracy from the target address information.

[0057] Furthermore, when the first address information and the second address information are different and the matching address information is specified by the user, the similarity prediction unit 103 may predict the similarity based on the first address information or the second address information that is specified more frequently. For example, when the number of times the second address information is specified is greater than the number of times the first address information is specified among the predetermined number of times that the matching address information is specified by the user among the first address information and the second address information, it can be said that the reliability of the second address information is high. Therefore, in this case, for the new target address information, the similarity prediction unit 103 may predict the similarity using the corrected token sequence instead of the token sequence.

[0058] [Variation 3] In the above embodiment, the similarity prediction unit 103 predicted the similarity between the entire token string or corrected token string and the address information included in the address information list 20. Alternatively, the similarity prediction unit 103 may divide the token string and predict the similarity for each divided token. The following explanation takes a token string as an example, but the same explanation can be applied to the corrected token string.

[0059] As a first example of this modification, the similarity prediction unit 103 divides the token string into an address text token string including multiple tokens before the code token and a code token string including one or more tokens after the code token. FIG. 8 shows an example of an address text token string and a code token string according to this modification. As in FIG. 3A, the token string 31 is composed of "Tokyo", "BC Ward", "D Town", "1-2-3", "Company", and "abc". In the token string 31, the code token is "1-2-3". Therefore, the similarity prediction unit 103 divides the token string 31 into an address text token string 80 before "1-2-3" (a token string consisting of "Tokyo", "BC Ward", and "D Town") and a code token string 81 after "1-2-3" (a token string consisting of "1-2-3", "Company", and "abc"). When the corrected token string 60 shown in FIG. 6 is used, the similarity prediction unit 103 divides the corrected token string 60 into an address text token string before "1-2-3" and a code token consisting of "1-2-3".

[0060] The similarity prediction unit 103 predicts the similarity between the address information included in the address information list 20 and the address text token string as the address text similarity. Furthermore, the similarity prediction unit 103 predicts the similarity between the address information included in the address information list 20 and the code token string as the code similarity. Here, since the address information included in the address information list 20 is configured so that each address division is distinguished (identified), the address text similarity and the code similarity may be predicted by dividing the address information included in the address information list 20 into address text information before the code division and address code information after the code division. In other words, the similarity prediction unit 103 may predict the address text similarity between the address text information and the address text token string, and the code similarity between the address code information and the code token string.

[0061] Next, the matching address identifying unit 104 identifies matching address information based on the address text similarity and the code similarity from the address information list 20. Here, the matching address identifying unit 104 may identify address information with the highest text similarity and code similarity as the matching address information from the address information list 20. Alternatively, the matching address identifying unit 104 may select a predetermined number of address information with high text similarity from the address information list 20, and identify address information with the highest code similarity to the selected address information as the matching address information.

[0062] As a second example of this modified example, the similarity prediction unit 103 may predict the similarity for each token included in the token string. For example, the similarity measurement unit 103 may use a token string 31 consisting of "Tokyo", "ABC Ward", "D Town", "1-2-3", "Inc.", and "abc" as shown in FIG. 3A, and predict and acquire the similarity for each token in this order (hereinafter referred to as "token similarity"). In this case, the matching address identification unit 104 may identify matching address information based on the token similarity. For example, when the lower the numerical value (similarity score), the higher the similarity, the matching address identification unit 104 may identify the address information with the lowest total value of the token similarity as the matching address information.

[0063] As described above, according to this modification, the token sequence (or the corrected token sequence) generated by the token sequence generating unit 102 is divided, and the similarity with the address information included in the address information list 20 is predicted for each divided unit. This makes it possible to evaluate the similarity in smaller units, and to identify matching address information from the target address information with high accuracy.

[0064] Although the above describes certain embodiments, the embodiments are merely illustrative and are not intended to limit the scope of the present invention. The apparatus and method described in this specification can be embodied in forms other than those described above. Furthermore, the above-described embodiments can be appropriately omitted, substituted, and modified without departing from the scope of the present invention. Such omitted, substituted, and modified forms are included in the scope of the claims and their equivalents, and belong to the technical scope of the present invention.

[0065] The disclosure of this embodiment includes the following configuration. [1] An information processing device having: a generation unit that generates a token sequence by tokenizing target address information into a plurality of tokens corresponding to hierarchical address classifications; a prediction unit that predicts the similarity between address information included in a list of predetermined address information and the token sequence; an identification unit that identifies address information from the list that has the highest similarity to the token sequence as matching address information; and a determination unit that determines, from a database that stores the list and latitude and longitude information corresponding to the address information included in the list, the latitude and longitude information corresponding to the matching address information as the latitude and longitude information of the target address information.

[0066] [2] The information processing device according to [1], wherein the generation unit generates the token sequence by inputting the target address information into a machine learning model trained for natural language processing.

[0067] [3] The information processing device described in [2], wherein the machine learning model is configured to output the multiple tokens and the position of each of the multiple tokens in the target address information from the target address information, and the generation unit generates the token sequence from the multiple tokens and the position of each token.

[0068] [4] The information processing device according to any one of [1] to [3], wherein the prediction unit predicts the similarity between the address information included in the list and the token sequence using a trained machine learning model.

[0069] [5] The information processing device described in [4], wherein the prediction unit selects multiple matching address information candidates from the list based on the token sequence and predicts the similarity between the matching address information candidates and the token sequence, and the identification unit identifies the candidate among the matching address information candidates that has the highest similarity to the token sequence as the matching address information.

[0070] [6] An information processing device according to any one of [1] to [5], wherein the token string includes, from the beginning, tokens for prefectures, cities, wards, counties and villages, tokens for towns or villages, and tokens for codes representing at least one of block codes and house numbers.

[0071] [7] The information processing device described in [6], wherein, when the token string includes a building name token following the code token, the generation unit generates a corrected token string by deleting one or more tokens following the building name token from the token string, the prediction unit predicts the similarity between address information included in the list and the corrected token string, and the identification unit identifies address information from the list that has the highest similarity to the corrected token string as the matching address information.

[0072] [8] The information processing device described in [7], further comprising a first selection unit that selects from the list address information having the highest similarity to the token sequence as first address information, and a second selection unit that selects from the list address information having the highest similarity to the corrected token sequence as second address information, wherein the identification unit identifies the second address information as the matching address information when the first address information and the second address information are different.

[0073] [9] The information processing device described in [7] further comprises a first selection unit that selects from the list address information having the highest similarity to the token sequence as first address information, and a second selection unit that selects from the list address information having the highest similarity to the corrected token sequence as second address information, wherein the identification unit identifies the first address information or the second address information selected by a user as the matching address information when the first address information and the second address information are different.

[0074]

[10] The information processing device described in [6], further comprising a splitting unit that splits the token string into an address text token string including multiple tokens before the code token and a code token string including one or more tokens after the code token, wherein the prediction unit predicts the similarity between each address information in the list and the address text token string and the code token as address text similarity and code similarity, respectively, and the identification unit identifies the matching address information from the list based on the address text similarity and the code similarity.

[0075]

[11] The information processing device according to

[10] , wherein the identification unit identifies, from the list, address information having the highest address text similarity and the highest code similarity as the matching address information. [Explanation of symbols]

[0076] 10: information processing device, 101: address information acquisition unit, 102: token string generation unit, 103: similarity prediction unit, 104: matching address identification unit, 105: geocoding unit, 110: learning model storage unit, 111: tokenization model, 112: similarity prediction model, 120: address point database

Claims

1. a generation unit that generates a token string by tokenizing the target address information into a plurality of tokens corresponding to hierarchical address classifications; a prediction unit that predicts a similarity between address information included in a list of predetermined address information and the token sequence; an identification unit that identifies, from the list, address information having the highest similarity to the token string as matching address information; a determination unit that determines, from a database in which the list and latitude and longitude information corresponding to the address information included in the list are stored, the latitude and longitude information corresponding to the matched address information as the latitude and longitude information of the target address information; An information processing device having the above configuration.

2. The information processing device according to claim 1 , wherein the generation unit generates the token sequence by inputting the target address information into a machine learning model trained for natural language processing.

3. The machine learning model is configured to output, from the target address information, the plurality of tokens and a position of each of the plurality of tokens in the target address information; the generating unit generates the token sequence from the plurality of tokens and the positions of each of the tokens. The information processing device according to claim 2 .

4. The information processing device according to claim 1 , wherein the prediction unit predicts a similarity between the address information included in the list and the token sequence by using a trained machine learning model.

5. the prediction unit selects a plurality of matching address information candidates from the list based on the token sequence, and predicts a similarity between the matching address information candidates and the token sequence; The information processing apparatus according to claim 4 , wherein the identifying unit identifies, from among the matching address information candidates, a candidate having a highest similarity to the token string as the matching address information.

6. 2. The information processing device according to claim 1, wherein the token string includes, from the beginning, tokens for prefectures, cities, wards, counties, and villages, tokens for towns or villages, and tokens for codes representing at least one of block codes and house numbers.

7. When the token string includes a building name token following the code token, the generating unit generates a corrected token string by deleting one or more tokens following the building name token from the token string, The prediction unit predicts a similarity between address information included in the list and the corrected token sequence, The information processing apparatus according to claim 6 , wherein the identifying unit identifies, from the list, address information having the highest similarity to the corrected token string as the matching address information.

8. a first selection unit that selects, from the list, address information having the highest similarity to the token sequence as first address information; a second selection unit that selects, from the list, address information having the highest similarity to the corrected token sequence as second address information, The information processing apparatus according to claim 7 , wherein the identifying unit identifies the second address information as the matching address information when the first address information and the second address information are different.

9. a first selection unit that selects, from the list, address information having the highest similarity to the token sequence as first address information; a second selection unit that selects, from the list, address information having the highest similarity to the corrected token sequence as second address information, The information processing device according to claim 7, wherein the identification unit identifies the first address information or the second address information selected by a user as the matching address information when the first address information and the second address information are different.

10. a splitting unit that splits the token string into an address text token string including a plurality of tokens preceding the code token and a code token string including one or more tokens following the code token, the prediction unit predicts a similarity between each piece of address information in the list and the address text token string and a similarity between each piece of address information in the list and the code token as an address text similarity and a code similarity, respectively; The information processing apparatus according to claim 6 , wherein the identifying unit identifies the matching address information from the list based on the address text similarity and the code similarity.

11. The information processing apparatus according to claim 10 , wherein the identifying unit identifies, from the list, address information having the highest address text similarity and the highest code similarity as the matching address information.

12. An information processing method executed by an information processing device, tokenizing the target address information into a plurality of tokens corresponding to hierarchical address segments to generate a token sequence; predicting a similarity between address information included in a list of predetermined address information and the token sequence; identifying, from the list, address information having the highest similarity to the token sequence as matching address information; determining, from a database in which the list and latitude and longitude information corresponding to the address information included in the list are stored, the latitude and longitude information corresponding to the matched address information as the latitude and longitude information of the target address information; An information processing method comprising:

13. An information processing program for causing a computer to execute information processing, the program including: A generation process of tokenizing the target address information into a plurality of tokens corresponding to hierarchical address classifications to generate a token string; A prediction process for predicting a similarity between address information included in a list of predetermined address information and the token sequence; a process of identifying, from the list, address information having the highest similarity to the token sequence as matching address information; and a determination process for determining, from a database in which the list and latitude and longitude information corresponding to the matching address information are stored, as the latitude and longitude information of the target address information. Information processing program.

Citation Information

Patent Citations

  • Chinese address element analysis method and device based on vocabulary enhancement and storage medium

    CN114792091A

  • Query parsing for map search

    JP2012532388A

  • Method and apparatus for providing address geo-coding

    US20140330865A1

  • System and method for prediction of geo-coordinates for a geographical element

    US20220065654A1

  • Address-latitude / longitude conversion apparatus and geographical information system using the same

    JP2016091315A