Address data processing method and geocoding method
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-07
- Publication Date
- 2026-08-11
AI Technical Summary
[0003]现有技术中,地址文本数据包括多样化描述的文本数据,导致获取的地址文本数据的准确率偏低
本申请提供的地址数据处理方法,根据样本目标地址数据和样本目标地址数据对应的多个样本候选地址数据之间分别对应的关联级别数据,在多个样本候选地址数据中选择属于不同关联级别数据的多个指定样本候选地址数据,并获得多个指定样本候选地址数据对应的样本标签数据。根据多个指定样本候选地址数据与样本标签数据,构建样本目标地址数据的偏序关系样本数据。而且,偏序关系样本数据用于对初始模型进行训练,以得到用于分析待查询地址数据与召回地址数据之间的目标关联指标数据的地址关联指标模型。该方法对样本数据进行处理,获得针对样本目标地址数据的偏序关系样本数据,以及偏序关系样本数据包含的多个指定样本候选地址数据与样本目标地址数据之间的关联指标数据,从而获得与样本目标地址数据的关联指标数据大于预设关联指标数据的匹配样本候选地址数据。通过偏序关系样本数据对初始模型进行训练,可以使得训练后的地址关联指标模型在待查询地址数据的多个召回地址数据中预测的目标召回地址数据与待查询地址数据之间的关联指标数据大于预设关联指标数据,提升模型获得的目标召回地址数据的准确率。
Smart Images

Figure CN121478832B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of geographic information technology, specifically to an address data processing method, apparatus, electronic device, and computer storage medium. This application also relates to a geocoding method, apparatus, electronic device, and computer storage medium. Background Technology
[0002] Geocoding technology converts textual address data into geographic coordinate data. It is crucial in location-based services; for example, in logistics, delivery resources require address lookup devices to obtain the corresponding geographic coordinates of configured addresses, thus providing navigation information to the delivery location. Advances in artificial intelligence and big data technologies can further enhance the accuracy of obtaining accurate address text data, thereby improving the accuracy of determining corresponding geographic coordinates in location-based services.
[0003] In existing technologies, address text data includes diverse descriptive text data, resulting in low accuracy of the obtained address text data.
[0004] Therefore, improving the accuracy of the obtained address text data is a technical problem that needs to be solved. Summary of the Invention
[0005] This application provides an address data processing method to improve the accuracy of acquired address text data. This application also provides an address data processing apparatus, electronic device, and computer storage medium. This application further provides a geocoding method, apparatus, electronic device, and computer storage medium.
[0006] The specific plan is as follows: In a first aspect, embodiments of this application provide an address data processing method, comprising: acquiring sample target address data and multiple sample candidate address data corresponding to the sample target address data; selecting multiple specified sample candidate address data belonging to different association levels from the multiple sample candidate address data according to the association level data corresponding to the sample target address data and the multiple sample candidate address data, and obtaining sample label data corresponding to the multiple specified sample candidate address data, wherein the sample label data is used to characterize the matching sample candidate address data in the multiple specified sample candidate address data whose sample association index data with the sample target address data is greater than a preset association index data; constructing partial order relation sample data of the sample target address data according to the multiple specified sample candidate address data and the sample label data; wherein the partial order relation sample data is used to train an initial model to obtain an address association index model for analyzing the target association index data between the query address data and the recalled address data.
[0007] Secondly, embodiments of this application provide a geocoding method, comprising: obtaining query address data and recall address data corresponding to the query address data; providing the query address data and the recall address data to an address association index model trained using partial order relation sample data, and obtaining target association index data output by the address association index model for the relationship between the query address data and the recall address data; sorting the multiple recall address data according to the target association index data between the multiple recall address data and the query address data; and using the geographic coordinate data corresponding to the recall address data whose sorting order is located at a second preset position as the geographic coordinate data corresponding to the query address data; wherein the partial order relation sample data is obtained by the address data processing method described in the first aspect.
[0008] Thirdly, embodiments of this application provide an address data processing apparatus, comprising: a sample acquisition unit, configured to acquire sample target address data and multiple sample candidate address data corresponding to the sample target address data; a first obtaining unit, configured to select multiple specified sample candidate address data belonging to different association levels from the multiple sample candidate address data according to the association level data corresponding to the sample target address data and the multiple sample candidate address data respectively, and obtain sample label data corresponding to the multiple specified sample candidate address data, wherein the sample label data is used to characterize matching sample candidate address data in the multiple specified sample candidate address data whose sample association index data with the sample target address data is greater than a preset association index data; and a partial order relation sample data construction unit, configured to construct partial order relation sample data of the sample target address data according to the multiple specified sample candidate address data and the sample label data; wherein the partial order relation sample data is used to train an initial model to obtain an address association index model for analyzing the target association index data between the query address data and the recalled address data.
[0009] Fourthly, embodiments of this application provide a geocoding device, comprising: a second obtaining unit, configured to obtain query address data and recall address data corresponding to the query address data; a third obtaining unit, configured to provide the query address data and the recall address data to an address association index model trained using partial order relation sample data, to obtain target association index data output by the address association index model for the relationship between the query address data and the recall address data; a sorting unit, configured to sort the multiple recall address data according to the target association index data between the multiple recall address data and the query address data; and a geocoding unit, configured to use the geographic coordinate data corresponding to the recall address data whose sorting order is located at a second preset position as the geographic coordinate data corresponding to the query address data; wherein the partial order relation sample data is obtained by the address data processing method described in the first aspect.
[0010] Fifthly, embodiments of this application provide an electronic device, including: a memory and a processor; the memory is used to store one or more computer instructions; the processor is used to execute the one or more computer instructions to implement the method described in the first aspect or the second aspect.
[0011] In a sixth aspect, embodiments of this application provide a computer-readable storage medium having stored thereon one or more computer instructions that are executed by a processor to implement the method described in the first or second aspect.
[0012] Compared with the prior art, this application has the following advantages: The address data processing method provided in this application selects multiple specified candidate address data belonging to different association levels from the target address data and the corresponding candidate address data, and obtains sample label data corresponding to the multiple specified candidate address data. Based on the multiple specified candidate address data and the sample label data, partial order relation sample data of the target address data is constructed. Furthermore, the partial order relation sample data is used to train an initial model to obtain an address association index model for analyzing the target association index data between the query address data and the recalled address data. This method processes the sample data to obtain partial order relation sample data for the target address data, and association index data between the multiple specified candidate address data contained in the partial order relation sample data and the target address data, thereby obtaining matching candidate address data whose association index data with the target address data is greater than a preset association index data. By training the initial model with partial order relation sample data, the association index model can predict that the association index between the target recall address data and the target query address data in multiple recall address data of the query address data is greater than the preset association index data, thereby improving the accuracy of the target recall address data obtained by the model. Attached Figure Description
[0013] Figure 1 This is the process of geocoding address data provided in the embodiments of this application.
[0014] Figure 2 This is a flowchart of an address data processing method provided in the first embodiment of this application.
[0015] Figure 3 A flowchart of a geocoding method provided in the second embodiment of this application.
[0016] Figure 4 This is a schematic diagram of an address data processing device provided in the third embodiment of this application.
[0017] Figure 5 This is a schematic diagram of a geocoding device provided in the fourth embodiment of this application. Detailed Implementation
[0018] Many specific details are set forth in the following description to provide a full understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this application; therefore, this application is not limited to the specific embodiments disclosed below.
[0019] It should be noted that the terms "first," "second," "third," etc., in the claims, specification, and drawings of this application are used to distinguish similar objects and are not used to describe a specific order or sequence. Such data are interchangeable where appropriate so that the embodiments of this application described herein can be implemented in a sequence other than that shown or described herein. Furthermore, the terms "comprising," "having," and their variations are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses.
[0020] It should be understood that in the embodiments of this application, "at least one" means one or more, and "more than one" means two or more. "And / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. The character " / " generally indicates that the related objects before and after it are in an "or" relationship. "Contains A, B and / or C" means containing any one, two, or three of A, B, and C.
[0021] It should be understood that in the embodiments of this application, "B corresponding to A", "B corresponding to A", "A corresponds to B" or "B corresponds to A" means that B is associated with A, and B can be determined based on A. Determining B based on A does not mean that B is determined solely based on A; B can also be determined based on A and / or other information.
[0022] This application provides an address data processing method, apparatus, electronic device, and computer storage medium. This application also provides a geocoding method, apparatus, electronic device, and computer storage medium. These will be described in detail in the following embodiments.
[0023] Before describing the implementation methods of this application in detail, the prior art will be further explained first.
[0024] Geocoding technology converts textual address data into geographic coordinate data. It is crucial in location-based services; for example, in logistics, delivery resources require address lookup devices to obtain the corresponding geographic coordinates of configured addresses, thus providing navigation information to the delivery location. Advances in artificial intelligence and big data technologies can further enhance the accuracy of obtaining accurate address text data, thereby improving the accuracy of determining corresponding geographic coordinates in location-based services.
[0025] In existing technologies, address text data includes diverse descriptive text data, resulting in low accuracy of the obtained address text data.
[0026] Therefore, improving the accuracy of the obtained address text data is a technical problem that needs to be solved.
[0027] To more clearly demonstrate the address data processing method provided in this application embodiment, the application scenario of the address data processing method provided in this application embodiment is first introduced. The partial order relation sample data obtained by this method is used to train an initial model to obtain an address association index model. This model is used to determine the corresponding standard address text based on the address text provided by the user, thereby using the geographic coordinate data corresponding to the standard address as the geographic coordinate data corresponding to the address text provided by the user, thus realizing the process of geocoding the address text provided by the user.
[0028] For details, please refer to the following: Figure 1 The process includes the following steps: Step S101: Obtain address text, which is to obtain the address data to be queried by the user; Step S102: Address segmentation, which involves segmenting the address data to be queried into words, for example, by segmenting by address level to obtain the provincial address phrases, district address phrases, street address phrases, community address phrases, building number and target location name phrases corresponding to the address data; Step S103: Address recall, which uses the geocoding method provided in the second embodiment to obtain multiple recall address data corresponding to the address data to be queried; Step S104: Address sorting, which uses the partial order relation sample data obtained by the method provided in the first embodiment to train the initial model to obtain the address association index data model. The target association index data corresponding to the query address data and the recall address data are analyzed by the address association index data model. The multiple recall address data are sorted according to the target association index data to obtain the target recall address data with the second preset position. The second preset position value is less than the other position values in the sorting result above. For example, the second preset position value is 1. Step S105: Location information, the target recall address data is used as the standard address of the query address data, and the location information of the target recall address data in the electronic map is used as the location information of the query address data in the electronic map.
[0029] The partial order relation sample data involved in the above process is obtained through the method provided in the first embodiment. Specifically, the sample data is processed to obtain partial order relation sample data for the target address data, as well as correlation index data between multiple specified candidate address data and the target address data contained in the partial order relation sample data. This results in matching candidate address data whose correlation index data with the target address data is greater than a preset correlation index data. Training the initial model using the partial order relation sample data allows the trained address correlation index model to predict a correlation index data between the target recall address data and the query address data that is greater than the preset correlation index data among multiple recall address data of the query address data, thus improving the accuracy of the target recall address data obtained by the model. Based on this, the accuracy of obtaining the geographic coordinate data corresponding to the query address data can be improved.
[0030] The following combination Figure 2 This application describes the address data processing method provided in the first embodiment.
[0031] First Embodiment Figure 2 This is a flowchart of an address data processing method provided in an embodiment of this application. Figure 2 The address data processing method shown is executed by a server, and the method includes steps S201 to S203.
[0032] Step S201: Obtain the target address data of the sample and the multiple candidate address data of the sample target address data.
[0033] This step is used to acquire sample data, including target address data and candidate address data. The candidate address data includes positive candidate addresses whose association level with the target address data is greater than or equal to a preset association level, and negative candidate addresses whose association level with the target address data is less than a preset association level. This data will then be processed in subsequent steps to obtain partially ordered relation sample data used to train the initial model.
[0034] Geocoding address data refers to converting text address data into geographic coordinate data, which means obtaining the location information of the text address data on an electronic map, thereby obtaining the corresponding POI (Point of Interest) information. To improve the accuracy of the location information of the obtained text address data on the electronic map, accurate text address data is first required. Therefore, the first embodiment of this application processes sample data to obtain partially ordered relation sample data. The partially ordered relation sample data is then used to train a model to obtain an address association index model, and accurate text address data is obtained through the address association index model.
[0035] The target address data in the sample is usually address data with different forms of expression. For example, address data obtained by using non-standard description methods or local aliases, such as X1 City X3 Street X4 Community X6 Floor YY No., or X7 City X3 Street X6 Floor YY No.; or address data with missing information, such as address data with missing address elements or address misspellings.
[0036] The sample candidate address data consists of standard addresses stored in the address database, such as: Unit YY, 6th Floor, Building X5, X3 Street, X4 Community, X1 City. The sample candidate address data includes standard address text data and the corresponding geographic coordinates.
[0037] By using the similarity between positive candidate address data and target address data, the model learns information from standard positive candidate address data that is similar to the target address data. Similarly, by using the similarity between negative candidate address data and target address data, the model learns information from standard negative candidate address data whose similarity to the target address data is below a preset similarity threshold. This enables the model to learn the ability to analyze target correlation metrics between the query address data and the recalled address data.
[0038] Step S202: Based on the correlation level data corresponding to the target address data and the multiple candidate address data, select multiple specified candidate address data belonging to different correlation levels from the multiple candidate address data, and obtain the sample label data corresponding to the multiple specified candidate address data. The sample label data is used to characterize the matching candidate address data in the multiple specified candidate address data whose sample correlation index data with the target address data is greater than the preset correlation index data.
[0039] This step is used to process the sample target address data and sample candidate address data obtained in step S201. Specifically, based on the association level data between multiple sample candidate address data and sample target address data, multiple specified sample candidate address data and sample label data corresponding to multiple specified sample candidate address data are obtained, so as to construct the partial order association sample data of sample target address data in subsequent steps.
[0040] The association level data includes correlation levels used to characterize the similarity between the sample target address data and the sample candidate address data.
[0041] For example, the target address data is 'a', and the corresponding candidate address data are a1, a2, a3, and a4. The correlation levels between a1, a2, a3, and a4 and 'a' are 2, 1, 1, and 0, respectively. Here, 2 indicates a correlation level of 2 between the candidate address data and the target address data, meaning they are perfectly correlated. For example, the first candidate address data a1 is perfectly correlated with the target address data a. 1 indicates a correlation level of 1 between the candidate address data and the target address data, meaning they are moderately correlated. For example, the second candidate address data a2 is moderately correlated with the target address data a. The third candidate address data a3 is moderately correlated with the target address data a. 0 indicates a correlation level of 0 between the candidate address data and the target address data, meaning they are not correlated. For example, the third candidate address data a3 is not correlated with the target address data a.
[0042] Therefore, the positive sample candidate address data mentioned in step S201 refers to the sample candidate address data whose association level data is greater than the preset association level data, where the preset association level data can be 1. That is, the first sample candidate address data a1, the second sample candidate address data a2, and the third sample candidate address data a3 are positive sample candidate address data, and the fourth sample candidate address data a4 is negative sample candidate address data.
[0043] The following describes the process of selecting multiple specified sample candidate address data from multiple sample candidate address data corresponding to the sample target address data.
[0044] During model training, to improve the accuracy of the model's analysis of the correlation indicators between candidate address data and target address data, multiple candidate address data were analyzed and processed to construct a partial-order relation sample data for the target address data. Specifically: Based on the correlation levels corresponding to the target address data and the plurality of candidate address data, matching candidate address data belonging to the target level and other candidate address data belonging to non-target levels are selected from the plurality of candidate address data. The non-target level is at least one correlation level lower than the target level. One matching candidate address data and at least one other candidate address data are used as the plurality of specified candidate address data.
[0045] The correlation levels between multiple candidate address data and the target address data are manually labeled correlation data. Based on the correlation levels between the multiple candidate address data and the target address data, multiple specified candidate address data included in the sample data for constructing the partial order relation are obtained.
[0046] Matching candidate address data belonging to the target level refers to the specified candidate address data among multiple specified candidate address data whose relevance level is higher than that of other candidate address data. Multiple specified candidate address data include one matching candidate address data and at least one other specified candidate address data.
[0047] When acquiring matching sample candidate address data and other sample candidate address data, the target relevance level and non-target relevance level are determined by classifying and sorting the relevance levels corresponding to the sample candidate address data, and then the matching sample candidate address data belonging to the target relevance level are determined.
[0048] Specifically, based on the correlation levels between the target address data and the multiple candidate address data, the candidate address data are classified to obtain candidate address data belonging to different correlation levels; based on the sorting results corresponding to the multiple different correlation levels, the target correlation level and non-target correlation levels are determined; among the candidate address data belonging to at least one candidate address data belonging to the target correlation level, one candidate address data is selected as the matching candidate address data; among the candidate address data belonging to at least one candidate address data belonging to the non-target correlation level, at least one candidate address data is selected as the other candidate address data.
[0049] Continuing with the above example, the candidate address data with a relevance level of 2 is the first candidate address data a1; the candidate address data with a relevance level of 1 are the second candidate address data a2 and the third candidate address data a3; and the candidate address data with a relevance level of 0 is the fourth candidate address data a4. The higher the relevance level, the greater the similarity between the corresponding candidate address data and the target address data.
[0050] The relevance ranking of the aforementioned candidate address data is as follows: first relevance level 2, second relevance level 1, and third relevance level 0. Therefore, the target level can be 2, and the non-target relevance levels can be 1 and 0. The resulting first set of multiple specified candidate address data can be a1, a2, a3, and a4. This multiple specified candidate address data is used to construct a partial order relation sample data, which represents the matching candidate address data whose correlation index with the target address data is greater than a preset correlation index.
[0051] The following describes the process of obtaining sample label data corresponding to multiple specified sample candidate address data in the first group.
[0052] Specifically, the data is obtained as follows: based on the relevance levels corresponding to the multiple specified candidate address data, sample association index data corresponding to the multiple specified candidate address data is obtained; based on the sample association index data corresponding to the multiple specified candidate address data, sample label data corresponding to the multiple specified candidate address data is obtained.
[0053] A higher relevance level indicates a greater similarity between the specified candidate address data and the target address data. Since the sample label data is used to characterize the matching candidate address data among multiple specified candidate address data in a set of partially ordered sample data where the sample association index data with the target address data is greater than the preset association index data, that is, the sample association index data corresponding to the matching candidate address data is greater than the sample association index data corresponding to other specified candidate address data.
[0054] Therefore, the sample correlation index data corresponding to the multiple specified sample candidate address data can be obtained in the following way: Based on the relevance levels corresponding to the multiple specified candidate address data, the multiple specified candidate address data are sorted; the number of sample correlation indicators corresponding to the specified candidate address data whose sorting order is in the first preset position is called the first correlation indicator data, and the number of sample correlation indicator data corresponding to the specified candidate address data whose sorting order is after the first preset position is called the second correlation indicator data.
[0055] The higher the relevance level of the candidate address data, the greater the similarity between the candidate address data and the target address data. Therefore, based on the ranking results of the relevance levels of the candidate address data, the corresponding sample association index data should be determined for each of the specified candidate address data.
[0056] The first set of multiple specified sample candidate address data obtained above are a1, a2, a3, a4, and their corresponding relevance levels are 2, 1, 1, 0, and their corresponding sorting results are a1, a2, a3, a4.
[0057] A partially ordered sample dataset contains only one matching candidate address data. Therefore, the first preset position value here is, for example, 1. The sample correlation index data corresponding to the specified candidate address data ranked in the first preset position in the correlation level ranking result is called the first correlation index data, and the sample correlation index data corresponding to the remaining specified candidate address data is called the second correlation index data.
[0058] Based on the sample association index data corresponding to the multiple specified sample candidate address data obtained above, and combined with the matching sample candidate address data and other candidate sample address data in the multiple specified sample candidate address data determined in the above process, the sample label data corresponding to the multiple specified sample address data is obtained in the following way.
[0059] The specified sample candidate address data with a sorting order in the first preset position is used as the matching sample candidate address data for constructing the partial order relationship sample data, and the sample association index data corresponding to the matching sample candidate address data is the first association index data; the specified sample candidate address data with a sorting order after the first preset position is used as other sample candidate address data for constructing the partial order relationship sample data, and the sample association index data corresponding to the other sample candidate address data is the second association index data; the first association index data and the second association index data are used as the sample label data corresponding to the plurality of specified sample candidate address data.
[0060] For example, the first correlation index data can be 1, and the second correlation index data can be 0. Therefore, the sample data of the first partial order relation in the above example are (a1, a2, a3, a4), and the corresponding sample correlation index data are (1, 0, 0, 0).
[0061] Based on the sample association index data corresponding to multiple specified sample candidate address data in the first partial order relation sample data mentioned above, it can be seen that the sample association index data corresponding to the matching sample candidate address data is greater than the sample association index data corresponding to other sample candidate address data.
[0062] Step S203: Based on the multiple specified candidate address data and the sample label data, construct the partial order relation sample data of the target address data; wherein, the partial order relation sample data is used to train the initial model to obtain an address association index model for analyzing the target association index data between the query address data and the recalled address data.
[0063] This step involves constructing a partial order relation sample data from the specified candidate address data and sample label data obtained in the previous steps, thereby training and obtaining an address association index model.
[0064] For example, the first group of multiple specified sample candidate address data is (a1, a2, a3, a4), and their corresponding sample association index data is (1, 0, 0, 0). The first partial order relation sample data constructed by them is (a1, a2, a3, a4) = (1, 0, 0, 0).
[0065] However, the first partially ordered relation sample data contains three other candidate address data. The magnitude relationship between the sample correlation index data corresponding to the other three candidate address data cannot be obtained from the first partially ordered relation sample data. Therefore, the second partially ordered relation sample data (a2, a4) = (1, 0) can be obtained, which indicates that the similarity between the second specified candidate address data and the target address data is greater than the similarity between the fourth specified candidate address data and the target address. The third partially ordered relation sample data (a3, a4) = (1, 0) indicates that the similarity between the third specified candidate address data and the target address data is greater than the similarity between the fourth specified candidate address data and the target address. Since the correlation level corresponding to the second specified sample candidate address data is the same as the correlation level corresponding to the third specified sample candidate address data, there is no need to construct partial order relation sample data that only contains the second specified sample candidate address data and the third specified sample candidate address data.
[0066] The target address data mentioned above corresponds to multiple candidate address data, a1, a2, a3, and a4. The partially ordered relation sample data that can be constructed include the first partially ordered relation sample data ((a1, a2, a3, a4) = (1, 0, 0, 0)), the second partially ordered relation sample data ((a2, a4) = (1, 0)), and the third partially ordered relation sample data ((a3, a4) = (1, 0)). The partially ordered relation sample data is used to train the initial model, enabling the initial model to learn the sample association index data between multiple specified candidate address data and the target address data. The initial model is then trained using both positive and negative candidate address data to improve the accuracy of the trained address association index model in predicting the target association index data between the query address data and the recalled address data.
[0067] The above describes the process of obtaining the partial order relation sample data. The following describes the process of training the initial model using the partial order relation sample data to obtain the address association index model. The initial model is one capable of analyzing whether two address data are related.
[0068] We utilize the address pair relevance labeling function of a pre-trained teacher model to label a large number of address pairs. We then train the model using these labeled address pairs, enabling it to predict whether a pair of addresses contains relevant or irrelevant addresses, thus obtaining the initial model.
[0069] Furthermore, by training the initial model using partial order relation sample data, the accuracy of the target correlation index data predicted by the obtained address correlation index model between the query address data and the recalled address data is improved.
[0070] Specifically as follows: The sample target address data and the multiple specified sample candidate address data are provided to the initial model. Based on the return result of the initial model, the predicted label data corresponding to the multiple specified sample candidate address data is obtained. Based on the loss value between the sample label data and the predicted label data, the parameters of the initial model are adjusted to obtain the address association index model.
[0071] The initial model is trained for each partial order relation sample data. The predicted label data can be obtained using the first method: All specified candidate address data of the sample contained in the partial order relation sample data and the target address data of the sample are simultaneously provided to the initial model: The target address data and the plurality of specified candidate address data are provided to the initial model to predict the predictive correlation index data between each specified candidate address data and the target address data; the predictive label data is constructed based on the predicted correlation index data corresponding to the plurality of specified candidate address data.
[0072] After obtaining all specified candidate address data contained in the partial order relation sample data, the initial model calculates the predicted association index data between each specified candidate address data and the target address data. For example, the predicted association index data corresponding to the multiple specified candidate address data (a1, a2, a3, a4) contained in the first partial order relation sample data are (0.1, 0.5, 0.5, 0.8). Based on the obtained multiple predicted association index data, predicted label data is constructed.
[0073] The initial model is trained for each partial order relation sample data. The second method can be used to obtain the predicted label data: For each specified candidate address data, the specified candidate address data and the target address data are provided to the initial model to obtain the predicted correlation index data between the specified candidate address data and the target address data predicted by the initial model; based on the predicted correlation index data corresponding to the multiple specified candidate address data obtained, the predicted label data is constructed.
[0074] For example, the first partial order relation sample data contains four specified candidate address data. The first specified candidate address data and the target address data are used as a first address data pair and provided to the initial model to obtain the first predicted association index data (e.g., 0.1). The second specified candidate address data and the target address data are used as a second address data pair and provided to the initial model to obtain the second predicted association index data (e.g., 0.5). The third specified candidate address data and the target address data are used as a third address data pair and provided to the initial model to obtain the third predicted association index data (e.g., 0.5). The fourth specified candidate address data and the target address data are used as a fourth address data pair and provided to the initial model to obtain the fourth predicted association index data (e.g., 0.8). Then, based on the first, second, third, and fourth predicted association index data, predicted label data (0.1, 0.5, 0.5, 0.8) are constructed.
[0075] Then, the parameters of the initial model are adjusted by the loss value between the predicted label data and the sample label data. Specifically, the cross-entropy loss value is calculated between the sample association index data contained in the sample label data and the predicted association index data contained in the predicted label data. Based on the cross-entropy loss value, the parameters of the initial model are adjusted to obtain the trained address association index model.
[0076] In the example above, the sample label data corresponding to the multiple specified candidate address data (a1, a2, a3, a4) in the first partial order relation sample data is (1, 0, 0, 0), while the predicted label data corresponding to the initial model is (0.1, 0.5, 0.5, 0.8). Here, the predicted association index data corresponding to each specified candidate address data in the predicted label data does not conform to the distribution relationship represented by the sample label data (1, 0, 0, 0) corresponding to the first partial order relation sample data. That is, the predicted label data here indicates that the fourth specified candidate address data a4 is a matching candidate address data, while the sample label data corresponding to the first partial order relation sample data indicates that the first specified candidate address data is a matching candidate address data. Therefore, there is a deviation between the predicted label data corresponding to the initial model and the sample label data corresponding to the first partial order relation sample data. This indicates that the cross-entropy loss value between the predicted label data and the sample label data is too large, requiring adjustment of the parameters of the initial model. The adjusted model should then be further trained using the partial order relation sample data.
[0077] For example, after the first round of parameter adjustments to the initial model, the resulting first model is further trained using the first partial order relation sample data. The predicted correlation indices between each specified candidate address data and the target address data are 0.9, 0.5, 0.5, and 0.1, respectively. Therefore, the predicted label data is (0.9, 0.5, 0.5, 0.1). Here, the predicted correlation indices corresponding to each specified candidate address data in the predicted label data conform to the distribution relationship represented by the sample label data. That is, the predicted label data here indicates that the first specified candidate address data is a matching candidate address data, which matches the result indicated in the first partial order relation sample data.
[0078] To improve the accuracy of the model in predicting the target association index data between the address data to be processed and the recalled address data, multiple partial order relation samples were used to train the initial model for multiple rounds to obtain the address association index model. Here, after one round of training of the initial model and adjustment of the model parameters, the trained model is tested online using a test set to obtain the result index value of the trained model predicting the target association index data. This value is used to characterize whether the parameters adjusted by the model after this training meet the preset conditions for predicting the target association index data. To avoid overfitting the model parameters, the result index values corresponding to the model after multiple rounds of training are compared, and the parameters adjusted by the trained model corresponding to the target result index value are selected, where the target result index value is greater than the result index values corresponding to the model after other rounds of training.
[0079] The above describes the process of obtaining partial-order relation sample data for the target address data based on the association level data between the target address data and the candidate address data. This involves obtaining the sample association index data between multiple specified candidate address data and the target address data contained in the partial-order relation sample data, thereby obtaining matching candidate address data whose association index data with the target address data is greater than a preset association index data. Training the initial model using the partial-order relation sample data obtained in this way can improve the accuracy of the target recall address data obtained by the model, ensuring that the association index data between the predicted target recall address and the target address data in the multiple recall address data of the query address data is greater than the preset association index data.
[0080] Second Embodiment Based on the first embodiment, Figure 3 A geocoding method is provided in the second embodiment of this application. The executing entity of this method can be a client used to query address coordinate data corresponding to address data. The method includes steps S301 to S304.
[0081] Step S301: Obtain the address data to be queried and the recall address data corresponding to the address data to be queried.
[0082] This step is used to obtain the address data to be queried and the corresponding recall address data. The address data to be queried is the address data provided by the user for querying. It is usually address data with different expressions, such as address data obtained using non-standard descriptions or local aliases, like "X1 City, X3 Street, X4 Community, X6th Floor, YY Number", or "X7 City, X3 Street, X6th Floor, YY Number"; or address data with missing information, such as address data with missing address elements or address misspellings.
[0083] The recall address data consists of standard address data existing in the address database, such as: Unit YY, 6th Floor, Building X5, X3 Street, X4 Community, X1 City. The recall address data includes standard address text data and the corresponding geographic coordinates.
[0084] The recall address data refers to address data whose correlation index with the address data to be queried is greater than the preset correlation index.
[0085] The method provided in the second embodiment obtains target recall address data corresponding to the address data to be queried from the recall address data, uses the target recall address data as the address data to be queried, and uses the geographic coordinate data of the target recall address data as the geographic coordinate data of the address data to be queried.
[0086] First, the query address data is processed by word segmentation, as follows: The address data to be queried and the candidate address data in the address database are segmented into words according to address hierarchy elements. For example, X1 City, X2 District, X3 Street, X4 Community, X5 Building, X6 Floor, YY Number is segmented into multiple words according to address hierarchy. The first level is the province / municipality level: X1 City; the second level is the district / county level: X2 District; the third level is the street / township / town level: X3 Street; the fourth level is the AOI (Area of Interest) level: X4 Community; and the fifth level is the building level: X5 Building, X6 Floor, YY Number.
[0087] To increase the number of retrieved address data corresponding to the query address data, multiple retrieved address data corresponding to the query address data are obtained through the following method: The recall address data corresponding to the query address data is obtained through at least one of the following methods: obtaining the recall address data corresponding to the query address data through address hierarchy matching; obtaining the recall address data corresponding to the query address data through address vector matching; obtaining the recall address data corresponding to the query address data through address word frequency matching.
[0088] To improve search efficiency, this embodiment extracts the segmentation information corresponding to the first level from the address after segmentation processing, thereby narrowing the scope of searching for and recalling address data in the address database, and obtaining selectable address data within the same province, city, and district as the address data to be queried.
[0089] The first method is address level matching: The word segments contained in each optional address data in the address database are compared with the word segments contained in the address data to be queried to obtain the address data to be matched that contains the word segments in the address data to be queried.
[0090] In the address data to be matched, address data with a number of matching address levels greater than a preset threshold, and including address data whose level is less than or equal to a preset threshold, are selected as the first recall address data set corresponding to the address data to be queried. For example, the preset threshold for the number of levels is 3, and the threshold for the level of the matching address is 4. Wherein, the address data to be queried is YY, Building X6, Community X4, Street X3, City X1. Two candidate recall address data are found in the address database, for example, the first candidate recall address data is YY, Building X6, Community X4, Street X3, District X1, City X2, and the second candidate recall address data is YY, Building X7, Community X4, Street X3, District X2, City X1.
[0091] It can be seen that the number of matching address levels between the query address data and the first candidate recall address data is 4 (X1 city, X3 street, X4 community, X5 building, X6 floor, YY number). Moreover, the matching address level of "X5 building, X6 floor, YY number" is the 5th level, which is less than the threshold of 4.
[0092] The number of matching address levels between the query address data and the second candidate recall address data is 3 (X1 city, X3 street, X4 community). Moreover, the matching address level "X4 community" is the 4th level, which is equal to the threshold of 4.
[0093] Therefore, the first candidate recall address data and the second candidate recall address data are used as the address data in the first recall address data set corresponding to the address data to be queried.
[0094] The second method is address vector matching: Calculate the vector of the address data to be queried, and the vector corresponding to each optional address data in the address database. By calculating the similarity between the vector of the optional address data and the vector of the address data to be queried, a second recall address data set with a similarity greater than a preset similarity threshold is obtained.
[0095] The third method is address segmentation matching: The step of obtaining the recall address data corresponding to the query address data through address word frequency matching includes: Preprocess all address data in the address database by word segmentation to obtain candidate word segments included in each address data; perform word segmentation on the to-be-query address data to obtain the to-be-query word segments included in the to-be-query address data; query the to-be-matched address data that includes the to-be-query word segments in the address database, and use the to-be-query word segments included in the to-be-matched address data as the hit word segments; obtain the recalled address data corresponding to the to-be-query address data based on the obtained multiple to-be-matched address data that include the hit word segments.
[0096] In order to further increase the number of types of possible word segments of the obtained address data, perform word segmentation on the address data through at least one of the following methods: Segment the text corresponding to the address data according to adjacent specified number of characters. For example, if the text information of the address data is "X5 Company" which contains 6 characters, namely "one two three four five six", and the adjacent specified number of characters is 2, then the word segments after word segmentation include "one two", "two three", "three four", "four five", "five six".
[0097] Segment the text corresponding to the address data according to the specified interval number of characters. For example, if the specified interval number of characters is 1, then the obtained word segments include "one three", "two four", "three five", "four six". For example, some to-be-query address data are described in abbreviated forms. For example, the full text address information is "one two university", and its corresponding abbreviation is "one big".
[0098] Perform word segmentation on the to-be-query address data and the address text data in the address database according to the above methods. For example, if the to-be-query address data is the above "one two three four five six", its corresponding word segments, called to-be-query word segments, include: "one two", "two three", "three four", "four five", "five six", "one three", "two four", "three five", "four six".
[0099] Query the optional address data that includes the to-be-query word segments among the optional address data at the same provincial level as the to-be-query address data in the address database, which is called the to-be-matched address data.
[0100] Specifically, obtain the candidate word segments that are the same as the to-be-query word segments included in the to-be-query address data among the candidate word segments included in the address database as the hit word segments; use the optional address data in the address database that includes the hit word segments as the to-be-matched address data. Based on the word frequency weights corresponding to the hit words in the address database, the sum of the word frequency weights of all hit words contained in the address data to be matched is calculated; based on the sum of the word frequency weights of the hit words corresponding to multiple address data to be matched, at least one address data to be matched whose sum of the word frequency weights of the hit words is greater than a preset weight threshold is obtained, and used as the recall address data corresponding to the address data to be queried.
[0102] After segmenting all address data in the address database, the word frequency weight corresponding to each segment is calculated using Formula 1: Formula Expression 1 in, This represents the word frequency weight of the i-th word; N represents the total number of addresses contained in the address database; n represents the number of address data points in the address database that contain the i-th word.
[0103] Formula 1 can be used to obtain the word frequency weight of each segmented word in the address database after all address data has been segmented. The more addresses a segmented word corresponds to in the address database, the more it is biased towards general terms, and thus the higher its word frequency weight. The lower the number of addresses corresponding to a segment, the more the segment is biased towards a specific word. For example, if a word is located at an address level higher than the preset level, such as X4 cell, then the word frequency weight corresponding to that segment will be lower. The lower.
[0104] After obtaining the word frequency weights of each segmented word in the address database using the above method, the word frequency weights corresponding to the hit segments in each address data to be matched are added together to obtain the sum of the word frequency weights corresponding to each address data to be matched.
[0105] This allows us to obtain addresses where the sum of the word frequency weights of the hit words is greater than a preset threshold, which are then used as the third recall address data.
[0106] Step S302: Provide the address data to be queried and the recall address data to the address association index model obtained by training through partial order relation sample data, and obtain the target association index data output by the address association index model for the relationship between the address data to be queried and the recall address data.
[0107] The partial order relation sample data is obtained using the method provided in the first embodiment.
[0108] This step is used to form an address data pair with each recall address data obtained in step S301 above and the address data to be queried. The address association index model obtained by training with the partial order relation sample data provided in the first embodiment is used for prediction, which can improve the accuracy of the target association index data of the obtained address data pairs.
[0109] Step S303: Sort the multiple recall address data according to the target correlation index data between the multiple recall address data and the address data to be queried.
[0110] Step S304: Use the geographic coordinate data corresponding to the recall address data that is ranked in the second preset position as the geographic coordinate data corresponding to the address data to be queried.
[0111] After sorting multiple recall address data according to their corresponding target association index data, the recall address data with a sorting position in the second preset position is selected as the target recall address data, and the geographic coordinate data corresponding to the target recall address data is selected as the geographic coordinate data corresponding to the address data to be queried. The second preset position value is less than the other position values in the sorting result; for example, the second preset position value is 1.
[0112] Based on the partial order relation sample data provided in the first embodiment, the address association index model obtained by training the initial model can improve the accuracy of analyzing the target association index data between the query address data and the recalled address data. Therefore, after obtaining the target association index data between multiple recalled address data and the query address data through the address association index model, the multiple recalled address data are sorted to improve the accuracy of the target recalled address data corresponding to the query address data, thereby improving the accuracy of obtaining the address coordinate data corresponding to the query address data.
[0113] Third Embodiment Based on the first embodiment, Figure 4 An address data processing apparatus provided in the third embodiment of this application includes: The sample acquisition unit 401 is used to acquire sample target address data and multiple sample candidate address data corresponding to the sample target address data; The first obtaining unit 402 is used to select multiple specified sample candidate address data belonging to different association levels from the multiple sample candidate address data according to the association level data corresponding to the sample target address data and the multiple sample candidate address data respectively, and obtain sample label data corresponding to the multiple specified sample candidate address data. The sample label data is used to characterize the matching sample candidate address data in the multiple specified sample candidate address data whose sample association index data with the sample target address data is greater than the preset association index data. The partial order relation sample data construction unit 403 is used to construct the partial order relation sample data of the sample target address data based on the plurality of specified sample candidate address data and the sample label data; The partial order relation sample data is used to train the initial model to obtain an address association index model for analyzing the target association index data between the query address data and the recall address data.
[0114] Fourth embodiment Based on the second embodiment, Figure 5 A geocoding device provided in the fourth embodiment of this application includes: The second obtaining unit 501 is used to obtain the address data to be queried and the recall address data corresponding to the address data to be queried. The third obtaining unit 502 is used to provide the address data to be queried and the recall address data to the address association index model obtained by training through partial order relation sample data, and obtain the target association index data between the address data to be queried and the recall address data output by the address association index model. The sorting unit 503 is used to sort the multiple recall address data according to the target correlation index data between the multiple recall address data and the address data to be queried; The geocoding unit 504 is used to use the geographic coordinate data corresponding to the recall address data whose sorting order is located in the second preset position as the geographic coordinate data corresponding to the address data to be queried. The partial order relation sample data is obtained by the method described in the first embodiment.
[0115] Fifth embodiment The fifth embodiment of this application also provides an electronic device, including: a processor; and a memory for storing a computer program. After the electronic device is powered on and runs the computer program through the processor, it performs the above-described method.
[0116] Sixth Embodiment The sixth embodiment of this application also provides a computer storage medium storing a computer program, which is executed by a processor to perform the above-described method.
[0117] The above-described device embodiments, electronic device embodiments, and storage medium embodiments correspond to the above-described method embodiments; please refer to the method embodiments for details. Although this application discloses preferred embodiments as described above, it is not intended to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the scope of protection of this application should be determined by the scope defined in the claims of this application. In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory. Memory may include non-permanent storage in computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media. 1. Computer-readable media includes permanent and non-permanent, removable and non-removable media. Information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media do not include non-transitory computer-readable media, such as modulated data signals and carrier waves. 2. Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. It should be noted that the embodiments of this application may involve the use of user data. In practical applications, user-specific personal data may be used in the scheme described herein within the scope permitted by applicable laws and regulations, provided that it complies with the applicable laws and regulations of the country (e.g., with the user's explicit consent, with the user being properly notified, etc.).It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
Claims
1. An address data processing method, characterized in that, include: Obtain sample target address data and multiple sample candidate address data corresponding to the sample target address data. The sample candidate address data includes standard address text data that has an association level data relationship with the sample target address data and the geographic coordinate data corresponding to the address text data. Based on the correlation levels between the target address data and the multiple candidate address data, multiple specified candidate address data belonging to different correlation levels are selected from the multiple candidate address data. This includes: selecting matching candidate address data belonging to the target correlation level and other candidate address data belonging to non-target correlation levels, and obtaining sample label data corresponding to the multiple specified candidate address data. The sample label data is used to characterize the matching candidate address data among the multiple specified candidate address data whose sample correlation index data with the target address data is greater than a preset correlation index data. The sample correlation index data corresponding to the matching candidate address data is the first correlation index data, and the sample correlation index data corresponding to the other candidate address data is the second correlation index data. The first correlation index data and the second correlation index data are used as the sample label data. The sorting order of the first correlation index data is located in the first preset position, and the sorting order of the second correlation index data is located after the first preset position. Based on the plurality of specified candidate address data and the sample label data, a partial order relation sample data of the target address data is constructed; wherein, the step of selecting the plurality of specified candidate address data belonging to different relevance levels from the plurality of candidate address data includes: classifying the plurality of candidate address data according to the relevance level to obtain candidate address data belonging to different relevance levels; determining the target relevance level and non-target relevance levels based on the sorting results corresponding to the obtained plurality of different relevance levels, wherein the non-target relevance level is at least one relevance level lower than the target relevance level, the target relevance level includes a first relevance level, and the non-target relevance level includes a second relevance level and a third relevance level; in belonging to One sample candidate address is selected from at least one sample candidate address data of the target relevance level as the matching sample candidate address data; at least one sample candidate address is selected from at least one sample candidate address data belonging to the non-target relevance level as the other sample candidate address data; one matching sample candidate address data and at least one other sample candidate address data are used as the plurality of specified sample candidate address data; wherein, the partial order relation sample data is used to train the initial model to obtain an address association index model for analyzing the target association index data between the query address data and the recalled address data, and the recalled address data includes address text data whose association index data with the query address data is greater than the preset association index data and the geographic coordinate data corresponding to the address text data.
2. The method according to claim 1, characterized in that, Obtaining the sample label data corresponding to the plurality of specified sample candidate address data includes: Based on the correlation levels corresponding to the multiple specified candidate address data, obtain the sample association index data corresponding to the multiple specified candidate address data; Based on the sample association index data corresponding to the multiple specified sample candidate address data, obtain the sample label data corresponding to the multiple specified sample candidate address data.
3. The method according to claim 2, characterized in that, The step of obtaining sample correlation index data corresponding to the multiple specified sample candidate address data according to the correlation levels corresponding to the multiple specified sample candidate address data respectively includes: Sort the multiple specified sample candidate address data according to the relevance level corresponding to each of the multiple specified sample candidate address data; The number of sample correlation indicators corresponding to the specified sample candidate address data whose sorting order is in the first preset position is called the first correlation indicator data, and the number of sample correlation indicator data corresponding to the specified sample candidate address data whose sorting order is after the first preset position is called the second correlation indicator data.
4. The method according to claim 3, characterized in that, The step of obtaining sample label data corresponding to the multiple specified sample candidate address data based on the sample association index data corresponding to the multiple specified sample candidate address data includes: The specified sample candidate address data with the sorting order located at the first preset position is used as the matching sample candidate address data for constructing the partial order relation sample data; The specified sample candidate address data whose sorting order is after the first preset position is used as other sample candidate address data for constructing the partial order relation sample data.
5. The method according to claim 1, characterized in that, The partial order relation sample data is used to train the initial model to obtain the address association index model in the following manner: The sample target address data and the multiple specified sample candidate address data are provided to the initial model, and the predicted label data corresponding to the multiple specified sample candidate address data are obtained based on the return result of the initial model. Based on the loss value between the sample label data and the predicted label data, the parameters of the initial model are adjusted to obtain the address association index model.
6. A geocoding method, characterized in that, include: Obtain the address data to be queried and the corresponding recall address data; The query address data and the recall address data are provided to the address association index model trained by the partial order relation sample data to obtain the target association index data between the query address data and the recall address data output by the address association index model. Based on the target correlation index data between the multiple retrieved address data and the address data to be queried, the multiple retrieved address data are sorted. The geographic coordinates of the recall address data that are ranked in the second preset position are used as the geographic coordinates of the address data to be queried. The partial order relation sample data is obtained by the method described in any one of the technical solutions of claims 1-5.
7. The method according to claim 6, characterized in that, Also includes: Obtain the recall address data corresponding to the query address data; The recall address data corresponding to the query address data is obtained through at least one of the following methods: The recall address data corresponding to the query address data is obtained by address hierarchy matching. The recall address data corresponding to the query address data is obtained by address vector matching. The recall address data corresponding to the query address data is obtained by address word frequency matching.
8. The method according to claim 7, characterized in that, The step of obtaining the recall address data corresponding to the query address data through address word frequency matching includes: The address data to be queried is segmented into words to obtain the segmented words to be queried contained in the address data to be queried; Query the address database for matching address data containing the query terminology, and use the query terminology contained in the matching address data as the hit terminology; Based on the obtained multiple address data to be matched containing the hit words, the recall address data corresponding to the query address data is obtained.
9. The method according to claim 8, characterized in that, Also includes: All address data in the address database are pre-processed into word segments to obtain candidate word segments for each address data. The step of querying the address database for matching address data containing the query term includes: Among the candidate segments contained in the address database, the candidate segment that is the same as the segment to be queried contained in the address data to be queried is obtained as the hit segment; The optional address data containing the hit word in the address database is used as the address data to be matched.
10. An address data processing device, characterized in that, include: The sample acquisition unit is used to acquire sample target address data and multiple sample candidate address data corresponding to the sample target address data. The sample candidate address data includes standard address text data that has an association level data relationship with the sample target address data and the geographic coordinate data corresponding to the address text data. The first obtaining unit is configured to select multiple specified sample candidate address data belonging to different correlation levels from the multiple sample candidate address data based on the correlation levels corresponding to the target address data and the multiple sample candidate address data, including: selecting matching sample candidate address data belonging to the target correlation level and other sample candidate address data belonging to non-target correlation levels, and obtaining sample label data corresponding to the multiple specified sample candidate address data. The sample label data is used to characterize the matching sample candidate address data among the multiple specified sample candidate address data whose sample correlation index data with the target address data is greater than a preset correlation index data. The sample correlation index data corresponding to the matching sample candidate address data is the first correlation index data, and the sample correlation index data corresponding to the other sample candidate address data is the second correlation index data. The first correlation index data and the second correlation index data are used as the sample label data. The sorting order of the first correlation index data is located in a first preset position, and the sorting order of the second correlation index data is... The sorting order is after the first preset position; the step of selecting multiple specified sample candidate address data belonging to different relevance levels from the multiple sample candidate address data includes: classifying the multiple sample candidate address data according to the relevance level to obtain sample candidate address data belonging to different relevance levels; determining the target relevance level and non-target relevance level according to the sorting results corresponding to the multiple different relevance levels obtained, wherein the non-target relevance level is at least one relevance level lower than the target relevance level, the target relevance level includes a first relevance level, and the non-target relevance level includes a second relevance level and a third relevance level; selecting one sample candidate address data belonging to the target relevance level as a matching sample candidate address data; selecting at least one sample candidate address data belonging to the non-target relevance level as the other sample candidate address data; and using one matching sample candidate address data and at least one other sample candidate address data as the multiple specified sample candidate address data. A partial order relation sample data construction unit is used to construct partial order relation sample data of the sample target address data based on the plurality of specified sample candidate address data and the sample label data; The partial order relation sample data is used to train the initial model to obtain an address association index model for analyzing the target association index data between the query address data and the recall address data. The recall address data includes standard address text data whose association index data with the query address data is greater than the preset association index data, and the geographic coordinate data corresponding to the address text data.
11. A geocoding device, characterized in that, include: The second obtaining unit is used to obtain the address data to be queried and the recall address data corresponding to the address data to be queried. The third obtaining unit is used to provide the address data to be queried and the recall address data to the address association index model obtained by training through partial order relation sample data, and obtain the target association index data between the address data to be queried and the recall address data output by the address association index model. The sorting unit is used to sort the multiple recall address data according to the target correlation index data between the multiple recall address data and the address data to be queried; A geocoding unit is used to use the geographic coordinate data corresponding to the recall address data whose sorting order is located in the second preset position as the geographic coordinate data corresponding to the address data to be queried. The partial order relation sample data is obtained by the method described in any one of the technical solutions of claims 1-5.
12. An electronic device, characterized in that, include: Memory and processor; The memory is used to store one or more computer instructions; The processor is used to execute one or more computer instructions to implement the method described in any one of claims 1-9.
13. A computer-readable storage medium storing one or more computer instructions thereon, characterized in that, The instruction is executed by the processor to implement the method described in any one of claims 1-9.
Citation Information
Patent Citations
Non-standard address standardization method and device
CN119917539A