Geographic Coordinate Optimization Method for a Multi-Source Network Geocoding Platform Constrained by Points of Interest

By building the optimization method of the minimum geographic encirclement and multi-source geocoding platform, using CRF and BERT models to process address features and POI data, the problem of large output error of the geocoding platform is solved, and more accurate geographic coordinate services are achieved.

CN119669381BActive Publication Date: 2025-07-08HOHAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510189077.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-07-08
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

Due to the differences in database sources and semantic similarity algorithms, the existing geocoding platform has large errors in the geographic coordinate output of the same text address, which affects the accuracy of the geographic coordinate service.

Method used

By building a minimum geographical encirclement, the address features are extracted using the CRF model and filtering POI data, vectorized similarity calculations are performed in combination with the BERT model, theoretical coordinates are generated, and the encoding results of multiple geocoding platforms are optimized to output the final optimized geographical coordinates.

Benefits of technology

The geographic coordinate error output by different geocoding platforms is effectively reduced, and the accuracy of geographic coordinate services is improved, especially the error within the error range greater than 500m is significantly reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119669381B_ABST
    Figure CN119669381B_ABST
Patent Text Reader

Abstract

The present invention provides a method for optimizing geographical coordinates of a multi-source network geocoding platform based on point-of-interest constraints, belonging to the field of processing geocoding result data, including S1: a user uses a CRF model to divide the input text address into rough address features and precise address features; S2: generating a minimum geographical bounding circle according to the rough address features; S3: calculating the theoretical coordinates according to the precise address features, and recording the theoretical coordinates as ( X m , Y m ); S4: forming a candidate set from the coding results of each geocoding platform; S5: if the coding result in the candidate set is not within the minimum geographical bounding circle, output the theoretical coordinates ( X m , Y m ); if at least one coding result in the candidate set is within the minimum geographical bounding circle, calculate the optimized geographical coordinates ( X 优 , Y 优 ) and output. This method reduces the coding result error of each geocoding platform and improves the accuracy of the coding result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of geocoding result data processing, and particularly relates to a method for optimizing geographical coordinates of a multi-source network geocoding platform based on point of interest constraints. Background Art

[0002] Geocoding is a technology that converts a text description of an address into geographical coordinates. Currently, online geocoding platforms (such as Baidu, Tencent, Amap, and Sogou) all rely on this geocoding technology to provide geographical coordinate services.

[0003] When using the above existing geocoding platforms, for example, typing "No. 17, Focheng West Road, Jiangning District, Nanjing City, Jiangsu Province" into the above multiple geocoding platforms respectively, the above geocoding platforms will, according to the text address of "No. 17, Focheng West Road, Jiangning District, Nanjing City, Jiangsu Province", query the corresponding coding result in their respective databases through their respective semantic similarity algorithms, and then feedback the corresponding coding result to the user. The coding result is the geographical coordinates corresponding to the text address of "No. 17, Focheng West Road, Jiangning District, Nanjing City, Jiangsu Province".

[0004] Due to reasons such as different data sources in the databases of the above multiple geocoding platforms or different designs of semantic similarity algorithms in each geocoding platform, the coding results output by each geocoding platform are also different. For example, for the same text address, the coding result output by geocoding platform K1 is P1, and the coding result output by geocoding platform K2 is P2. There are relatively large errors in the geographical coordinates output by multiple geocoding platforms, with an average error of more than 500m, which makes it impossible to provide accurate geographical coordinate services to users and causes inconvenience to users. For example, delivery personnel cannot accurately reach the location according to the geographical coordinates provided by a certain coding platform.

[0005] In order to reduce the error of geographical coordinates output by multiple geocoding platforms for the same text address and provide more accurate geographical coordinate services to users, therefore, it is necessary to propose a method for optimizing geographical coordinates of a multi-source network geocoding platform based on point of interest constraints in this solution. Summary of the Invention

[0006] The present invention proposes a method for optimizing geographical coordinates of a multi-source network geocoding platform based on point of interest constraints, which is used to reduce the error of the coding results output by each geocoding platform.

[0007] To solve the above problems, the present invention proposes the following solutions:

[0008] The method for optimizing geographical coordinates of a multi-source network geocoding platform based on point of interest constraints includes the following steps:

[0009] S1: The user inputs the text address in the CRF model. The CRF model extracts the address features from the input text address and classifies the address features into rough address features and precise address features;

[0010] S2: According to the rough address features, filter the POI data in the POI address database to generate the smallest geographical bounding circle;

[0011] S3: Take all the POI data within the smallest geographical bounding circle as the new POI database; query the new POI database, calculate the theoretical coordinates according to the precise address features, and record the theoretical coordinates as ( X m , Y m );

[0012] S4: Input the text address in S1 into each geocoding platform respectively, obtain the coding results of each geocoding platform, and form a candidate set with the coding results of each geocoding platform;

[0013] S5: If none of the coding results in the candidate set are within the smallest geographical bounding circle, output the theoretical coordinates ( X m , Y m ); if at least one coding result in the candidate set is within the smallest geographical bounding circle, calculate the optimized geographical coordinates ( X 优 , Y 优 ), and output the optimized geographical coordinates ( X 优 , Y 优 ).

[0014] Furthermore, step S2 includes the following steps:

[0015] S2.1: Input both the rough address features and all the POI data in the POI address database into the BERT model; the BERT model vectorizes both the rough address features and the text addresses of all the POI data;

[0016] S2.2: Calculate the cosine similarity between the rough address features and the text addresses of each POI data;

[0017] The formula is:

[0018]

[0019] In formula (1), A is the vector representation of the rough address features; Bis the vector representation of one of the POI data text addresses in the POI address database; represents the dot product of two vectors; and are respectively A the vector sum B and the L2 norm of the vector; CS is the cosine similarity;

[0020] S2.3: Set the cosine similarity threshold R, and mark the POI data with a cosine similarity greater than or equal to R with respect to the vector A as the POI data to be surrounded; Use the method of the minimum polygon to construct a polygon area from all the POI data to be surrounded, and the constructed polygon area is the minimum geographical surrounding area.

[0021] Furthermore, step S3 includes two cases,

[0022] Case 1: If there is at least one house number in the new POI database that is the same as the house number corresponding to the exact address feature, the calculation formula for the theoretical coordinates is:

[0023]

[0024] In formula (2), X m and Y m are the longitude and latitude of the theoretical coordinates; X i and Y i are the longitude and latitude of the POI data with the same exact address feature, , i is an integer; n is the number of POI data with the same exact address feature, n≥ 1;

[0025] Case 2: If there is no house number in the new POI database that is the same as the exact address feature, query the two house numbers in the new POI database that are close to the exact address feature house number, and denote them as C 1 and C 2,

[0026] The calculation formula for the theoretical coordinates is:

[0027]

[0028] In formula (3), X m and Y m are respectively the longitude and latitude of the theoretical coordinates;C x is the house number corresponding to the precise address feature; and are respectively the longitude and latitude of house number C 1; and are respectively the longitude and latitude of house number C 2.

[0029] Furthermore, in step S4, record the coding results of each geocoding platform as: ( X z1 , Y z1 ), ( X z2 , Y z2 ), ……, ( X zu , Y zu ), and the composed candidate set is: {( X z1 , Y z1 ), ( X z2 , Y z2 ), ……, ( X zu , Y zu )}, u is the number of coding results.

[0030] Furthermore, step S5 includes the following steps:

[0031] S5.1: If none of the coding results in the candidate set are within the minimum geographical bounding circle, output the theoretical coordinates ( X m , Y m );

[0032] S5.2: If at least one coding result in the candidate set is within the minimum geographical bounding circle, calculate the optimized geographical coordinates, and record the optimized geographical coordinates as ( X 优 , Y 优 ), and output the optimized geographical coordinates ( X 优 , Y 优 );

[0033] S5.2 specifically includes the following steps:

[0034] S5.2.1: Calculate the Euclidean distance between each coding result and the theoretical coordinates. The formula is:

[0035]

[0036] In formula (4), ; d k represents the Euclidean distance between the theoretical coordinates ( X m , Y m ) and the k th coding result ( X zk , Y zk );

[0037] S5.2.2: The coordinate weight of each coding result is:

[0038]

[0039] In formula (5), w k represents the coordinate weight of the k th coding result; u is the number of coding results; d k represents the Euclidean distance between the theoretical coordinates ( X m , Y m ) and the k th coding result ( X zk , Y zk );

[0040] S5.2.3: Calculate the optimized geographical coordinates:

[0041]

[0042] In formula (6), w k represents the coordinate weight of the k th coding result; u is the number of coding results; X zk and Y zk respectively represent the longitude and latitude of the k th coding result; ( X 优 , Y 优 ) is the optimized geographical coordinate.

[0043] Further, in step S2.3, the cosine similarity threshold R is 0.95.

[0044] Adopting the above technical solution, the beneficial effects that this solution can achieve are as follows:

[0045] This solution constructs a minimum geographical encirclement as a spatial constraint condition, calculates the theoretical coordinates based on the POI data constraints in the minimum geographical encirclement, generates encoding results using multiple geocoding platforms respectively, and outputs the theoretical coordinates or the optimized geographical coordinates through the relationship between multiple encoding results and the minimum geographical encirclement. It effectively considers the encoding results of multiple geocoding platforms and improves the accuracy of the encoding results. It reduces the encoding result errors caused by different databases of different geocoding platforms and different algorithms for similar semantics. Description of the Drawings

[0046] Figure 1 is the logic flow chart of this optimization method;

[0047] Figure 2 is a schematic diagram of POI data in the POI address database;

[0048] Figure 3 is a schematic diagram after the POI data in the POI address database forms a minimum geographical encirclement;

[0049] Figure 4 is the error histogram of this optimization method, other encoding platforms and the real address respectively;

[0050] Figure 5 is the process instance diagram of this optimization method. Detailed Embodiments

[0051] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0052] See Figure 1 , a geographical coordinate optimization method for a multi-source network geocoding platform based on point of interest (POI) constraints, includes the following steps:

[0053] S1: A user inputs a text address in the CRF model, and the CRF model extracts address features from the input text address and classifies the address features into rough address features and precise address features.

[0054] The explanation of the CRF model is: the Conditional Random Field (CRF) model.

[0055] Suppose the text address input into the CRF model is: "No. 17, Focheng West Road, Jiangning District, Nanjing City, Jiangsu Province".

[0056] After the CRF model extracts the address features, the text address "No. 17, Focheng West Road, Jiangning District, Nanjing City, Jiangsu Province" is divided into rough address features and precise address features. Among them, the rough address features are: "Focheng West Road, Jiangning District, Nanjing City, Jiangsu Province". The precise address features are: "No. 17".

[0057] S2: According to the rough address features, screen the POI data in the POI address database to generate the smallest geographical encirclement.

[0058] The POI address database is obtained by downloading from an open-source POI address website.

[0059] Among them, the explanation of POI is: Point of Interest.

[0060] There are several POI data in the POI address database. The POI address database is as Figure 2 shown.

[0061] Step S2 specifically includes the following steps:

[0062] S2.1: Input both the rough address features and all the POI data in the POI address database into the BERT model. Using the BERT model, vectorize the rough address features and the text addresses of all POI data.

[0063] The POI data includes longitude, latitude, and text address. Step S2.1 only vectorizes the text address of the POI data.

[0064] The explanation of the BERT model is: BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language model based on the Transformer architecture.

[0065] S2.2: Calculate the cosine similarity between the rough address features and the text address of each POI data.

[0066] The formula is:

[0067]

[0068] In formula (1), A is the vector representation of the rough address features; Bis the vector representation of one of the POI data text addresses in the POI address database; and are respectively A the L2 norm of the vector sum B and the vector. CS is the cosine similarity. represents the dot product of two vectors.

[0069] S2.3: Set the cosine similarity threshold R, and filter the POI data according to the cosine similarity threshold R. Mark the POI data whose cosine similarity with the vector A is greater than or equal to R as the POI data to be used to construct the surrounding circle; Use the POI data to be used to construct the surrounding circle to construct the smallest geographical surrounding circle.

[0070] S2.3 specifically includes the following steps:

[0071] S2.3.1: Mark in the POI address database the POI data whose cosine similarity with the vector A is greater than or equal to R; The marked POI data is the POI data to be used to construct the surrounding circle; R is selected as 0.95 in this embodiment.

[0072] S2.3.2: Adopt the method of the smallest polygon to construct all the POI data to be used to construct the surrounding circle into a polygon area. The constructed polygon area is shown in Figure 3 . The constructed polygon area is the smallest geographical surrounding circle.

[0073] For the text addresses of some POI data inside the constructed polygon area, their cosine similarity with the rough address features may be less than R, which does not affect the overall solution.

[0074] S3: Take all the POI data within the smallest geographical surrounding circle as the new POI database; Query the new POI database, calculate the theoretical coordinates according to the precise address features, and record the theoretical coordinates as ( X m , Y m ).

[0075] S3 specifically includes two cases:

[0076] Case 1: If there is at least one house number in the POI data text address in the new POI database that is the same as the house number corresponding to the precise address feature, the calculation formula for the theoretical coordinates is:

[0077]

[0078] In formula (2), X m andY m are the longitude and latitude of the theoretical coordinates respectively; X i and Y i are the longitude and latitude of the POI data with the same precise address feature respectively, ; n is the number of POI data with the same precise address feature, n≥ is 1 and is an integer.

[0079] Explanation of the situation where multiple POI data text addresses in the new POI database have the same house number:

[0080] For a building, different storefronts on its east and west sides may have the same house number. For example, for the building e the storefronts on the east and west sides are store S1 and store S2 respectively, but both store S1 and store S2 use the e house number of the building. Therefore, different stores have the same house number. However, store S1 and store S2 are different in geographical coordinates.

[0081] Situation 2: If there is no house number in the new POI database that is the same as the precise address feature, then query two house numbers in the new POI database that are close to the precise address feature house number, and denote them as C 1 and C 2 respectively. The calculation formula for the theoretical coordinates is:

[0082]

[0083] In formula (3), X m and Y m are the longitude and latitude of the theoretical coordinates respectively; C x is the house number corresponding to the precise address feature; and are the longitude and latitude of the house number C 1 respectively; and are the longitude and latitude of the house number C 2 respectively.

[0084] Illustrative example: Assume that the house number of the precise address feature C x is "No. 17", and in the new POI database, there are house numbers "No. 16", "No. 18", "No. 19", and "No. 21" respectively. Then the two house numbers close to "No. 17" are "No. 16" and "No. 18" respectively.

[0085] S4: Input the text addresses in S1 into each geocoding platform respectively, obtain the coding results of each geocoding platform, and form a candidate set with the coding results of each geocoding platform.

[0086] Specifically, input "No. 17, Focheng West Road, Jiangning District, Nanjing City, Jiangsu Province" into each geocoding platform respectively. The geocoding platforms include but are not limited to: AutoNavi, Baidu, Sogou, Tencent, etc. Suppose there are u geocoding platforms, and each geocoding platform will generate a coding result based on this text address. Therefore, a total of u coding results are generated. Let the respective coding results be: ( X z1 , Y z1 ), ( X z2 , Y z2 ), ……, ( X zu , Y zu ). The formed candidate set is: {( X z1 , Y z1 ), ( X z2 , Y z2 ), ……, ( X zu , Y zu )}.

[0087] S5: If none of the coding results in the candidate set are within the minimum geographical bounding circle, output the theoretical coordinates ( X m , Y m ), and provide the theoretical coordinates ( X m , Y m ) to the users as the final coding results; if at least one coding result in the candidate set is within the minimum geographical bounding circle, calculate the optimized geographical coordinates, denoted as ( X 优 , Y 优 ), output the optimized geographical coordinates ( X 优 , Y 优 ), and use the optimized geographical coordinates ( X 优 , Y优 ) is provided to the users as the final coding result.

[0088] S5 specifically includes the following steps:

[0089] S5.1: If none of the coding results in the candidate set are within the minimum geographical bounding circle, output the theoretical coordinates ( X m , Y m ); The theoretical coordinates ( X m , Y m ) are provided to the users as the final coding result.

[0090] S5.2: If at least one coding result in the candidate set is within the minimum geographical bounding circle, calculate the optimized geographical coordinates, denoted as ( X 优 , Y 优 ), and output the optimized geographical coordinates ( X 优 , Y 优 ).

[0091] S5.2 specifically includes the following steps:

[0092] S5.2.1: Calculate the Euclidean distance between each coding result and the theoretical coordinates. The formula is:

[0093]

[0094] In formula (4), .

[0095] S5.2.2: Calculate the coordinate weight of each coding result according to the Euclidean distance. The coordinate weight of each coding result is:

[0096]

[0097] In formula (5), w k represents the coordinate weight of the k th coding result. u is the number of coding results. d k represents the theoretical coordinates ( X m , Y m ) and the k th coding result ( X zk , Yzk ) Euclidean distance

[0098] S5.2.3: Calculate the optimized geographical coordinates:

[0099]

[0100] In formula (6), w k represents the coordinate weight of the k th coding result. u is the number of coding results. X zk and Y zk respectively represent the longitude and latitude of the k th coding result. ([[]] X 优 , Y 优 ) is the optimized geographical coordinate, and the optimized geographical coordinate ([[]] X 优 , Y 优 ) is provided to the users as the final coding result.

[0101] Case analysis:

[0102] To verify that this method can reduce the coding result errors of multiple coding platforms, 424 actual addresses are selected in this case, and the geographical coordinates of 424 actual addresses are accurately obtained. The 424 actual addresses all come from Gulou District, Nanjing City, Jiangsu Province, and the address types are diverse. Due to the large amount of data, only partial address data is provided, and partial address data is shown in Table 1. Each actual address contains rough address features and accurate address features.

[0103] Table 1 Actual address data

[0104]

[0105] The true geographical coordinates corresponding to each actual address are obtained through manual coding and manual correction.

[0106] Moreover, the coding results are respectively obtained through the open API interfaces of the existing three coding platforms of Baidu, Amap, and Tencent;

[0107] Then, the geographical coordinates are output through the optimization algorithm of this solution.

[0108] The statistical error data is shown in Table 2 and Figure 4 displayed.

[0109] Table 2 Errors between the coding results of each type of geographical coding platform and the true geographical coordinates

[0110]

[0111] As can be seen from Table 2, the optimization algorithm of this solution has the smallest error with the actual real geographical coordinates. Therefore, the optimization algorithm of this solution is beneficial to reducing the geographical coordinate error.

[0112] From Figure 4 it can be seen that the error results of this optimization algorithm mostly concentrate in the range of 0 - 50m. The number of errors decreases as the error range increases, and the number of errors greater than 500m is 0, which indicates that this optimization algorithm can significantly reduce the large - range errors greater than 500m.

[0113] Figure 5 is the process schematic diagram of this optimization algorithm.

[0114] Taking the ideal embodiments based on the present invention described above as an inspiration, through the above - described description, relevant staff can completely make various changes and modifications within the scope not deviating from the technical idea of this invention. The technical scope of this invention is not limited to the content in the specification, and its technical scope must be determined according to the scope of the claims.

Claims

1. A method for optimizing geographical coordinates of a multi-source network geocoding platform based on point-of-interest constraints, characterized in that, It includes the following steps: S1: The user inputs a text address into the CRF model. The CRF model extracts address features from the input text address and classifies the address features into rough address features and precise address features; S2: According to the rough address features, filter POI data in the POI address database to generate the smallest geographical bounding circle; S3: All POI data within the minimum geographic encirclement are used as a new POI database; query the new POI database, calculate theoretical coordinates based on precise address features, and mark the theoretical coordinates as (X m , Y m ); S4: Enter the text address in S1 into each geocoding platform respectively, obtain the coding results of each geocoding platform, and form a candidate set with the coding results of each geocoding platform; S5: If none of the coding results in the candidate set is within the minimum geographical bounding circle, output the theoretical coordinates (X m , Y m ); if at least one coding result in the candidate set is within the minimum geographical bounding circle, calculate the optimized geographical coordinates (X 优 , Y 优 ), and output the optimized geographical coordinates (X 优 , Y 优 ); Among them, step S2 further includes: S2.1: Input both the rough address features and all POI data in the POI address database into the BERT model; the BERT model vectorizes both the rough address features and the text addresses of all POI data; S2.2: Calculate the cosine similarity between the rough address features and the text addresses of each POI data; The formula is: In formula (1), A is the vector representation of the rough address features; B is the vector representation of the text address of one of the POI data in the POI address database; A·B represents the dot product of the two vectors; ||A|| and ||B|| are the L2 norms of vector A and vector B respectively; CS is the cosine similarity; S2.3: Set a cosine similarity threshold R. Mark the POI data whose cosine similarity with vector A is greater than or equal to R as the POI data to be used for constructing the bounding circle; use the method of the smallest polygon to construct all the POI data to be used for constructing the bounding circle into a polygon area, and the constructed polygon area is the smallest geographical bounding circle; Among them, step S5 includes the following steps: S5.1: If none of the encoding results in the candidate set is within the minimum geographic bounding circle, output the theoretical coordinates (X m , Y m ); S5.2: If at least one of the encoding results in the candidate set is within the minimum geographical bounding circle, calculate the optimized geographical coordinates, denoted as (X 优 , Y 优 ), and output the optimized geographical coordinates (X 优 , Y 优 ); S5.2 specifically includes the following steps: S5.2.1: Calculate the Euclidean distance between each coding result and the theoretical coordinates. The formula is: In formula (4), k ∈ [1, u]; d k represents the Euclidean distance between the theoretical coordinates (Xm, Ym) and the k-th coding result (X zk , Y zk ). S5.2.2: The coordinate weight of each coding result is: In formula (5), w k represents the coordinate weight of the k-th coding result; u is the number of coding results; d k represents the Euclidean distance between the theoretical coordinates (X m , Y m ) and the k-th coding result (X zk , Y zk ); S5.2.3: Calculate the optimized geographical coordinates: In formula (6), w k represents the coordinate weight of the k-th coding result; u is the number of coding results; X zk and Y zk respectively represent the longitude and latitude of the k-th coding result; (X 优 , Y 优 ) is the optimized geographical coordinate.

2. The geographic coordinate optimization method of the multi-source network geocoding platform based on point-of-interest constraints according to claim 1, characterized in that Step S3 includes two cases, Case 1: If there is at least one house number in the text address of the POI data in the new POI database that is the same as the house number corresponding to the precise address features, the calculation formula of the theoretical coordinates is: In formula (2), X m and Y m are the longitude and latitude of the theoretical coordinates; X i and Y i are the longitude and latitude of the POI data with the same exact address features, i ∈ [1, n], where i is an integer; n is the number of POI data with the same exact address features, and n ≥ 1; Case 2: If there is no house number in the new POI database that is the same as the precise address features, query two house numbers in the new POI database that are close to the house number of the precise address features, and record them as C1 and C2 respectively, The calculation formula of the theoretical coordinates is: In formula (3), X m and Y m are the longitude and latitude of the theoretical coordinates respectively; C x is the house number corresponding to the precise address feature; and are the longitude and latitude of the house number C1 respectively; and are the longitude and latitude of the house number C2 respectively.

3. The method for optimizing the geographical coordinates of a multi-source network geocoding platform based on point-of-interest constraints according to claim 2, characterized in that, In step S4, record the coding results of each geocoding platform as: (X z1 , Y z1 ), (X z2 , Y z2 ), ……, (X zu , Y zu ). The candidate set formed is: {(X z1 , Y z1 ), (X z2 , Y z2 ), ……, (X zu , Y zu )}, where u is the number of coding results.

4. The method for optimizing geographical coordinates of a multi-source network geocoding platform based on point-of-interest constraints according to claim 1, wherein The cosine similarity threshold R in step S2.3 is 0.95.

Citation Information

Patent Citations

  • Interest point address data processing method and device, server and medium

    CN110413904A

  • Systems and methods for location processing and geocoding optimization

    US20210256040A1