Automatic address cleaning method based on geographic information system weight analysis model
Through the geographic information system weight analysis model, combined with linear regression and map tool reverse analysis, the problem of inaccurate address information was solved, the correction and accuracy judgment of address information were achieved, and the authenticity and integrity of address information were improved.
Patent Information
- Application Number
- CN202310214957.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-08
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2043-03-08
AI Technical Summary
When judging the address information of patent applicants, the existing technology has problems such as non-standard, incomplete or inaccurate filling, which makes it impossible to accurately match the address information with the real address, especially in the case of complex address information, which is prone to misjudgment.
A weight analysis model based on the geographic information system is used to query the latitude and longitude coordinates through the standard coordinate system, establish a linear regression equation, construct a weighted model, use map tools to reversely analyze and verify the address, and adjust the weight value through manual comparison to achieve the correction of address information and accuracy judgment.
It improves the accuracy of address information matching, can correct false, erroneous or missing address information, and ensure the authenticity and integrity of address information.
Smart Images

Figure CN116257703B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to an automatic address cleaning method, in particular to an automatic address cleaning method based on a geographic information system weight analysis model. Background Art
[0002] There are more and more application scenarios for intellectual property data information. For example, address information can be used to monitor the maintenance of intellectual property rights in the target area and real-time changes in migration in and out, but this depends on the authenticity and accuracy of the address information.
[0003] However, in the application process of trademarks and patents, the address information related to applicants and other entities is often filled out in an irregular, unstable, and incomplete manner. When the data is entered into the database, the address information cannot be accurately matched with the actual address.
[0004] Currently, the method commonly used to determine the regional classification of patent applicants is to perform keyword segmentation on the applicant's address and match the applicant's address with the province, city, district, and county to achieve regional matching. For example, "11th Floor, Block B, Twin Towers, New District, Zhenjiang City, Jiangsu Province" is segmented to form a vocabulary of "Jiangsu", "Jiangsu Province", "Zhenjiang", "Zhenjiang City" [major elements] and "11th Floor, Block B, Twin Towers, New District" [minor elements]. Probability calculation is performed to form effective regional information. This regional matching method is commonly used in national regional sorting and pre-screening in the express delivery industry. However, blindly using statistical methods based on probability analysis cannot solve the problem of confirming the ownership of multiple address information. At the same time, it is easy to make mistakes in judgment under complex and interfering address information. For example, "No. 118, Nanjing Street, Xizang Road, Shanghai" will be seriously interfered with during segmentation and matching. Summary of the Invention
[0005] Purpose of the invention: The purpose of the present invention is to provide an automatic address cleaning method, which determines the accuracy of address information and performs corrections through a weight analysis model constructed based on a geographic information system.
[0006] Technical solution: A method for automatic address cleaning based on a geographic information system weight analysis model, comprising the following steps:
[0007] Step 100: Query the longitude and latitude coordinates of the subject's storage name through the standard coordinate system to obtain the coordinate set A{A1,…,A n}, n≥1, query the longitude and latitude coordinates of the main storage address through the standard coordinate system, and obtain the coordinate set B{B1,…,B m}, m≥1, determine whether the coordinate set A is completely consistent with the coordinate set B. If not, determine that there is a deviation in the storage address information;
[0008] Step 200: Select the coordinate set A and / or the coordinate set B,
[0009] Step 201: Using one of the longitude and latitude coordinates in the selected coordinate set, reverse parsing and verification are performed using the map tool of the standard coordinate system to obtain a verification address, extracting the verification address into large elements, and obtaining a matching address by combining the large elements with the stored name. The matching address is then queried for longitude and latitude coordinates using the standard coordinate system to obtain a coordinate set C{C1,…,C p}, p≥1,
[0010] Step 202: Execute step 201 for each other latitude and longitude coordinate of the selected coordinate set; obtain a total of i coordinate sets C, where i = n or m or not greater than n + m;
[0011] Step 300: Establish a linear regression equation based on the range of all longitude and latitude coordinates of coordinate set A, coordinate set B, and i groups of coordinate set C;
[0012] Step 400: Use coordinate set A, coordinate set B, and i groups of coordinate sets C as three weight items and assign weight values to build a weighted model. Then calculate the confidence value of each longitude and latitude coordinate and the regression line, and take the longitude and latitude coordinate corresponding to the minimum confidence value as the valid address.
[0013] Specifically, the subject is non-individual.
[0014] Specifically, the standard coordinate system includes the GCJ-02 coordinate system, the Martian coordinate system, and the Earth coordinate system.
[0015] Furthermore, the confidence value is W*|K|, where |K| is the absolute value of the discrete value of the longitude and latitude coordinates and the regression line, and W is the weight value of the weight item corresponding to the longitude and latitude coordinates.
[0016] Furthermore, when constructing the weighted model, the initial weight value of each weight item is assigned as 1 / the number of weight items. Then, through manual comparison of the regression line consistency of voting and / or reverse comparison of small elements of the storage address, the calculation results are verified to improve the weighted model and achieve model accuracy training for the weight value.
[0017] Furthermore, the manual comparison voting refers to the manual accuracy judgment of the latitude and longitude coordinates of the storage name, the real address of the subject, the large element + the storage name.
[0018] Beneficial effects: The advantages of the present invention are: the proposed automatic address cleaning method based on the geographic information system weight analysis model divides and compares the large elements in the address information based on the standard coordinate system, and constructs a weight analysis model with the results of the subject's name, address one-way query and large element joint query. The weight value can be dynamically improved to judge the accuracy of the address information, perform corrections to obtain accurate address information, and improve the accuracy of address information matching. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 The longitude and latitude coordinates A1 position map in Example 1;
[0020] Figure 2 The longitude and latitude coordinates B1 position map in Example 1;
[0021] Figure 3 The latitude and longitude coordinates C1 position map in Example 1;
[0022] Figure 4 is the linear regression graph in Example 1;
[0023] Figure 5 The longitude and latitude coordinates A1 position map in Example 2;
[0024] Figure 6 The latitude and longitude coordinates B1 position map in Example 2;
[0025] Figure 7 The latitude and longitude coordinates C1 position map in Example 2;
[0026] Figure 8 is the linear regression graph in Example 2;
[0027] Figure 9 The longitude and latitude coordinates A1 position map in Example 3;
[0028] Figure 10 The longitude and latitude coordinates B1 position map in Example 3;
[0029] Figure 11 The latitude and longitude coordinates C1 position map in Example 3;
[0030] Figure 12 is the linear regression graph in Example 3;
[0031] Figure 13 The longitude and latitude coordinates A1 position map in Example 4;
[0032] Figure 14 The latitude and longitude coordinates A2 position map in Example 4;
[0033] Figure 15 The latitude and longitude coordinates C1 position map of the first coordinate set C in Example 4;
[0034] Figure 16 The latitude and longitude coordinates C1 position map of the third coordinate set C in Example 4;
[0035] Figure 17 This is the linear regression graph in Example 4. DETAILED DESCRIPTION
[0036] The present invention will be further explained below with reference to the accompanying drawings and specific embodiments.
[0037] An automatic address cleaning method based on a geographic information system weight analysis model specifically comprises the following steps:
[0038] Step 100: Query the longitude and latitude coordinates of the subject's storage name through the standard coordinate system to obtain the coordinate set A{A1,…,A n}, n≥1, query the longitude and latitude coordinates of the main storage address through the standard coordinate system, and obtain the coordinate set B{B1,…,B m}, m≥1, determine whether the coordinate set A and the coordinate set B are completely consistent. If not, it is determined that there is a deviation in the storage address information.
[0039] Since the name of an individual subject, i.e., a name, does not have the attribute of an address, in this technical solution, the subject refers to an enterprise or institution, etc., and the names of these subjects have one or more address attributes when indexed by GIS. When data information is stored in the database, data with individuals as the subject is first filtered.
[0040] Standard coordinate systems include the GCJ-02 coordinate system, the Martian coordinate system, and the Earth coordinate system.
[0041] Step 200: Select one of the coordinate set A and / or the coordinate set B,
[0042] Step 201: Use one of the longitude and latitude coordinates in the selected coordinate set to perform reverse parsing verification using a map tool in the standard coordinate system to obtain a verification address, extract the verification address to obtain a large element, and obtain a matching address by combining the large element and the stored name. Query the longitude and latitude coordinates of the matching address using the standard coordinate system to obtain a coordinate set C{C1,…,C p}, p≥1,
[0043] Step 202: Execute step 201 for each other longitude and latitude coordinate of the selected coordinate set; obtain a total of i coordinate sets C, where i = n or m or not greater than n + m.
[0044] Map tools that use standard coordinate systems: Map tools that use the GCJ-02 coordinate system include Amap, and map tools that use the Martian coordinate system include Baidu Map.
[0045] Step 300: Establish a linear regression equation based on the range of all longitude and latitude coordinates of coordinate set A, coordinate set B, and i groups of coordinate sets C.
[0046] Step 400: Use coordinate set A, coordinate set B, and i groups of coordinate sets C as three weight items and assign weight values to build a weighted model. Then calculate the confidence value of each longitude and latitude coordinate and the regression line, and take the longitude and latitude coordinate corresponding to the minimum confidence value as the valid address.
[0047] When constructing the weighted model, the weight W of each weight item is initially assigned a value of 1 / number of weight items. The weighted model is then refined and trained for accuracy through manual comparison voting and / or regression line consistency of reverse comparison of small elements of the entry address. Manual comparison voting involves manually determining the accuracy of the entry name, the subject's actual address, and the large element + entry name, along with their respective latitude and longitude coordinates.
[0048] The confidence value is W*|K|, where |K| is the absolute value of the discrete value between the latitude and longitude coordinates and the regression line, and W is the weight value of the weight item corresponding to the latitude and longitude coordinates. For example, when calculating the confidence value of the latitude and longitude coordinate A2, W is the dynamic parameter that can be used to weight the weight item of the coordinate set A.
[0049] The following examples illustrate the application of the automatic address cleaning method of the present invention by using several cases of deviation in the stored data information.
[0050] Example 1
[0051] Taking the case where the storage address of the patent applicant (i.e., the subject) is false address information as an example, the automatic address cleaning method of the present invention is used to correct the error.
[0052] Patent applicant's name in the database: Nanjing University of Technology
[0053] The address of the patent applicant: Box 8020, No. 30, Puzhu South Road, Pukou District, Nanjing, Jiangsu Province
[0054] Step 100: Query the name in the database through the GCJ-02 coordinate system and only get a longitude and latitude coordinate (see attached Figure 1 ), then the coordinate set A{A1(118.640081,32.082496)}, the storage address is queried through the GCJ-02 coordinate system, and only one longitude and latitude coordinate is obtained (see attached Figure 2 ), then the coordinate set B{B1(118.657928,32.082914)}, due to the inconsistency between A1 and B1, by appending Figure 3 Manual judgment also reveals that B1 is a transceiver station near "Nanjing University of Technology", and it is determined that there is a deviation in the storage address information.
[0055] Step 200: Select coordinate set A from coordinate set A and coordinate set B.
[0056] Step 201: Using the only longitude and latitude coordinate A1 in the coordinate set A, reverse parsing verification is performed through AutoNavi Maps to obtain the verification address: No. 30, Puzhu South Road, Nanjing, Jiangsu Province. The verification address is segmented and extracted to obtain the large element: Nanjing, Jiangsu Province. The matching address is obtained by combining the large element and the stored name: Nanjing University of Technology, Nanjing, Jiangsu Province. The matching address is queried through the GCJ-02 coordinate system and only one longitude and latitude coordinate is obtained (see attached). Figure 3 ), then the coordinate set C{C1(118.640081,32.082496)}.
[0057] Step 300: Establish a linear regression equation (see Appendix) based on the range of all longitude and latitude coordinates of coordinate set A, coordinate set B, and a set of coordinate set C. Figure 4 ).
[0058] Step 400: With coordinate set A, coordinate set B, and coordinate set C as the three weight items, the initial weight values are all assigned to 1 / 3, and an equal-weighted model is constructed. Then, the confidence value of each longitude and latitude coordinate and the regression line is calculated. The longitude and latitude coordinates (118.640081, 32.082496) are closest to the regression line. It is determined that the confidence values of the longitude and latitude coordinates A1 and C1 are equal and both are the smallest, and both are valid addresses.
[0059] The results of the above methods are consistent with those of actual manual judgment and verification.
[0060] Example 2
[0061] Taking the case where the storage address of the patent applicant (i.e., the subject) is erroneous address information as an example, the automatic address cleaning method of the present invention is used to correct the error.
[0062] Patent applicant's registered name: Jiangsu Mingyue Optical Glasses Co., Ltd.
[0063] The address of the patent applicant: No. 200, Qiliang Road, Danyang Economic Development Zone, Jiangsu Province
[0064] Step 100: Query the name in the database through the GCJ-02 coordinate system and only get a longitude and latitude coordinate (see attached Figure 5 ), the location information is in Beijing, then the coordinate set A{A1(116.462689,39.878249)}, the storage address is queried through the GCJ-02 coordinate system, and only one longitude and latitude coordinate is obtained (see attached Figure 6 ), the location information is in Danyang, Zhenjiang, Jiangsu, then the coordinate set B{B1(119.617603,32.036547)}, due to the inconsistency between A1 and B1, it is determined that there is a deviation in the storage address information.
[0065] Step 200: Select coordinate set B from coordinate set A and coordinate set B.
[0066] Step 201: Using the only longitude and latitude coordinate B1 in the coordinate set B, reverse parsing verification is performed through the AutoNavi map to obtain the verification address: No. 200, Qiliang Road, Danyang City, Zhenjiang City, Jiangsu Province. The verification address is segmented and extracted to obtain the large element: Danyang City, Zhenjiang City, Jiangsu Province. The matching address is obtained by combining the large element and the stored name: Jiangsu Mingyue Optical Glasses Co., Ltd., Danyang City, Zhenjiang City, Jiangsu Province. The matching address is queried through the GCJ-02 coordinate system and only one longitude and latitude coordinate is obtained (see attached). Figure 7 ), the positioning information is in Danyang, Zhenjiang, Jiangsu, then the coordinate set is C{C1(119.617603,32.036547)}.
[0067] Step 300: Establish a linear regression equation (see Appendix) based on the range of all longitude and latitude coordinates of coordinate set A, coordinate set B, and a set of coordinate set C. Figure 8 ).
[0068] Step 400: With coordinate set A, coordinate set B, and coordinate set C as the three weight items, the initial weight values are all assigned to 1 / 3, and an equal-weighted model is constructed. Then, the confidence value of each longitude and latitude coordinate and the regression line is calculated. The longitude and latitude coordinates (119.617603, 32.036547) are closest to the regression line. It is determined that the confidence values of the longitude and latitude coordinates B1 and C1 are equal and both are the smallest, and both are valid addresses.
[0069] In this embodiment, "Danyang Province" in the storage address is erroneous information. Only through the present invention can the error of writing "Danyang City" as "Danyang Province" be corrected.
[0070] The results of the above methods are consistent with those of actual manual judgment and verification.
[0071] Example 3
[0072] Taking the case where the storage address of the patent applicant (ie, the subject) is missing address information as an example, the automatic address cleaning method of the present invention is used to correct the address information.
[0073] Patent applicant's registered name: Nanjing Ruisen Fuel Co., Ltd.
[0074] The address of the patent applicant: Qinfeng Village, Ma'an Town, Liuhe District, Nanjing City, Jiangsu Province
[0075] Step 100: Query the name in the database through the GCJ-02 coordinate system and only get a longitude and latitude coordinate (see attached Figure 9 ), the location information is in Nanjing, Jiangsu, then the coordinate set A{A1(118.802459,32.414463)}, the storage address is queried through the GCJ-02 coordinate system, and only one longitude and latitude coordinate is obtained (see attached Figure 10), the location information is in Nanjing, Jiangsu, then the coordinate set B{B1(118.804505,32.40086)}, due to the inconsistency between A1 and B1, it is determined that there is a deviation in the storage address information.
[0076] Step 200: Select coordinate set B from coordinate set A and coordinate set B.
[0077] Step 201: Using the only longitude and latitude coordinate B1 in the coordinate set B, reverse parsing verification is performed through the AutoNavi map to obtain the verification address: Qinfeng Village Committee, Nanjing City, Jiangsu Province. The verification address is segmented and extracted to obtain the large element: Nanjing City, Jiangsu Province. The matching address is obtained by combining the large element and the stored name: Nanjing Ruisen Fuel Co., Ltd., Nanjing City, Jiangsu Province. The matching address is queried through the GCJ-02 coordinate system and only one longitude and latitude coordinate is obtained (see attached). Figure 11 ), the positioning information is in Nanjing, Jiangsu, then the coordinate set is C{C1(119.617603,32.036547)}.
[0078] Step 300: Establish a linear regression equation (see Appendix) based on the range of all longitude and latitude coordinates of coordinate set A, coordinate set B, and a set of coordinate set C. Figure 12 ).
[0079] Step 400: With coordinate set A, coordinate set B, and coordinate set C as the three weight items, and the initial weight values are all assigned to 1 / 3, an equal-weighted model is constructed, and then the confidence value of each longitude and latitude coordinate and the regression line is calculated. The longitude and latitude coordinates (119.617603, 32.036547) are closest to the regression line and have the smallest confidence value, so the longitude and latitude coordinates C1 are judged to be a valid address.
[0080] In this embodiment, the missing detailed small element information in the storage address leads to inaccurate positioning range, which can be corrected by the present invention.
[0081] The results of the above methods are consistent with those of actual manual judgment and verification.
[0082] Example 4
[0083] Taking the case where the storage address of a contact (i.e., subject) is erroneous address information as an example, the automatic address cleaning method of the present invention is used to correct the error.
[0084] Contact person's database name: School of Economics and Management, Jiangsu University of Science and Technology
[0085] Contact person's storage address: No. 666, Changhui Road, Dantu District, Zhenjiang City
[0086] Step 100: Query the stored name through the GCJ-02 coordinate system and obtain two longitude and latitude coordinates, A1 (119.468078, 32.197314), and the location information is in Zhenjiang City, Jiangsu Province (see attached Figure 13 ), A2 (116.357342, 39.993032), the location information is in Haidian District, Beijing (see attached Figure 14 ), then the coordinate set A{A1(119.468078,32.197314),A2(116.357342,39.993032)}, the storage address is queried through the GCJ-02 coordinate system, and only one longitude and latitude coordinate is obtained, then the coordinate set B{B1(119.360779,32.109739)}, since A1, A2, and B1 are inconsistent, it is judged that there is a deviation in the storage address information.
[0087] Step 200: Select coordinate set A and coordinate set B,
[0088] Step 201: Using the only longitude and latitude coordinate B1 in coordinate set B, reverse parsing verification is performed through Amap, and the verification address is obtained: No. 666, Changhui Road, Dantu District, Zhenjiang City, Jiangsu Province. The verification address is segmented and extracted to obtain the large element: Zhenjiang City, Jiangsu Province. The matching address is obtained by combining the large element and the stored name: School of Economics and Management, Jiangsu University of Science and Technology, Zhenjiang City, Jiangsu Province. The matching address is queried through the GCJ-02 coordinate system, and only one longitude and latitude coordinate is obtained (see attached). Figure 15 ), the positioning information is in Zhenjiang, Jiangsu, then the first set of coordinates is C{C1(119.468078,32.197307)}.
[0089] In step 202, one of the longitude and latitude coordinates A1 in the coordinate set A is reverse parsed and verified using AutoNavi Maps to obtain a verification address. The verification address is segmented and extracted to obtain a large element: Zhenjiang City, Jiangsu Province. The large element plus the stored name is used to obtain a matching address: School of Economics and Management, Jiangsu University of Science and Technology, Zhenjiang City, Jiangsu Province. The matching address is queried using the GCJ-02 coordinate system, obtaining only one longitude and latitude coordinate. The positioning information is in Zhenjiang City, Jiangsu Province. Therefore, the second coordinate set C is {C1(119.468078,32.197307)}.
[0090] Using another latitude and longitude coordinate A2 of coordinate set A, reverse parsing verification is performed through Amap to obtain the verification address. The verification address is segmented and extracted to obtain the large element: School of Economics and Management, Jiangsu University of Science and Technology, Haidian District, Beijing. The matching address is queried through the GCJ-02 coordinate system and only one latitude and longitude coordinate is obtained (see attached Figure 16 ), the positioning information is in Haidian District, Beijing, then the third coordinate set is C{C1(116.357342,39.993028)}.
[0091] Step 300: Establish a linear regression equation (see Appendix) based on the range of all longitude and latitude coordinates of coordinate set A, coordinate set B, and three sets of coordinate sets C. Figure 17 ).
[0092] Step 400: With coordinate set A, coordinate set B, and three sets of coordinate sets C as three weight items, and the initial weight values are all assigned to 1 / 3, an equal-weighted model is constructed, and then the confidence value of each longitude and latitude coordinate and the regression line is calculated. The longitude and latitude coordinates (119.468078, 32.197314) are closest to the regression line and have the smallest confidence value, so the longitude and latitude coordinates A1 are judged to be a valid address.
[0093] In this embodiment, the address of the branch campus without "School of Economics and Management" is mistakenly written as the main campus address of "School of Economics and Management of Jiangsu University of Science and Technology", resulting in inaccurate positioning range. The present invention can correct the error.
[0094] The results of the above methods are consistent with those of actual manual judgment and verification.
Claims
1. An automatic address cleaning method based on a geographic information system weight analysis model, characterized in that The following steps are involved: Step 100: Query the longitude and latitude coordinates of the subject's storage name through the standard coordinate system to obtain the coordinate set A{A1,…,A n }, n≥1, query the longitude and latitude coordinates of the main storage address through the standard coordinate system, and obtain the coordinate set B{B1,…,B m }, m≥1, determine whether the coordinate set A is completely consistent with the coordinate set B. If not, determine that there is a deviation in the storage address information; Step 200: Select the coordinate set A and / or the coordinate set B, Step 201: Using one of the longitude and latitude coordinates in the selected coordinate set, reverse parsing and verification are performed using the map tool of the standard coordinate system to obtain a verification address, extracting the verification address into large elements, and obtaining a matching address by combining the large elements with the stored name. The matching address is then queried for longitude and latitude coordinates using the standard coordinate system to obtain a coordinate set C{C1,…,C p }, p≥1, Step 202: Execute step 201 for each other latitude and longitude coordinate in the selected coordinate set; A total of i coordinate sets C are obtained, where i = n or m or not greater than n + m; Step 300: Establish a linear regression equation based on the range of all longitude and latitude coordinates of coordinate set A, coordinate set B, and i groups of coordinate set C; Step 400: Using coordinate set A, coordinate set B, and i sets of coordinate sets C as three weight items and assigning weight values, a weighted model is constructed. The confidence value between each longitude and latitude coordinate and the regression line is calculated, and the longitude and latitude coordinate corresponding to the minimum confidence value is regarded as the valid address. The confidence value is W*|K|, |K| is the absolute value of the longitude and latitude coordinates and the discrete value of the regression line, and W is the weight value of the weight item corresponding to the longitude and latitude coordinates; when constructing the weighted model, the initial weight value assignment of each weight item is 1 / the number of weight items, and then the calculation results are verified and improved through manual comparison voting and / or reverse comparison of the regression line consistency of the small elements of the storage address, and the model accuracy training of the weight value is achieved; the manual comparison voting refers to the manual accuracy judgment of the longitude and latitude coordinates of the storage name, the real address of the subject, the large element + the storage name.
2. The automatic address cleaning method based on the geographic information system weight analysis model according to claim 1 is characterized in that: The subject is non-personal.
3. The automatic address cleaning method based on the geographic information system weight analysis model according to claim 1 is characterized in that: The standard coordinate systems include the GCJ-02 coordinate system, the Martian coordinate system, and the Earth coordinate system.
Citation Information
Patent Citations
Airborne infrared moving target detection method based on geographical homologous point registration
CN106056625A
Maintaining method and system for hotel longitude-latitude information on OTA website
CN106991185A