A mapping system for IP geolocation data
Through reference point screening and dynamic clustering analysis based on geolocation entropy, combined with network topological similarity calculation and shortest path correction, the problem of insufficient geolocation accuracy of IP in the existing technology is solved, and high-precision positioning of reference point and non-reference point IP is achieved.
Patent Information
- Application Number
- CN202411828796.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-12
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-12-12
Smart Images

Figure CN119729763B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of IP geolocation, and particularly relates to a mapping system for IP geolocation data. Background Art
[0002] Currently, the IP geolocation technology has the problem of insufficient accuracy. For example, the existing technologies mainly rely on IP address allocation rules, routing detection, and intelligence analysis methods, which can only achieve the positioning accuracy at the city level and are difficult to meet the requirements of high-precision geolocation. Especially in scenarios such as cyber space asset supervision, map-based operations, asset exposure surface sorting, and industry analysis, the city-level accuracy cannot meet the precise requirements for IP address positioning, resulting in problems such as incomplete data coverage and large positioning errors in actual applications. In addition, the existing technologies have obvious deficiencies in terms of the utilization rate of reference point data, algorithm dynamic optimization, real-time positioning of non-reference point IPs, and adaptability to complex scenarios. For example, the geographical distribution data of reference point IPs has not been effectively mapped and utilized. Traditional methods are mostly based on static rules or simple algorithms and cannot cope with the multi-dimensional data requirements in complex network environments. For non-reference point IPs without a clearly bound geographical location, there is a lack of accurate dynamic positioning ability, resulting in accuracy far lower than the requirements of actual applications. In dynamic network scenarios or complex network topologies, the accuracy rate of the existing technologies is significantly reduced, and it is impossible to achieve street-level or even meter-level positioning accuracy. Therefore, there is an urgent need for a high-precision IP positioning method that combines data mapping and network measurement technologies and can integrate multiple data sources, so as to still be able to achieve street-level or even meter-level geolocation when both reference point data and non-reference point data are applicable, thereby greatly improving the accuracy and applicability of IP positioning and meeting the high-precision positioning requirements in multiple scenarios. Summary of the Invention
[0003] Aiming at the above-mentioned technical deficiencies, the purpose of the present invention is to propose a mapping system for IP geolocation data, aiming to solve the technical problems in the existing technologies that only rely on IP allocation rules, static routing detection, or single algorithms to achieve city-level positioning, especially the inability to fully utilize reference point data or accurately position non-reference point IP addresses in complex network environments.
[0004] To solve the above technical problems, the present invention adopts the following technical solutions: The present invention provides a mapping system for IP geolocation data, including:
[0005] A data acquisition module, configured to obtain IP basic information from the World Wide Web and mobile platforms to form an open data set, where the open data set includes IP addresses, BGP routing information, Whois registration data, open WIFI hotspot information, host-domain name association information, user device records, and geographical distribution statistical data;
[0006] Standardize the obtained open dataset, and calculate the geographical location entropy value H for the geographical data point i in the standardized open dataset to obtain the geographical location entropy value H i , and retain the geographical location entropy value H i , retain the geographical data points whose geographical location entropy value H is greater than the preset geographical location entropy value threshold to form a benchmark point dataset; retain the geographical location entropy value H i , retain the geographical data points whose geographical location entropy value H is less than or equal to the preset geographical location entropy value threshold to form a non-benchmark point dataset;
[0007] A scenario division module, which is used to extract the feature vector F and the IP address from the open dataset. The feature vector F includes access frequency features, port distribution features, domain name association information features, and geographical distribution information features. Based on the feature vector F, the K-means clustering algorithm is used to divide the IP addresses into different application scenarios, including enterprise dedicated lines, residential users, and school institutions;
[0008] Use the division results of different application scenarios as classification labels, and mark the benchmark point dataset to generate a scenario-specific benchmark point dataset;
[0009] A clustering analysis module, which is used to set dynamic clustering parameters for the scenario-specific benchmark point dataset. The dynamic clustering parameters include the distance threshold ∈ d and the minimum number of points N min , and perform dynamic clustering analysis on the scenario-specific benchmark point dataset using the density clustering algorithm DBSCAN. The density clustering algorithm DBSCAN outputs the first-level positioning result P u , including the street-level benchmark point IP geographical location range of different application scenarios;
[0010] A topology analysis module, which is used to obtain the network path information of the data points in the non-benchmark point dataset through a network detection tool for the data points in the non-benchmark point dataset, extract the first adjacent set E t of the data points in the non-benchmark point dataset, preset the target IP for the data points in the non-benchmark point dataset and construct the second adjacent set E b , and calculate the network topology similarity S t between the target IP and the data points in the non-benchmark point dataset according to the first adjacent set E b and the second adjacent set E t , bind the target IP to the data points in the non-benchmark point dataset based on the calculation result of the network topology similarity S t to obtain the non-benchmark point IP geographical location range, and output the second-level positioning result Q t after correcting the non-benchmark point IP geographical location range through the shortest path algorithm, including the street-level non-benchmark point IP geographical location range;
[0011] A location output module, which is used to fuse the positioning results of the clustering analysis module and the topology analysis module, and generate and output the final IP geographical positioning data after verifying the consistency of the positioning results.
[0012] Preferably, in the data acquisition module, for the step of standardizing the obtained open dataset, the normalization formula is adopted:
[0013]
[0014] where x' is the normalized open dataset, x is the open dataset, μ is the mean of the open dataset, and σ is the standard deviation of the open dataset.
[0015] Preferably, in the clustering analysis module, in the density clustering algorithm DBSCAN, the formula for judging the geographical distance between points in the scene-specific reference point dataset is:
[0016]
[0017] where φ m and λ m are the latitude and longitude of the geographical data point m, and φ n and λ n are the latitude and longitude of the geographical data point n.
[0018] Preferably, in the clustering analysis module, the dynamic clustering parameters are adjusted according to the optimization objective function F, and the formula for the optimization objective function F is:
[0019] F(∈ d , N min ) = α @ R + β · P
[0020] where F(∈ d , N min ) is the optimization objective function; ∈ d is the distance threshold in the dynamic clustering parameters; N min is the minimum number of points in the dynamic clustering parameters; α and β are weight coefficients; R is the recall rate of reference points, which is used to measure the coverage of reference points in the clustering results; P is the geographical positioning accuracy, which is used to measure the matching degree between the clustering range and the actual geographical location.
[0021] Preferably, in the network topology module, the calculation formula for the network topology similarity S t is:
[0022]
[0023] where |E t ∩ E p | is the number of common adjacent nodes between the target IP and the data points in the non-reference point dataset.
[0024] Preferably, in the position output module, the formula for verifying the consistency of the positioning result is:
[0025]
[0026] where C is the consistency index.
[0027] Preferably, in the data acquisition module, the geographical location entropy value H i is calculated by the formula:
[0028]
[0029] where H i is the geographical location entropy value of the geographical data point i, n is the number of geographical data points divided within the preset geographical area, and p i is the occurrence probability of the IP address of the data point in the open dataset within the geographical data point i.
[0030] The beneficial effects of the present invention are as follows: Compared with the prior art that only relies on IP allocation rules, static route detection or a single algorithm to achieve city-level positioning, especially the technical problems that in a complex network environment, the benchmark point data cannot be fully utilized or the non-benchmark point IP addresses cannot be accurately located, this application realizes the street-level geographical positioning with a meter-level accuracy for benchmark points and non-benchmark point IPs through technical means such as benchmark point screening based on geographical location entropy, scene-specific dynamic clustering analysis, network topology similarity calculation, and shortest path correction. Thus, it avoids the problems of low utilization rate of benchmark point data and insufficient positioning accuracy in the prior art, and greatly improves the accuracy rate and application scope of IP positioning. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0032] Figure 1 It is a schematic flowchart of the first embodiment of a mapping system for IP geographical location data provided by the present invention.
[0033] Figure 2 It is a schematic diagram of the equipment of a mapping system for IP geographical location data provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0035] Embodiment 1: As Figure 1 shown, it is a schematic flowchart of the first embodiment of the IP geolocation data mapping system of the present invention, and the first embodiment of the IP geolocation data mapping system of the present invention is proposed.
[0036] In the first embodiment, the IP geolocation data mapping system includes:
[0037] A data acquisition module, which is used to obtain IP basic information from the World Wide Web and mobile platforms to form an open data set. The open data set includes IP addresses, BGP routing information, Whois registration data, open WIFI hotspot information, host-domain name association information, user device records, and geographical distribution statistical data;
[0038] Perform standardization processing on the obtained open data set, calculate the geographical location entropy value H for the geographical data point i in the standardized open data set, i , and retain the geographical location entropy value H i Greater than the preset geographical location entropy value threshold of the geographical data points, to form a reference point data set; retain the geographical location entropy value H i Less than or equal to the preset geographical location entropy value threshold of the geographical data points, to form a non-reference point data set;
[0039] It should be noted that the BGP routing information included in the open data set in the data acquisition module is used to provide the approximate network path of the IP address, the Whois registration data is used to extract the ownership information and management agency information of the IP address, the open WIFI hotspot information provides geographical marking points such as longitude and latitude and MAC addresses in public places, the host-domain name association information can help further identify the usage environment of the IP address such as enterprise-specific IP or server IP, and the user device records and geographical distribution statistical data provide auxiliary data support for subsequent positioning from the perspectives of behavior and spatial distribution.
[0040] It can be understood that the standardization processing performed by the data acquisition module on the open data set is to unify the data from different sources into a consistent format and numerical range.
[0041] It should be understood that the calculation of the geographical location entropy value is based on the probability of the IP data distribution within each geographical region, and the geographical location entropy value is calculated through the following formula:
[0042]
[0043] Among them, H i is the geographical location entropy value of geographical data point i, n is the number of geographical data points divided within the preset geographical area, and p i is the occurrence probability of the IP address of the data point in the open dataset within the geographical data point i.
[0044] For example, an open dataset contains access records of an IP address, distributed in five regions: Region A (50 times), Region B (30 times), Region C (20 times), Region D (0 times), Region E (0 times). Then the corresponding probability distribution is: pA = 0.5, pB = 0.3, pC = 0.2, pD = 0, pE = 0. Through the entropy value formula calculation, we get: H = -(0.5log(0.5) + 0.3log(0.3) + 0.2 * log(0.2)) ≈ 1.485 bits. If the entropy value threshold is set to 1.2, then the entropy value H = 1.485 of this point exceeds the threshold and is retained as part of the reference point dataset. Through the cooperation of the above modules and algorithms, data points with geographical location stability can be effectively screened out, providing reliable reference point data support for subsequent high-precision IP geolocation.
[0045] The scenario division module is used to extract the feature vector F and the IP address from the open dataset. The feature vector F includes access frequency features, port distribution features, domain name association information features, and geographical distribution information features. Based on the feature vector F, the K-means clustering algorithm is used to divide the IP addresses into different application scenarios, including enterprise dedicated lines, residential users, and school institutions;
[0046] Take the division results of different application scenarios as classification labels and mark the reference point dataset to generate a scenario-specific reference point dataset;
[0047] It should be noted that the extraction of the feature vector F is to reflect different usage patterns of the IP address. For example, the access frequency can be measured by counting the daily access volume of the IP address, the port distribution can represent the types and proportions of ports opened by the IP address, the domain name association information can reflect the binding relationship between the IP address and a specific domain name, and the geographical distribution information reflects the usage range and concentration degree of this IP address.
[0048] It can be understood that the core of the K-means clustering algorithm is to iteratively optimize so that each IP address belongs to the application scenario class with the closest features.
[0049] It should be understood that after the application scenarios are divided, it is necessary to associate the benchmark point dataset with the classification results of different scenarios to generate scenario-specific benchmark point datasets. The purpose of this operation is to refine the classification of benchmark points so that more accurate location analysis can be carried out in specific application scenarios.
[0050] For example, assume that the feature vector of an IP address is F = [access frequency = 100 times / day, port distribution = 80 / 443 / 22, domain name association = example.com, geographical distribution = concentrated in a certain office area]. After K-means clustering, this IP address is classified into the enterprise dedicated line scenario. Further combining the geographical location entropy value of this IP address and the benchmark point dataset, it is labeled with the enterprise dedicated line label and incorporated into the benchmark point dataset of the enterprise dedicated line scenario to provide support for subsequent location analysis. In this way, more accurate IP location results can be achieved under different application scenarios.
[0051] The clustering analysis module is used to set dynamic clustering parameters for the scenario-specific benchmark point dataset. The dynamic clustering parameters include a distance threshold ∈ d and a minimum number of points N min , and the density-based spatial clustering of applications with noise (DBSCAN) algorithm is used to perform dynamic clustering analysis on the scenario-specific benchmark point dataset. The DBSCAN algorithm outputs the first-level location result P u , including the street-level benchmark point IP geographical location ranges of different application scenarios;
[0052] It should be noted that the distance threshold in the dynamic clustering parameters is used to limit the maximum geographical distance between benchmark points. Benchmark points exceeding this threshold will not belong to the same cluster; the minimum number of points is used to define the minimum number of benchmark points required to form a valid cluster, avoiding misclustering caused by isolated points or sparse areas. The DBSCAN algorithm analyzes the scenario-specific benchmark point dataset through these two parameters and can automatically identify the number and distribution of clusters.
[0053] It can be understood that the working principle of DBSCAN is to calculate the number of neighborhood points of each point with the distance threshold as the radius. If the number of neighborhood points of a certain point is greater than or equal to the minimum number of points, then this point is regarded as a core point, and all points within its neighborhood are grouped into the same cluster; for points that do not reach the minimum number of points, they are regarded as noise points or isolated points and do not participate in clustering.
[0054] It should be understood that by dynamically adjusting the distance threshold and the minimum number of points, the clustering effect can be optimized for different application scenarios. For example, in the enterprise dedicated line scenario, since the IP distribution is relatively concentrated, the distance threshold can be set smaller (such as 50 meters), while in the residential user scenario, the IP distribution is more dispersed, so the distance threshold can be appropriately relaxed (such as 200 meters) to improve the clustering coverage rate.
[0055] For example, assume that the reference point dataset in a certain enterprise dedicated line scenario contains 10 reference points, and their geographical locations are relatively concentrated. Set the distance threshold to 50 meters and the minimum number of points to 3. Through analysis, the DBSCAN algorithm finds that the maximum distance between three points (A, B, and C) is less than 50 meters, and the number of points in the neighborhood of each point is greater than or equal to 3. Then, A, B, and C are classified into the same cluster, and the street-level geographical range is output as the clustering result of the reference points; in the residential user scenario, if the distance threshold is set to 200 meters, the same method can cover a larger range of reference point clustering results.
[0056] Through the above clustering analysis module, it is possible to generate the street-level reference point IP geographical location range for different application scenarios, providing key input data support for the subsequent positioning of non-reference points.
[0057] The topology analysis module is used to obtain the network path information of the data points in the non-reference point dataset through a network detection tool for the data points in the non-reference point dataset, and extract the first adjacency set E of the data points in the non-reference point dataset t , preset a target IP for the data points in the non-reference point dataset and construct a second adjacency set E b , according to the first adjacency set E t and the second adjacency set E b calculate the network topology similarity S between the target IP and the data points in the non-reference point dataset t , based on the calculation result of the network topology similarity S t bind the target IP to the data points in the non-reference point dataset to obtain the non-reference point IP geographical location range, and output the second-level positioning result Q after correcting the non-reference point IP geographical location range through the shortest path algorithm t , including the street-level non-reference point IP geographical location range;
[0058] It should be noted that the network topology similarity is used to measure the similarity degree of the network path structure between the target IP and the data points in the non-reference point dataset, and is defined as the proportion of the common adjacent points between the target IP and the non-reference point; the network detection tool can provide the network path information of the target IP, including the path hop count and intermediate nodes. By extracting the adjacency set, the topology similarity between the target IP and the non-reference point data points can be calculated.
[0059] It should be understood that by calculating the network topology similarity, the target IP can be bound to the non-reference point data point with the most similar path. After the binding is completed, combined with the street-level reference point IP geographical location range and the shortest path algorithm, the positioning error of the target IP can be further corrected to ensure that the positioning accuracy of the non-reference point IP meets the street-level requirements.
[0060] For example, assume that the adjacent set of a target IP is {A, B, C, D}, and the adjacent sets of non-reference points b1 and b2 are E_b1 = {B, C, D, E} and E_b2 = {A, C, F, G} respectively. According to the formula, the topological similarity is calculated as S_t,b1 = 2 * 3 / (4 + 4) = 0.75, S_t,b2 = 2 * 2 / (4 + 4) = 0.5; because S_t,b1 > S_t,b2, the target IP is bound to the non-reference point b1. Subsequently, the positioning range of b1 is corrected by the shortest path algorithm, and finally the street-level positioning result of the target IP is output.
[0061] Through the topological analysis module, the positioning problem of non-reference point IPs can be effectively solved, providing accurate input data for subsequent result fusion.
[0062] The position output module is used to fuse the positioning results of the clustering analysis module and the topological analysis module, and generate and output the final IP geolocation data after verifying the consistency of the positioning results.
[0063] It should be noted that the clustering analysis module and the topological analysis module respectively handle the geolocation of reference point IPs and non-reference point IPs. There may be intersections or deviations in the results generated by the two in some areas. The role of the position output module is to fuse and verify the consistency of the two types of results to ensure that the finally output geolocation data has high accuracy and reliability. The core method for verifying consistency is to calculate the coincidence degree between the positioning results from different sources to quantify its consistency level.
[0064] It should be understood that when the consistency verification result is lower than the preset threshold, the parameters of the clustering analysis and topological analysis modules such as the distance threshold, minimum number of points, or topological similarity threshold need to be readjusted to improve the matching degree of the positioning results. After the consistency verification passes, the two types of results are fused to output the final positioning data.
[0065] In addition, the present invention also provides a mapping device for IP geolocation data. Please refer to Figure 2, A mapping device for IP geolocation data includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a mapping system for IP geolocation data in the first embodiment above. A mapping device for IP geolocation data in the embodiments of the present invention may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistant), PADs (Portable Application Description: tablet computers), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. A mapping device for IP geolocation data is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present invention. A mapping device for IP geolocation data may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of a mapping device for IP geolocation data are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow a mapping device for IP geolocation data to communicate with other devices wirelessly or wiredly to exchange data. Although a mapping device for IP geolocation data with various systems is shown in the figure, it should be understood that it is not required to implement or have all the shown systems. Instead, more or fewer systems may be implemented or had.
[0066] The present invention also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of a mapping system for IP geolocation data as described above. The computer program product provided by the present invention can solve the technical problem of mapping IP geolocation data. Compared with the prior art, the beneficial effects of the computer program product provided by the present invention are the same as those of the mapping system for IP geolocation data provided in the above embodiments, and will not be elaborated herein.
[0067] Specifically, according to the embodiments disclosed by the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed by the present invention include a computer program product which includes a computer program carried on a computer-readable medium. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by a processing device 1001, it implements the above functions defined in the system of the embodiments disclosed by the present invention.
[0068] It should be understood that the various parts disclosed by the present invention can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in a suitable manner in any one or more of the embodiments or examples.
[0069] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A system for mapping IP geolocation data, characterized in that: The system includes: A data collection module is used to obtain IP basic information from the World Wide Web and mobile platforms to form an open data set, which includes IP addresses, BGP routing information, Whois registration data, open WIFI hotspot information, host and domain name association information, user device records and geographical distribution statistics; The obtained open data set is standardized, and the geographic location entropy value of the geographic data point i in the standardized open data set is calculated to obtain the geographic location entropy value H i , retain the geographic location entropy value H i Geographic data points with entropy values greater than a preset geographic location threshold form a benchmark point data set; Keep the geographic location entropy value H i Geographic data points that are less than or equal to a preset geographic location entropy value threshold form a non-reference point data set; The scenario division module is used to extract feature vectors F and IP addresses from open data sets. Feature vector F includes access frequency features, port distribution features, domain name association information features, and geographic distribution information features. Based on feature vector F, the K-means clustering algorithm is used to divide IP addresses into different application scenarios, including enterprise dedicated lines, residential users, and school institutions. The division results of different application scenarios are used as classification labels, and the benchmark point datasets are marked to generate scenario-specific benchmark point datasets; The clustering analysis module is used to set dynamic clustering parameters for scene-specific benchmark data sets. The dynamic clustering parameters include the distance threshold ∈ d and the minimum number of points N min , the density clustering algorithm DBSCAN is used to perform dynamic clustering analysis on the scene-specific benchmark data set, and the first-level positioning result P is output. u , including the IP geographic location range of street-level benchmarks for different application scenarios; The topology analysis module is used to obtain the network path information of the data points of the non-reference point data set through the network detection tool, and extract the first adjacent set E of the data points of the non-reference point data set. t , preset the target IP for the data points of the non-reference point dataset and construct the second adjacent set E b , according to the first adjacency set E t and the second adjacent set E b Calculate the network topology similarity S between the target IP and the data points of the non-reference point dataset t , based on the network topology similarity S t The calculation result binds the target IP to the data point of the non-reference point data set to obtain the non-reference point IP geographical location range, and corrects the non-reference point IP geographical location range through the shortest path algorithm to output the second-level positioning result Q t , including street-level non-reference point IP geographic location ranges; The location output module is used to integrate the positioning results of the cluster analysis module and the topology analysis module, verify the consistency of the positioning results, and then generate and output the final IP geolocation data.
2. A system for mapping IP geographic location data as claimed in claim 1, characterized in that: In the data collection module, the step of standardizing the obtained open data set adopts the normalization formula: Among them, x' is the normalized open dataset, x is the open dataset, μ is the mean of the open dataset, and σ is the standard deviation of the open dataset.
3. The system for mapping IP geographic location data according to claim 1, characterized in that: In the cluster analysis module, in the density clustering algorithm DBSCAN, the formula used to determine the geographic distance between points in the scene-specific benchmark data set is: Among them, φ m and λ m is the latitude and longitude of geographic data point m, φ n and λ n is the latitude and longitude of geographic data point n.
4. The system for mapping IP geographic location data according to claim 1, characterized in that: In the cluster analysis module, the dynamic clustering parameters are adjusted according to the optimization objective function F. The formula of the optimization objective function F is: F(∈ d ,N min )=α·R+β@P Among them, F(∈ d ,N min ) is the optimization objective function; ∈ d is the distance threshold in the dynamic clustering parameters; N min is the minimum number of points in the dynamic clustering parameters; α and β are weight coefficients; R is the benchmark recall rate, which is used to measure the coverage of benchmark points in the clustering results; P is the geographic positioning accuracy, which is used to measure the matching degree between the clustering range and the actual geographic location.
5. The system for mapping IP geographic location data according to claim 1, characterized in that: In the network topology module, the network topology similarity S t The calculation formula is: Among them, |E t ∩E p | is the number of common adjacent nodes between the target IP and the data points of the non-benchmark point dataset.
6. The system for mapping IP geographic location data according to claim 1, characterized in that: In the position output module, the consistency of the positioning results is verified using the formula: Among them, C is the consistency index.
7. The system for mapping IP geographic location data according to claim 1, characterized in that: In the data collection module, the geographic location entropy value H i The calculation formula is: Among them, H i is the geographic location entropy value of geographic data point i, n is the number of geographic data points divided in the preset geographic area, p i is the probability of the IP address of a data point in the open dataset appearing in geographic data point i.
Citation Information
Patent Citations
Systems, methods, and devices for payment recovery platform
CA3052163A1
Geospatial data grading evaluation method and device based on decision tree and medium
CN119046394A