Web access geographic position anomaly detection method based on density estimation
By using DBSCAN spatial clustering algorithm and density estimation technology in Web access log exception detection, the problem of difficulty in utilizing IP address geolocation information in the existing technology is solved, and accurate detection and identification of Web access geolocation abnormalities is achieved.
Patent Information
- Application Number
- CN202411983830.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-16
AI Technical Summary
In the prior art, in the abnormal detection of web access logs, it is difficult to effectively utilize the geographical location information of IP addresses, resulting in the inability to accurately identify the abnormal access behavior of geolocation.
The Web access geographic location anomaly detection method based on density estimation is adopted, and the DBSCAN spatial clustering algorithm is used to obtain the geographical location through IP addresses, train the geographical location distribution model, and perform clustering and anomaly detection.
It realizes flexible description and abnormal detection of the geographic location of Web access, can identify regionally-oriented network attacks, and improves the accuracy and effectiveness of abnormal detection.
Smart Images

Figure CN120017310A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of access detection, and in particular to a method for detecting abnormal geographic location of Web access based on density estimation. Background Art
[0002] In the anomaly detection of Web access logs, the detection of anomalies in access geographic location information (latitude and longitude) mainly refers to the detection of access requests from different geographic locations to discover abnormal access behaviors. This usually involves analyzing IP addresses, geographic location information, etc. to determine whether the access request comes from the expected geographic location. For example, if a website suddenly receives a large number of access requests from overseas, this may indicate unauthorized access or hacker attacks; or if an account is usually logged in and accessed in a certain place in China, if the account suddenly appears to log in and access overseas, there is a high risk of account theft, so further detection and analysis are required.
[0003] The existing technical solution extracts the features related to IP, builds a model for the IP access feature vector using a linear regression model, and detects abnormal access behavior based on the model. However, the features in this method do not include the geographical location of the IP address, and the geographical location is a binary, which is not suitable for integration into the linear regression model. Summary of the invention
[0004] In view of the above technical problems, the present invention provides a method for detecting abnormal geographic location of Web access based on density estimation.
[0005] The present invention is implemented by adopting the following technical scheme: a method for detecting abnormal geographic location of Web access based on density estimation, based on DBSCAN spatial clustering algorithm, specifically comprising the following steps: Step S1: Obtain the corresponding geographic location according to the IP address and pre-process the geographic location data set; Step S2: Determine and input the DBSCAN algorithm parameters neighborhood radius ε and minimum point density number MinPts; Step S3: traverse the IP geographical location data set, train the geographical location distribution model of Web application access, and cluster the geographical locations of IP addresses; Step S4: Calculate the center point and radius of each cluster after clustering to detect whether the geographic location of Web access is abnormal.
[0006] Specifically, the conversion and preprocessing of the geographic location in step S1 specifically includes: querying the geographic location according to the IP address, converting the geographic location into a longitude and latitude coordinate format, and counting the same longitude and latitude data.
[0007] Specifically, the method for determining the parameters of the DBSCAN algorithm in step S2 includes: using a trial and error method, a data distribution method, an elbow rule, and a silhouette coefficient method to determine the optimal values of the neighborhood radius ε and the minimum point density number MinPts.
[0008] Specifically, step S3 includes the following sub-steps: Step S31: Mark all data points in the IP geographic location data set as unvisited, randomly select an unvisited data point P, and calculate the point set of the ε neighborhood of P, which is recorded as N(P); Step S32: If the sum of the frequencies of each point in N(P) is not less than MinPts, the point is judged to be a core point; if it is a core point, a new cluster C is created and P is added to the cluster; if it is not a core point, the point is temporarily retained to check whether it can be connected to an existing cluster through density reachability; another unvisited point is selected from the data set to continue processing; Step S33: Select an unvisited point Q from N(P) and mark it as visited, check whether it is a core point; if so, add it to cluster C and update N(Q); recursively check each unvisited point in N(Q) until there are no more unvisited points that can be added to cluster C; Step S34: Repeat step S32 until all data points have been visited.
[0009] Specifically, the step S31 further includes: for the points with geographic coordinates on the spherical surface, using the haversine formula to calculate the distance between two points on the spherical surface.
[0010] Specifically, the step S34 further includes: if a point is still in an unvisited state at the end, or has never been included in the ε neighborhood of any core point, it is regarded as a noise point.
[0011] Specifically, the step S4 of calculating the center point and radius of each cluster after clustering includes: according to all points in each cluster after clustering, combining the frequency of each longitude and latitude, considering it as the weight of the point, and approximately calculating the center point and radius of each cluster.
[0012] Specifically, the step S4 detects whether the geographical location of the Web access is abnormal and specifically includes: first using the latitude and longitude point closest to the center point of all historical clusters as the target cluster, and then using the latitude and longitude point to calculate the distance with all points in the cluster to determine whether there is a point that satisfies the domain radius ε; if so, the latitude and longitude point is determined to be normal, otherwise it is determined to be an abnormal point.
[0013] Specifically, the detecting whether the Web access geographical location is abnormal also includes: calculating the distance from the latitude and longitude point to the cluster center, if the distance exceeds the sum of the radius of this cluster and the neighborhood radius ε, and the distances to the remaining clusters all exceed the corresponding radius of the cluster, then the data point is determined to be abnormal.
[0014] Specifically, the detecting whether the geographical location of the Web access is abnormal further includes: calculating the distances between the latitude and longitude point and all points of the cluster, and if all of them exceed the neighborhood radius ε, determining that the point is an abnormal point.
[0015] The beneficial effects of the present invention are as follows: the present invention converts the visitor's IP address into a corresponding geographic location, which is represented by a latitude and longitude tuple; uses the geographic location history access records of Web applications to train a user access geographic location model, and uses the model to perform anomaly detection; through density estimation, the geographic location distribution of Web application visitors can be more flexibly described; the geographic location can reflect the regional attributes of Web application visitors, and can identify regionally inclined network attacks. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying creative work.
[0017] Figure 1 This is a flow chart of an abnormality detection method in an embodiment of the present invention; Figure 2 A schematic diagram of data point types in an embodiment of the present invention; Figure 3 It is a schematic diagram of density connection relationship in another embodiment of the present invention. DETAILED DESCRIPTION
[0018] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0019] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.
[0020] The following is combined with Figures 1 to 3, some embodiments of the present invention are described in detail. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.
[0021] The present invention proposes a method for detecting abnormal geographic location of Web access based on density estimation, based on the DBSCAN spatial clustering algorithm, such as Figure 1 As shown, the specific steps include: Step S1: Obtain the corresponding geographic location according to the IP address and pre-process the geographic location data set; Step S2: Determine and input the DBSCAN algorithm parameters neighborhood radius ε and minimum point density number MinPts; Step S3: traverse the IP geographical location data set, train the geographical location distribution model of Web application access, and cluster the geographical locations of IP addresses; Step S4: Calculate the center point and radius of each cluster after clustering to detect whether the geographic location of Web access is abnormal.
[0022] In this embodiment, step S1 of converting and preprocessing the geographic location specifically includes: querying the geographic location according to the IP address, converting the geographic location into a longitude and latitude coordinate format, and counting the same longitude and latitude data.
[0023] In this embodiment, the method for determining the DBSCAN algorithm parameters in step S2 includes: using trial and error method, data distribution method, elbow rule and silhouette coefficient method to determine the optimal values of the neighborhood radius ε and the minimum point density number MinPts.
[0024] In this embodiment, step S3 includes the following sub-steps: Step S31: Mark all data points in the IP geographic location data set as unvisited, randomly select an unvisited data point P, and calculate the point set of the ε neighborhood of P, which is recorded as N(P); Step S32: If the sum of the frequencies of each point in N(P) is not less than MinPts, the point is judged to be a core point; if it is a core point, a new cluster C is created and P is added to the cluster; if it is not a core point, the point is temporarily retained to check whether it can be connected to an existing cluster through density reachability; another unvisited point is selected from the data set to continue processing; Step S33: Select an unvisited point Q from N(P) and mark it as visited, check whether it is a core point; if so, add it to cluster C and update N(Q); recursively check each unvisited point in N(Q) until there are no more unvisited points that can be added to cluster C; Step S34: Repeat step S32 until all data points have been visited.
[0025] In this embodiment, step S31 further includes: for points with geographic coordinates on the spherical surface, using the haversine formula to calculate the distance between two points on the spherical surface.
[0026] In this embodiment, step S34 further includes: if a point is still in an unvisited state at the end, or has never been included in the ε neighborhood of any core point, it is regarded as a noise point.
[0027] In this embodiment, step S4 of calculating the center point and radius of each cluster after clustering includes: according to all points in each cluster after clustering, combining the frequency of each longitude and latitude, considering it as the weight of the point, and approximately calculating the center point and radius of each cluster.
[0028] In this embodiment, step S4 detects whether the geographical location of the Web access is abnormal and specifically includes: first using the latitude and longitude point closest to the center point of all historical clusters as the target cluster, and then using the latitude and longitude point to calculate the distance with all points in the cluster to determine whether there is a point that satisfies the domain radius ε; if so, the latitude and longitude point is determined to be normal, otherwise it is determined to be an abnormal point.
[0029] In this embodiment, detecting whether the geographic location of Web access is abnormal also includes: calculating the distance from the latitude and longitude point to the cluster center. If the distance exceeds the sum of the radius of this cluster and the neighborhood radius ε, and the distances to the remaining clusters all exceed the corresponding radius of the cluster, then the data point is determined to be abnormal.
[0030] In this embodiment, detecting whether the geographical location of the Web access is abnormal further includes: calculating the distance between the latitude and longitude point and all points in the cluster, and if all of them exceed the neighborhood radius ε, then it is determined to be an abnormal point. In one embodiment, the present invention uses the DBSCAN algorithm concept to detect abnormal longitude and latitude access of IP addresses. DBSCAN (Density-Based Spatial Clustering of Applications with Noise) is a density-based spatial clustering algorithm. The algorithm can identify any number of clusters and can find clusters of any shape, including noise and outliers. In the detection of geographic location anomalies in Web access, the algorithm can be used to cluster the geographic location of IP addresses to find abnormal geographic location access. In the clustering process, points with a domain radius less than a threshold are clustered in a cluster, and those clusters with a number of nodes exceeding the threshold are the clustering results.
[0031] The technical solution steps of the present invention are as follows: S1: IP address data conversion and preprocessing: Check the geographic location based on the IP address, convert it into longitude and latitude coordinate format, and count the same longitude and latitude data.
[0032] S2: Determine algorithm parameters: The DBSCAN algorithm needs to determine two parameters, namely the neighborhood radius (ε) and the minimum point density (MinPts). The optimal values of these two parameters can be determined using trial and error, data distribution-based methods, elbow rule, silhouette coefficient, and other methods. Finally, MinPts was determined to be 300 and ε to be 50 km.
[0033] S3: Training the geographic location distribution model for web application access: (1) Initialization: Mark all data points in the IP's geographic location data set as unvisited; randomly select an unvisited data point P and calculate the point set of P's ε neighborhood, denoted as N(P). The distance between geographic coordinates on the sphere does not meet the requirements of Euclidean distance, so the distance between two points on the sphere is calculated using the haversine formula; (2) Check whether it is a core point: If the sum of the frequencies of each point in N(P) is not less than MinPts, the point is considered a core point (that is, there are at least MinPts points in the ε neighborhood); create a new cluster: If it is a core point, create a new cluster C and add P to the cluster; if it is not a core point, temporarily retain the point and check whether it can be connected to an existing cluster through density accessibility. Select another "unvisited" point from the data set to continue processing; (3) Density reachability query: Select an “unvisited” point Q from N(P) and mark it as “visited”. Check whether it is a core point. If so, add it to cluster C and update N(Q). Recursively perform the above steps for each “unvisited” point in N(Q) until there are no more “unvisited” points that can be added to cluster C.
[0034] (4) Repeat step (2) until all data points have been visited. If a point is still “unvisited” at the end of the algorithm, or it has never been included in the ε-neighborhood of any core point, it is considered a noise point.
[0035] (5) Based on all the points in each cluster after clustering, combined with the frequency of each longitude and latitude, which is regarded as the weight of the point, the center point and radius of each cluster are approximately calculated.
[0036] S4: Web access geographic location anomaly detection, you can use one of the following methods: (1) First, the latitude and longitude point closest to the center point of all historical clusters is used as the target cluster, and then the latitude and longitude point is used to calculate the distance with all points in the cluster to see if there is a point that is less than the neighborhood radius. If so, then this latitude and longitude point is considered normal, otherwise it is judged as an abnormal point.
[0037] (2) Calculate the distance from the latitude and longitude point to the cluster center. If the distance exceeds the radius of this cluster + the neighborhood radius ε = 50 kilometers, and exceeds the corresponding radius of the cluster with the remaining clusters, then the data point can be regarded as an outlier. (Outliers are usually considered to be data points that are far away from all cluster centers) (3) Calculate the distance between the longitude and latitude point and all the points in the cluster. If they all exceed the neighborhood radius, they are considered outliers.
[0038] In another embodiment, when performing outlier detection analysis on the results after DBSCAN clustering, it is necessary to understand the following three types of data points: core points, boundary points, and noise points, such as Figure 2 As shown in the figure, the points whose number of sample points within the neighborhood radius R is greater than or equal to MinPts are called core points, the points that are not core points but within the neighborhood of a core point are called boundary points, and the points that are neither core points nor boundary points are noise points.
[0039] On this basis, it is also necessary to combine the four relationships of these points: density direct, density reachable, density connected, and non-density connected, to finally determine the outliers.
[0040] If P is a core point and Q is in the R neighborhood of P, then P is said to be density-reachable from Q. Any core point is density-reachable from itself, and density-reachable is not symmetric. If P is density-reachable from Q, then Q is not necessarily density-reachable from P.
[0041] If there are core points P2, P3, ..., Pn, and P1 is density-reachable from P2, P2 is density-reachable from P3, ..., P(n-1) is density-reachable from Pn, and Pn is density-reachable from Q, then P1 is density-reachable from Q, and density-reachable is not symmetric.
[0042] If there is a core point S, so that S is density-reachable to both P and Q, then P and Q are density-connected. Density connection is symmetric. If P and Q are density-connected, then Q and P must also be density-connected. Two density-connected points belong to the same cluster.
[0043] If two points are not density-connected, then the two points are non-density-connected, and the two non-density-connected points belong to different clusters, or there are noise points between them. Figure 3 shown.
[0044] The algorithm logic of this solution is as follows: Input: Source IP address of the client side {is}; Calculation steps: 1. Establish clusters of historical IP latitude and longitude: 2. Call the address conversion API to convert the IP address into longitude and latitude; 3. Use the haversine formula to convert the longitude and latitude data into plane coordinates; 4. Window task, count the number of visits to each IP in the window in Flink, remove duplicate IPs and record the number of visits when storing. The storage table name is assumed to be t1; 5. Scheduled tasks, use DBSACAN density clustering algorithm to calculate historical data. Set MinPts=1, neighborhood radius ε=50 kilometers. Use the deduplicated data for DBSACAN density clustering; 6. Filter out noise points, sum the number of times the cluster midpoint corresponds to the set midpoint in table t1. If it is less than 10, it is considered an outlier and the cluster is discarded; otherwise, it is retained; 7. Calculate the cluster center point and the maximum radius of the cluster, and calculate the longitude and latitude of the center by taking the weighted average of the number of times the cluster center point corresponds to the set midpoint in table t1; calculate the maximum distance from all points in the cluster to the center point and set it as the radius.
[0045] Output storage: cluster center point, radius, and longitude and latitude points without duplication within the cluster In this embodiment, outlier detection includes multiple detection methods: Method 1: Calculate the distance from the latitude and longitude point to the cluster center. If the distance exceeds the radius of this cluster + the neighborhood radius ε = 50 kilometers, and exceeds the corresponding radius of the cluster with the rest of the clusters, then the data point can be regarded as an outlier. (Outliers are usually considered to be data points that are far away from all cluster centers).
[0046] Method 2: Calculate the distance between the longitude and latitude point and all the points in the cluster. If they all exceed the neighborhood radius, they are considered abnormal points. (It can be improved by first using the longitude and latitude point closest to the center point of all historical clusters as the target cluster, and then using the longitude and latitude point to calculate the distance with all the points in the cluster to see if there is a point that is less than the neighborhood radius. If so, then this longitude and latitude point is considered normal, otherwise it is considered an abnormal point).
[0047] Output: event name, longitude, latitude, and source IP address of the abnormal geographic location access {is}.
[0048] For the aforementioned embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the order of the actions described, because according to the present application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily required by the present application.
[0049] The above embodiments describe the basic principles and main features of the present invention and the advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the changes and modifications made by those skilled in the art shall be within the scope of protection of the appended claims of the present invention without departing from the spirit and scope of the present invention.
Claims
1. A method for detecting abnormal geographic location of Web access based on density estimation, characterized in that: Based on the DBSCAN spatial clustering algorithm, the following steps are specifically included: Step S1: Obtain the corresponding geographic location according to the IP address and pre-process the geographic location data set; Step S2: Determine and input the DBSCAN algorithm parameters neighborhood radius ε and minimum point density number MinPts; Step S3: traverse the IP geographical location data set, train the geographical location distribution model of Web application access, and cluster the geographical locations of IP addresses; Step S4: Calculate the center point and radius of each cluster after clustering to detect whether the geographic location of Web access is abnormal.
2. A method for detecting abnormal Web access geographic location based on density estimation as claimed in claim 1, characterized in that: The step S1 of geographic location conversion and preprocessing specifically includes: querying the geographic location according to the IP address, converting the geographic location into longitude and latitude coordinate format, and counting the same longitude and latitude data.
3. A method for detecting abnormal Web access geographic location based on density estimation as claimed in claim 1, characterized in that: The method for determining the DBSCAN algorithm parameters in step S2 includes: using trial and error method, data distribution method, elbow rule and silhouette coefficient method to determine the optimal values of the neighborhood radius ε and the minimum point density number MinPts.
4. A method for detecting abnormal Web access geographic location based on density estimation as claimed in claim 1, characterized in that: The step S3 comprises the following sub-steps: Step S31: Mark all data points in the IP geographic location data set as unvisited, randomly select an unvisited data point P, and calculate the point set of the ε neighborhood of P, which is recorded as N(P); Step S32: If the sum of the frequencies of the points in N(P) is not less than MinPts, the point is determined to be a core point; if it is a core point, a new cluster C is created and P is added to the cluster; If it is not a core point, keep the point temporarily and check whether it can be connected to an existing cluster through density reachability; select another unvisited point from the data set to continue processing; Step S33: Select an unvisited point Q from N(P) and mark it as visited, check whether it is a core point; if so, add it to cluster C and update N(Q); recursively check each unvisited point in N(Q) until there are no more unvisited points that can be added to cluster C; Step S34: Repeat step S32 until all data points have been visited.
5. A method for detecting abnormal Web access geographic location based on density estimation as claimed in claim 4, characterized in that: The step S31 further includes: for the points with geographic coordinates on the spherical surface, using the haversine formula to calculate the distance between two points on the spherical surface.
6. A method for detecting abnormal Web access geographic location based on density estimation as claimed in claim 4, characterized in that: The step S34 also includes: if a point is still in an unvisited state at the end, or has never been included in the ε neighborhood of any core point, it is regarded as a noise point.
7. A method for detecting abnormal Web access geographic location based on density estimation as claimed in claim 1, characterized in that: The step S4 of calculating the center point and radius of each cluster after clustering includes: according to all the points in each cluster after clustering, combining the frequency of each longitude and latitude, considering it as the weight of the point, and approximately calculating the center point and radius of each cluster.
8. A method for detecting abnormal Web access geographic location based on density estimation as claimed in claim 7, characterized in that: The step S4 detects whether the geographical location of the Web access is abnormal and specifically includes: firstly taking the latitude and longitude point closest to the center point of all historical clusters as the target cluster, then calculating the distance between the latitude and longitude point and all points in the cluster to determine whether there is a point that satisfies the domain radius ε; if so, the latitude and longitude point is determined to be normal, otherwise it is determined to be an abnormal point.
9. A method for detecting abnormal Web access geographic location based on density estimation as claimed in claim 8, characterized in that: The detecting whether the Web access geographical location is abnormal also includes: calculating the distance from the latitude and longitude point to the cluster center, if the distance exceeds the sum of the radius of this cluster and the neighborhood radius ε, and the distances to the remaining clusters all exceed the corresponding radius of the cluster, then the data point is determined to be abnormal.
10. The method for detecting abnormal Web access geographic location based on density estimation according to claim 8, characterized in that: The detecting whether the geographical location of the Web access is abnormal further includes: calculating the distances between the latitude and longitude point and all points of the cluster, and if all of them exceed the neighborhood radius ε, determining that the point is an abnormal point.