An IP address profiling method based on spatio-temporal statistics
By using the IP address portrait method of spatiotemporal statistics in large-scale network flow data, combined with Count-Min Sketch and statistical methods, the spatiotemporal patterns and access frequency rules of IP addresses are extracted, and the problem that the existing technology cannot characterize these features in fine-grained manner is solved, and efficient IP address portrait generation and behavior abnormality detection are achieved.
Patent Information
- Application Number
- CN202111308488.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-05
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-11-05
AI Technical Summary
The prior art is difficult to fine-grained the spatio-temporal mode and access frequency rules of IP addresses in large-scale network streaming data, resulting in the inability to effectively detect IP addresses with abnormal behavior.
The IP address portrait method based on spatiotemporal statistics is adopted. By setting up global and local Count-Min Sketch for each IP address, recording its access mode, and extracting feature information in combination with statistical methods and data dimensionality reduction technology, classifying using hierarchical clustering to generate an IP address portrait.
It realizes fine-grained representation of the spatio-temporal mode and access frequency rules of IP addresses, greatly saving space overhead, and can observe the connection mode, activity mode and semantic mode of each IP address in real time.
Smart Images

Figure CN114037009B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of associated data mining and portrait technology, and particularly relates to an IP address portrait method based on spatio-temporal statistics. Background Art
[0002] IP portraits are very important in real data center networks. Facing a huge amount of data streams, it is not only necessary to correctly count the number of accesses and accessed times of each IP address, but also necessary to generate an IP address portrait according to its spatio-temporal access pattern, so as to detect IP addresses with abnormal behaviors and avoid causing adverse effects on other service functions within the network.
[0003] The related technologies mainly focus on the counting of huge data streams and the mining of associated data. Each IP address accessing other IP addresses and ports has spatio-temporal patterns and frequency rules. It is hoped that by extracting the feature information of the access patterns of IP addresses, IP address portraits can be formed through clustering and automatically labeled with pseudo-labels.
[0004] However, in real scenarios, the scale of network data streams is very large. How to correctly record the information of each data stream with as little space overhead as possible, complete the search of stream records in constant time, and how to extract the spatio-temporal patterns and access frequency rules of each IP address from the stored information are problems that need to be considered and solved. Summary of the Invention
[0005] Object of the Invention: Aiming at the above problems, the present invention proposes an IP address portrait method based on spatio-temporal statistics, which can fundamentally solve the problem that the existing portrait algorithms cannot finely depict the spatio-temporal patterns and access frequency rules of IP addresses under large-scale network flow data.
[0006] Technical Solution: In order to achieve the above object of the invention, the technical solution of the present invention is as follows:
[0007] An IP address portrait method based on spatio-temporal statistics, comprising the following steps:
[0008] (1) Set a global Count-Min Sketch (hereinafter referred to as Sketch) for each IP address. After receiving a network flow data packet, parse the source IP address and destination IP address information, and update this information into the global Sketch corresponding to each IP address;
[0009] (2) Additionally, divide a day into several time periods. Besides maintaining the global Sketch, each time period also needs to maintain a corresponding local Sketch for each IP address. When a new time period starts, save and clear the local Sketch of the previous time period. In this way, over the course of a day, the access and being-accessed patterns of each IP address within each time period are saved;
[0010] (3) Based on the obtained global Sketch and single-time-period Sketch containing spatio-temporal statistical information, obtain the feature information of each IP address through statistical methods and data dimensionality reduction methods;
[0011] (4) According to the feature information of the IP addresses, use hierarchical clustering to classify the IP addresses so that each IP address gets a corresponding class label, completing the profiling of the IP address population;
[0012] (5) According to the frequently and frequently-being-accessed objects of each IP address recorded in the global Sketch, parse to obtain the activity pattern, connection pattern, and semantic pattern of each IP address, completing the profiling of individual IP addresses.
[0013] Furthermore, in step (1), the Count-Min Sketch data structure is used to determine the required storage space according to the range of hash values, which can greatly reduce the storage overhead.
[0014] Furthermore, Count-Min Sketch is a two-dimensional array with w columns and d rows. The parameters w and d are determined at the time of creation and are related to the query error rate. Each row is associated with a hash function, and there are d mutually independent hash functions in total. When a new event arrives, use the d hash functions to obtain d corresponding column indices, and increment the count at the corresponding position in each row. During the query phase, when it is necessary to count a certain event i, d corresponding column indices can be obtained similarly, and then take the minimum value at the corresponding positions.
[0015] Furthermore, in step (1), in order to facilitate the recording and query of high-frequency items, a corresponding min-heap is designed for each Sketch. The min-heap is updated together with the Sketch each time the Sketch is updated, and finally the Top-K items in the stream data can be obtained through the min-heap.
[0016] Furthermore, in step (1), in order to better obtain information, the cells in the Sketch are designed as vectors with a length of 4, and each bit stores the frequency, traffic, Session number, and port information of the current record respectively, so as to facilitate subsequent reverse lookup of the Sketch.
[0017] Further, in step (1), in order to count global information, five Sketches are constructed: CS_SIP, CS_DIP, CS_DIP_Port, CS_IP_Pair, and CS_DIP_Pair_Port, which record the total number of accesses initiated by each source IP in all data streams, the total number of accesses received by each destination IP, the total number of accesses received by each destination IP port, the total number of accesses between hosts, and the total number of accesses by each source IP to server applications, respectively.
[0018] Further, in step (2), the data stream is segmented into a Session every 15 minutes. At the same time, two Sketches are created for each IP address to record the IP address and its frequency when it initiates and receives accesses. When each new Session arrives, based on the global information statistics, the access and being accessed situations of each IP in the current time period are further recorded (each IP address is further divided into client and server for recording), and they are saved in the Sketches corresponding to each IP address under this Session.
[0019] Further, in step (3), a high-dimensional spatio-temporal matrix is constructed using the Sketches recorded in each Session in step (2). One day is divided into 96 Sessions at 15-minute intervals. Each IP address maintains a Sketch in each Session. For each IP address, the corresponding 96 Sketches in its corresponding Sessions are extracted. Each Sketch is a matrix of w×d×4. Combining these 96 Sketches forms a tensor of w×d×4×96, which is the original feature vector of this IP address and contains spatio-temporal pattern and access frequency information.
[0020] Further, in step (3), the PCA dimensionality reduction method is used to reduce the dimensionality of the high-dimensional spatio-temporal matrix to obtain the feature vector of each IP address. The dimensionality reduction method is as follows: The original data is formed into a matrix X by columns, and the covariance matrix The eigenvalues and eigenvectors of C are obtained. The eigenvectors are arranged in a matrix row by row in descending order according to the corresponding eigenvalue magnitudes. The first k rows are taken to form a matrix P, and let F = PX be the data after dimensionality reduction.
[0021] Further, in step (4), the K-Means clustering method is used to complete clustering based on the feature vectors of each IP address to achieve the group portrait of IP addresses.
[0022] Further, in step (5), according to the global record information, the portrait analysis of a single IP address is performed from multiple perspectives (IP address service information, spatio-temporal access habits, access frequency time series), and the connection mode, activity mode, and semantic mode of the single IP address are portraited.
[0023] Beneficial effects: The present invention first proposes to use an improved Count-Min Sketch data structure for network flow data measurement, and proposes a brand-new IP address portrait algorithm to extract the spatio-temporal patterns and access frequency time series features of IP address activities, and assigns pseudo-labels to each IP address through a hierarchical clustering method. Its advantages are that it greatly saves space overhead, and the connection mode, activity mode, and semantic mode of each IP address can be observed in real time and intuitively through IP address portraits. Brief Description of the Drawings
[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0025] Figure 1 is a flowchart of an IP address portrait method based on spatio-temporal statistics;
[0026] Figure 2 is a schematic framework diagram of an IP address portrait method based on spatio-temporal statistics;
[0027] Figure 3 is a schematic diagram of the Count-Min Sketch data structure according to the embodiment of the present invention. Detailed Embodiments
[0028] The technical solutions of the present invention will be further described below with reference to the drawings. It should be understood that the embodiments provided below are only for disclosing the present invention in detail and completely, and fully conveying the technical concept of the present invention to those skilled in the art. The present invention can also be implemented in many different forms and is not limited to the embodiments described herein. The terms in the exemplary embodiments shown in the drawings are not limitations to the present invention.
[0029] As Figure 1-2 shown, the present invention proposes a new network data flow statistics and IP address portrait algorithm based on Count-Min Sketch (hereinafter referred to as Sketch). The entire model framework consists of three parts: a data flow statistics module, a feature extraction and clustering module, and a portrait module.
[0030] The data flow statistics module uses Count-Min Sketch to record the information of each data flow in the network. When a data flow arrives, the key information in it is first extracted: source IP address, destination IP address, source port, destination port, protocol type, traffic, etc. Then, according to the above information, the record of this data flow is updated to the Sketch in the global and single time period. In this way, each Sketch saves the access pattern of the IP address, and the attribute of time is also taken into account by dividing the time period.
[0031] The feature extraction and clustering module uses the Sketch obtained from the previous module to construct the feature vector of the IP address. Since the day is divided into fine-grained intervals (15 minutes), a corresponding Sketch is maintained within each time slice. Therefore, by stacking these Sketches, a high-dimensional feature vector of each IP address can be obtained. However, the dimension of the original feature tensor is too high, so PCA dimensionality reduction is also required. Using the dimensionality-reduced feature vector for hierarchical clustering can complete the division of the IP address, thus completing the group portrait of the IP address.
[0032] The portrait module uses spatio-temporal statistical information and clustering results to comprehensively portrait the IP address. For each IP, its activity pattern and connection pattern can be depicted through the global Sketch and the corresponding Sketch within a single time period. The semantic pattern can be analyzed using the protocol number information, and combined with the group portrait of the IP address, the comprehensive portrait of a single IP address can be obtained.
[0033] Figure 3Schematic diagram of the Count-Min Sketch data structure according to an embodiment of the present invention. Count-Min Sketch is a two-dimensional array with w columns and d rows. In the present invention, w is fixed at 100 and d is fixed at 10. Each row is associated with a hash function, and there are d mutually independent hash functions in total. The present invention adopts BKDRhash. When a new event arrives, d corresponding column indices are obtained by using the d hash functions, and the count at the corresponding position in each row is incremented by one. In the query phase, to count a certain event i, d corresponding column indices can be obtained similarly, and then the minimum value among the corresponding positions is taken. In the present invention, a global Sketch and a time-segmented Sketch are used. The global Sketch includes CS_SIP (Key = source IP, Value = total number), CS_DIP (Key = destination IP, Value = total number), CD_DIP_Port (Key = source IP + port, Value = total number), CS_IP_Pair (Key = source IP, Value = destination IP), CS_IP_Pair_Port (Key = source IP, Value = destination IP + port), which respectively record the total number of accesses initiated by each source IP, the total number of accesses received by each destination IP, the total number of accesses received by each destination IP port, the total number of accesses between hosts, and the total number of accesses by each source IP to the server application. At the same time, for each IP address, four Sketches need to be maintained within each time period: CS_DIP (Key = destination IP, Value = total number), CS_DIP_Port (Key = destination IP + port, Value = total number), CS_SIP (Key = source IP, Value = total number), CS_SIP_Port (Key = source IP + port, Value = total number), which respectively record the total number of accesses received by each destination IP, the total number of accesses received by each destination IP port, the total number of accesses by each source IP, and the total number of accesses from each source IP to the port.
[0034] Algorithm 1 is a network data stream statistical algorithm according to an embodiment of the present invention. For each data stream, the algorithm first analyzes the information in its header, generates key-value pairs according to information such as its source IP address, and inserts them into the global Sketch and the time-segmented Sketch respectively, and then updates the minimum heap corresponding to each Sketch. After all data streams are received, the traffic of each IP address and the frequently accessed objects can be queried through the Sketch.
[0035] Algorithm 1: Network Flow Data Statistical Algorithm
[0036] Input: N-tuple (source IP, source port, destination IP, destination port, protocol type, traffic, etc.)
[0037] Initialize TotalInfo (including 5 global Sketches) and SessionInfo (including 4 Sketches corresponding to each IP address)
[0038]
[0039]
[0040] Save the statistical results
[0041] Output: Global statistical result TotalInfo and per-period statistical results SessionInfo0 to n
[0042] Algorithm 2 is the IP address feature extraction and clustering algorithm according to the embodiment of the present invention. For each IP address, the present invention selects the Sketches corresponding to 96 Sessions of it. For each element in these 96 Sketches, calculate the mean, variance, minimum value, and maximum value at its corresponding position respectively, so as to generate four new two-dimensional matrices. Stacking these new matrices can obtain the original high-dimensional feature tensor with a size of (4, 10, 1000). Then use the PCA dimensionality reduction method to reduce the original feature tensor to (1, 10), and then use the K-Means clustering method to complete the clustering of IP addresses. Finally, analyze the spatio-temporal patterns and semantic information of each class of IP addresses according to the clustering results to obtain the group portrait of IP addresses.
[0043] Algorithm 2: IP Address Feature Extraction and Clustering Algorithm
[0044] Input: Set of IP addresses to be classified
[0045] Traverse the input set of IP addresses:
[0046] Initialize the feature vector feature(4, 10, 1000)
[0047] Traverse SessionInfo:
[0048] Extract the Sketch corresponding to the IP name under SessionInfo
[0049] Calculate the average value, variance, minimum value, and maximum value at the corresponding positions of this group of Sketches
[0050] Store the results in feature
[0051] Use PCA to reduce the dimension of feature to get new_feature
[0052] Use the K-Means algorithm to perform clustering according to new_feature
[0053] Traverse each type of IP set:
[0054] Randomly select 10 IP addresses
[0055] Obtain its corresponding feature[0,:,:]
[0056] Calculate the average of these 10 matrices and visualize them using a heatmap
[0057] Calculate the frequently accessed objects of 10 IP addresses
[0058] Visualize using a word cloud
[0059] Output: Each type of IP address set, and its corresponding spatio-temporal matrix and word cloud
[0060] In an embodiment of the present application, for a single IP address, its ego network and the trend graph of access frequency over time can be depicted through a global Sketch. Combining the protocol number and service content corresponding to each data stream, its semantic information is presented in the form of a word cloud.
[0061] The method of the present invention uses two data structures, Count-Min Sketch and min-heap, to save the access and being accessed situations of each IP in the network. When receiving a network data stream, information such as the source IP address, destination IP address, source port number, destination port number, and protocol number of the stream is obtained, and the above information is updated to the Sketch of the corresponding IP address. After parsing all network streams, a set of spatio-temporal matrices is generated using the Sketch of each IP address, so as to obtain the feature information of each IP address. On this basis, hierarchical clustering is performed on the IP addresses, and pseudo-labels are assigned to each IP address according to the clustering results to form a portrait of group IP addresses. Then, based on the frequently accessed and being accessed records of IP addresses under the same label, the analysis of this type of IP address is completed, and the portrait of individual IP addresses is completed. The present invention uses probability data structures based on Count-Min Sketch and min-heap, which greatly reduces the storage space while ensuring the accuracy of data records in the face of real-time massive stream data, and cooperates with the IP address portrait algorithm based on spatio-temporal statistics to complete the portrait of group IP addresses and individual IP addresses using the global information of network streams and IP access pattern information respectively.
[0062] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all belong to the protection scope of the present invention.
Claims
1. An IP address profiling method based on spatio-temporal statistics, characterized in that, the method comprises the following steps: (1) Set a global Count-Min Sketch for each IP address. After receiving a network flow datagram, parse the source IP address and destination IP address information, and update this information into the global Count-Min Sketch corresponding to each IP address; (2) Divide a day into several time periods. In addition to maintaining the global Count-Min Sketch, each time period also needs to maintain a corresponding local Count-Min Sketch for each IP address. When a new time period starts, save and clear the local Count-Min Sketch of the previous time period. In this way, the access and being accessed patterns of each IP address in each time period are saved; (3) According to the obtained global Count-Min Sketch and single-time period Count-Min Sketch containing spatio-temporal statistical information, obtain the characteristic information of each IP address through statistical methods and data dimensionality reduction methods; (4) According to the characteristic information of the IP address, use hierarchical clustering to classify the IP addresses, so that each IP address obtains a corresponding class label, and complete the profiling of the IP address group; (5) According to the frequently and frequently accessed objects of each IP address recorded in the global Count-Min Sketch, combine the local Count-Min Sketch to parse the activity pattern, connection pattern and semantic pattern of each IP address, and complete the profiling of the individual IP address; wherein, use the Count-Min Sketch recorded in each Session to construct a high-dimensional spatio-temporal matrix. Divide a day into 96 Sessions at 15-minute intervals. Each IP address maintains a Count-Min Sketch in each Session. For each IP address, extract the corresponding Count-Min Sketch in its corresponding 96 Sessions. Each Count-Min Sketch is a w×d×4 matrix. Combine these 96 Count-Min Sketches to form a w×d×4×96 tensor, which is the original feature vector of this IP address, containing spatio-temporal patterns and access frequency information. Then use the PCA dimensionality reduction method to reduce the dimensionality of the high-dimensional spatio-temporal matrix to obtain the feature vector of each IP address. The dimensionality reduction method is: form a matrix X by the original data column by column, calculate the covariance matrix C, calculate the eigenvalues and eigenvectors of C, arrange the eigenvectors in a matrix row by row in descending order according to the corresponding eigenvalue sizes, take the first k rows to form a matrix P, and let F = PX be the data after dimensionality reduction.
2. The IP address profiling method based on spatio-temporal statistics according to claim 1, characterized in that, different from the storage form of key-value pairs used in traditional databases, The Count-Min Sketch data structure determines the required storage space according to the range of hash values, which can greatly reduce the storage overhead.
3. A method for IP address profiling based on spatio-temporal statistics according to claim 1, characterized in that Count-Min Sketch is a two-dimensional array with w columns and d rows. The parameters w and d are determined at the time of creation and are related to the query error rate. Each row is associated with a hash function, and there are d independent hash functions in total. When a new event arrives, d corresponding column indexes are obtained using the d hash functions, and the count at the corresponding position in each row is incremented by one. In the query phase, to count the occurrences of a certain event i, d corresponding column indexes can be obtained, and then the minimum value at the corresponding positions is taken.
4. A method for IP address profiling based on spatio-temporal statistics according to claim 1, characterized in that To facilitate the recording and query of high-frequency items, a corresponding min-heap is designed for each Count-Min Sketch. The min-heap is updated whenever the Count-Min Sketch is updated, and finally the TopK items in the stream data can be obtained through the min-heap.
5. A method for IP address profiling based on spatio-temporal statistics according to claim 1, characterized in that Each cell in the basic Count-Min Sketch only stores the current recorded frequency. To better obtain information, the cell is designed as a vector of length 4, and each bit stores the current recorded frequency, traffic, Session number, and port information respectively, so as to perform reverse lookup on the Count-Min Sketch later.
6. A method for IP address profiling based on spatio-temporal statistics according to claim 1, characterized in that In step (1), to count the global information, five Count-Min Sketches are constructed: CS_SIP, CS_DIP, CS_DIP_Port, CS_IP_Pair, CS_DIP_Pair_Port, which record the total number of accesses initiated by each source IP, the total number of accesses received by each destination IP, the total number of accesses received by each destination IP port, the total number of accesses between hosts, and the total number of accesses by each source IP to the server application respectively.
7. A method for IP address profiling based on spatio-temporal statistics according to claim 1, characterized in that In step (2), the data stream is segmented into a Session every 15 minutes. At the same time, two Count-Min Sketches are created for each IP address to record the IP address and its frequency when it initiates and receives accesses respectively. When each new Session arrives, based on the global information statistics, the access and being accessed situations of each IP in the current time period are further recorded. Each IP address is further divided into client and server for recording, and it is saved in the Count-Min Sketch corresponding to each IP address under this Session.
8. A method for IP address profiling based on spatio-temporal statistics according to claim 1, characterized in that, in step (4), the K-Means clustering method is used to complete clustering based on the feature vectors of each IP address, and the group profiling of IP addresses is realized.
9. A method for IP address profiling based on spatio-temporal statistics according to claim 1, characterized in that, in step (5), according to the global record information, the profiling analysis of a single IP address is carried out from multiple perspectives, and the connection pattern, activity pattern, and semantic pattern of a single IP address are profiled.
Citation Information
Patent Citations
Big data environment-oriented summary information dynamic constructing and querying method and device
CN104657450A
Method and system for constructing user portrait
CN111210326A