Network service access node wire expanding method based on IP fingerprint multi-view clustering
Through the multi-view clustering method based on IP fingerprint and the adaptive IP scanning algorithm, the IP fingerprint of the software-defined network service access point is extracted and upgraded, and the problem of difficult to discover the domain network access node of the same software-defined network service provider in the prior art is solved, and efficient evaluation of the security and reliability of software-defined network service is achieved.
Patent Information
- Application Number
- CN202411914941.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-24
AI Technical Summary
It is difficult to accurately discover enough domain network access nodes of the same software-defined network service provider, affecting the evaluation of the security and reliability of software-defined network services.
The multi-view clustering method based on IP fingerprint is adopted to extract the IP fingerprint of the network service access point through breadth search and adaptive IP scanning algorithms, use machine learning model to upgrade the dimensions, and expand the line through multiple clustering algorithms to accurately discover the domain network access node of the same software-defined network service provider.
It improves the breadth and accuracy of discovering the same network service access node from massive Internet data, helps users evaluate the security and reliability of software-defined network services, and thus selects high-quality service providers.
Smart Images

Figure CN119945936A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of software-defined network services, and relates to a network service access node extension method based on IP fingerprint multi-view clustering. Background Art
[0002] With the advancement of digital transformation in various vertical industries in my country, the digitization, computing power and intelligence of information and communication networks will become an inevitable choice for the high-quality development of the industry. Driven by the demand for services and responsiveness of the underlying network from emerging businesses such as cloud services, AI, big data, and the Internet of Things, the concept of software-defined networking has gradually fermented in the ICT field and has been introduced into the enterprise network service supply market, prompting the emergence of software-defined network solutions.
[0003] However, choosing a suitable software-defined network service provider is not an easy task. Currently, most software-defined network service providers claim to be able to provide high-quality services, but in fact their services may have problems with poor performance, security, and privacy protection. In the software-defined network architecture, the access point of the data center plays a key hub role. The status of the domain network access node not only affects the operating efficiency of the network, but is also directly related to the security and stability of the network. Therefore, discovering enough domain network access nodes of the same software-defined network service provider is crucial to evaluating the security and reliability of the software-defined network services provided by the provider. Summary of the invention
[0004] In view of the problems existing in the prior art, the purpose of the present invention is to provide a network service access node expansion method based on IP fingerprint multi-view clustering, which can accurately discover a sufficient number of domain network access nodes of the same software-defined network service provider, thereby helping users evaluate the security and reliability of software-defined network service providers, and thus choose the high-quality software-defined network service that best suits their needs.
[0005] The present invention extracts the IP fingerprint of the access point of the network service based on the traffic characteristics, combines the breadth search and adaptive IP scanning algorithm, uses the machine learning model to upgrade the dimension of the feature to obtain sufficiently robust fingerprint features, and then through multiple clustering layers, it can effectively expand the line from the known network access nodes, and improve the breadth and accuracy of discovering the same network service access node from the massive data of the Internet.
[0006] The main difficulties that the present invention needs to solve are: 1. Efficient and wide-range detection and collection of network access point seed IP (i.e., target IP) data; 2. Extracting IP feature fingerprints based on seed IP data so that it can effectively distinguish software-defined network service network access point traffic from general node traffic; 3. Based on the detected IP data, efficiently expanding the domain network access nodes that may exist under the same software-defined network service.
[0007] In order to overcome the above difficulties, the present invention uses a number of innovative technologies to complete the predetermined requirements. The main related technologies are: ① Adaptive IP scanning technology based on breadth search and combined with domain network access node traffic characteristics: Research host inventory discovery technology based on ICMP, TCP and other means, adaptively select efficient breadth scanning strategy for network access point traffic to perform IP scanning, ensure scanning breadth and efficiency, and complete the detection and collection of basic seed data; ② Software-defined network access point IP fingerprint extraction technology based on adaptive weight adjustment strategy: Analyze the port service, certificate, banner, web page and other feature information of the IP address of the software-defined network service network access point, and clean and pre-process the data to reduce the adverse effects of noise on the features, and then dynamically adjust the weight of the features to dynamically provide more attention to different high-quality features in different software-defined network services; ③ Domain network access node extension technology based on multi-view clustering: Use three clustering methods, K-means, hierarchical clustering and DBSCAN, to cluster fingerprint feature vectors, divide similar IPs into the same cluster, and thus form IP clusters with the same software-defined network service provider as the unit.
[0008] The specific steps of the present invention include:
[0009] (1) Select the traffic characteristics of the domain network access node;
[0010] (2) Perform breadth-first adaptive IP scanning based on the traffic characteristics of the domain network access nodes;
[0011] (3) Obtain scanning results and form an IP library;
[0012] (4) Clean and preprocess the scan results to obtain standardized traffic characteristics and form a primary IP fingerprint library;
[0013] (5) Upgrading the dimension of IP fingerprints through machine learning models to obtain sufficiently robust feature vectors, and storing them in the primary IP fingerprint library;
[0014] (6) Select the features that need to be weighted for the software-defined network service and initialize their weights;
[0015] (7) Select three clustering algorithms according to the principle of progressive complexity to form a clustering layer. Generate a fingerprint vector for each IP according to the primary IP fingerprint library and input it into the clustering algorithm of the clustering layer in parallel to obtain the primary IP clustering result.
[0016] (8) Obtain the IP clustering result, reversely map it to the software-defined network weighted fingerprint, obtain the updated weight of the weighted fingerprint, input the updated fingerprint vector into the clustering layer for clustering again, and obtain the updated IP clustering result;
[0017] (9) Repeat the above process (2) at different time periods and different nodes, and store the incremental IPs into the IP library;
[0018] (10) Record dynamically changing data and repeat (3)-(8) to continuously and dynamically expand the network access nodes of the same software-defined network service domain. When the expansion results decrease sharply, it means that the optimal discovery effect has been achieved and clustering is stopped. IP classes with high density in the clustering results are more likely to have network service access points. Find the group of points with the highest density in the three-category clustering diagram and use the overlapping points of the three-category results as service access points to complete the node expansion.
[0019] The line extension results are used as an evaluation method. By performing asset detection on the line extension nodes, extracting their registration information, flow information, etc., analyzing security and reliability, and finally integrating all the line extension results of each software-defined network service domain as the basis for security and reliability evaluation of the corresponding software-defined network service domain.
[0020] The technical solution of the present invention is:
[0021] A network service access node extension method based on IP fingerprint multi-view clustering, the steps of which include:
[0022] 1) Selecting several traffic characteristics of the domain network access node and setting a corresponding screening condition according to each selected traffic characteristic;
[0023] 2) Filtering nodes that meet the filtering conditions from the software defined network as network service access nodes;
[0024] 3) The IP of each network service access node obtained in step 2) is used as the target IP to form an IP library;
[0025] 4) Obtaining the traffic characteristics corresponding to each of the target IPs as the IP fingerprint of the corresponding target IP to form a primary IP fingerprint library;
[0026] 5) Upgrading each IP fingerprint in the primary IP fingerprint library through a machine learning model;
[0027] 6) Selecting features that need to be weighted from the IP fingerprints in the primary IP fingerprint library and initializing their weights;
[0028] 7) Selecting multiple clustering algorithms to form a clustering layer, using each clustering algorithm in the clustering layer to cluster the IP fingerprints processed in step 6), dividing the network service access nodes of the same software-defined network service domain into the same cluster, and generating a primary clustering graph corresponding to each clustering algorithm; wherein different clustering algorithms focus on different features;
[0029] 8) updating the weights of corresponding features in the IP fingerprint according to each of the primary clustering graphs, and then clustering the updated IP fingerprints using each clustering algorithm in the clustering layer to obtain updated IP clustering results;
[0030] 9) The IP class with the highest density in each IP clustering result finally obtained in step 8) is used as the network service access node of the corresponding software-defined network service domain to complete the node extension.
[0031] Furthermore, the method for filtering out nodes that meet the filtering conditions from the software-defined network is as follows: first, passive detection is performed using the traffic data of known software-defined network services to determine whether the traffic characteristics of each node meet the filtering conditions, and the node IP that meets the filtering conditions is used as the target IP; then, active detection is performed using a breadth-first algorithm, with each target IP as the starting node and expanding outward layer by layer to determine whether the traffic characteristics of each node actively detected meet the filtering conditions. If the traffic characteristics of the node meet the filtering conditions, the corresponding node is used as a network service access node.
[0032] Furthermore, the method for filtering out nodes that meet the filtering conditions from the software defined network is:
[0033] 21) Divide the nodes corresponding to each target IP into multiple queues according to the software-defined service to which they belong; and perform IP scanning on the multiple queues simultaneously;
[0034] 22) Maintain a detection depth parameter n for determining the range of each round of scanning; perform steps 23) to 25) during each round of scanning;
[0035] 23) Select a target node b from the current queue and obtain the node with n hops connected to the target node b according to the detection depth n, and send a setting message to detect the reachability and activity of the obtained node;
[0036] 24) If the activity of the detected node is greater than the set threshold, the depth n value corresponding to the target node b is gradually increased to detect more layers of nodes and gradually expand the detection range;
[0037] 25) When the detection depth n corresponding to the target node b reaches a set maximum value or satisfies a set detection condition, stop scanning the target node b.
[0038] Furthermore, according to the reachability metric Determine the reachability of the node; where u is the detection server node, v is the reachability metric node, and d(i,v) is the Boolean value of the reachability from node i to node v.
[0039] Furthermore, according to the activity metric Determine the activity of the node; where f(v,t i )=Tr(v,t i )+Re(v,t i )+Fl(v,t i ), v is the activity measurement node, T is the measurement time period, N is the number of discrete time points in the time window Δt, t i is the i-th time point, f(v,t i ) is the time point t i The amount of activity in the body, Tr(v,t i ) is at time t i The number of bytes of traffic passing through node v, Re(v,t i ) is at time t i The number of requests received by internal node v, Fl(v,t i ) is at time t i The number of TCP connections within node v.
[0040] Furthermore, the method for updating the weight of the corresponding feature in the IP fingerprint according to each of the primary clustering graphs is as follows: extracting a batch of points with the highest density from each of the primary clustering graphs, obtaining points with the same IP from the extracted points as candidate service access points; for each selected feature x to be weighted, calculating the consistency index of the feature x based on the feature value of the feature x in each candidate service access point, and when the consistency index of the feature x exceeds a set threshold, Adjust the weight of the feature x; where, represents the weight of feature x after updating, represents the weight of feature x before updating; α is the learning rate, which is responsible for controlling the adjustment amplitude; C(x) is the contribution of feature x.
[0041] Furthermore, the selected traffic characteristics include traffic density characteristics, traffic distribution characteristics, traffic peak characteristics, upstream and downstream traffic asymmetry characteristics, traffic distribution difference characteristics of different applications, traffic and geographic location correlation characteristics, and traffic encapsulation and forwarding overhead characteristics.
[0042] Furthermore, the primary IP fingerprint library uses a hash table to store each IP fingerprint.
[0043] Furthermore, the IP fingerprints in the primary IP fingerprint library include TCP fingerprints, UDP fingerprints, ICMP fingerprints and page banner fingerprints; the features that need to be weighted include ASN, TTL and port.
[0044] Furthermore, the clustering algorithms in the clustering layer include K-means clustering algorithm, hierarchical clustering algorithm and DBSCAN clustering algorithm.
[0045] The advantages of the present invention are as follows:
[0046] Based on the breadth search and the adaptive IP scanning algorithm combined with the traffic characteristics of the domain network access node, the network access point IP seed data can be efficiently captured, and then the network access point IP fingerprint of the software-defined network service can be analyzed and extracted. The machine learning model is used to upgrade the feature dimension to obtain sufficiently robust fingerprint features, and then it is input into multiple clustering layers. It can effectively expand the line from the known software-defined network domain network access nodes and accurately discover enough domain network access nodes provided by the same software-defined network service provider. The expansion line results can be further used as the evaluation data source for the domain network of the same software-defined network service provider. By performing asset detection on the expansion line nodes, extracting their filing information, flow information, etc., analyzing security and reliability, and finally integrating all the expansion line results of the domain as the evaluation basis for the domain. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 Schematic diagram of the method for extending the network access nodes for the same software-defined network service area. DETAILED DESCRIPTION
[0048] The present invention is further described in detail below in conjunction with the accompanying drawings. The examples given are only used to explain the present invention but not to limit the scope of the present invention.
[0049] Figure 1The flowchart of the method for extending the domain network access node of the same software defined network service provided by the present invention is shown. The method is based on the adaptive IP scanning algorithm of breadth search and combined with the traffic characteristics of the domain network access node, which can efficiently capture the IP seed data of the network access point, and then analyze and extract the IP fingerprint of the network access point of the software defined network service, use the machine learning model to upgrade the feature dimension to obtain sufficiently robust fingerprint features to form a primary IP fingerprint library, select the features that need to be weighted for the software defined network service to form a software defined network weighted fingerprint, generate the fingerprint vector of each IP according to the primary IP fingerprint library and input it into the clustering algorithm of the clustering layer, the IP class with high density in the clustering result is more likely to have a network service access point, find a batch of points with the highest density in the three types of clustering graphs respectively, and use the overlapping points of the three types of results as service access points, so as to complete the node extension, so as to effectively extend the line from the known software defined network domain network access node, and use the newly added extension data to reversely iterate the feature weight of the software defined network weighted fingerprint, which is helpful to more accurately discover enough domain network access nodes provided by the same software defined network service provider in continuous detection. The main steps include:
[0050] 1. Determine traffic characteristics of domain network access nodes
[0051] A network access point refers to a series of edge facilities and infrastructure nodes deployed by network service providers in the network. They are used to connect networks in different geographical locations and act as imports and exports of traffic so that data can be transmitted and exchanged between different networks. The traffic at a domain network access node has the following characteristics:
[0052] 1) Traffic concentration
[0053] As a network edge node, the domain network access node usually concentrates a large amount of user traffic. User traffic from different regions is aggregated and forwarded through the nearest domain network access node, resulting in very concentrated and high-density traffic at the domain network access node. The corresponding characteristic index is expressed by calculating the traffic density within a certain period of time.
[0054] 2) Uneven traffic distribution
[0055] The number of users at different LAN access nodes varies greatly, so the traffic distribution of each LAN access node is also uneven. The traffic of LAN access nodes in densely populated cities may be very high, while the traffic of LAN access nodes in remote areas is relatively low. The corresponding characteristic index is expressed by calculating the degree of traffic dispersion within a certain area.
[0056] 3) Sudden traffic peaks
[0057] Due to the differences in the usage patterns of network applications, the traffic of domain network access nodes will show obvious sudden peak characteristics, such as video live broadcast period, online live events, etc., which will bring short-term but extremely high traffic peaks to the domain network access nodes. The corresponding characteristic indicators are expressed by calculating the traffic peak within a certain period of time.
[0058] 4) Asymmetric uplink and downlink traffic
[0059] Upstream traffic (user->network access point) is usually relatively small, mainly request data; while downstream traffic (network access point->user) is large, including response data content, so there will be a certain asymmetry between upstream and downstream traffic. The corresponding characteristic index is expressed by calculating the difference between the average upstream and downstream traffic of the user.
[0060] 5) Obvious differences in application distribution
[0061] Different types of network applications (such as web browsing, online video, file downloading, etc.) are distributed in different proportions in each domain network access node, and the corresponding traffic patterns are quite different. The corresponding characteristic indicators are expressed by calculating the variance of the user's traffic type. 6) Geographic location correlation
[0062] The traffic patterns and characteristics of domain network access nodes in the same geographical location are often similar, which is closely related to the Internet habits and application usage of netizens in the region. The corresponding characteristic indicators are calculated by calculating the degree of correlation between traffic and application information in the region.
[0063] 7) Encapsulation and forwarding overhead
[0064] Since the domain network access node needs to perform operations such as encapsulation, decapsulation, and forwarding on the traffic, a certain delay and bandwidth overhead will be generated, affecting the delay and throughput performance of the traffic. The corresponding characteristic indicators are expressed by calculating the average delay and throughput of the traffic.
[0065] The rule for selecting features is to calculate the index of each feature, and then set a threshold for each traffic feature as the corresponding screening condition.
[0066] (II) Breadth-first IP detection
[0067] In terms of detection breadth, a combination of active and passive strategies is adopted. First, passive detection is performed using the traffic data of known software-defined network services, and the characteristics of domain network access nodes in the traffic are analyzed to identify possible target IPs. This passive detection method is not only effective, but also does not interfere with normal network communications. Secondly, active detection is performed using the breadth-first algorithm, with the passively discovered target IP as the starting node, expanding outward layer by layer, first visiting all neighbor nodes of the starting node, and then visiting their neighbors layer by layer, and so on, until the target node that meets the traffic characteristics of the domain network access node is found or the entire graph is traversed. This method can effectively improve the detection efficiency while increasing the detection range as much as possible.
[0068] (III) Adaptive IP scanning strategy
[0069] The adaptive IP scanning strategy is designed to intelligently select the scanning depth in the breadth-first IP detection scenario to better balance the efficiency and scope of detection. The details of the adaptive IP scanning strategy are as follows:
[0070] 1) Breadth search basis: The adaptive IP scanning strategy is based on the basic idea of breadth search. The software-defined network IP nodes (i.e., each target IP) obtained in the previous step are divided into multiple queues according to the software-defined services they belong to; multiple queues perform IP scanning simultaneously.
[0071] 2) Adaptive depth selection: Maintain a detection depth parameter n, which will determine the scope of each round of scanning. Initially, the depth can be set to 1, indicating that only the neighboring nodes of the target IP address are scanned.
[0072] 3) Detect target nodes: select an IP node b from the current queue and obtain the nodes with n hop connections to the IP node b according to the detection depth n, send specific messages to detect the reachability and activity of these nodes, and perform detection (such as sending ICMP requests) to determine their reachability and activity.
[0073] 4) Depth increasing strategy: The update of the depth parameter n refers to a certain strategy. If the target node is active, the algorithm can gradually increase the depth parameter to detect more levels of nodes and gradually expand the scope.
[0074] 5) Adaptive stop condition: Set the stop condition, such as stopping the scan of node b when the detection depth reaches a certain level or meets certain success conditions, to prevent excessive detection and waste of resources.
[0075] Regarding the quantification of reachability and activity, we use the following formula to complete the calculation:
[0076]
[0077] Where D(u,v) is the reachability metric, u is the detection server node, v is the reachability metric node, and d(i,v) is the reachability Boolean value between nodes i and v;
[0078]
[0079] f(v,t i )=Tr(v,t i )+Re(v,t i )+Fl(v,t i )
[0080] Where A(v,T) is the activity measurement, v is the activity measurement node, T is the measurement time period, N is the number of discrete time points in the time window Δt, and t i is a time point, f(v,t i ) is the time point t i The amount of activity within.
[0081] f(v,t i ) is obtained by summing the three activity quantity representation methods, Tr(v,t i ) is at time t i The number of bytes of traffic passing through node v, Re(v,t i ) is at time t i The number of requests received by internal node v, Fl(v,t i ) is at time t i The number of TCP connections within node v.
[0082] Through the strategy, the depth is dynamically and intelligently selected, and the scanning range is expanded according to the detection results. The target node can be quickly located at the beginning, reducing the scanning time and resource consumption. Different stop conditions can also be set as needed, such as finding the target node or reaching the maximum depth, to control the scanning process and reasonably allocate detection resources.
[0083] (IV) Intelligent extraction of IP fingerprints based on machine learning models
[0084] Traditional traffic feature extraction mainly targets Internet traffic, which is characterized by routine access, poor confidentiality, and rich samples. For traffic passing through network access points, the access methods are complex, and different encryption protocols and implementation methods make the traffic features extremely diverse, making it difficult to obtain samples, and the applicability of traditional identification methods is limited. This method uses traditional IP fingerprint extraction technology and combines it with machine learning to achieve intelligent dimensionality upgrade of IP fingerprints.
[0085] First, use TCP, ICMP and UDP protocols to detect the seed IP, and its main fingerprint details are:
[0086] 1) IP fingerprint generation and detection based on TCP:
[0087] a) TCP connection flags (SYN, ACK, FIN, etc.) and sequence numbers.
[0088] b) Parameters of the TCP header such as window size and option field.
[0089] c) Connection timeout and delay.
[0090] d) The process of establishing and terminating a TCP connection.
[0091] 2) IP fingerprint generation and detection based on ICMP:
[0092] a) ICMP echo request and response (Ping).
[0093] b) ICMP error message (ICMP error report).
[0094] c) ICMP timestamp request and response, etc.
[0095] 3) UDP-based IP fingerprint generation and detection:
[0096] a) Source port and destination port.
[0097] b) The length and content of the UDP data packet.
[0098] c) Frequency and traffic pattern of packets.
[0099] Secondly, use machine learning technology to train the model to automatically extract the characteristics of the target IP. These models can be classified based on factors such as the statistical characteristics of IP attributes and protocol identifiers, thereby achieving efficient IP fingerprint identification. First, meaningful features are extracted from the network data for feature engineering, so that different types of encrypted traffic can be effectively distinguished. Usually, these features include protocol identifiers, packet sizes, delay characteristics of the traffic to which the IP belongs, frequency distribution, etc. Then, the data set construction, model selection, training optimization and model evaluation are carried out in sequence to realize the construction of the learning model for IP fingerprint detection. When applied, the traffic data corresponding to the sample IP needs to be input into the learning model, and then the model performs adaptive learning based on the software-defined network service traffic and the characteristics of the network access point itself, and finally outputs an IP fingerprint that can effectively represent the network access point.
[0100] Finally, it is combined with traditional detection methods to achieve efficient fingerprint extraction of the target seed IP.
[0101] (V) Construction of primary IP fingerprint library
[0102] The construction of the fingerprint library is aimed at storing and managing a large number of feature fingerprints to facilitate subsequent clustering analysis.
[0103] When building a fingerprint library, the first task is to select an appropriate data structure and storage format to effectively organize and store feature fingerprints. The selection of an appropriate data structure should take into account storage requirements, query performance, and data update requirements. The present invention selects a hash table for storage, and the characteristics of a linear table improve the efficiency of retrieval.
[0104] There are many software-defined network service providers, and the corresponding detected IP fingerprints also change frequently, which leads to constant changes in the data of the fingerprint library. Therefore, it is necessary to establish an effective data update and maintenance strategy, including data cleaning, incremental update, data backup and transaction management, to ensure the accuracy and consistency of the data. In addition, an efficient fingerprint library requires continuous performance tuning and query optimization. Methods such as query caching, database partitioning, index optimization and load balancing can be used to improve query performance. Finally, it is crucial to ensure the data security of the fingerprint library. Appropriate security measures can be taken, including access control, authentication, encrypted transmission, etc., to protect the fingerprint library from unauthorized access and attacks.
[0105] (VI) Selecting features that need to be weighted from the IP fingerprints in the primary IP fingerprint library for software-defined network services
[0106] The present invention combines the characteristics of the domain network access node itself and the characteristics of the traffic flowing through the software-defined network, and optimizes the performance of the clustering algorithm by adjusting the weight of the feature. The adjustment of the feature weight is crucial for accurately clustering the IP addresses related to the domain network access node of the software-defined network service. Adjusting the weight can highlight the differences between different software-defined network service providers, ensuring that the unique features of different software-defined network service providers receive higher attention during the clustering process, effectively increasing the distance between classes, and thus optimizing the generation results of IP clusters.
[0107] These features include specific data transmission protocols, specific data encryption methods, and specific data flow patterns, etc. At the same time, reducing the weights of features that are irrelevant or less important to the domain network access node to reduce their interference with the clustering results helps filter out noise or irrelevant features and improve the accuracy of clustering.
[0108] In general, the feature weights are optimized as follows:
[0109] 1) Specific ASN (Autonomous System Number): The same SDN service domain network access node is usually associated with a specific ASN, which is used to identify the network or service provider to which the node belongs. Therefore, the weight of the AS to which the source IP belongs and the network segment to which the source IP belongs is increased.
[0110] 2) Network topology: The access nodes of the same software-defined network service domain have specific positions in the network topology, so the weights of the average packet routing forwarding times and the average TTL are increased.
[0111] 3) Specific protocols and ports: The same SDN service domain network access nodes usually use specific network protocols and ports to provide services and communicate with users, so the weight of IP open ports and closed ports is increased.
[0112] 4) Traffic pattern: The same SDN service domain network access nodes usually generate specific traffic patterns, such as a large amount of upstream or downstream traffic in a specific period, or a specific packet frequency. Therefore, the weights of TCP window size, packet frequency and traffic pattern are increased.
[0113] The weights of each fingerprint are set to be the same as the initialization of the weighted fingerprint of the software-defined network.
[0114] In addition, the present invention also adopts an adaptive weight adjustment strategy to dynamically adjust the weight of the feature according to the dynamic changes in the distribution of data and the clustering process. Each iteration will recalculate the centrality and density of the graph and use it as the basis for weight adjustment, which makes the method more flexible and can better adapt to the changing network traffic situation.
[0115] (VII) Select three clustering algorithms according to the principle of progressive complexity to form a clustering layer
[0116] IP fingerprints include: ① Primary IP fingerprint library, mainly including TCP fingerprint, UDP fingerprint, ICMP fingerprint and page banner fingerprint; ② Software Defined Network weighted fingerprint, mainly including ASN, TTL, port and other features. The key task of clustering technology is to select different sub-features from these two types of fingerprints and combine them to construct multi-dimensional feature vectors. These feature vectors can be used to characterize the behavior and characteristics of the access nodes of the software defined network service domain network.
[0117] To ensure the scientific nature of the clustering results, three clustering algorithms, K-means, hierarchical clustering and DBSCAN, are used to cluster the feature vectors according to the principle of progressive complexity, and the network service access nodes of the same software-defined network service domain are divided into the same cluster. The three algorithms focus on different aspects of the node characteristics. K-means considers the numerical characteristics of the node, hierarchical clustering considers the relationship structure of the node, and DBSCAN considers the density characteristics of the node. Finally, the coincidence points of the three clustering algorithms are used as the extension nodes.
[0118] The following are the technical routes of three clustering algorithms:
[0119] 1) Use K-means algorithm to control the final number of IP clusters
[0120] a) Data preparation: Collect and prepare IP fingerprint datasets, and generate a feature vector based on the fingerprint features of each IP. The vector contains multiple feature dimensions such as the primary IP fingerprint library and software-defined network weighted fingerprint, such as TCP fingerprint, UDP fingerprint, ICMP fingerprint, page banner fingerprint, ASN, TTL, and port.
[0121] b) Initialization: Randomly select K IP addresses as the initial cluster centers, where K is the predefined number of clusters.
[0122] c) Iterative clustering: multiple rounds of iterations are performed, each round of iteration is divided into two steps:
[0123] i. Assignment phase: For each IP address, calculate the distance between it and each cluster center and assign it to the nearest cluster.
[0124] ii. Update phase: Update the center of each cluster so that the average feature vector of the IP addresses in the cluster is closest to the new cluster center.
[0125] d) Convergence judgment: When the cluster center no longer changes significantly or reaches the predetermined number of iterations, the algorithm converges.
[0126] 2) Use hierarchical clustering algorithm to control the distance within IP clusters
[0127] a) Data preparation: Similar to K-means, first prepare the IP fingerprint dataset and generate a feature vector based on the fingerprint features of each IP.
[0128] b) Hierarchical structure construction: Starting from the dataset, each IP address is treated as an independent cluster, and then a hierarchical dendrogram (dendroclustering) is built by merging the nearest clusters.
[0129] c) Hierarchical division: Specific clusters can be divided according to the structure of the dendrogram and the threshold of the cluster distance.
[0130] 3) Use DBSCAN algorithm to control the distance between IP clusters
[0131] a) Data preparation: Similarly, prepare the IP fingerprint dataset and generate a feature vector based on the fingerprint features of each IP.
[0132] b) Core point selection: Select core points based on the specified radius ε and the minimum number of neighbors MinPts. A core point is a point that contains at least MinPts IP addresses within the radius ε.
[0133] c) Density reachability: Clusters are constructed by density reachability between core points. If a point can reach another core point through a series of jump connections starting from a core point, they belong to the same cluster.
[0134] d) Noise points: Points that are not connected to core points are considered noise points.
[0135] (VIII) Reverse iterative weighted features to update IP clustering results
[0136] Get the IP clustering results and reversely map them to the software-defined network weighted fingerprint, that is, count the weights of different parameters in the software-defined network weighted fingerprint according to the network service access node IP associated with each software-defined network service domain in the primary IP clustering results, and get the updated weight of the weighted fingerprint of each software-defined network service domain. Input the IP library, primary IP fingerprint library and updated software-defined network weighted fingerprint into the clustering layer to get the updated IP clustering results.
[0137] For the weight adjustment method, after the clustering results of the three graphs, extract the highest density of the points, and take the points with the same IP in the three batches as candidate service access points. For each feature x selected in step (six) that needs to be weighted, calculate the consistency index of the feature x based on the feature value of the feature x in each candidate service access point. When the index exceeds a certain threshold, make a targeted weight adjustment. The consistency index of the feature refers to the inverse of the variance of the field value corresponding to the feature of all nodes.
[0138] This means re-evaluating the importance of certain features based on the clustering results. For example, if some IP addresses are correctly clustered into the same category, and they are very similar in certain specific features, the weights of these features may be enhanced because they play a key role in distinguishing IP behavior. On the contrary, if some features have no obvious impact in the clustering process, their weights can be appropriately reduced. The specific calculation formula is as follows:
[0139]
[0140] in, represents the weight of feature x after updating, represents the weight before update; α is the learning rate, which is responsible for controlling the adjustment amplitude; C(x) is the contribution of feature x, which is obtained by the consistency of the feature in similar IP nodes. The consistency index is obtained by the inverse of the variance of the field values corresponding to all nodes under the feature; C is the total contribution of all features, which is obtained by summing the consistency index of each feature.
[0141] (IX) Repeat the breadth-first adaptive IP scanning at different time periods and different nodes, and store the incremental IPs into the IP library;
[0142] (10) Record dynamically changing data and repeat the above process to continuously and dynamically expand the network access nodes of the same software defined network service domain.
[0143] Figure 1 The figure is a flow chart of the method for expanding the access node of the same software defined network service domain network provided by the present invention. The overall process is divided into three parts:
[0144] 1) IP detection: Determine the traffic characteristics of the domain network access node. If the traffic of a certain node is very concentrated and high-density, has obvious sudden peak characteristics, there is a stable asymmetry in the upstream and downstream traffic, and the traffic patterns of domain network access nodes in the same geographical location are significantly similar, then this IP is considered to be a target domain network access node. Based on the selected domain network access node traffic characteristics, in a network where software-defined network services are known to exist, the domain network access node IP is obtained through passive network traffic. Then, a breadth-first adaptive IP scan is performed, with the passively discovered target IP as the starting node of the network segment to be detected, and it expands outward layer by layer, first visiting all neighbor nodes of the starting node, and then visiting their neighbors layer by layer, and so on, until the target node that meets the traffic characteristics of the domain network access node is found or the entire graph is traversed. In this process, the reachability and activity of the nodes in the network segment can be determined by sending ICMP requests, and the detection depth of accessing the next node layer by layer can be set accordingly. If the activity of the network segment nodes exceeds 90%, the detection depth is set to 1 (each neighboring node is traversed). When the network segment activity is 50%, the detection depth is set to 3 (every two nodes are detected, that is, only the third node adjacent to it is detected). Then the detection results can be obtained to form an IP library.
[0145] 2) Fingerprint extraction: Use TCP, ICMP and UDP protocols to detect the target IP and obtain the corresponding fingerprint features, such as the TCP connection flags (SYN, ACK, FIN, etc.) and sequence numbers, ICMP echo requests and responses (Ping), the length and content of UDP packets, etc., to form a primary IP fingerprint library. Then use machine learning technology to train the model to intelligently upgrade the target features. First, the data needs to be preprocessed, including standardization and normalization, to reduce the difference between data and the influence of noise. The model can choose a multi-layer perceptron (MLP) neural network for feature dimensionality upgrade. Unlike general linear models that can only obtain global features, MLP obtains local features through its deep architecture, which increases the data expression ability. Each layer of neurons receives the output of the neurons in the previous layer and performs weighted summation on it. Then, a nonlinear activation function such as Sigmoid or ReLU function is used to obtain the output of the neurons. The output layer usually obtains the probability value of the access node of the domain network through the Softmax function. When the probability value is set to be above 90%, the feature vector output by the hidden layer of the MLP network is the result of the feature dimensionality upgrade at this stage, which can be stored in the primary IP fingerprint library. At the same time, in order to avoid overfitting, regularization terms such as L1 or L2 regularization, or dropout technology can be added during the training process. In addition, in order to optimize the network weights, the gradient descent method or its variants such as Adam are usually used to update the weights.
[0146] 3) Clustering extension line: For software-defined network services, select features that need to be weighted, such as ASN features, network topology features, protocol and port features, traffic pattern features, etc., and initialize them. Initialize them with the same weight, all set to 1. Then select three clustering algorithms according to the principle of progressive complexity to form a clustering layer, such as using the K-means algorithm to control the final number of IP clusters, using the hierarchical clustering algorithm to control the distance within the IP cluster, and using the DBSCAN algorithm to control the distance between IP clusters. Then generate the fingerprint vector of each IP based on the primary IP fingerprint library and input it into the layer to obtain the primary IP clustering result. Obtain the IP clustering result, reversely map it to the software-defined network weighted fingerprint, obtain the updated weight of the weighted fingerprint, input the updated fingerprint vector into the clustering layer, and obtain the updated IP clustering result.
[0147] Repeat the IP detection process at different time periods and different nodes, and store the incremental IPs in the IP library. Record dynamically changing data, repeat fingerprint extraction and clustering to continuously and dynamically expand the access nodes of the same software-defined network service domain network.
[0148] Although the specific embodiments of the present invention are disclosed for the purpose of illustration, the purpose is to help understand the content of the present invention and implement it accordingly, those skilled in the art will understand that various substitutions, changes and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the best embodiment, and the scope of the present invention is subject to the scope defined in the claims.
Claims
1. A network service access node extension method based on IP fingerprint multi-view clustering, the steps comprising: 1) Selecting several traffic characteristics of the domain network access node and setting a corresponding screening condition according to each selected traffic characteristic; 2) Filtering nodes that meet the filtering conditions from the software defined network as network service access nodes; 3) The IP of each network service access node obtained in step 2) is used as the target IP to form an IP library; 4) Obtaining the traffic characteristics corresponding to each of the target IPs as the IP fingerprint of the corresponding target IP to form a primary IP fingerprint library; 5) Upgrading each IP fingerprint in the primary IP fingerprint library through a machine learning model; 6) Selecting features that need to be weighted from the IP fingerprints in the primary IP fingerprint library and initializing their weights; 7) Selecting multiple clustering algorithms to form a clustering layer, using each clustering algorithm in the clustering layer to cluster the IP fingerprints processed in step 6), dividing the network service access nodes of the same software-defined network service domain into the same cluster, and generating a primary clustering graph corresponding to each clustering algorithm; wherein different clustering algorithms focus on different features; 8) updating the weights of corresponding features in the IP fingerprint according to each of the primary clustering graphs, and then clustering the updated IP fingerprints using each clustering algorithm in the clustering layer to obtain updated IP clustering results; 9) The IP class with the highest density in each IP clustering result finally obtained in step 8) is used as the network service access node of the corresponding software-defined network service domain to complete the node extension.
2. The method according to claim 1, characterized in that The method for filtering out nodes that meet the filtering conditions from the software-defined network is as follows: first, passive detection is performed using the traffic data of known software-defined network services to determine whether the traffic characteristics of each node meet the filtering conditions, and the node IP that meets the filtering conditions is used as the target IP; then, active detection is performed using a breadth-first algorithm, with each target IP as the starting node and expanding outward layer by layer to determine whether the traffic characteristics of each actively detected node meet the filtering conditions. If the traffic characteristics of the node meet the filtering conditions, the corresponding node is used as a network service access node.
3. The method according to claim 2, characterized in that The method for filtering out nodes that meet the filtering conditions from the software defined network is: 21) Divide the nodes corresponding to each target IP into multiple queues according to the software-defined service to which they belong; and perform IP scanning on the multiple queues simultaneously; 22) Maintain a detection depth parameter n for determining the range of each round of scanning; perform steps 23) to 25) during each round of scanning; 23) Select a target node b from the current queue and obtain the node with n hops connected to the target node b according to the detection depth n, and send a setting message to detect the reachability and activity of the obtained node; 24) If the activity of the detected node is greater than the set threshold, the depth n value corresponding to the target node b is gradually increased to detect more layers of nodes and gradually expand the detection range; 25) When the detection depth n corresponding to the target node b reaches a set maximum value or satisfies a set detection condition, stop scanning the target node b.
4. The method according to claim 3, characterized in that According to the reachability metric Determine the reachability of the node; where u is the detection server node, v is the reachability metric node, and d(i,v) is the Boolean value of the reachability from node i to node v.
5. The method according to claim 3, characterized in that: According to the activity measurement Determine the activity of the node; where f(v,t i )=Tr(v,t i )+Re(v,t i )+Fl(v,t i ), v is the activity measurement node, T is the measurement time period, N is the number of discrete time points in the time window Δt, t i is the i-th time point, f(v,t i ) is the time point t i The amount of activity in the body, Tr(v,t i ) is at time t i The number of bytes of traffic passing through node v, Re(v,t i ) is at time t i The number of requests received by internal node v, Fl(v,t i ) is at time t i The number of TCP connections within node v.
6. The method according to claim 1, characterized in that The method for updating the weight of the corresponding feature in the IP fingerprint according to each of the primary clustering graphs is as follows: extracting a batch of points with the highest density from each of the primary clustering graphs, and obtaining points with the same IP from the extracted points as candidate service access points; for each selected feature x that needs to be weighted, calculating the consistency index of the feature x based on the feature value of the feature x in each candidate service access point, and when the consistency index of the feature x exceeds a set threshold, Adjust the weight of the feature x; where, represents the weight of feature x after updating, represents the weight of feature x before updating; α is the learning rate, which is responsible for controlling the adjustment amplitude; C(x) is the contribution of feature x.
7. The method according to claim 1, characterized in that The selected traffic characteristics include traffic density characteristics, traffic distribution characteristics, traffic peak characteristics, upstream and downstream traffic asymmetry characteristics, traffic distribution difference characteristics of different applications, traffic and geographic location correlation characteristics, and traffic encapsulation and forwarding overhead characteristics.
8. The method according to claim 1, characterized in that: The primary IP fingerprint library uses a hash table to store each IP fingerprint.
9. The method according to claim 1 or 8, characterized in that: The IP fingerprints in the primary IP fingerprint library include TCP fingerprints, UDP fingerprints, ICMP fingerprints and page banner fingerprints; the features that need to be weighted include ASN, TTL and port.
10. The method according to claim 1, characterized in that The clustering algorithms in the clustering layer include K-means clustering algorithm, hierarchical clustering algorithm and DBSCAN clustering algorithm.
Citation Information
Patent Citations
Model construction method, malicious code identification method, storage medium and terminal
CN115983342A
Network topology generation method and system
CN116915620A
POP node discovery method and device based on similarity correlation analysis
CN118487938A
Self organizing learning topologies
US20170310691A1
Cited By
A data processing system for network service fingerprinting
CN122741387A