Network service access node topology method based on ip fingerprint multi-view clustering

By employing an IP fingerprint-based multi-view clustering method, and combining breadth-first search and adaptive IP scanning techniques with a machine learning model, the problem of discovering LAN access nodes in software-defined networks was solved. This enabled accurate assessment and efficient network expansion for the same software-defined network service provider, thereby improving network security and reliability assessment.

CN119945936BActive Publication Date: 2025-11-25INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411914941.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-11-25
Estimated Expiration
2044-12-24

AI Technical Summary

Technical Problem

In software-defined network architecture, it is difficult to efficiently and accurately discover LAN access nodes of the same software-defined network service provider, which affects the security and reliability assessment of the network.

Method used

The network service access node topology method based on IP fingerprint multi-view clustering extracts traffic features through breadth-first search and adaptive IP scanning technology. It combines machine learning models and multiple clustering algorithms to analyze and separate the IP fingerprints of network access points of software-defined network services, forming a clustering layer to discover domain network access nodes of the same service provider.

Benefits of technology

It improves the breadth and accuracy of discovering the same network service access nodes from massive amounts of internet data, effectively assesses the security and reliability of software-defined network services, and provides high-quality network service options.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119945936B_ABST
    Figure CN119945936B_ABST
Patent Text Reader

Abstract

The application discloses a network service access node topology method based on IP fingerprint multi-view clustering, and steps of the method comprise the following steps: 1) selecting traffic characteristics of domain network access nodes; 2) performing IP scanning based on the traffic characteristics of the domain network access nodes; 3) obtaining scanning results to form an IP library; 4) forming a primary fingerprint library according to the scanning results; 5) upgrading IP fingerprint characteristics and storing the IP fingerprint characteristics in the primary IP fingerprint library; 6) selecting features needing weighting for a software-defined network service, and initializing weights of the features; 7) inputting the IP library, the primary fingerprint library and the initialized software-defined network weighted fingerprint into a clustering layer to obtain primary IP clustering results; 8) obtaining the IP clustering results, and reversely mapping the IP clustering results to the software-defined network weighted fingerprint to obtain updated weights of the weighted fingerprint; and 9) repeating the steps 3) to 8) to continuously and dynamically perform topology on the same software-defined network service domain network access nodes.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of software-defined network service and relates to a network service access node mapping method based on IP fingerprint multi-view clustering. BACKGROUND

[0002] With the promotion of the digital transformation of various vertical industries in China, the digitization, computing power and intelligence of information communication networks will become an inevitable choice for high-quality development of industries. Under the demand for service and response ability of bottom-layer networks driven by emerging businesses such as cloud services, AI, big data and the Internet of Things, the concept of software-defined networks gradually fermented in the ICT field and was introduced into the enterprise network service supply market, prompting the emergence of software-defined network solutions.

[0003] However, it is not easy to choose a suitable software-defined network service provider. At present, most software-defined network service providers claim to be able to provide high-quality services, but in fact, their services may have problems of poor performance, security and privacy protection ability. In the software-defined network architecture, the access point of the data center plays a key role, and the state of the domain network access node not only affects the running efficiency of the network, but also directly relates to the security and stability of the network, therefore, discovering enough domain network access nodes of the same software-defined network service provider is crucial to evaluating the security and reliability of the software-defined network service provided by the provider. SUMMARY

[0004] In view of the problems in the prior art, the purpose of the application is to provide a network service access node mapping method based on IP fingerprint multi-view clustering, which can accurately discover enough domain network access nodes of the same software-defined network service provider, thereby helping users to evaluate the security and reliability of the software-defined network service provider and to select the most suitable high-quality software-defined network service for their own needs.

[0005] The application extracts the IP fingerprint of the access point of the network service based on the traffic characteristics, combines the breadth search and the adaptive IP scanning algorithm, uses the machine learning model to upgrade the features to obtain a robust fingerprint feature, and then through multiple clustering layers, the network service access node can be effectively mapped from the known network access nodes, thereby improving the breadth and accuracy of discovering the same network service access node from the massive data on the Internet.

[0006] The main difficulties to be solved by the present application are: ①efficient and wide-range detection and collection of network access point seed IP (i.e. target IP) data; ②extraction of IP feature fingerprints based on seed IP data, so as to effectively distinguish software-defined network service network access point traffic from general node traffic; and ③efficient mapping of domain network access nodes that may exist under the same software-defined network service based on the IP data that has been probed.

[0007] To overcome the above difficulties, the present application uses a plurality of innovative technologies to complete the predetermined requirements. The main related technologies are: ①adaptive IP scanning technology based on breadth search and combined with domain network access node traffic characteristics: research on host inventory discovery technology based on ICMP, TCP and other means, adaptive selection of efficient breadth scanning strategy for network access point traffic for IP scanning, to ensure scanning breadth and efficiency, and complete the detection and collection of basic seed data; ②software-defined network access point IP fingerprint extraction technology based on adaptive weight adjustment strategy: analyze the port service, certificate, banner, webpage and other characteristic information of the IP address of the software-defined network service network access point, and clean and preprocess the data to reduce the adverse effects of noise on the characteristics, and then dynamically adjust the weight of the characteristics to dynamically provide more attention to different high-quality characteristics in different software-defined network services; ③domain network access node mapping technology based on multi-view clustering: use K-means, hierarchical clustering and DBSCAN three clustering methods to cluster the fingerprint feature vectors, and divide similar IPs into the same cluster, thereby forming IP clusters in the same software-defined network service provider unit.

[0008] The specific steps of the present application include:

[0009] (1) selecting the traffic characteristics of the domain network access node;

[0010] (2) breadth-first adaptive IP scanning based on the traffic characteristics of the domain network access node;

[0011] (3) obtaining the scanning results to form an IP library;

[0012] (4) cleaning and preprocessing the scanning results to obtain standard traffic characteristics and thereby form a primary IP fingerprint library;

[0013] (5) dimensioning the IP fingerprint through a machine learning model to obtain a feature vector robust enough, and storing it in the primary IP fingerprint library;

[0014] (6) selecting the features that need to be weighted for the software-defined network service, and initializing the weight of the features;

[0015] (7) According to the complexity progressive principle, three clustering algorithms are selected to form a clustering layer, the fingerprint vector of each IP is generated according to the primary IP fingerprint library and is input into the clustering algorithm of the clustering layer in parallel, and the primary IP clustering result is obtained;

[0016] (8) The IP clustering result is obtained, is reversely mapped to the software defined network weighted fingerprint, the update weight of the weighted fingerprint is obtained, the updated fingerprint vector is input into the clustering layer to perform clustering again, and the updated IP clustering result is obtained;

[0017] (9) The above process (2) is repeated at different time periods and different nodes, and the incremental IP is stored in the IP library;

[0018] (10) The dynamic change data is recorded, and (3)-(8) are repeated, the same software defined network service domain network access node is continuously and dynamically mapped, when the mapping result is reduced, it means that the optimal discovery effect has been reached, and the clustering is stopped; the IP class with high density in the clustering result is more likely to exist in the network service access point, a batch of points with the highest density in the three clustering maps are found, and the coincident points of the three results are taken as the service access point, so that the node mapping is completed.

[0019] The mapping result is taken as an evaluation method, asset detection is performed on the mapped node, record information, flow information and the like are extracted, security and reliability are analyzed, and finally all the mapping results of each software defined network service domain are integrated as the security and reliability evaluation basis of the corresponding software defined network service domain.

[0020] The technical scheme of the application is:

[0021] A network service access node mapping method based on IP fingerprint multi-view clustering, comprising the following steps:

[0022] 1) Selecting a plurality of traffic characteristics of a domain network access node, setting a corresponding screening condition according to each selected traffic characteristic;

[0023] 2) Screening nodes meeting each screening condition from the software defined network as network service access nodes;

[0024] 3) Forming an IP library according to the IP of each network service access node obtained in step 2) as a target IP;

[0025] 4) Obtaining the traffic characteristic corresponding to each target IP as the IP fingerprint of the corresponding target IP to form a primary IP fingerprint library;

[0026] 5) Dimensionality of each IP fingerprint in the primary IP fingerprint library is increased through a machine learning model;

[0027] 6) selecting features from the IP fingerprints in the primary IP fingerprint library that need to be weighted and initializing the weights of the features;

[0028] 7) selecting multiple clustering algorithms to form a clustering layer, using each clustering algorithm in the clustering layer to cluster the IP fingerprints processed in step 6) respectively, dividing network service access nodes of the same software-defined network service domain into the same cluster, and generating a primary clustering graph corresponding to each clustering algorithm; wherein different clustering algorithms focus on different features;

[0029] 8) updating the weights of the corresponding features in the IP fingerprints according to the primary clustering graphs, and then using each clustering algorithm in the clustering layer to cluster the updated IP fingerprints respectively to obtain updated IP clustering results;

[0030] 9) taking the IP class with the highest density in each IP clustering result obtained in step 8) as the network service access node of the corresponding software-defined network service domain, and completing the node topology.

[0031] Further, the method for screening nodes meeting the screening conditions from the software-defined network is: first, using the traffic data of known software-defined network services for passive detection to determine whether the traffic characteristics of each node meet the screening conditions, and taking the node IP meeting the screening conditions as a target IP; then, using the breadth-first algorithm for active detection, expanding layer by layer outward from each target IP as a starting node, determining whether the traffic characteristics of each node detected actively meet the screening conditions, and if the traffic characteristics of the node meet the screening conditions, taking the corresponding node as a network service access node.

[0032] Further, the method for screening nodes meeting the screening conditions from the software-defined network is:

[0033] 21) dividing the nodes corresponding to each target IP into multiple queues according to the software-defined services to which they belong; and performing IP scanning on the multiple queues simultaneously;

[0034] 22) maintaining a detection depth parameter n for determining the range of each round of scanning; and performing steps 23) to 25) in each round of scanning;

[0035] 23) selecting a target node b from the current queue and obtaining nodes having n-hop connections with the target node b according to the detection depth n, and sending a set message to probe the reachability and activity of the obtained nodes;

[0036] 24) if the activity of the probed nodes is greater than a set threshold, gradually increasing the depth n value corresponding to the target node b so as to probe more levels of nodes and gradually expand the detection range;

[0037] 25) When the target node b corresponds to the detection depth n reaches the set maximum value or meet the set detection conditions, stop scanning the target node b.

[0038] Further, according to the reachability metric Determine the reachability of the node; wherein u is the probe server node, v is the reachability metric node, d(i, v) is the reachability Boolean value between node i and node v.

[0039] Further, according to the liveliness metric Determine the liveliness of the node; wherein f(v, t i ) = Tr(v, t i ) + Re(v, t i ) + Fl(v, t i ), v is the liveliness metric node, T is the metric time period, N is the number of discrete time points within the time window Δt, t i is the i th time point, f(v, t i ) is the activity amount within the time point t i , Tr(v, t i ) is the number of flow bytes passing through the node v within the time point t i , Re(v, t i ) is the number of requests accepted by the node v within the time point t i , Fl(v, t i ) is the number of TCP connections of the node v within the time point t i .

[0040] Further, the method for updating the weight of the corresponding feature in the IP fingerprint according to each said primary clustering graph is: extracting a batch of points with the highest density from each said primary clustering graph, respectively, and obtaining points with the same IP from the extracted points as candidate service access points; for each selected feature x that needs to be weighted, calculating a consistency index of the feature x based on the feature value of the feature x in each candidate service access point, and when the consistency index of the feature x exceeds a set threshold, adjusting the weight of the feature x by ; wherein, represents the weight of the feature x after updating, represents the weight of the feature x before updating; α is the learning rate, responsible for controlling the adjustment range; C(x) is the contribution degree of the feature x.

[0041] Further, the selected traffic features include traffic density features, traffic distribution features, traffic peak features, uplink and downlink traffic asymmetry features, traffic distribution difference features of different applications, traffic and geographical location correlation features, and traffic encapsulation and forwarding overhead features.

[0042] Further, the primary IP fingerprint library stores each IP fingerprint by using a hash table.

[0043] Further, the IP fingerprint in the primary IP fingerprint library includes TCP fingerprint, UDP fingerprint, ICMP fingerprint and page banner fingerprint; and the selected features to be weighted include ASN, TTL and port.

[0044] Further, the clustering algorithm in the clustering layer includes K-means clustering algorithm, hierarchical clustering algorithm and DBSCAN clustering algorithm.

[0045] The advantages of the present application are as follows:

[0046] Based on the breadth-first search and the adaptive IP scanning algorithm combining with the traffic characteristics of domain network access nodes, the network access point IP seed data can be efficiently captured, and then the network access point IP fingerprint of the software-defined network service can be analyzed and extracted, the machine learning model is used to upgrade the features to obtain a sufficient robust fingerprint feature, and then the fingerprint feature is input into multiple clustering layers, so that the network access nodes of the known software-defined network domain can be effectively expanded, and sufficient network access nodes of the same software-defined network service provider can be accurately found. BRIEF DESCRIPTION OF DRAWINGS

[0047] Figure 1 The figure is a schematic diagram of the method for expanding the network access nodes of the same software-defined network service domain. DETAILED DESCRIPTION

[0048] The present application will be further described in detail below with reference to the accompanying drawings, and the examples are only used to explain the present application, and are not used to limit the scope of the present application.

[0049] Figure 1A flowchart of a same software-defined network service domain network access node topology mapping method provided by the application is shown. Based on the breadth-first search and the adaptive IP scanning algorithm combined with the domain network access node traffic characteristics, the network access point IP seed data can be efficiently captured, and then the network access point IP fingerprint of the software-defined network service is analyzed and extracted. The machine learning model is used to upgrade the features to obtain a sufficient robust fingerprint feature to form a primary IP fingerprint library. The features that need to be weighted are selected for the software-defined network service to form a software-defined network weighted fingerprint. According to the primary IP fingerprint library, the fingerprint vector of each IP is generated and input into the clustering algorithm of the clustering layer. The IP class with high density in the clustering result is more likely to exist in the network service access point. Find the highest density of points in the three types of clustering graphs, and the overlapping points of the three types of results are used as the service access point, so as to complete the node topology mapping. Thus, the topology mapping from the known software-defined network domain network access node is effectively realized, and the feature weight of the software-defined network weighted fingerprint is iterated in reverse by using the newly mapped data, which helps to more accurately find enough domain network access nodes provided by the same software-defined network service provider in continuous detection. The main steps include:

[0050] (I) Determine the domain network access node traffic characteristics

[0051] Network access points refer to a series of edge facilities and infrastructure nodes deployed by network service providers in the network. They are used to connect different geographical locations of the network and act as the import and export of traffic so that data can be transmitted and exchanged between different networks. The domain network access node traffic has the following characteristics:

[0052] 1) Traffic concentration

[0053] As a network edge node, the domain network access node will usually concentrate a large amount of user traffic. The user traffic in different regions will be aggregated and forwarded through the nearest domain network access node, resulting in very concentrated and high-density traffic at the domain network access node. The corresponding feature index is represented by calculating the traffic density within a certain period of time.

[0054] 2) Uneven traffic distribution

[0055] The number of users in different locations of the domain network access node is very different, so the traffic distribution of each domain network access node is also uneven. The domain network access node in a densely populated city may have high traffic, while the domain network access node in a remote area may have low traffic. The corresponding feature index is represented by calculating the traffic dispersion within a certain region.

[0056] 3) High peak of burst traffic

[0057] Due to the difference in the use mode of network applications, the domain network access node traffic will show obvious burst peak characteristics, such as video live broadcast period, online live event, etc., bringing short but extremely high traffic flood to the domain network access node. The corresponding characteristic index is represented by calculating the traffic peak value in a certain period of time.

[0058] 4) Upstream and downstream traffic asymmetry

[0059] The upstream traffic (user -> network access point) is usually relatively small, mainly for requesting data; while the downstream traffic (network access point -> user) is larger, containing the response data content, so there is a certain asymmetry between the upstream and downstream traffic. The corresponding characteristic index is represented by calculating the difference between the average upstream and downstream traffic of the user.

[0060] 5) Significant difference in application distribution

[0061] Different types of network applications (such as web browsing, online video, file download, etc.) have different distribution ratios in each domain network access node, and the corresponding traffic mode difference is large. The corresponding characteristic index is represented by calculating the traffic type variance of the user.

[0062] The traffic mode and characteristics of the domain network access nodes in the same geographical location are often similar, which is closely related to the online habits and application use of the netizens in the region. The corresponding characteristic index is represented by calculating the correlation degree of the traffic in the region and the application information.

[0063] 7) Encapsulation and forwarding overhead

[0064] Since the domain network access node needs to perform encapsulation, decapsulation, forwarding and other operations on the traffic, it will produce certain delay and bandwidth overhead, affecting the delay and throughput performance of the traffic. The corresponding characteristic index is represented by calculating the average delay and throughput of the traffic.

[0065] The rule for selecting the characteristics is to calculate the index for each characteristic, and then set a threshold value for each traffic characteristic as the corresponding screening condition.

[0066] (II) Breadth-first IP probe

[0067] In terms of the detection range, a combination of active and passive strategies is adopted. First, passive detection is performed using the traffic data of known software-defined network services to analyze the domain network access node features in the traffic to identify possible target IPs. This passive detection method is not only effective but also does not interfere with normal network communication. Second, active detection is performed using a breadth-first algorithm, starting from the target IPs discovered passively, expanding layer by layer outward, first accessing all neighbor nodes of the starting node, then layer by layer accessing their neighbors, and so on, until the target nodes that meet the domain network access node traffic features are found or the entire graph is traversed. This method can effectively improve the detection efficiency while maximizing the detection range.

[0068] (III) Adaptive IP scanning strategy

[0069] The adaptive IP scanning strategy aims to intelligently select the scanning depth in the breadth-first IP detection scenario, better balancing the efficiency and range of detection. The details of the adaptive IP scanning strategy are as follows:

[0070] 1) Breadth search basis: The adaptive IP scanning strategy is based on the basic idea of breadth search, dividing the software-defined network IP nodes (i.e., each target IP) obtained in the previous step into multiple queues according to the software-defined services they belong to; multiple queues simultaneously perform IP scanning.

[0071] 2) Adaptive depth selection: Maintain a detection depth parameter n, which will determine the range of each round of scanning. Initially, the depth can be set to 1, indicating that only the neighbor nodes of the target IP address are scanned.

[0072] 3) Detection of target nodes: Select an IP node b from the current queue and obtain nodes with n-hop connectivity from the IP node b according to the detection depth n, send specific messages to probe the reachability and liveliness of these nodes, and perform detection (such as sending ICMP requests) to determine their reachability and liveliness.

[0073] 4) Depth increment strategy: The update of the depth parameter n refers to a certain strategy. If the target node is active, the algorithm can gradually increase the depth parameter to detect more levels of nodes and gradually expand the range.

[0074] 5) Adaptive stopping condition: Set a stopping condition, such as stopping scanning for node b when the detection depth reaches a certain level or meets certain success conditions, to prevent excessive detection and resource waste.

[0075] Regarding the quantification of reachability and liveliness, the following formula is used to complete the calculation:

[0076]

[0077] Wherein, D(u, v) is the reachability metric, u is the probe server node, v is the reachability metric node, and d(i, v) is the reachability Boolean value between node i and v.

[0078]

[0079] f(v, t i ) = Tr(v, t i ) + Re(v, t i ) + Fl(v, t i )

[0080] Wherein, A(v, T) is the activity metric, v is the activity metric node, T is the metric time period, N is the number of discrete time points within the time window Δt, t i is the time point, and f(v, t i ) is the activity amount within the time point t i .

[0081] f(v, t i ) is summed by three activity amount representation methods, Tr(v, t i ) is the number of flow bytes passing through node v within the time point t i , Re(v, t i ) is the number of requests accepted by node v within the time point t i , and Fl(v, t i ) is the number of TCP connections of node v within the time point t i .

[0082] The depth is dynamically and intelligently selected according to the strategy, and the scanning range is expanded according to the probe result. The target node can be quickly located at the beginning, and the scanning time and resource consumption are reduced. Different stop conditions can also be set according to the needs, such as finding the target node or reaching the maximum depth, so as to control the scanning process and reasonably allocate the probe resources.

[0083] (Four) Intelligent extraction of IP fingerprint based on machine learning model

[0084] Traditional traffic feature extraction is mainly aimed at the traffic in the Internet, and this part of traffic is characterized by regular access, poor confidentiality and rich samples. For the traffic through the network access point, the access means is complex, and different encryption protocols and implementation methods make the traffic feature diversity extremely high, and the sample acquisition is difficult, so the applicability of traditional identification method is limited. This method follows the traditional IP fingerprint extraction technology, and on this basis, the intelligent dimensionality of IP fingerprint is realized by combining machine learning.

[0085] First, the seed IP is probed by using TCP, ICMP and UDP protocols, and the main fingerprint details are as follows:

[0086] 1) TCP-based IP fingerprint generation and detection:

[0087] a) Flag bits (SYN, ACK, FIN, etc.) and sequence numbers of TCP connections.

[0088] b) Parameters of TCP header such as window size, option field, etc.

[0089] c) Timeout and delay of connections.

[0090] d) Establishment and termination process of TCP connections.

[0091] 2) ICMP-based IP fingerprint generation and detection:

[0092] a) ICMP echo request and response (Ping).

[0093] b) ICMP error messages (ICMP error reports).

[0094] c) ICMP timestamp request and response, etc.

[0095] 3) UDP-based IP fingerprint generation and detection:

[0096] a) Source port and destination port.

[0097] b) Length and content of UDP packets.

[0098] c) Frequency and traffic pattern of packets.

[0099] Secondly, use machine learning techniques to train models to automatically extract features of target IP. These models can be based on statistical properties of IP attributes, protocol identifiers, etc. to classify, so as to achieve efficient IP fingerprint recognition. First, extract meaningful features from network data for feature engineering, so as to effectively distinguish different types of encrypted traffic. Usually, these features include protocol identifiers, packet size, delay characteristics of IP belonging to traffic, frequency distribution, etc. Then go through the links of dataset construction, model selection, training optimization and model evaluation, etc. to realize the construction of learning model for IP fingerprint detection. When applied, the traffic data corresponding to the sample IP needs to be input into the learning model, and then the model learns adaptively according to the software-defined network service traffic and the characteristics of the network access point itself, and finally outputs the IP fingerprint that can effectively represent the traffic through the network access point.

[0100] Finally, combined with traditional detection methods, efficient fingerprint extraction of target seed IP is realized.

[0101] (Five) Construction of primary IP fingerprint library

[0102] The construction of the fingerprint library aims to store and manage a large number of feature fingerprints to facilitate subsequent clustering analysis.

[0103] When building the fingerprint library, the first task is to select appropriate data structures and storage formats to effectively organize and store feature fingerprints. The selection of appropriate data structures should consider storage requirements, query performance and data update requirements. The present application selects a hash table for storage, and the characteristics of the linear table improve the efficiency of retrieval.

[0104] There are many software-defined network service providers, and the corresponding detected IP fingerprints also change frequently, which leads to the continuous change of the data of the fingerprint library. Therefore, an effective data update and maintenance strategy needs to be established, including data cleaning, incremental update, data backup and transaction management, to ensure the accuracy and consistency of the data. In addition, efficient fingerprint library needs to be continuously optimized in performance and query, which can use query caching, database partitioning, index optimization and load balancing, etc. to improve the query performance. Finally, it is essential to ensure the data security of the fingerprint library. Appropriate security measures can be taken, including access control, identity verification, encrypted transmission, etc. to protect the fingerprint library from unauthorized access and attacks.

[0105] (VI) Selecting features that need to be weighted from the IP fingerprints in the primary IP fingerprint library for software-defined network services

[0106] The present application combines the characteristics of the domain network access node itself and the traffic characteristics of the software-defined network, and optimizes the performance of the clustering algorithm by adjusting the weight of the features. The adjustment of the feature weight is crucial for accurate clustering of IP addresses related to the domain network access node of the software-defined network service. Adjusting the weight can highlight the differences between different software-defined network service providers, ensuring that unique features of different software-defined network service providers receive more attention during the clustering process, effectively increasing the inter-class distance, and thus optimizing the generation result of the IP cluster.

[0107] These features include specific data transmission protocols, specific data encryption methods, and specific data flow patterns. At the same time, reducing the weight of features unrelated to the domain network access node or less important to reduce their interference with the clustering results helps to filter noise or irrelevant features and improve the accuracy of clustering.

[0108] Overall, the following optimizations are mainly made to the feature weight:

[0109] 1) Specific ASN (Autonomous System Number): The same software-defined network service domain network access node is usually associated with a specific ASN, which is used to identify the network or service provider to which the node belongs. Therefore, the weight of the source IP belonging to the AS and the source IP belonging to the network segment is increased.

[0110] 2) Network topology: The same software-defined network service domain network access node has a specific location in the network topology. Therefore, increase the weight of the average message routing forwarding times and the average TTL.

[0111] 3) Specific protocol and port: The same software-defined network service domain network access node usually uses specific network protocols and ports to provide services and communicate with users. Therefore, increase the weight of IP open ports and closed ports.

[0112] 4) Traffic pattern: The same software-defined network service domain network access node usually generates a specific traffic pattern, such as a large amount of uplink or downlink traffic at a specific time, or a specific packet frequency. Therefore, increase the weight of TCP window size and packet frequency and traffic pattern.

[0113] Set the same weight for each fingerprint as the initialization of the software-defined network weighted fingerprint.

[0114] In addition, the present application also adopts an adaptive weight adjustment strategy, which dynamically adjusts the weight of the feature according to the distribution of the data and the dynamic changes in the clustering process. The centrality and density of the graph will be recalculated each iteration, which will be used as the basis for weight adjustment, making the method more flexible and better able to adapt to changing network traffic conditions.

[0115] (Seven) Select three clustering algorithms to form a clustering layer according to the complexity progression principle

[0116] IP fingerprint includes: ① Primary IP fingerprint library, mainly including TCP fingerprint, UDP fingerprint, ICMP fingerprint and page banner fingerprint; ② Software-defined network weighted fingerprint, mainly including ASN, TTL, port, etc. The key task of clustering technology is to select different sub-features from these two types of fingerprints for combination to build multi-dimensional feature vectors, which can be used to represent the behavior and characteristics of software-defined network service domain network access nodes.

[0117] To ensure the scientificity of the clustering results, according to the complexity progression principle, three clustering algorithms, K-means, hierarchical clustering and DBSCAN, are used to cluster the feature vectors, and the network service access nodes of the same software-defined network service domain are divided into the same cluster. The three algorithms focus on different aspects of the features of the nodes, k-means considers the numerical features of the nodes, hierarchical clustering considers the relationship structure of the nodes, and DBSCAN considers the density features of the nodes, and finally the overlapping points of the three clustering algorithms are used as the extension line nodes.

[0118] The technical route of the three clustering algorithms is as follows:

[0119] 1) Use K-means algorithm to control the number of final IP clusters

[0120] a) Data Preparation: Collect and prepare the dataset of IP fingerprints, generate a feature vector for each IP based on its fingerprint characteristics, which includes multiple feature dimensions such as TCP fingerprint, UDP fingerprint, ICMP fingerprint, page banner fingerprint, ASN, TTL, and port, etc.

[0121] b) Initialization: Randomly select K IP addresses as initial cluster centers, where K is the predefined number of clusters.

[0122] c) Iterative Clustering: Perform multiple rounds of iteration, each round consisting of two steps:

[0123] i. Assignment Phase: For each IP address, calculate its distance to each cluster center and assign it to the nearest cluster.

[0124] ii. Update Phase: Update the center of each cluster to make the average feature vector of IP addresses in that cluster closest to the new cluster center.

[0125] d) Convergence Judgment: When the cluster centers no longer change significantly or reach the predetermined number of iterations, the algorithm converges.

[0126] 2) Control IP cluster intra-cluster distance using hierarchical clustering algorithm

[0127] a) Data Preparation: Similar to K-means, first prepare the dataset of IP fingerprints, generate a feature vector based on the fingerprint characteristics of each IP.

[0128] b) Hierarchical Structure Construction: Start from the dataset, each IP address is considered as an independent cluster, then build a hierarchical tree structure (tree clustering) by merging the nearest clusters.

[0129] c) Hierarchical Division: Specific cluster division can be made according to the structure of the tree and the threshold of clustering distance.

[0130] 3) Control IP cluster inter-cluster distance using DBSCAN algorithm

[0131] a) Data Preparation: Similarly, prepare the dataset of IP fingerprints, generate a feature vector based on the fingerprint characteristics of each IP.

[0132] b) Core Point Selection: Select core points according to the specified radius ε and minimum neighbor number MinPts. Core points are points that contain at least MinPts IP addresses within the radius ε.

[0133] c) Density accessibility: Clusters are constructed by density accessibility between core points. If a point can be reached from another core point by a series of hop connections, they belong to the same cluster.

[0134] d) Noise points: Points not connected to core points are considered as noise points.

[0135] (Eight) Reverse iterative weighted features, update IP clustering results

[0136] Get the IP clustering result, and map it back to the software-defined network weighted fingerprint, that is, according to the IP associated with each software-defined network service domain in the primary IP clustering result, the weight of each parameter in the software-defined network weighted fingerprint is calculated, and the updated weight of each software-defined network service domain weighted fingerprint is obtained. The IP library, the primary IP fingerprint library and the updated software-defined network weighted fingerprint are input into the clustering layer to obtain the updated IP clustering result.

[0137] For weight adjustment method, after the clustering results of the three graphs, extract the points with the highest density, and the points with the same IP in the three results are selected as candidate service access points. For each feature x selected in step (six) that needs to be weighted, the consistency index of the feature x is calculated based on the feature value of the feature x in each candidate service access point. When the index exceeds a certain threshold, the weight is adjusted accordingly. The consistency index of the feature is the reciprocal of the variance of the field value corresponding to the feature of all nodes.

[0138] This means that the importance of certain features is re-evaluated according to the clustering results. For example, if some IP addresses are correctly clustered into the same class, and they are very similar in some specific features, the weight of these features may be enhanced, because they play a key role in distinguishing IP behavior. Conversely, if some features have no obvious effect in the clustering process, their weight can be appropriately reduced. The specific calculation formula is as follows:

[0139]

[0140] wherein, represents the weight of feature x after updating, represents the weight before updating; α is the learning rate, which is responsible for controlling the adjustment amplitude; C(x) is the contribution degree of feature x, which is obtained by the consistency of the feature in the same IP node, and the consistency index is obtained by the reciprocal of the variance of the field value corresponding to the feature of all nodes; C is the total contribution of all features, which is obtained by summing the consistency index of each feature.

[0141] (Nine) Repeat the breadth-first adaptive IP scan based on different time periods and different nodes, and store the incremental IP into the IP library;

[0142] (X) record dynamic change data, repeat the above process, will continue to dynamically map the same software-defined network service domain network access node.

[0143] Figure 1 The same software-defined network service domain network access node mapping method provided by the present application is shown in the flow chart, the overall process is divided into three parts:

[0144] 1) IP detection: determine the traffic characteristics of the domain network access node, such as a node traffic is very concentrated and high density, has obvious burst peak characteristics, uplink and downlink traffic exists stable asymmetry, the same geographical location of domain network access node traffic pattern similarity is significant and so on, then consider this IP for a target domain network access node. Based on the selected domain network access node traffic characteristics in the network known to exist software-defined network service, through passive network traffic to obtain domain network access node IP. Then breadth-first adaptive IP scanning, passive discovery of target IP as the starting node of the to-be-probed network segment, layer by layer outward expansion, first access to all neighbor nodes of the starting node, and then layer by layer access their neighbors, and so on, until the target node that meets the domain network access node traffic characteristics is found or the entire graph is traversed. In this process, the reachability and activity of the nodes in the network segment can be determined by sending ICMP requests, and the probe depth for layer-by-layer access to the next node can be set accordingly, such as when the network segment node activity is more than 90%, the probe depth is set to 1 (each neighbor node is traversed), and when the network segment activity is 50%, the probe depth is set to 3 (every two nodes are probed, that is, only the third node adjacent to them is probed). Then the probe results can be obtained to form an IP library.

[0145] 2) Fingerprint extraction: Use TCP, ICMP and UDP protocols to probe the target IP and obtain the corresponding fingerprint characteristics, such as TCP connection flag bits (SYN, ACK, FIN, etc.) and sequence numbers, ICMP echo request and response (Ping), UDP packet length and content, etc., to form a preliminary IP fingerprint library. Then use machine learning techniques to train the model to intelligently upgrade the target features. First, the data needs to be preprocessed, including standardization, normalization, etc., to reduce the differences and noise effects between data. The model can choose a multi-layer perception (MLP) neural network for feature upgrading. Unlike general linear models that can only obtain global features, MLP can obtain local features through its deep architecture, increasing the data representation capability. Each layer of neurons receives the output of the previous layer of neurons, and performs weighted summation, then passes it through a nonlinear activation function such as Sigmoid or ReLU function to get the neuron output. The output layer usually gets the probability value of the decision region access node through the Softmax function. When the probability value is 90% or more, the feature vector output by the MLP network hidden layer is the result of feature upgrading in this stage, which can be stored in the preliminary IP fingerprint library. At the same time, in order to avoid overfitting, a regularization term such as L1 or L2 regularization can be added during training, or the dropout technique can be used. In addition, in order to optimize the network weights, the gradient descent method or its variants such as Adam are usually used for weight update.

[0146] 3) Cluster topology: Select features that need to be weighted for software-defined network services, such as ASN features, network topology features, protocol and port features, traffic pattern features, etc., and initialize them. The same weight is selected, all set to 1. Then select three clustering algorithms according to the complexity progression principle to form a clustering layer, such as using K-means algorithm to control the number of final IP clusters, using hierarchical clustering algorithm to control the intra-class distance of IP clusters, and using DBSCAN algorithm to control the inter-class distance of IP clusters. Then generate the fingerprint vector of each IP according to the preliminary IP fingerprint library and input it into the layer to get the preliminary IP clustering result. Get the IP clustering result and map it back to the software-defined network weighted fingerprint to get the updated weight of the weighted fingerprint. Input the updated fingerprint vector into the clustering layer to get the updated IP clustering result.

[0147] Repeat the IP probing process at different times and different nodes, and store the incremental IP in the IP library. Record the dynamically changing data, repeat the fingerprint extraction and clustering topology, and continuously and dynamically map the same software-defined network service domain network access node.

[0148] While specific embodiments of the application have been disclosed in order to illustrate the application and to assist those skilled in the art in practicing the application, it is to be understood that various substitutions, modifications and changes can be made by those skilled in the art without departing from the spirit of the application and the scope of the appended claims. Accordingly, it is intended that the application not be limited, except by the scope of the claims.

Claims

1. A method for extending network service access nodes based on IP fingerprint multi-view clustering, comprising the following steps: 1) Select several traffic characteristics of the domain network access node, and set a corresponding filtering condition for each selected traffic characteristic. The selected traffic characteristics include traffic density characteristics, traffic distribution characteristics, traffic peak characteristics, uplink and downlink traffic asymmetry characteristics, traffic distribution differences between different applications, traffic-geographical location correlation characteristics, and traffic encapsulation and forwarding overhead characteristics. 2) Select nodes from the software-defined network that meet the selection criteria as network service access nodes; The method for selecting nodes that meet the filtering conditions from a software-defined network is as follows: First, passive probing is performed using known traffic data of the software-defined network service to determine whether the traffic characteristics of each node meet the filtering conditions, and the IPs of nodes that meet the filtering conditions are taken as target IPs; then, active probing is performed using a breadth-first search algorithm, starting from each target IP and expanding outwards layer by layer, to determine whether the traffic characteristics of each actively probed node meet the filtering conditions. If the traffic characteristics of a node meet the filtering conditions, the corresponding node is taken as a network service access node. 3) Based on the IPs of each network service access node obtained in step 2), an IP database is formed. 4) Obtain the traffic characteristics corresponding to each target IP as the IP fingerprint of the corresponding target IP to form a primary IP fingerprint database; 5) Upgrade the dimensionality of each IP fingerprint in the primary IP fingerprint database using a machine learning model; 6) Select the features to be weighted from the IP fingerprints in the primary IP fingerprint database and initialize their weights; 7) Select multiple clustering algorithms to form a clustering layer. Use each clustering algorithm in the clustering layer to cluster the IP fingerprint processed in step 6), and divide the network service access nodes of the same software-defined network service domain into the same cluster, generating a primary clustering graph corresponding to each clustering algorithm; where different clustering algorithms focus on different features. 8) Update the weights of corresponding features in the IP fingerprint according to each of the primary clustering graphs, and then use each clustering algorithm in the clustering layer to cluster the updated IP fingerprint to obtain the updated IP clustering results; wherein, the method for updating the weights of corresponding features in the IP fingerprint according to each of the primary clustering graphs is as follows: extract the batch of points with the highest density from each of the primary clustering graphs, and obtain points with the same IP from the extracted points as candidate service access points; for each selected feature x that needs to be weighted, calculate the consistency index of feature x based on the feature value of feature x in each candidate service access point, and when the consistency index of feature x exceeds a set threshold, then... Adjust the weight of feature x; This represents the weight of feature x after the update. The weights of feature x before the update are represented; α is the learning rate, which controls the adjustment range; C(x) is the contribution of feature x. 9) Take the IP class with the highest density in each IP clustering result obtained in step 8) as the network service access node for the corresponding software-defined network service domain, and complete the node expansion.

2. The method according to claim 1, characterized in that, The method for selecting nodes from a software-defined network that meet the aforementioned selection criteria is as follows: 21) Divide the nodes corresponding to each target IP into multiple queues according to their respective software-defined services; and perform IP scanning on the multiple queues simultaneously; 22) Maintain a detection depth parameter n to determine the range of each scan round; execute steps 23) to 25) during each scan round; 23) Select a target node b from the current queue and obtain the nodes with n-hop connections to the target node b based on the detection depth n, and send a set message to detect the reachability and activity of the obtained nodes; 24) If the activity of the detected node is greater than the set threshold, the depth n value corresponding to the target node b is gradually increased in order to detect nodes at more levels and gradually expand the detection range. 25) Stop scanning target node b when the detection depth n corresponding to the target node b reaches the set maximum value or meets the set detection conditions.

3. The method according to claim 2, characterized in that, Based on accessibility metrics Determine the reachability of nodes; where u is the probe server node, v is the reachability measurement node, and d(i,v) is the Boolean value for reachability between node i and node v.

4. The method according to claim 2, characterized in that, Based on activity metrics Determine the activity of a node; where f(v,t) i ) = Tr(v,t i )+Re(v,t i )+Fl(v,t i ), where v is the activity measurement node, T is the measurement time period, N is the number of discrete time points within the time window Δt, and t i It is the i-th time point, f(v,t) i ) is the time point t i The amount of activity within, Tr(v,t) i ) is at time point t i The number of bytes of traffic passing through node v, Re(v,t) i ) is at time point t i The number of requests received by internal node v, Fl(v,t) i ) is at time point t i The number of TCP connections for internal node v.

5. The method according to claim 1, characterized in that, The primary IP fingerprint database uses a hash table to store each IP fingerprint.

6. The method according to claim 1 or 5, characterized in that, The primary IP fingerprint database includes TCP fingerprints, UDP fingerprints, ICMP fingerprints, and page banner fingerprints; the features to be weighted include ASN, TTL, and port.

7. The method according to claim 1, characterized in that, The clustering algorithms in the clustering layer include K-means clustering, hierarchical clustering, and DBSCAN clustering.

Citation Information

Patent Citations

  • Model construction method, malicious code identification method, storage medium and terminal

    CN115983342A

  • Network topology generation method and system

    CN116915620A