Large-scale-oriented IPv6 target address generation method and system
By constructing IPv6 seed address sets and using MDHC clustering and address pattern spatial intersection feedback strategy, the problem of small total number of IPv6 active address predictions and long time is solved, and the effect of efficiently generating more active addresses in IPv6 address space is achieved.
Patent Information
- Application Number
- CN202510545119.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-12
AI Technical Summary
The existing IPv6 target address generation algorithm has the problem of small total prediction amount of IPv6 active addresses and long prediction time. Especially when the IPv6 address space is huge and the distribution is sparse, the detection efficiency of the existing methods is inefficient.
The IPv6 seed address set is constructed by using the IPv6 expansion strategy based on sampling distribution and the IPv6 classification strategy based on address structure. The IPv6 address space tree is constructed through clustering, and the low-dimensional and high-dimensional address pattern collection is generated. The MDHC clustering strategy is used for address mining. Combined with the Bloom filter and the longest prefix matching algorithm to eliminate alias addresses, a feedback strategy for detection results based on the spatial intersection of IPv6 address pattern is proposed.
It significantly improves the prediction number and efficiency of IPv6 active addresses, and can generate more IPv6 active addresses in a short time, improving the workspace and hit rate of the IPv6 target address generation algorithm.
Smart Images

Figure CN120475013A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer networks, and in particular to a method and system for generating large-scale IPv6 target addresses. Background Art
[0002] The next-generation Internet Protocol (IPv6) is entering a period of rapid global development. Compared to IPv4 (Internet Protocol version 4), IPv6 offers a wider address space, built-in security, simplified configuration, and optimized routing efficiency. This significantly improves network scalability and performance, making it more adaptable to the demands of modern network environments.
[0003] Efficient internet-wide (global) IPv6 address scanning is a key technology for IPv6 network asset discovery and cyberspace mapping. Due to the vast size of the IPv6 address space and the sparse distribution of active addresses, detecting live internet-wide IPv6 addresses is extremely difficult, and existing methods are highly inefficient. Using the asynchronous scanning tool ZMap, a full scan of the IPv4 address space can be completed in 45 minutes on a gigabit network. However, scanning the entire IPv6 address space using this traversal address scanning method would take tens of millions of years.
[0004] To address the challenges of rapidly scanning IPv6 addresses posed by the vast IPv6 address space, IPv6 target address generation (also known as IPv6 address prediction) has become a research hotspot. This technology generates highly active IPv6 target addresses (IPv6 predicted address sets) by learning from known IPv6 seed address sets (living or previously living IPv6 addresses). This reduces the IPv6 address detection space and allows for the rapid discovery of active IPv6 addresses.
[0005] One of the core goals of IPv6 target address generation algorithms is to predict and obtain as many active IPv6 addresses as possible in a short period of time. Most existing IPv6 target address generation algorithms suffer from a low number of predicted active IPv6 addresses and a long prediction time. With the exception of AddrMiner, few other IPv6 target address generation algorithms have predicted and collected over 100 million active IPv6 addresses. AddrMiner also suffers from a long prediction time for active IPv6 addresses, with the entire target address generation process taking one month. Summary of the Invention
[0006] Aiming at the problem that most existing IPv6 target address generation algorithms have a small total amount of predicted IPv6 active addresses and a long prediction time, the present invention proposes a large-scale IPv6 target address generation method and system.
[0007] In a first aspect, the present invention provides a large-scale IPv6 target address generation method, comprising:
[0008] Constructing IPv6 seed address sets based on IPv6 expansion strategy based on sampling distribution and IPv6 classification strategy based on address structure;
[0009] Clustering the IPv6 seed address set to construct an IPv6 address space tree, thereby generating a low-dimensional IPv6 address pattern set and a high-dimensional IPv6 address pattern set;
[0010] Generate an IPv6 target address set in a low-dimensional IPv6 address pattern space, and remove the IPv6 alias addresses therein to obtain a non-alias target address set;
[0011] Detect the non-alias target address set in the low-dimensional IPv6 address pattern space, filter out the active high-dimensional IPv6 address pattern from the high-dimensional IPv6 address pattern set based on the detected IPv6 active address, generate an IPv6 target address set in the active high-dimensional IPv6 address pattern space, remove the IPv6 alias addresses therein, obtain the non-alias target address set, and then detect it and collect the corresponding IPv6 active addresses;
[0012] The IPv6 active addresses detected in all seed address sets are merged to form the final IPv6 active address set.
[0013] Furthermore, the IPv6 capacity expansion strategy based on sampling distribution specifically includes: downsampling a preset seed address source to obtain multiple IPv6 seed address sets of the same size.
[0014] Furthermore, the IPv6 classification strategy based on address structure specifically includes: dividing the seed address in the preset seed address source into IPv6 seed address sets of low byte, EUI-64 (64-bit Extended Unique Identifier), port embedding, IPv4 embedding, mode byte and random categories.
[0015] Furthermore, clustering the IPv6 seed address set to construct an IPv6 address space tree, thereby generating a low-dimensional IPv6 address pattern set and a high-dimensional IPv6 address pattern set, specifically including:
[0016] The IPv6 seed address set is clustered using the leftmost variable dimension clustering algorithm, the most complete coverage clustering algorithm, the minimum entropy clustering algorithm and the rightmost variable dimension clustering algorithm respectively to construct four IPv6 address space trees, thereby obtaining a low-dimensional IPv6 address pattern set and a high-dimensional IPv6 address pattern set in the four IPv6 address space trees; wherein, the clustering process of the rightmost variable dimension clustering algorithm includes: when constructing the IPv6 address space tree, splitting from right to left in sequence according to the variable dimension in the pattern until the number of IPv6 seed addresses in the node is lower than a threshold or the pattern dimension of the node is lower than a threshold.
[0017] Furthermore, the method further includes: adding a wildcard function to the deterministic finite automaton DFA, using the modified DFA to filter the patterns in the four constructed IPv6 address space trees, and eliminating repeated sub-pattern spaces.
[0018] Furthermore, an IPv6 target address set is generated in the low-dimensional IPv6 address pattern space, specifically including: when generating a new IPv6 target address, using a Bloom filter to determine whether the address is unique; if it is a unique address, adding it to the IPv6 target address list and updating the Bloom filter; otherwise, ignoring the address.
[0019] Furthermore, removing the IPv6 alias addresses specifically includes: using a longest prefix matching algorithm to remove the alias addresses in the IPv6 target address set.
[0020] In a second aspect, the present invention provides a large-scale IPv6 target address generation device, comprising:
[0021] A seed address set construction module is used to construct an IPv6 seed address set based on an IPv6 expansion strategy based on sampling distribution and an IPv6 classification strategy based on address structure;
[0022] An address pattern mining module is used to cluster the IPv6 seed address set to construct an IPv6 address space tree, thereby generating a low-dimensional IPv6 address pattern set and a high-dimensional IPv6 address pattern set;
[0023] A target address prediction module is used to generate an IPv6 target address set in a low-dimensional IPv6 address pattern space and remove IPv6 alias addresses therein to obtain a non-alias target address set;
[0024] The detection and feedback module is used to detect the non-alias target address set in the low-dimensional IPv6 address pattern space, filter out the active high-dimensional IPv6 address pattern from the high-dimensional IPv6 address pattern set based on the detected IPv6 active address, generate an IPv6 target address set in the active high-dimensional IPv6 address pattern space, remove the IPv6 alias addresses therein, and then detect the non-alias target address set and collect the corresponding IPv6 active addresses;
[0025] The post-processing module is used to merge the IPv6 active addresses detected in all seed address sets to form a final IPv6 active address set.
[0026] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect when executing the program.
[0027] In a fourth aspect, the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method described in the first aspect.
[0028] The beneficial effects of the present invention are:
[0029] This paper proposes an IPv6 seed address expansion strategy based on sampling distribution and an IPv6 seed address classification strategy based on address structure to generate IPv6 seed address sets, fully utilizing the seed address source. 6Massive uses the expansion strategy to construct multiple seed address sets of equal size, a classification strategy to generate multiple sub-seed address sets, and cluster learning across all address sets, expanding the target address generation workspace. The expansion and classification strategies enable the IPv6 target address generation algorithm to detect a greater number of active IPv6 addresses even on machines with hardware limitations.
[0030] This paper proposes Merge Divisive Hierarchical Clustering (MDHC), which can fully exploit the structural information in IPv6 seed addresses. Using the MDHC clustering strategy, the seed addresses are clustered four times to increase the learning frequency of the seed addresses. Four IPv6 address space trees are constructed, and subspace regions of these space trees are probed to collect active IPv6 addresses. Theoretically, it is demonstrated that the MDHC clustering strategy can generate a larger number of geo-dimensional patterns, thereby improving the efficiency of IPv6 target address generation.
[0031] This paper proposes an IPv6 address detection result feedback strategy based on the intersection of IPv6 address pattern spaces, which can improve the scale of active IPv6 addresses predicted by the IPv6 target address generation algorithm. After completing the detection of the low-dimensional IPv6 address pattern space, this strategy enables 6Massive to continue generating IPv6 target addresses in the high-dimensional IPv6 address pattern space without significantly reducing the hit rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 Provided for embodiments of the present invention: (a) Problems with existing algorithms such as 6Gen, 6Tree, 6Scan, 6Hit, 6Forest, and AddrMiner, and (b) a comparison of the present invention;
[0033] Figure 2 A flowchart of a large-scale IPv6 target address generation method provided by an embodiment of the present invention;
[0034] Figure 3 The problem of repeated sub-patterns and intersecting patterns of the MDHC clustering strategy provided by the embodiment of the present invention;
[0035] Figure 4 An example of using a low-dimensional IPv6 address pattern to filter active high-dimensional IPv6 address patterns provided by an embodiment of the present invention;
[0036] Figure 5 The seed address set category provided in the embodiment of the present invention is H 10w ,Prediction performance of different IPv6 target address generation algorithms under low-dimensional IPv6 address mode;
[0037] Figure 6 The number of active addresses and hit rate performance of different IPv6 target address generation algorithms when the feedback policy is not applied according to the embodiment of the present invention;
[0038] Figure 7 The overlap between active IPv6 addresses predicted by different DHC clustering strategies provided in the embodiments of the present invention;
[0039] Figure 8 The number of active IPv6 addresses predicted by different DHC clustering strategies provided in the embodiment of the present invention; where MaxCoverUniq = MaxCover-Common(Left, MaxCover), MinEntropyUniq = MinEntropy-Common(Left, MaxCover, MinEntropy), RightUniq = Right-Common(Left, MaxCover, MinEntropy, Right)
[0040] Figure 9 The overlap between active IPv6 addresses detected by different protocols provided by the embodiment of the present invention;
[0041] Figure 10 A structural diagram of a large-scale IPv6 destination address generation device provided by an embodiment of the present invention;
[0042] Figure 11 This is a structural block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0043] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0044] The present invention specifically relates to a software system for predicting active IPv6 addresses, which can be applied to fields such as cyberspace mapping, network security analysis, and network optimization.
[0045] IPv6 seed addresses: These are known, active IPv6 addresses. These addresses serve as initial data for further generation and detection of more target addresses.
[0046] IPv6 target address: This is an address generated based on the seed address. The target address is generated to discover more potentially active IPv6 addresses to expand the scope of network scanning.
[0047] Based on the different ways of learning seed addresses during the IPv6 address pattern generation phase, this paper divides IPv6 target address generation algorithms into machine learning-based and hierarchical clustering-based IPv6 target address generation algorithms. Table 1 shows the performance of existing IPv6 target address generation algorithms in terms of the number of IPv6 target addresses, the number of active IPv6 addresses, and the total algorithm time.
[0048] Table 1 IPv6 address prediction performance of existing IPv6 target address generation algorithms
[0049]
[0050] Through analysis, the inventors found that the existing IPv6 target address generation algorithm has the following problems: (1) Insufficient utilization of seed address sources. Figure 1 As shown, in Figure 1In (a), algorithms such as 6Scan only use a small area (blue sector area) in the seed address source, which limits the target address generation space at the seed address source usage level. (2) The seed address structure information is not learned sufficiently. Figure 1 In (a), algorithms such as 6Scan learn the seed address set C1 once and construct an IPv6 address space tree. The generated address patterns are limited, which limits the target address generation space at the seed address information utilization level.
[0051] In response to the above problems, Figure 2 As shown, the embodiment of the present invention provides a large-scale IPv6 target address generation method (abbreviated as 6Massive), which includes the following five steps:
[0052] IPv6 seed address set construction phase: The IPv6 seed address set is constructed based on the IPv6 expansion strategy based on sampling distribution and the IPv6 classification strategy based on address structure, thereby making full use of the active IPv6 addresses in the seed address source to generate more IPv6 target addresses.
[0053] Pattern mining stage: clustering the IPv6 seed address set to construct an IPv6 address space tree, thereby generating a low-dimensional IPv6 address pattern set and a high-dimensional IPv6 address pattern set;
[0054] Target address prediction stage: Generate an IPv6 target address set in the low-dimensional IPv6 address pattern space and remove the IPv6 alias addresses to obtain a non-alias target address set;
[0055] Detection and feedback phase: Detect the non-alias target address set in the low-dimensional IPv6 address pattern space, filter out the active high-dimensional IPv6 address pattern from the high-dimensional IPv6 address pattern set based on the detected IPv6 active address, generate an IPv6 target address set in the active high-dimensional IPv6 address pattern space, remove the IPv6 alias addresses, obtain the non-alias target address set, and detect it and collect the corresponding IPv6 active addresses.
[0056] Specifically, 6Massive prioritizes IPv6 address generation in the low-dimensional IPv6 address pattern space, then filters out active high-dimensional IPv6 address patterns based on the active IPv6 addresses in the low-dimensional IPv6 address pattern space, and continues address generation in the active high-dimensional IPv6 address pattern space, thereby increasing the number of active IPv6 addresses predicted by 6Massive.
[0057] Post-processing stage: Merge the IPv6 active addresses detected in all seed address sets to form the final IPv6 active address set.
[0058] This paper proposes an IPv6 seed address expansion strategy based on sampling distribution and an IPv6 seed address classification strategy based on address structure to improve the utilization rate of seed address sources and the degree of seed address mining. By integrating the hierarchical divisive clustering strategy (MDHC), this paper fundamentally expands the IPv6 address prediction space and increases the scale of IPv6 active address prediction. Furthermore, to further expand the scale of IPv6 active address prediction, this paper proposes an IPv6 address detection result feedback strategy based on the intersection of IPv6 address pattern spaces. This strategy can screen out active high-dimensional address patterns without pre-scanning, thereby increasing the number of predicted IPv6 active addresses.
[0059] In one embodiment, a method for constructing an IPv6 seed address set is provided. This embodiment first introduces the categories of IPv6 seed address sets, and then introduces an IPv6 expansion strategy based on sampling distribution and an IPv6 classification strategy based on address structure, thereby solving the problem of insufficient utilization of seed address sources.
[0060] (1) IPv6 seed address set size division
[0061] The downsampling method can make the prefix distribution of the IPv6 addresses in the seed address set as consistent as possible with the seed address source IPv6 Hitlist. Therefore, the present invention uses the downsampling method to construct the seed address set for the seed address source Hitlist, and randomly extracts a certain number of addresses from the seed address source as the IPv6 seed address set.
[0062] According to the division rules of the seed address sets of 6Subpattern, 6Forest, and 6Graph, as shown in Table 2, 6Massive of the present invention sets the sampling scale of the seed address source Hitlist to six types: 100,000, 200,000, 500,000, 1 million, 2 million, and 5 million.
[0063] Table 2 IPv6 seed address set categories and sizes
[0064]
[0065]
[0066] (2) IPv6 expansion strategy based on sampling distribution and IPv6 classification strategy based on address structure
[0067] To address the problem of insufficient utilization of seed address sources in algorithms such as 6Scan, this embodiment proposes a seed address expansion and classification strategy, including an IPv6 expansion strategy based on sampling distribution and an IPv6 classification strategy based on address structure.
[0068] The IPv6 expansion strategy based on sampling distribution constructs multiple IPv6 seed address sets of the same size by downsampling the seed address source, and generates IPv6 target addresses for each seed address set to increase the predicted number of IPv6 active addresses.
[0069] There are unique seed addresses between different seed address sets. Since the addresses in the seed address source are extracted with equal probability to form the IPv6 seed address set, when the size of the seed address set is smaller than the seed address source, there are unique seed addresses between different seed address sets. Let the number of seed address sources be N source , the size of the seed address set to be constructed is N set , then the probability of an address in the seed address source being selected is
[0070] 6Massive can generate its unique IPv6 address pattern based on different seed address sets. Assume that the pattern dimension of all leaf nodes in the IPv6 address space tree is m, then the leaf node is at the (33-m)th layer of the IPv6 address space tree. According to Theorem 1 below, when there are enough seed addresses, the number of address patterns in the IPv6 address space tree is N. pattern =16 (32-m) , the number of seed addresses included is at least 2N pattern seed addresses. The probability that two addresses belong to different address patterns is The pattern dimension of the pattern used to generate IPv6 target addresses is generally less than 10, so the probability that different addresses belong to different address patterns is p ≈ 1. Since there are unique addresses between different seed address sets, 6Massive can generate unique patterns on different seed addresses.
[0071] The IPv6 address structure classification strategy, guided by RFC 7707, divides the seed address set into sub-seed address sets of different categories and combines detection results across all seed address sets to increase the predicted number of active IPv6 addresses. The classification strategy uses the Addr6 tool to categorize seed addresses into sub-seed address sets with low byte, EUI-64, port-embedded, IPv4-embedded, pattern byte, and random categories.
[0072] Let P be the set of active addresses in the entire IPv6 address space that respond to a protocol, and A be the set of all alias addresses in the IPv6 address space. S is the source of seed addresses. The seed address set of category T is C T-i , where i∈{1,n}, that is, a total of n categories of seed address sets are constructed. The output of the algorithm г given a budget of b can be defined as г(C T-i ,b), then after the algorithm г applies the expansion and classification strategies, the number of active IPv6 addresses predicted is N active It can be expressed as:
[0073]
[0074] In one embodiment, this embodiment provides an IPv6 address pattern mining method, called fused divisive hierarchical clustering (MDHC), which can fully mine seed address structure information. This embodiment first designs a new DHC clustering segmentation strategy: the rightmost variable dimension clustering algorithm; then, it fully introduces MDHC; and finally, it introduces the theoretical support, existing problems, and solutions of MDHC.
[0075] (1) Rightmost variable dimension clustering algorithm
[0076] In order to further improve the utilization rate of seed addresses and increase the number of IPv6 target addresses that 6Massive can generate, this paper designs a new DHC clustering segmentation strategy: the rightmost variable dimension clustering algorithm.
[0077] When constructing the IPv6 address space tree, the IPv6 target address generation algorithm using the rightmost variable dimension clustering algorithm splits from right to left according to the variable dimension in the pattern until the number of IPv6 seed addresses in the node is lower than the threshold or the node pattern dimension is lower than the threshold.
[0078] Assume that the node ω in the IPv6 address space tree contains q IPv6 seed addresses, and the possible values of the node ω in dimension X are {x1,…,x k}, with value x in dimension X i The number of seed addresses included is q i . Then the value in dimension X is x i The probability of The entropy value of node ω in dimension X can be expressed as:
[0079]
[0080] If H(X)=0, the dimension X is a fixed dimension, otherwise the dimension X is a variable dimension.
[0081] The dimension of node ω is {X1,…,X 32}, when performing the rightmost variable dimension splitting algorithm, find the first variable dimension from right to left for clustering. Calculate the dimension X in descending order based on the dimension subscripts. j The corresponding entropy value H(X j ). If H(X j )=0, then the dimension is fixed, and the calculation of dimension X is continued. j-1 , until the first one that satisfies H(X j )≠0 for clustering.
[0082] (2) Fusion of hierarchical clustering technology
[0083] Existing IPv6 target address generation algorithms all use a single DHC clustering algorithm to construct an IPv6 address space tree, which results in a limited number of low-dimensional IPv6 address patterns. 6Massive of the present invention uses MDHC to generate multiple IPv6 address space trees to generate more low-dimensional IPv6 address patterns, thereby obtaining more active IPv6 addresses.
[0084] The pseudocode for IPv6 address pattern generation is shown in Algorithm 1. The algorithm inputs are the seed address set C and the maximum pattern dimension D of the low-dimensional IPv6 address pattern. The output is the low-dimensional IPv6 address pattern set LowDimPatterns and the high-dimensional IPv6 address pattern set HighDimPatterns. 6Massive uses the MDHC clustering algorithm to learn the address structure information of the seed address set C. It clusters the seed address set four times, constructs four IPv6 address space trees, and returns the low-dimensional IPv6 address pattern set and high-dimensional IPv6 address pattern set from the four space trees (Line 3).
[0085] MDHC consists of four DHC clustering algorithms: Left, Right, MaxCover, and MinEntropy (line 5). 6Massive sequentially executes the Left, MaxCover, MinEntropy, and MinEntropy algorithms on the seed address set C, generating a greater number of low- and high-dimensional IPv6 address patterns and expanding the workspace of the IPv6 target address generation algorithm (lines 6-8).
[0086]
[0087] 6Massive classifies all generated IPv6 address patterns into low-dimensional IPv6 address patterns and high-dimensional IPv6 address patterns based on a pattern dimension threshold. Algorithms such as HMap6 have demonstrated experimentally that IPv6 addresses in the low-dimensional IPv6 address pattern space are more likely to be active. 6Massive prioritizes IPv6 address generation in the low-dimensional IPv6 address pattern space. It then selects active high-dimensional IPv6 address patterns based on the active IPv6 addresses in the low-dimensional IPv6 address pattern space and continues address generation in the active high-dimensional IPv6 address pattern space, thereby increasing the number of active IPv6 addresses predicted by 6Massive.
[0088] (3) MDHC algorithm theoretical proof
[0089] Theorem 1. When there are sufficient seed addresses, the total number of nodes in the i-th layer of the IPv6 address space tree is 16. i-1 .
[0090] Proof. The process of constructing the IPv6 address space tree by DHC clustering is the process of constructing a complete 16-way tree. Each node splits a variable dimension and specifies [0-f] with a total of 16 values, that is, each node contains 16 adjacent child nodes. The first layer has only the root node, that is, the number of nodes in the first layer can be expressed as 16. 0 The second layer nodes are the adjacent child nodes of the root node, with a total of 16 nodes, that is, the number of second layer nodes can be expressed as 16 1 Similarly, the number of nodes in the i-th layer is 16. i-1 .
[0091] Theorem 2. When the number of IPv6 addresses is unevenly distributed, there are a large number of loss patterns in different IPv6 address space trees.
[0092] Proof. Let the total number of active IPv6 addresses be N0, the root node has m dimensions to be split (concretized), each dimension has 16 possible values (0-f), p d is the probability of the data of the dth dimension at a specific value (for the convenience of calculation, set p d is uniform, that is, p d =1 / 16), according to the variable dimension subscript array {d0,d1,…,d m-1}. For node v, at the kth split (child node of node v), the amount of data of the node is in It can be calculated using the following formula:
[0093]
[0094] If a node is a loss node, the number of IPv6 addresses contained in the node is lower than the threshold, that is, (where T is 2). The number of seed sets in existing IPv6 address prediction algorithms is usually less than 1 million. Let N0 = 10 6 ≈2 20 Then the condition for node v to be a loss node can be expressed by the following formula:
[0095]
[0096] The solution is k≥5.75. This means that if the node v is at a level greater than or equal to 6 in the IPv6 address space tree, the node is likely to be a lost node. When algorithms such as 6Probe and 6Scan construct the IPv6 address space tree, the root node dimension is D root=32, the node pattern of the low-dimensional node is less than or equal to the threshold D1 = 4. The level of the low-dimensional node in the IPv6 address space tree is at least (D root -D1+1)=29≥6. As the formula shows, the number of IPv6 addresses a node contains decreases exponentially as its level in the IPv6 address space tree increases. Therefore, the nodes in the IPv6 address space tree constructed by algorithms like 6Probe and 6Scan are often loss nodes, meaning that the IPv6 address space tree contains a large number of loss patterns.
[0097] Theorem 3. The existence of loss patterns guarantees that different IPv6 address space trees constructed by the MDHC clustering algorithm have unique target addresses.
[0098]
[0099] Proof. The IPv6 seed set is C, and the dimension of the node is £. Then the node has £ types of segmentation strategies. The sub-node generation pattern according to the i-th variable dimension can be expressed as A i[j] (0≤j≤f), that is, all seed addresses in the child node have the value j in the i-th dimension. If the pattern space and the seed set have an intersection, then the pattern is a node pattern or a seed pattern, otherwise it is a loss pattern. If the pattern dimension of the seed pattern is less than or equal to the threshold D1, then the pattern is a low-dimensional address pattern, otherwise it is a high-dimensional address pattern. As shown in the above formula, if the i-th IPv6 address space tree has a loss pattern A i[k] , the jth IPv6 address space tree has node pattern A j[l] , there is a common sub-pattern A between the two i[k]j[l] Detection Mode A j[l] When the address in the space is detected, the loss mode A will be detected at the same time i[k] As long as there is a loss pattern in the IPv6 address space tree, the IPv6 address space tree formed according to different segmentation strategies can have unique IPv6 destination addresses.
[0100] (4) Problems and solutions of MDHC
[0101] like Figure 3 As shown in the figure, while MDHC provides more low-dimensional IPv6 address patterns, it also brings the challenge of repeated sub-pattern space. In order to fully learn the seed address set, 6Massive uses the MDHC clustering strategy to repeatedly learn the IPv6 seed address 4 times, resulting in a single address belonging to multiple patterns, thus causing the problem of repeated sub-pattern space. The phenomenon of repeated sub-patterns will not only waste budget, resulting in duplicate addresses, but also repeatedly detect the same IPv6 target address, thus affecting the network ecology. Figure 3In the example, mode B∈mode A, that is, the IPv6 address space of mode A covers the IPv6 address space of mode B.
[0102] To address the challenge of repeated subpattern spaces, 6Massive uses Deterministic Finite Automaton (DFA) to perform preliminary filtering of patterns in multiple IPv6 address space trees, initially eliminating repeated subpattern spaces. DFA has a high matching speed, and the time complexity of detecting whether an IPv6 address pattern is a subpattern of a known pattern is O(1).
[0103] To adapt to IPv6 patterns and detect repeated sub-patterns, 6Massive modified DFA and added a wildcard function that can filter out patterns containing wildcards.
[0104] In one embodiment, the IPv6 target address prediction stage provided in this embodiment includes three sub-stages: target address set construction, alias address filtering, and active address scanning. The input is the low-dimensional IPv6 address pattern of the IPv6 pattern mining stage and the high-dimensional IPv6 address pattern of the feedback stage, and the output is the IPv6 target address set.
[0105] In the IPv6 pattern mining phase, DFA can only filter out completely matching sub-patterns, and it is difficult to filter out patterns that only partially intersect with the pattern space. Figure 3 In the example, mode C∩mode And mode Mode D, Mode When the DFA detects pattern C, it is difficult for the DFA to detect pattern D. In other words, scenario 1: pattern B can be detected by the DFA (B is a sub-pattern of A); scenario 2: C or D cannot be detected by the DFA (C or D is not a sub-pattern of the other party).
[0106] To address this issue, the 6Massive algorithm uses a Bloom filter to eliminate duplicate IPv6 addresses when generating IPv6 destination addresses. The Bloom filter eliminates duplicate addresses from all pattern spaces at a low cost (in terms of false positive rate and memory usage). Using a uniformly distributed hash function and increasing the number of hash functions can reduce the false positive rate. The 6Massive Bloom filter uses MurmurHash3, with two hash functions.
[0107] When generating a new IPv6 address, 6Massive uses a Bloom filter to determine whether the address is unique. If the address is determined to be unique, it is added to the IPv6 destination address list and the Bloom filter is updated. Otherwise, the process is skipped.
[0108] Specifically, let the IPv6 address be x and the Hash function be h i (x), the index position is Index(h i (x)), the bit set storing element information is BitSet. The condition for 6Massive to determine whether an IPv6 address is an existing address is
[0109]
[0110] Otherwise, 6Massive determines that the address is a new unique address, adds it to the IPv6 target address list, and sets the BitSet response index position to 1.
[0111] When the number of IPv6 addresses reaches 100 million, the Bloom filter has a false positive rate of 1%, meaning that 1 million unique addresses are misidentified as duplicates. The Bloom filter uses 2 bits of memory for each IPv6 address, which is much smaller than the size of the IPv6 address itself, saving the overhead of filtering out duplicate IPv6 addresses.
[0112] Alias address filtering is a critical stage in the IPv6 destination address generation algorithm. All addresses with alias prefixes point to a small number of devices, wasting detection budget. Therefore, alias addresses must be eliminated. 6Massive uses a longest prefix match algorithm to eliminate alias addresses from the IPv6 destination address set. 6Massive first reads the alias prefix list published by Gasser et al., then iterates through the IPv6 destination address set. If a destination address belongs to an alias prefix, it is removed from the IPv6 destination address set. Finally, all non-alias addresses are merged into the IPv6 non-alias destination address set.
[0113] In one embodiment, this embodiment provides a feedback strategy based on pattern space fusion. After exploring all low-dimensional IPv6 address pattern spaces, to obtain more active IPv6 addresses, it is necessary to explore the high-dimensional IPv6 address pattern space. The IPv6 address space of a pattern grows exponentially with increasing pattern dimensionality, while the number of IPv6 seed addresses contained in high-dimensional IPv6 address patterns is similar to that of low-dimensional IPv6 address patterns. Based on prior knowledge, IPv6 target addresses generated by high-dimensional IPv6 address patterns tend to be less active. To address the low activity of high-dimensional IPv6 address patterns, algorithms such as 6Forest use pre-scan feedback to screen for highly active high-dimensional IPv6 address patterns. However, allocating a reasonable budget for pre-scanning is difficult. A small pre-scan budget makes it difficult to screen for relatively active high-dimensional IPv6 address patterns, while a large pre-scan budget wastes the detection budget. Therefore, 6Massive proposes a feedback strategy based on pattern space fusion, which can screen for relatively active high-dimensional IPv6 address patterns without pre-scanning. 6Massive divides the patterns in the IPv6 address space tree into low-dimensional IPv6 address patterns and high-dimensional IPv6 address patterns based on the pattern dimension. Based on the uniqueness of the MDHC strategy, it uses low-dimensional IPv6 address patterns to filter out active high-dimensional IPv6 address patterns.
[0114] The core of the feedback strategy is that there is an intersection between the low-dimensional IPv6 address pattern space and the high-dimensional IPv6 address pattern space, and this is used to screen out active high-dimensional IPv6 address patterns. 6Massive uses the MDHC strategy to cluster the seed address set four times to construct four IPv6 address space trees, and each seed address exists in four IPv6 address space trees. There may be intersections between patterns belonging to different IPv6 address space trees, including seed addresses and some newly detected IPv6 active addresses in the low-dimensional pattern space. Figure 4 As shown, the low-dimensional IPv6 address patterns E and F intersect with the high-dimensional IPv6 address pattern G, including seed addresses (solid line addresses) and newly detected IPv6 active addresses (dashed line addresses).
[0115] 6Massive determines whether a high-dimensional IPv6 address pattern is active based on the number of active addresses in the low-dimensional IPv6 address pattern contained in the high-dimensional IPv6 address pattern. Based on prior knowledge of seed density, i.e., patterns with greater active address density are more likely to be active, larger pattern dimensions increase the screening threshold. Experiments have shown that when 6Massive sets the thresholds for high-dimensional IPv6 address patterns with pattern dimensions of 5 and 6 to 3000 and 1000, respectively, the address generation hit rate is close to that of low-dimensional IPv6 address patterns. Therefore, 6Massive sets the thresholds for high-dimensional IPv6 address patterns with pattern dimensions of 5 and 6 to 3000 and 1000, respectively.
[0116] To fully utilize seed address sources, we propose an IPv6 seed address expansion strategy based on sampling distribution and an IPv6 seed address classification strategy based on address structure to generate IPv6 seed address sets. 6Massive uses the expansion strategy to construct multiple seed address sets of equal size, a classification strategy to generate multiple sub-seed address sets, and cluster learning across all address sets, expanding the target address generation workspace. The expansion and classification strategies enable the IPv6 target address generation algorithm to detect a greater number of active IPv6 addresses even on machines with hardware limitations.
[0117] To fully exploit the structural information in IPv6 seed addresses, a fused split hierarchical clustering (MDHC) strategy was proposed. The seed addresses were clustered four times using the MDHC clustering strategy to increase the learning frequency of the seed addresses. Four IPv6 address space trees were constructed, and active IPv6 addresses were collected by probing subspaces within these trees. Theoretically, it was demonstrated that the MDHC clustering strategy can generate a larger number of geo-dimensional patterns, thereby improving the efficiency of IPv6 target address generation.
[0118] Furthermore, to increase the number of active IPv6 addresses predicted by the IPv6 target address generation algorithm, a feedback strategy for IPv6 address detection results based on the intersection of IPv6 address pattern spaces is proposed. After completing the detection of the low-dimensional IPv6 address pattern space, 6Massive can continue generating IPv6 target addresses in the high-dimensional IPv6 address pattern space without significantly reducing the hit rate.
[0119] To evaluate the performance of 6Probe in predicting active IPv6 addresses, we collected live IPv6 addresses from public data sources to construct a data source. We conducted experiments in a real IPv6 network environment and compared it with the current five mainstream IPv6 network topology detection technologies: 6Gen, 6Tree, 6Hit, HMap6, and 6Scan in terms of hit rate, number of predicted active IPv6 addresses, and efficiency of generating active IPv6 addresses.
[0120] In all the charts of this invention, the scales of K (thousand), M (million), and B (billion) are 10 3 , 10 6 , 10 9 .
[0121] Gasser et al. obtained an IPv6 hitlist containing server, router, and client addresses from multiple public data sources, including domain lists and Bitnodes, as well as traceroute measurements. The present invention used data published on July 20, 2024, as a seed address source, containing approximately 21.49M active IPv6 addresses worldwide.
[0122] (a) Problem statement and measurement metrics
[0123] The core goal of the IPv6 target address generation algorithm is to predict as many active IPv6 addresses as possible in a short time. This section explains the goal of the IPv6 target address generation algorithm. Let P be the set of active addresses that respond to a certain protocol in the entire IPv6 address space, and A be the set of all alias addresses in the IPv6 address space. S is the seed address source, and C is the seed address set. Where C∈S∈P. Given a budget of b, the output of the algorithm г can be defined as г(C,b). Then the number of active IPv6 addresses predicted by the algorithm г is N. active It can be expressed as:
[0124] N active =(г(C,b)-A)∩P
[0125] in The hit rate of algorithm г can be expressed as:
[0126]
[0127] Assume that the running time of algorithm г and the detection time of active addresses are t gen and t probe , then the time cost of algorithm г can be expressed as:
[0128] t=t gen +t probe
[0129] Then the efficiency of the algorithm in generating active addresses is E active Expressed as:
[0130]
[0131] The core objectives of the IPv6 destination address generation algorithm can be interpreted as three elements:
[0132] (1) High hit rate, that is, maximizing the hitrate value. A larger hitrate value means more active IPv6 addresses are predicted under a fixed budget.
[0133] (2) The predicted number of active IPv6 addresses is large, i.e., N active As many as possible. N active It means the ability of the IPv6 destination address generation algorithm to predict the total size of active IPv6 addresses based on the seed address set or seed address source.
[0134] (3) High efficiency of active address generation, i.e. maximizing E active Whether an IPv6 active address survives at a certain moment is related to time, so the time cost t of the algorithm г should be as short as possible, so that the active address generation efficiency E active As high as possible.
[0135] 6Massive's goal is to maximize hitrate, N active and E active The value of improves the algorithm's ability to generate as many active IPv6 addresses as possible under a fixed budget, to generate as many active IPv6 addresses as possible under a seed address set or seed address source, and to generate as many active IPv6 addresses as possible in a short period of time.
[0136] (b) Hit rate and number of active addresses of different IPv6 target address generation algorithms
[0137] This section verifies the superiority of 6Massive technology by constructing two scenarios by controlling two parameters: the IPv6 address detection mode dimension and the IPv6 target address generation budget.
[0138] (1) Low-dimensional IPv6 address model scenario
[0139] Due to their incorporation of dynamic feedback strategies, algorithms like 6Scan do not clearly separate low-dimensional and high-dimensional IPv6 address patterns. The HMap algorithm supports a dimensionality threshold, limiting IPv6 target address generation to below a set dimension. Due to its integration of both AHC and DHC clustering, HMap constructs more low-dimensional IPv6 address patterns than algorithms like 6Scan, which use only DHC clustering. Therefore, for experimental comparisons, the HMap budget was applied to algorithms like 6Scan.
[0140] like Figure 5 As shown in the figure, the five IPv6 target address generation algorithms such as 6Massive and 6Scan are used in the IPv6 seed address set category H 10w, comparison scenarios in low-dimensional IPv6 address mode. HMap6 has high memory requirements. Under high budget, the memory required is higher than the server memory. Therefore, in low-dimensional scenarios, only HMap6 10w Compare the performance of different IPv6 destination address generation algorithms under different seed address sets.
[0141] Using the same seed address set, 6Massive generates a greater number of IPv6 target addresses, with the resulting low-dimensional IPv6 address pattern space being 1.63 times larger than that of HMap6. With a seed address scale of 100,000, HMap6 generates approximately 260 million IPv6 target addresses in the low-dimensional IPv6 address pattern space, while 6Massive generates approximately 420 million IPv6 target addresses. Using the same seed address set, 6Massive, using the MDHC clustering algorithm, generates a larger low-dimensional IPv6 address pattern space than HMap6, which uses both the DHC and AHC clustering algorithms.
[0142] 6Massive detects a large number of active IPv6 addresses in the low-dimensional IPv6 address pattern space while maintaining a high hit rate. In a low-dimensional IPv6 address pattern space constructed from a set of 100,000 seed addresses, 6Massive detected and collected 152 million active IPv6 addresses, representing a 268.27% to 516.83% increase compared to algorithms like HMap6. When targeting hundreds of millions of active addresses, 6Massive achieved a hit rate of 36.24%, an improvement of 86.25% to 211.41% compared to algorithms like HMap6.
[0143] The feedback strategy can alleviate the problem of the IPv6 target address generation algorithm constructing a small low-dimensional IPv6 address pattern space and a limited number of generated low-dimensional IPv6 address patterns. Both 6Gen and 6Scan generate fewer low-dimensional IPv6 address patterns than HMap6, while generating similar or even more active IPv6 addresses than HMap6 with the same budget. This is primarily due to the feedback strategy incorporated into the 6Gen and 6Scan algorithms, which dynamically guides the IPv6 target address generation algorithm's subsequent detection direction based on already detected active IPv6 addresses. Therefore, 6Massive technology also incorporates a feedback strategy, enabling it to obtain a larger number of active IPv6 addresses.
[0144] (2) Equivalent IPv6 target address prediction scenario
[0145] This section compares the performance of 6Massive and 6Scan, among other IPv6 destination address generation algorithms, before and after feedback, for scenarios with equal IPv6 destination address budgets. The HMap algorithm does not support a custom budget for IPv6 destination address generation. Therefore, in scenarios with equal budgets, only 6Massive is compared with 6Scan, 6Tree, 6Gen, and 6Hit.
[0146] 6Massive only in H 10w 、H 20w The two categories were compared with 6Scan and other algorithms in the feedback strategy scenario. This is because as the seed size increases, the number of target addresses generated by the algorithm also increases. 20w Under the seed address set of the category, the number of target addresses generated by 6Massive is 2.779 billion, which is higher than the maximum target address generation number of 2 billion by algorithms such as 6Scan.
[0147] Under the same budget, 6Massive can obtain more active IPv6 addresses than algorithms such as 6Scan, and the effect becomes more obvious as the budget increases.
[0148] like Figure 6 As shown, no feedback strategy is applied. 10w Under the seed address set of the category, the number of active IPv6 addresses detected by 6Massive is 2.58 to 3.3 times that of algorithms such as 6Scan. 100w Under the seed address set of the category, this effect is 3.45 to 6.97 times.
[0149] As shown in Table 3, after applying the feedback strategy, 10w Under the seed address set of the category, the number of active IPv6 addresses detected by 6Massive is 3.9 to 5.84 times that of algorithms such as 6Scan. 20w Under the seed address set of the category, this effect is 3.61 to 7.17 times.
[0150] The feedback strategy can increase the number of active IPv6 addresses detected by 6Massive. 10w Category to H 20w For the seed address set of this category, the number of active IPv6 addresses detected by 6Massive increased by 198.64% to 375.16% compared with the low-dimensional IPv6 address pattern budget.
[0151] Table 3 Comparison of target addresses, active addresses, and hit rates of different IPv6 target address generation algorithms after applying the feedback strategy
[0152]
[0153] (c) Comparison of active address generation efficiency of different IPv6 target address generation algorithms
[0154] This section tests the third indicator of different IPv6 target address generation algorithms: IPv6 active address generation efficiency. Table 4 shows the H 100w Comparison of the algorithm time overhead and IPv6 active address generation efficiency of different IPv6 address generation algorithms in three scenarios on the category dataset.
[0155] 6Massive's active address generation efficiency surpasses that of algorithms like 6Scan in various scenarios. In low-dimensional IPv6 address patterns, the same target address budget scenario (without feedback), and the same target address budget scenario (with feedback), 6Massive achieved IPv6 active address generation efficiencies of 14,188.0318 addresses / second, 14,188.0318 addresses / second, and 9,852.79 addresses / second, respectively. These are 1.87 to 3.72 times, 3.34 to 5.08 times, and 3.41 to 5.19 times higher than algorithms like 6Scan, respectively.
[0156] As the target address budget increases, the efficiency of IPv6 active address generation by algorithms like 6Massive gradually decreases. After exploring all low-dimensional address pattern spaces with high seed address density, 6Scan uses a feedback strategy to filter out active high-dimensional address patterns (pattern spaces containing a large number of active IPv6 addresses already explored by the algorithm) to obtain a larger number of active IPv6 addresses while maintaining high IPv6 active address generation efficiency. It then continues to generate IPv6 target addresses within these IPv6 address pattern spaces. However, the number of these active high-dimensional address patterns is limited. After exploring all high-dimensional address pattern spaces, the efficiency of IPv6 active address generation by algorithms like 6Scan decreases rapidly.
[0157] Because 6Massive fully mines the IPv6 seed address structure through the MDHC clustering strategy and generates a larger number of IPv6 address patterns, in the three scenarios, 6Massive can predict a larger number of active IPv6 addresses and its IPv6 active address generation efficiency is much higher than that of algorithms such as 6Scan.
[0158] Table 4 Comparison of algorithm time overhead and IPv6 active address generation efficiency of different IPv6 address generation algorithms in three scenarios
[0159]
[0160] (d) Total number of active IPv6 addresses
[0161] This paper verifies the versatility of the 6Probe algorithm by constructing three scenarios, controlling the quality of the IPv6 seed set, the dimensionality of the IPv6 address detection pattern, and the IPv6 target address generation budget. Table 4 shows the specific test data for the 6Probe and 6Scan algorithms in different scenarios.
[0162] The goal of the IPv6 target address generation algorithm is to discover as many active IPv6 addresses as possible in a short period of time. Therefore, it is necessary to test 6Massive's ability to obtain the total number of active IPv6 addresses.
[0163] The total number of active IPv6 addresses consists of three parts: low-dimensional active addresses, feedback active addresses, and classified and expanded active addresses. IPv6 patterns are categorized into low-dimensional IPv6 address patterns and high-dimensional IPv6 address patterns. 6Massive first generates IPv6 target addresses within the low-dimensional IPv6 address pattern space to obtain low-dimensional active addresses.
[0164] 6Massive then uses feedback to obtain a larger number of active IPv6 addresses from a single seed address set. 6Massive uses the generated IPv6 target addresses and detected IPv6 active addresses to filter out active high-dimensional IPv6 address patterns, and then generates IPv6 target addresses within the active high-dimensional IPv6 address pattern space to obtain feedback active addresses.
[0165] To obtain a larger number of active IPv6 addresses, it is necessary to learn more active IPv6 addresses from IPv6 seed address sources. This paper generates multiple IPv6 seed address sets through classification and expansion, expanding the prediction range of the IPv6 target address generation algorithm. 6Massive combines the active IPv6 addresses detected across all seed address sets to form the final active IPv6 address set.
[0166] In the tested optimal seed address set category H 100w The details of each category of IPv6 active addresses detected by 6Massive are shown in Table 5. 100w Based on the seed address set, 6Massive detected 888 million active IPv6 addresses in the low-dimensional IPv6 address pattern space. Through feedback and expansion strategies, 6Massive detected a total of 1.644 billion active IPv6 addresses within 1.76 days (42.25 hours), generating an active address efficiency of 10,813.31 addresses per second.
[0167] Table 5 Details of the three types of 6Massive addresses under different data sets
[0168] Scene Category Destination Address Active addresses Hit rate Time cost Low Dimension 2.3B 887.96M 38.6% 16:41:8 feedback 1.67B 615.26M 36.81% 7:26:2 Capacity expansion 1.73B 583.1M 33.68% 18:7:57 Total 4.22B 1.64B 39.01% 42:15:8
[0169] (e) Correlation between pattern dimensions and hit rate
[0170] This section first uses experiments to illustrate the relationship between address pattern dimensions and hit rate. Then, it experimentally compares the four DHC clustering strategies included in the MDHC strategy. It then tests the optimal sampling size of the hitlist. Finally, it tests the correlation between the three protocols, ICMPv6, TCP, and UDP, and active addresses.
[0171] (1) Relationship between pattern dimension and hit rate: In the dataset H 10w The correlation between the pattern dimension and the hit rate of the IPv6 destination address generation algorithm was tested, as shown in Table 6.
[0172] Table 6H 10K Relationship between different dimensional patterns and hit rate
[0173]
[0174]
[0175] As the pattern dimension increases, the average pattern hit rate generally shows a downward trend. When the pattern dimension is less than or equal to 4, the average pattern hit rate is higher than 30%. When the pattern dimension increases to 5 and 6, the average pattern hit rate drops rapidly to 4.53% and 2.22% respectively. Therefore, in this experiment, the maximum dimension of the low-dimensional IPv6 address pattern is set to 4.
[0176] Target addresses generated by patterns with higher seed density are more likely to be active. As the pattern dimension increases, the size of the pattern space grows exponentially, and thus the cost of detecting high-dimensional IPv6 address patterns also grows exponentially.
[0177] (2) Performance comparison of the four DHC clustering strategies included in the MDHC strategy: Figure 7 and Figure 8 It is 6Massive in the dataset H 10w Above, the IPv6 active address prediction performance of different DHC clustering strategies in MDHC.
[0178] Figure 7 This is a heat map showing the percentage of common active IPv6 addresses relative to the total number of active addresses for different DHC clustering strategies. The Left and MaxCover clustering strategies are similar. Both select variable dimensions for splitting from left to right, resulting in a higher ratio of common addresses. The Right clustering strategy has a lower ratio of common addresses than the other DHC clustering strategies. This is because the Right strategy selects the first variable dimension for splitting from right to left each time, resulting in a lower similarity with the other strategies.
[0179] Figure 8 This is the distribution of 152.28M active IPv6 addresses predicted by 6Massive in the low-dimensional pattern space. Based on the three DHC clustering strategies (Left, MaxCover, and MinEntropy), the Right strategy predicted 46M unique active IPv6 addresses, accounting for 30.2% of the low-dimensional address pattern space. This demonstrates that the Right strategy constructs a larger number of IPv6 address patterns than Left, MaxCover, and MinEntropy, and thus predicts a greater number of active IPv6 addresses.
[0180] (3) Optimal sampling size of seed address sets: The correlation between the size of seed address sets in the low-dimensional address pattern space and the predicted number of active IPv6 addresses was tested on six categories of seed address sets, as shown in Table 7.
[0181] Table 7 Prediction performance in 6Massive low-dimensional address pattern space on different seed address sets
[0182] Seed address set category Number of IPv6 destination addresses Number of active IPv6 addresses Hit rate <![CDATA[H 10w ]]> 420M 152.28M 36.24% <![CDATA[H 20w ]]> 800M 310.22M 39.13% <![CDATA[H 50w ]]> 1.6B 638.31M 39.78% <![CDATA[H 100w ]]> 2.3B 887.96M 38.6% <![CDATA[H 200w ]]> 2.86B 826.15M 28.86% <![CDATA[H 500w ]]> 3.18B 719.45M 22.62%
[0183] As the size of the seed address set increases, the number of IPv6 target addresses generated by 6Massive in the low-dimensional pattern space also increases. 500w ), the number of IPv6 target addresses generated by 6Massie is 3.18 billion, which is 7.57 times that generated by the 100,000-scale seed address set.
[0184] However, the predicted number of active IPv6 addresses does not increase with the increase in the size of seed addresses. 100w ), 6Massive predicted the largest number of IPv6 active addresses, which is 888 million. Therefore, when testing the total predicted scale of IPv6 active addresses of 6Massive, the seed address set category used in the expansion strategy is H 100w .
[0185] (4) Correlation between protocols and active addresses: Table 8 shows the detection results of different protocols on the 420M IPv6 target addresses generated by 6Massive. The number of IPv6 active addresses obtained through the ICMPv6 protocol is much higher than that obtained through the TCP and UDP protocols. Figure 9 This chart compares the overlap between active IPv6 addresses detected by different protocols. ICMPv6 can detect over 50% of the active addresses detected by other protocols. Therefore, 6Massive uses ICMPv6 to detect IPv6 destination addresses to obtain active IPv6 addresses.
[0186] Table 8 Effects of different protocols on 6Massive 10w Probe results for IPv6 target addresses generated from the seed address set
[0187] protocol Active IPv6 addresses Hit rate protocol Active IPv6 addresses Hit rate ICMPv6 152.28M 36.24% UDP443 79.11K 0.02% TCP443 874K 0.21% UDP53 134.42K 0.03% TCP53 309.94K 0.07% UDP80 124 0% TCP80 906.01K 0.22% - - -
[0188] Due to insufficient utilization of seed address sources and insufficient mining of seed address structure information, existing IPv6 target address generation algorithms find it difficult to achieve the goal of predicting and obtaining 1 billion active IPv6 addresses in a short period of time.
[0189] To this end, 6Massive optimizes the seed address set construction, address pattern mining, and feedback phases of the IPv6 target address generation process to predict as many active IPv6 addresses as possible. 6Massive predicted 1.644 billion active IPv6 addresses in 1.76 days, generating an active address efficiency of 10,813.31 addresses per second.
[0190] This paper proposes a fusion split clustering technique: MDHC. 10w In the low-dimensional space of the category seed address set, the number of IPv6 target addresses predicted by MDHC is 1.63 times that of HMap6 (the IPv6 target address generation algorithm with the largest low-dimensional IPv6 address pattern space), and the number of predicted IPv6 active addresses is 2.57 to 3.28 times that of algorithms such as 6Scan.
[0191] This paper proposes two strategies to generate IPv6 seed address sets: an IPv6 expansion strategy based on sampling distribution and an IPv6 classification strategy based on address structure. 10w In the seed address concentration of the category, 6Massive newly detected 583 million active IPv6 addresses through the two strategies of expansion and classification, which is 65.65% of the low-dimensional IPv6 active addresses.
[0192] Based on the same inventive concept, Figure 10 As shown, an embodiment of the present invention provides a large-scale IPv6 target address generation device, including: a seed address set construction module, an address pattern mining module, a target address prediction module, a detection and feedback module and a post-processing module.
[0193] Among them, the seed address set construction module is used to construct the IPv6 seed address set based on the IPv6 expansion strategy of sampling distribution and the IPv6 classification strategy based on address structure; the address pattern mining module is used to cluster the IPv6 seed address set to construct the IPv6 address space tree, thereby generating a low-dimensional IPv6 address pattern set and a high-dimensional IPv6 address pattern set; the target address prediction module is used to generate the IPv6 target address set in the low-dimensional IPv6 address pattern space, and eliminate the IPv6 alias addresses therein to obtain the non-alias target address set; the detection and response module is used to detect and reverse the IPv6 address set. The feedback module is used to detect the non-alias target address set in the low-dimensional IPv6 address pattern space, filter out active high-dimensional IPv6 address patterns from the high-dimensional IPv6 address pattern set based on the detected IPv6 active addresses, generate an IPv6 target address set in the active high-dimensional IPv6 address pattern space, eliminate the IPv6 alias addresses therein, obtain the non-alias target address set, and then detect it and collect the corresponding IPv6 active addresses; the post-processing module is used to merge the IPv6 active addresses detected in all seed address sets to form the final IPv6 active address set.
[0194] It should be noted that the large-scale IPv6 target address generation device provided in the embodiment of the present invention is for implementing the above method. Its specific functions can be referred to the above method embodiment and will not be described again here.
[0195] Figure 11 An example of a physical structure diagram of an electronic device is shown below. Figure 11As shown, the electronic device may include: a processor (processor) 1101, a communication interface (Communications Interface) 1102, a memory (memory) 1103 and a communication bus 1104, wherein the processor 1101, the communication interface 1102, and the memory 1103 communicate with each other through the communication bus 1104. The processor 1101 can call the logic instructions in the memory 1103 to execute a large-scale IPv6 target address generation method, which includes: constructing an IPv6 seed address set based on an IPv6 expansion strategy based on sampling distribution and an IPv6 classification strategy based on address structure; clustering the IPv6 seed address set to construct an IPv6 address space tree, thereby generating a low-dimensional IPv6 address pattern set and a high-dimensional IPv6 address pattern set; generating an IPv6 target address set in the low-dimensional IPv6 address pattern space, and eliminating the IPv6 alias addresses therein to obtain a non-alias target address set; detecting the non-alias target address set in the high-dimensional IPv6 address pattern space and collecting the corresponding IPv6 active addresses. If the number of the corresponding IPv6 active addresses exceeds a threshold, generating an IPv6 target address set in the high-dimensional IPv6 address pattern space; merging the IPv6 active addresses detected in all seed address sets to form a final IPv6 active address set.
[0196] In addition, when the logic instructions in the above-mentioned memory 1103 are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0197] An embodiment of the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the large-scale IPv6 target address generation method provided by the above-mentioned method embodiments.
[0198] An embodiment of the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program implements the large-scale IPv6 target address generation method provided by the above-mentioned method embodiments.
[0199] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0200] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A large-scale IPv6 target address generation method, characterized in that: include: Constructing IPv6 seed address sets based on IPv6 expansion strategy based on sampling distribution and IPv6 classification strategy based on address structure; Clustering the IPv6 seed address set to construct an IPv6 address space tree, thereby generating a low-dimensional IPv6 address pattern set and a high-dimensional IPv6 address pattern set; Generate an IPv6 target address set in a low-dimensional IPv6 address pattern space, and remove the IPv6 alias addresses therein to obtain a non-alias target address set; Detect the non-alias target address set in the low-dimensional IPv6 address pattern space, filter out the active high-dimensional IPv6 address pattern from the high-dimensional IPv6 address pattern set based on the detected IPv6 active address, generate an IPv6 target address set in the active high-dimensional IPv6 address pattern space, remove the IPv6 alias addresses therein, obtain the non-alias target address set, and then detect it and collect the corresponding IPv6 active addresses; The IPv6 active addresses detected in all seed address sets are merged to form the final IPv6 active address set.
2. The method for generating a large-scale IPv6 target address according to claim 1, wherein: The IPv6 capacity expansion strategy based on sampling distribution specifically includes: downsampling a preset seed address source to obtain multiple IPv6 seed address sets of the same size.
3. The method for generating a large-scale IPv6 target address according to claim 1, wherein: The IPv6 classification strategy based on address structure specifically includes: dividing the seed addresses in the preset seed address source into IPv6 seed address sets of low byte, EUI-64, port embedding, IPv4 embedding, mode byte and random categories.
4. The method for generating a large-scale IPv6 target address according to claim 1, wherein: Clustering the IPv6 seed address set to construct an IPv6 address space tree, thereby generating a low-dimensional IPv6 address pattern set and a high-dimensional IPv6 address pattern set, specifically including: The IPv6 seed address set is clustered using the leftmost variable dimension clustering algorithm, the most complete coverage clustering algorithm, the minimum entropy clustering algorithm and the rightmost variable dimension clustering algorithm respectively to construct four IPv6 address space trees, thereby obtaining a low-dimensional IPv6 address pattern set and a high-dimensional IPv6 address pattern set in the four IPv6 address space trees; wherein, the clustering process of the rightmost variable dimension clustering algorithm includes: when constructing the IPv6 address space tree, splitting from right to left in sequence according to the variable dimension in the pattern until the number of IPv6 seed addresses in the node is lower than a threshold or the pattern dimension of the node is lower than a threshold.
5. The method for generating a large-scale IPv6 target address according to claim 4, wherein: Also includes: A wildcard function is added to the deterministic finite automaton (DFA). The modified DFA is used to filter the patterns in the four constructed IPv6 address space trees to eliminate repeated sub-pattern spaces.
6. The method for generating a large-scale IPv6 target address according to claim 1, wherein: An IPv6 target address set is generated in a low-dimensional IPv6 address pattern space, specifically including: when generating a new IPv6 target address, using a Bloom filter to determine whether the address is unique; if it is a unique address, adding it to the IPv6 target address list and updating the Bloom filter; otherwise, ignoring the address.
7. The method for generating a large-scale IPv6 target address according to claim 1, wherein: Eliminating the IPv6 alias addresses, specifically including: using a longest prefix matching algorithm to eliminate the alias addresses in the IPv6 target address set.
8. A large-scale IPv6 destination address generation device, characterized in that: include: A seed address set construction module is used to construct an IPv6 seed address set based on an IPv6 expansion strategy based on sampling distribution and an IPv6 classification strategy based on address structure; An address pattern mining module is used to cluster the IPv6 seed address set to construct an IPv6 address space tree, thereby generating a low-dimensional IPv6 address pattern set and a high-dimensional IPv6 address pattern set; A target address prediction module is used to generate an IPv6 target address set in a low-dimensional IPv6 address pattern space and remove IPv6 alias addresses therein to obtain a non-alias target address set; The detection and feedback module is used to detect the non-alias target address set in the low-dimensional IPv6 address pattern space, filter out the active high-dimensional IPv6 address pattern from the high-dimensional IPv6 address pattern set based on the detected IPv6 active address, generate an IPv6 target address set in the active high-dimensional IPv6 address pattern space, remove the IPv6 alias addresses therein, and then detect the non-alias target address set and collect the corresponding IPv6 active addresses; The post-processing module is used to merge the IPv6 active addresses detected in all seed address sets to form a final IPv6 active address set.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Efficient IPv6 address detection method based on neural network model
CN120675973A
IPv6 asset efficient discovery method and system with unified address and port mapping
CN121567675A