Network address prediction method and device based on prefix clustering

By using a network address prediction method based on prefix clustering, the problem of blind spots in the detection of IPv6 address prediction algorithms under non-seed prefix conditions is solved, achieving efficient and accurate IPv6 address detection and improving network security.

CN120915764APending Publication Date: 2025-11-07Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511116499.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing IPv6 address prediction algorithms have blind spots in non-seed prefix detection, leading to network security vulnerabilities. Current technologies are unable to efficiently scan for surviving IPv6 addresses in non-seed prefix environments.

Method used

A network address prediction method based on prefix clustering is adopted. A seed address pattern list is generated by split hierarchical clustering and then migrated to a non-seed prefix list. The address pattern migration is combined with Hamming distance and Whois information similarity to construct a two-level decision architecture across address prefixes and patterns, and the detection budget is dynamically allocated to improve the hit rate.

Benefits of technology

It improves the accuracy of IPv6 address mode migration and the utilization rate of detection resources, increases the IPv6 address hit rate, reduces detection blind spots, and enhances network security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120915764A_ABST
    Figure CN120915764A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a network address prediction method and device based on prefix clustering. The specific embodiment of the method comprises the following steps: migrating address modes in a seed address mode list to a non-seed prefix list according to a seed prefix list corresponding to the seed address mode list; performing detection budget allocation on each non-seed prefix in the non-seed prefix list; for each non-seed prefix, executing the following detection processing steps: determining an address mode set corresponding to the non-seed prefix; according to the initial detection total budget, the number of target format addresses generated by non-seed prefixes and the number of detected survival target format addresses are determined in the t-th round of detection, and the hit rate of the non-seed prefixes in the t-th round of detection is calculated; and outputting a survival target format address subset according to the detection budget and the address mode set. In order to improve the accuracy of address mode migration, the embodiment provides an address mode migration method based on prefix clustering.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present disclosure relate to the field of network address prediction, and in particular to a network address prediction method and device based on prefix clustering. BACKGROUND

[0002] With the rapid development of the Internet, IPv6 (Internet Protocol version 6) has gradually replaced IPv4 (Internet Protocol version 4) as the basis for network communication as the next generation of Internet protocol. The global IPv6 user support capability level is steadily improving. According to the statistical data of APNIC (Asia-Pacific Network Information Center), as of August 4, 2025, the global IPv6 user support capability has reached 39.86%, and this proportion is still growing overall. It is particularly noteworthy that in some Internet developed countries such as the United States, Germany and India, the deployment rate of IPv6 has exceeded 50%, which marks that the popularization of IPv6 has entered a new stage.

[0003] One of the significant technical advantages of IPv6 is its huge address space, with a 128-bit address space that theoretically provides 3.4 x 10 38 Such a huge address space not only fundamentally solves the problem of IPv4 address exhaustion, but also provides sufficient address resource guarantee for the development of emerging fields such as the Internet of Things, 5G (5th Generation Mobile Communication Technology) communication and smart cities.

[0004] However, the huge address space also brings new challenges to network security, asset management and service discovery. In the field of network security, address scanning as a basic technical means of asset detection plays a crucial role in network space mapping, vulnerability discovery and security protection. Through systematic address scanning, network security researchers can map network assets, discover potential security risks, and take timely protective measures.

[0005] Traditional IPv4 scanning, due to the limited 32-bit address space (about 429 million addresses), can use a relatively simple and direct traversal scanning method. Existing mature IPv4 traversal scanning tools such as Masscan, ZMap can complete scanning of the entire 32-bit IPv4 address space in 6 minutes at the fastest.

[0006] But in the IPv6 environment, due to the exponential growth of its address space, and the IPv6 address has the characteristics of sparse distribution, single point multi-address, short survival period, address change, etc., it is not feasible to scan all possible addresses comprehensively in time and computing resources on the basis of existing technology. For example, even at a scanning speed of one million addresses per second, scanning a 64-bit IPv6 subnet (about 1.8x10 19 Eighth addresses) also needs more than 580 billion years, which is obviously unrealistic. For this reason, researchers have begun to explore non-exhaustive IPv6 address fast scanning detection techniques and means.

[0007] As a new means of non-exhaustive IPv6 address fast scanning, in recent years, researchers have proposed a series of IPv6 address prediction algorithms, which generate potential live IPv6 addresses as detection targets by analyzing the structural characteristics and pattern mining of seed addresses (existing live or live IPv6 addresses). At present, scholars at home and abroad have carried out extensive research in the field of IPv6 address prediction. The existing IPv6 address prediction algorithms can be roughly divided into two categories: seed prefix-oriented IPv6 address prediction algorithms and non-seed prefix-oriented IPv6 address prediction algorithms.

[0008] Seed prefix-oriented IPv6 address prediction algorithms mainly generate target IPv6 addresses with high similarity to seed addresses, and are mostly limited to the prefix range of seed addresses. While the existing typical live IPv6 address dataset IPv6Hitlist has a BGP (Border Gateway Protocol) address prefix coverage rate of less than 40%, more than half of the IPv6 addresses under the address prefix are to be detected.

[0009] The existence of this detection blind area poses a major risk to network security, and threats may use these undiscovered addresses as a springboard to launch subsequent network threats, or these addresses may be running services with vulnerabilities that have not been discovered. Therefore, how to break through the limitations of existing seed address prefixes and efficiently scan live IPv6 addresses under non-seed prefixes has become an important research topic in the field of network security and network measurement that needs to be solved. SUMMARY

[0010] The summary portion of the disclosure is used to introduce the concepts in a brief form, which will be described in detail in the specific implementation part later. The summary portion of the disclosure is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0011] Some embodiments of the present disclosure propose a prefix clustering based network address prediction method, device, electronic equipment and computer readable medium to solve the technical problems mentioned in the background section.

[0012] In a first aspect, some embodiments of the present disclosure provide a prefix clustering based network address prediction method, which comprises: based on split hierarchical clustering, performing target format address pattern mining processing on a seed address set to be clustered to generate a seed address pattern list; according to a seed prefix list corresponding to the seed address pattern list, migrating the address patterns in the seed address pattern list to a non-seed prefix list to generate a target format address list; performing probe budget allocation for each non-seed prefix in the non-seed prefix list; for each non-seed prefix in the non-seed prefix list, performing the following probe processing steps: determining a set of address patterns corresponding to the non-seed prefix; allocating an initial total probe budget for each address pattern in the set of address patterns; according to the initial total probe budget, determining the number of target format addresses generated by the non-seed prefix and the number of live target format addresses detected in the first round of probe, and calculating the hit rate of the non-seed prefix in the first round of probe; according to the hit rate of the first round of probe, calculating the probe budget allocated to the non-seed prefix in the second round of probe; according to the probe budget and the set of address patterns, outputting a subset of live target format addresses; and merging each subset of live target format addresses into a list of live target format addresses. t t t t

[0013] In a second aspect, some embodiments of the present disclosure provide a prefix clustering based network address prediction device, which comprises: a mining unit configured to perform target format address pattern mining processing on a seed address set to be clustered based on split hierarchical clustering to generate a seed address pattern list; a migration unit configured to migrate the address patterns in the seed address pattern list to a non-seed prefix list according to a seed prefix list corresponding to the seed address pattern list to generate a target format address list; a budget allocation unit configured to perform probe budget allocation for each non-seed prefix in the non-seed prefix list; and an output unit configured to, for each non-seed prefix in the non-seed prefix list, perform the following probe processing steps: determining a set of address patterns corresponding to the non-seed prefix; allocating an initial total probe budget for each address pattern in the set of address patterns; according to the initial total probe budget, determining the number of target format addresses generated by the non-seed prefix and the number of live target format addresses detected in the first round of probe, and calculating the hit rate of the non-seed prefix in the first round of probe; according to the hit rate of the first round of probe, calculating the probe budget allocated to the non-seed prefix in the second round of probe; according to the probe budget and the set of address patterns, outputting a subset of live target format addresses; and merging each subset of live target format addresses into a list of live target format addresses. t t t ​​​​​​​t +1 round of probes assigned to the non-seed prefixes; outputting a subset of live target format addresses according to the probe budget and the address pattern set; a merging unit configured to merge each subset of live target format addresses into a list of live target format addresses.

[0014] In a third aspect, some embodiments of the present disclosure provide an electronic device, comprising: one or more processors; a storage device having one or more programs stored thereon, when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the first aspect.

[0015] In a fourth aspect, some embodiments of the present disclosure provide a computer readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any implementation manner of the first aspect.

[0016] The above various embodiments of the present disclosure have the following beneficial effects: to improve the accuracy of address pattern migration, the prefix clustering based network address prediction method of some embodiments of the present disclosure proposes a prefix clustering based IPv6 address pattern migration method, on the basis of Whois information similarity, a weighted Hamming distance is designed to measure the similarity of IPv6 address prefixes, and the accuracy of IPv6 address pattern migration is improved; to make full use of probe resources, a double-layer decision architecture across address prefixes and address patterns is constructed, the probe budget is tilted to the high hit rate address area, and the IPv6 address hit rate is improved. BRIEF DESCRIPTION OF DRAWINGS

[0017] The above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, same or similar reference numerals can represent same or similar elements. It should be understood that the drawings are schematic, and elements and elements are not necessarily drawn to scale.

[0018] Figure 1 is a flowchart of some embodiments of the prefix clustering based network address prediction method according to the present disclosure; Figure 2 is a schematic diagram of dividing each seed address to be clustered into different nodes; Figure 3 is a schematic diagram of migrating address patterns in the seed address pattern list to the non-seed prefix list; Figure 4 is a structural schematic diagram of some embodiments of the prefix clustering based network address prediction device according to the present disclosure; Figure 5 is a structural schematic diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0019] Embodiments of the present disclosure will be described below in greater detail with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure will be more thoroughly and completely understood. It should be understood that the drawings of the present disclosure are only for illustrative purposes and are not intended to limit the scope of protection of the present disclosure.

[0020] It should also be noted that only parts related to the present application are shown in the drawings for the purpose of description. The embodiments in the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0021] It should be noted that the terms "first", "second", and the like in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.

[0022] It should be noted that the terms "one", "multiple" in the present disclosure are illustrative and not restrictive, and those skilled in the art should understand that unless otherwise explicitly stated in the context, it should be understood as "one or more".

[0023] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0024] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.

[0025] Figure 1 Flow 100 of some embodiments of a prefix clustering based network address prediction method of some embodiments of the present disclosure. The prefix clustering based network address prediction method includes the following steps: Step 101, based on split hierarchical clustering, target format address pattern mining processing is performed on the set of addresses to be clustered to generate a seed address pattern list.

[0026] In some embodiments, the execution subject (e.g. a computing device) of the prefix clustering based network address prediction method can perform target format address pattern mining on the seed address set to be clustered based on the split hierarchical clustering to generate a seed address pattern list. IPv6 address pattern mining aims to discover high-density address regions from a large number of seed addresses for generating new target IPv6 addresses. After matching the seed addresses to the corresponding seed prefixes through the longest prefix matching, 6X can perform address pattern mining on the seed addresses under each seed prefix, respectively. To perform pattern mining on IPv6 addresses, it is necessary to represent them in a suitable manner first. An IPv6 address can be represented as a string composed of 32 hexadecimal characters. An address pattern is a string obtained by replacing some of the characters with wildcards “*”, for example, “200120e4************************”. 6X can use the split hierarchical clustering algorithm commonly used in academia to cluster the seed addresses and obtain address patterns. The target format address can refer to an IPv6 address.

[0027] In practice, the execution subject described above can perform target format address pattern mining on the seed address set to be clustered to generate a seed address pattern list through the following steps: First, use wildcards to represent the variable nibble bits in the seed address set to be clustered. Treat all seed addresses to be clustered as a node, and use wildcards to represent the variable nibble bits in them.

[0028] Second, select a target nibble bit from all nibble bits, and traverse all seed addresses to be clustered. According to the differences of the characters at the position of the target nibble bit, divide each seed address to be clustered into different nodes. Select a nibble bit from all nibble bits, traverse all seed addresses, and divide the addresses into different nodes according to the differences of the characters at the position of the nibble bit.

[0029] Among them, dividing each seed address to be clustered into different nodes includes: using the leftmost splitting strategy, i.e. selecting the nibble bits from left to right. In Figure 2 , the first variable nibble bit from left to right is the 14th nibble bit, i.e. the red nibble bit in node 0, which divides the node into node 1 and node 2 according to the value of the nibble bit. In this paper, a threshold Threhold 1 is defined, which represents the number of variable nibble bits in the address pattern represented by the node, and is used as the stopping condition for node splitting. Figure 2 In Threhold 1 is set to 1, the number of address pattern variable nibbles in node 1, node 3 and node 4 is reduced to 1, so the splitting is stopped. After the clustering of the surviving IPv6 addresses is completed, the leaf nodes are obtained, each of which represents an address pattern and the seed addresses contained therein.

[0030] Thirdly, the seed address pattern corresponding to each node is determined to obtain a seed address pattern list. One seed address pattern contains multiple seed addresses. Finally, multiple leaf nodes are obtained, each of which represents an address pattern and the seed addresses contained therein. Figure 2 The process of clustering 6 seed addresses in the range of IPv6 address prefix 2001:20e4:: / 32 is shown. The 4 nibbles of 14, 16, 31 and 32 have different characters, which are variable nibble bits, so the address pattern represented by node 0 is “2001 20e4 00 01 *0* 00 00 00 00 00**”.

[0031] The IPv6 address pattern mining algorithm is shown in Algorithm 1. The input of the algorithm includes a list of surviving IPv6 addresses AddrLis t and a node splitting threshold Threhold . Wherein Threhold The value range is, during the node splitting process, the number of variable nibbles in the address pattern gradually decreases, and when it is reduced to less than or equal to Threhold , the node stops splitting. The Group function is used to cluster seed addresses to obtain address patterns (seed address patterns).

[0032]

[0033] Step 102, according to the seed prefix list corresponding to the seed address pattern list, the address patterns in the seed address pattern list are migrated to the non-seed prefix list to generate a target format address list.

[0034] In some embodiments, the above execution subject can generate a target format address list by migrating the address patterns in the seed address pattern list to the non-seed prefix list according to the seed prefix list corresponding to the seed address pattern list. After the address prefix clustering is completed, how to accurately migrate the address patterns under the seed prefix to the non-seed prefix becomes a key challenge. According to the IPv6 address allocation strategy, on the one hand, the IPv6 addresses of the same country, autonomous system and organization usually have similar address patterns. On the other hand, the address patterns under the address prefixes of similar structures also have similarities.

[0035] This paper proposes an IPv6 address pattern migration method based on prefix clustering, which realizes address prefix association and pattern migration by combining bit pattern similarity and Whois information similarity. This paper introduces Hamming distance to measure the similarity of IPv6 address prefixes. The traditional Hamming distance only counts the number of different binary bits in the string. However, in IPv6 addresses, different positions of binary bits have different meanings.

[0036] In practice, the aforementioned execution entity can migrate the address patterns in the seed address pattern list to the non-seed prefix list through the following steps to generate the target format address list: The first step is to perform the following address pattern migration steps for each non-seed prefix in the non-seed prefix list: First, determine the Hamming distance between the non-seed prefixes and each seed prefix in the seed prefix list.

[0037] When calculating the Hamming distance of address prefixes, different weights are assigned to the binary bits at different positions. For two 128-bit IPv6 address vectors v1 (non-seed prefix) and v2 (seed prefix), the improved distance calculation formula is as follows:

[0038] in, For indicator functions, if x If true, the value is 1; otherwise, the value is 0. and They represent the first and second digits in v1 and v2 respectively. i 1 bit, weighting coefficient The bits are arranged hierarchically according to their importance, with higher weights for earlier bits. The specific settings are as follows: .

[0039] Second, using a hierarchical matching strategy, the information field values ​​between the aforementioned non-seed prefixes and each seed prefix in the aforementioned seed prefix list are determined. The calculation of the information field values ​​adopts the following hierarchical matching strategy formula: .

[0040] in, C, O, P These represent the country code, autonomous system number, and registrar, respectively. T , S These represent the target non-seed prefix and seed prefix, respectively. α , β , γ These are the hierarchical weight coefficients, representing the matching weights for the country code, autonomous system number, and organization, respectively, with default values ​​of 0.5, 0.3, and 0.2.

[0041] Third, for each seed prefix in the above seed prefix list, perform the following processing steps: 1. Determine the similarity score between the non-seed prefix and the seed prefix according to the Hamming distance corresponding to the seed prefix, the information field value. The algorithm uses inverted index to achieve fast retrieval. Inverted index is a database index structure used to quickly find documents containing a specific keyword. In this algorithm, the field value of each level (such as country code "US", autonomous system number "AS123") is used as the index key, and the prefix list containing the field is used as the index value. According to the Whois information field value of the non-seed prefix, the seed prefix list corresponding to the information field value is weighted and assigned, that is, the registration information similarity of the target non-seed prefix and various seed prefixes can be quickly calculated. The final similarity score is generated by linear combination, and the calculation method is as follows: .

[0042] Wherein, is the balance coefficient (β ), which can be customized by the user according to the actual use, and the larger the value represents the greater the weight of the prefix structure relative to the Whois information field value when clustering address prefixes.

[0043] 2. In response to determining that the similarity score is greater than or equal to the preset score, migrate the seed address pattern corresponding to the seed prefix to the non-seed prefix to obtain the target format address. After the address prefix association is completed, each non-seed prefix has several associated seed prefixes, and the next step is to migrate the address pattern under the seed prefix to the non-seed prefix to generate the target IPv6 address. As shown in Figure 3 , the associated seed prefix of the non-seed prefix 2001:20e5:: / 32 is 2001:20e4:: / 32, and there are three address patterns under this seed prefix. Migrate the last 96 bits to 2001:20e5:: / 32 to generate the target IPv6 address. When migrating the address pattern, there may be two problems to be solved, which are too many or too few migratable address patterns of the non-seed prefix.

[0044] For the case of too few migratable address patterns, the main reason is that the number of seed prefixes associated with the non-seed prefix is too small or the number of address patterns under the seed prefix is too small. To solve this problem, 6Xcan not only migrates the address pattern under the seed prefix to the non-seed prefix, but also uses a target IPv6 address generation method similar to HMap6 to traverse the 16 subnet prefixes under the non-seed prefix and generate a target IPv6 address with an interface identifier of "::1" under each subnet prefix.

[0045] For the case of too many migratable address patterns, the reason is that there are too many address patterns under the seed prefix. 6Xcan defines a threshold Threhold2. When performing address mode migration, 6X can sort all address modes according to the number of seed addresses, and select the top Threhold 2 address modes with the largest number of seed addresses. The IPv6 address mode migration algorithm is shown in Algorithm 2.

[0046]

[0047] Secondly, the obtained target format addresses are merged into a target format address list.

[0048] Due to the large IPv6 address space, it is not possible to directly determine which address prefixes have higher hit rates before detection, so it is necessary to dynamically generate target addresses and dynamically allocate detection budgets based on detection results to focus on detection in high hit rate address areas. Although AddrMiner-N uses a dynamic detection method, its budget allocation strategy is equal allocation between prefixes, and only address modes are allocated differently, resulting in a large amount of detection budget being wasted in low response rate address prefixes. To solve the shortcomings of AddrMiner-N, the budget allocation strategy is optimized in this paper, and a dual budget allocation space is constructed, with budget allocation across address prefixes and address modes, to improve detection hit rate.

[0049] Step 103, allocate detection budgets to each non-seed prefix in the above non-seed prefix list.

[0050] In some embodiments, the above execution subject can allocate detection budgets to each non-seed prefix in the above non-seed prefix list. Assuming that the total detection budget is Budget , the non-seed prefix set is PrefixSet ={ P 1, P 2, …, P n}, n is the number of non-seed prefixes, and the non-seed prefix P i corresponding address mode set is PatternSet i ={ p i1 , p i2 , …, p im}. In the first round of detection, equal detection budgets are allocated to each non-seed prefix. Assuming that the preset number of detection iterations is Epoch , then the detection budget allocated to each non-seed prefix is: .

[0051] Step 104, for each non-seed prefix in the above non-seed prefix list, the following detection processing steps are performed: Step 1041, determine the address pattern set corresponding to the above non-seed prefix.

[0052] In some embodiments, the above execution subject can determine the address pattern set corresponding to the above non-seed prefix. The non-seed prefix P i The corresponding address pattern set is PatternSet i ={ p i1 , p i2 , …, p im}.

[0053] Step 1042, assign an initial total detection budget for each address pattern in the above address pattern set.

[0054] In some embodiments, the above execution subject can assign an initial total detection budget for each address pattern in the above address pattern set.

[0055] In practice, the above execution subject can assign an initial total detection budget for each address pattern in the above address pattern set by the following steps: First, determine the number of addresses of the seed prefix in the above address pattern. That is, the number of seed prefixes corresponding to the address pattern s ij .

[0056] Second, according to the above address number and the initial total detection budget, assign an initial total detection budget for the address pattern. The initial total detection budget is .

[0057] For example, the number of seed addresses in the address pattern p ij is s ij , then p ij The assigned initial total detection budget b ij is: .

[0058] Step 1043, according to the initial total detection budget, determine the number of target format addresses generated by the above non-seed prefix in the first t round of detection, and the number of survived target format addresses detected, and calculate the hit rate of the above non-seed prefix in the first t round of detection.

[0059] In some embodiments, the execution subject can determine the number of target format addresses generated by the non-seed prefix according to the initial detection total budget, the number of detected surviving target format addresses in the first round of detection, and calculate the hit rate of the non-seed prefix in the first round of detection. Assuming that the number of target IPv6 addresses generated by the non-seed prefix in the first round of detection (the number of target format addresses) is N1, the number of detected surviving IPv6 addresses is N2, then the hit rate of the address prefix is N1 / N2. t t In some embodiments, the execution subject can determine the number of target format addresses generated by the non-seed prefix according to the initial detection total budget, the number of detected surviving target format addresses in the first round of detection, and calculate the hit rate of the non-seed prefix in the first round of detection. Assuming that the number of target IPv6 addresses generated by the non-seed prefix in the first round of detection (the number of target format addresses) is N1, the number of detected surviving IPv6 addresses is N2, then the hit rate of the address prefix is N1 / N2. t P i h i .

[0060] That is, according to the detection budget of each round, scan the target IPv6 address, and determine the responding target IPv6 address as the surviving IPv6 address.

[0061] Step 1044, according to the hit rate of the first round of detection, calculate the detection budget allocated to the non-seed prefix in the first round of detection. t t In some embodiments, the execution subject can calculate the detection budget allocated to the non-seed prefix in the first round of detection according to the hit rate of the first round of detection. According to the hit rate, calculate the detection budget allocated to the prefix in the first round of detection.

[0062] In some embodiments, the execution subject can calculate the detection budget allocated to the non-seed prefix in the first round of detection according to the hit rate of the first round of detection. According to the hit rate, calculate the detection budget allocated to the prefix in the first round of detection. t t In some embodiments, the execution subject can calculate the detection budget allocated to the non-seed prefix in the first round of detection according to the hit rate of the first round of detection. According to the hit rate, calculate the detection budget allocated to the prefix in the first round of detection. t P i .

[0063] Step 1045, output the surviving target format address subset according to the detection budget and the address mode set.

[0064] In some embodiments, the execution subject can output the surviving target format address subset according to the detection budget and the address mode set. That is, according to the detection budget, randomly construct the same number of target format addresses through the address mode set. Then, send instructions to each constructed target format address. Finally, determine the responding target format address as the surviving target format address. Here, the detection budget represents the number. The construction method of the target format address can refer to the construction method described above.

[0065] Step 105, combine each surviving target format address subset into a surviving target format address list.

[0066] ​​​​​​​​​​In some embodiments, the execution subject described above can combine each subset of the survived target format address into a list of survived target format addresses. Through the above method, the detection budget is tilted to the high hit rate address prefix and address mode, and the hit rate of the survived IPv6 address is improved. The IPv6 address dynamic detection algorithm is shown in Algorithm 3.

[0067]

[0068] Further referring to Figure 4 , as an implementation of the method shown in the above figures, the present disclosure provides some embodiments of a prefix clustering-based network address prediction device, which corresponds to the method embodiments shown in Figure 1 , and the prefix clustering-based network address prediction device can be applied in various electronic devices.

[0069] As shown in Figure 4 , the prefix clustering-based network address prediction device 400 of some embodiments includes a mining unit 401, a migration unit 402, a budget allocation unit 403, an output unit 404, and a merging unit 405. The mining unit 401 is configured to perform target format address mode mining processing on a set of seed addresses to be clustered based on a split hierarchical clustering to generate a list of seed address modes; the migration unit 402 is configured to migrate the address modes in the list of seed address modes to a list of non-seed prefixes according to a list of seed prefixes corresponding to the list of seed address modes, and generate a list of target format addresses; the budget allocation unit 403 is configured to allocate a detection budget to each non-seed prefix in the list of non-seed prefixes; the output unit 404 is configured to, for each non-seed prefix in the list of non-seed prefixes, perform the following detection processing steps: determine a set of address modes corresponding to the non-seed prefix; allocate an initial total detection budget to each address mode in the set of address modes; determine the number of target format addresses generated by the non-seed prefix and the number of survived target format addresses detected according to the initial total detection budget in the first round of detection; calculate the hit rate of the non-seed prefix in the first round of detection; calculate the detection budget allocated to the non-seed prefix in the first +1 round of detection according to the hit rate of the first round of detection; output a subset of survived target format addresses according to the detection budget and the set of address modes; and the merging unit 405 is configured to combine each subset of the survived target format addresses into a list of survived target format addresses. t t t t

[0070] It can be understood that the units described in the prefix clustering-based network address prediction device 400 are described with reference to Figure 1 ​​​​The various steps in the described method correspond. Thus, the operations, features and advantages described above for the method also apply to the network address prediction apparatus 400 based on prefix clustering and the units contained therein, which will not be repeated here.

[0071] Reference is made below in detail to Figure 5 which shows a structural schematic diagram of an electronic device (such as a computing device) suitable for implementing some embodiments of the present disclosure. Figure 5 The electronic device shown is merely an example and should not impose any limitation on the functions and use range of the embodiments of the present disclosure. As Figure 5 As shown, the computer device includes a processor, a memory and a network interface connected through a system bus, wherein the memory can include a non-volatile storage medium and an internal memory. The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions which, when executed, can cause the processor to perform any one of the network address prediction methods based on prefix clustering. The processor is used to provide computing and control capabilities to support the operation of the entire computer device. The internal memory provides an environment for the computer program in the non-volatile storage medium, which, when executed by the processor, can cause the processor to perform any one of the network address prediction methods based on prefix clustering. The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 5 The structure shown in the figure is only a block diagram of part of the structure related to the present disclosure scheme, and does not constitute a limitation on the computer device to which the present disclosure scheme is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0072] It should be understood that the processor can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0073] In one embodiment, the processor is configured to run a computer program stored in the memory to perform the following steps: based on the hierarchical clustering, performing target format address pattern mining on the set of clustering seeds to generate a seed address pattern list; according to a seed prefix list corresponding to the seed address pattern list, migrating the address patterns in the seed address pattern list to a non-seed prefix list to generate a target format address list; performing probe budget allocation for each non-seed prefix in the non-seed prefix list; for each non-seed prefix in the non-seed prefix list, performing the following probe processing steps: determining a set of address patterns corresponding to the non-seed prefix; allocating an initial total probe budget for each address pattern in the set of address patterns; according to the initial total probe budget, determining the number of target format addresses generated by the non-seed prefix and the number of live target format addresses detected in the first round of probing, and calculating the hit rate of the non-seed prefix in the first round of probing; according to the hit rate of the first round of probing, calculating the probe budget allocated to the non-seed prefix in the second round of probing; according to the probe budget and the set of address patterns, outputting a subset of live target format addresses; and merging each subset of live target format addresses into a list of live target format addresses. t t t t

[0074] The embodiments of the present disclosure further provide a computer readable storage medium, and the computer readable storage medium stores a computer program. The computer program includes program instructions. When the program instructions are executed, the method can refer to each embodiment of the network address prediction method based on prefix clustering of the present disclosure.

[0075] The computer readable storage medium can be an internal storage unit of the computer device, for example, a hard disk or a memory of the computer device. The computer readable storage medium can also be an external storage device of the computer device, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.

[0076] ​​​​It should be noted that, in the present document, the terms "comprising", "comprises" or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or system. Without further limitation, an element preceded by "comprises a" does not, without more constraints, foreclose the existence of additional identical elements in the process, method, article, or system that comprises the recited element.

[0077] The above description is merely exemplary of the present disclosure and the application principles of the technology used. Those skilled in the art should understand that the inventive scope of the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by the combinations of the above technical features or equivalent features without departing from the above inventive concept. For example, the above features and the technical features disclosed in the embodiments of the present disclosure (but not limited to) with similar functions are replaced with each other to form technical solutions.

Claims

1. A method for prefix-cluster-based network address prediction, characterized in that, The method comprises the following steps: Based on the split hierarchical clustering, the target format address pattern mining processing is performed on the seed address set to be clustered to generate a seed address pattern list; According to the seed prefix list corresponding to the seed address pattern list, the address patterns in the seed address pattern list are migrated to the non-seed prefix list to generate a target format address list; Each non-seed prefix in the non-seed prefix list is allocated a detection budget; For each non-seed prefix in the non-seed prefix list, the following detection processing steps are performed: Determine the address pattern set corresponding to the non-seed prefix; Each address pattern in the address pattern set is allocated an initial detection total budget; According to the initial detection total budget, the number of target format addresses generated by the non-seed prefix and the number of detected live target format addresses are determined in the tth round of detection, and the hit rate of the non-seed prefix in the tth round of detection is calculated; According to the hit rate of the tth round of detection, the detection budget allocated to the non-seed prefix in the t+1th round of detection is calculated; According to the detection budget and the address pattern set, the live target format address subset is output; Each live target format address subset is combined into a live target format address list.

2. The method of claim 1, wherein, The method comprises the following steps: The variable nibble bits in the seed address set to be clustered are represented by a wildcard; A target nibble bit is selected from all nibble bits, and all seed addresses to be clustered are traversed. According to the difference of the character at the position of the target nibble bit, each seed address to be clustered is divided into different nodes; Determine the seed address pattern corresponding to each node to obtain a seed address pattern list, wherein a seed address pattern contains multiple seed addresses.

3. The method of claim 2, wherein, The method comprises the following steps: For each non-seed prefix in the non-seed prefix list, the following address pattern migration steps are performed: Determine the Hamming distance between the non-seed prefix and each seed prefix in the seed prefix list; Through a hierarchical matching strategy, the information field values between the non-seed prefix and each seed prefix in the seed prefix list are determined; For each seed prefix in the seed prefix list, the following processing steps are performed: According to the Hamming distance and the information field value corresponding to the seed prefix, the similarity score between the non-seed prefix and the seed prefix is determined; In response to determining that the similarity score is greater than or equal to a preset score, the seed address pattern corresponding to the seed prefix is migrated to the non-seed prefix to obtain a target format address; The obtained target format addresses are combined into a target format address list.

4. The method of claim 3, wherein, The method comprises the following steps: Determine the number of addresses of the seed prefix in the address pattern; According to the address number and the initial total detection budget, the initial detection total budget is allocated to the address pattern.

5. A prefix clustering based network address prediction apparatus, characterized by, The method comprises the following steps: The mining unit is configured to perform target format address pattern mining processing on the set of clustering seed addresses based on the split hierarchical clustering to generate a seed address pattern list; The migration unit is configured to migrate the address patterns in the seed address pattern list to a non-seed prefix list according to a corresponding seed prefix list of the seed address pattern list to generate a target format address list; The budget allocation unit is configured to allocate a detection budget to each non-seed prefix in the non-seed prefix list; The output unit is configured to perform the following detection processing steps for each non-seed prefix in the non-seed prefix list: determining a set of address patterns corresponding to the non-seed prefix; allocating an initial total detection budget to each address pattern in the set of address patterns; According to the initial total detection budget, in the tthround of detection, determining the number of target format addresses generated by the non-seed prefix and the number of detected live target format addresses, and calculating the hit rate of the non-seed prefix in the tthround of detection; according to the hit rate of the tthround of detection, calculating the detection budget allocated to the non-seed prefix in the t+1thround of detection; according to the detection budget and the set of address patterns, outputting a subset of live target format addresses; The merging unit is configured to merge each subset of live target format addresses into a live target format address list.

6. An electronic device, comprising: comprise: one or more processors; a storage device having stored thereon one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 4.

7. A computer readable medium characterized by a computer program is stored thereon, wherein the computer program is executed by a processor to implement the method according to any one of claims 1 to 4.