Efficient IPv6 address detection method based on neural network model

By constructing an IPv6 address detection method based on a neural network model, self-supervised training is used to generate candidate addresses and perform activity detection, which solves the problems of low efficiency and low hit rate of IPv6 address detection in the existing technology and realizes efficient IPv6 address detection.

CN120675973AActive Publication Date: 2025-09-19NAT COMPUTER NETWORK & INFORMATION SECURITY MANAGEMENT CENT JIANGSU BRANCH

Patent Information

Application Number
CN202511183913.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-09-19
Estimated Expiration
2045-08-22

AI Technical Summary

Technical Problem

The existing technology is inefficient and has a low hit rate in IPv6 address detection, and is unable to effectively detect IPv6 addresses.

Method used

An IPv6 address detection method based on a neural network model is constructed. The neural network model is trained through self-supervision, candidate addresses are generated using IPv6 seed addresses, and activity detection is performed to remove duplicate and alias prefix addresses.

Benefits of technology

The hit rate of IPv6 address detection is improved. The hit rate of generated candidate addresses is between 56% and 74%, which is 1.1 to 292 times that of existing methods, significantly improving detection efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120675973A_ABST
    Figure CN120675973A_ABST
Patent Text Reader

Abstract

The invention discloses an efficient IPv6 address detection method based on a neural network model, and belongs to the technical field of IPv6 address detection. The method comprises the following steps: step 1, collecting an IPv6 seed address, and preprocessing the IPv6 seed address to obtain a plurality of IPv6 seed address data sets; 2, constructing a neural network model, and training the neural network model by using the plurality of IPv6 seed address data sets; using the trained neural network model to generate a candidate IPv6 address; and step 3, active address detection is carried out on the generated candidate IPv6 addresses, and an active target IPv6 address is obtained. According to the method, an IPv6 address is regarded as a text, a seed address is used for training a network in a self-supervision mode, and the hit rate is higher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of IPv6 address detection, and in particular relates to an efficient IPv6 address detection method based on a neural network model. Background Art

[0002] Network asset detection is a fundamental task in network security. Currently, IPv4 asset detection technology is highly mature. For example, the port scanning tool Masscan can scan designated ports across the entire IPv4 address space in just six minutes. As IPv4 addresses gradually become depleted, IPv6 addresses are becoming increasingly common, with IPv6 currently accounting for 44.96% of global traffic. However, due to the vast and sparse IPv6 address space, a brute-force scan of all IPv6 addresses would take millions of years. In other words, with current computing power, scanning the entire IPv6 address space is simply unfeasible. Therefore, developing effective IPv6 address detection methods remains a technical challenge.

[0003] For IPv4 / IPv6 dual-stack hosts, existing technologies, in an IPv4 network environment, induce the target host to actively send a request to a constructed IPv6 server by sending SSDP (Simple Service Discovery Protocol) packets, and then extract the IPv6 address from the server log; or enumerate the target host's service list and its corresponding AAAA records through the DNS-SD protocol to obtain the target host's IPv6 address. This method finds a small number of IPv6 addresses.

[0004] Existing techniques also include generating a set of predicted candidate addresses that may survive from a seed address set, and then performing liveness detection on these candidate target addresses to obtain valid target addresses. These methods can be divided into the following four categories.

[0005] (1) Based on address entropy, for example, the Entropy / IP method mines the IPv6 address structure implicit in a set of IPv6 addresses through entropy to generate predicted addresses, but the hit rate of these addresses is low.

[0006] (2) Based on address density: For example, based on the assumption that areas with dense seed addresses are more likely to contain other active hosts, similar seeds are clustered to form high-density areas, thereby generating target addresses with a high hit rate. Alternatively, a spatial tree is constructed using the dual-debiased heterogeneous co-training (DHC) algorithm to reveal the distribution characteristics of seed addresses. The search direction is then dynamically adjusted based on real-time scanning feedback, prioritizing scanning areas with dense active addresses. 6Graph, a graph-based IPv6 address pattern mining method, extracts high-density address patterns using an improved hierarchical clustering algorithm and a density-optimized minimum spanning tree (MST) clustering algorithm. These methods rely too much on seed addresses, and the generated target addresses lack diversity.

[0007] (3) Based on reinforcement learning, for example, the Internet IPv6 scanning method 6Sense combines reinforcement learning with online scanning: it generates candidate IPv6 addresses through iterative optimization and narrows the scanning range. A comprehensive global active IPv6 address discovery system, AddrMiner, divides the IPv6 address space into seedless areas (AddrMiner-N), seed-sparse areas (AddrMiner-F), and seed-sufficient areas (AddrMiner-S). In seed-sufficient areas, it dynamically adjusts the address generation direction through reinforcement learning to optimize the detection efficiency in high-density areas.

[0008] (4) Based on deep learning, for example, the gated convolutional variational autoencoder 6GCVAE for IPv6 target generation learns the seed address distribution through the gated convolution layer, forms the latent space parameters, and generates new predicted addresses from the latent space sampling.

[0009] The above-mentioned existing technologies use a certain number of known IPv6 seed addresses, analyze their entropy, structural characteristics, or address density distribution, generate a certain number of new predicted addresses, and then perform activity detection on these generated addresses to find hidden IPv6 addresses. However, these target generation algorithms (TGA) often suffer from low address generation efficiency and low hit rate. Summary of the Invention

[0010] Purpose of the invention: The technical problem to be solved by the present invention is to provide an efficient IPv6 address detection method based on a neural network model in response to the shortcomings of the existing technology.

[0011] In order to solve the above technical problems, the present invention discloses an efficient IPv6 address detection method based on a neural network model, comprising the following steps:

[0012] Step 1: Collect IPv6 seed addresses, pre-process the IPv6 seed addresses, and obtain multiple IPv6 seed address data sets;

[0013] Step 2: constructing a neural network model and training the neural network model using the multiple IPv6 seed address data sets; generating candidate IPv6 addresses using the trained neural network model;

[0014] Step 3: Perform active address detection on the generated candidate IPv6 addresses to obtain an active target IPv6 address.

[0015] Furthermore, step 1 includes:

[0016] Perform activity detection on the collected IPv6 seed addresses to obtain responsive active IPv6 seed addresses;

[0017] Downsampling the active IPv6 seed addresses to obtain multiple IPv6 seed address sets with different sampling numbers;

[0018] The IPv6 seed addresses in each IPv6 seed address set are expanded, the colon is removed, a start identifier and an end identifier are added to the beginning and end of the IPv6 seed address respectively to obtain a character string, each character of the character string is converted into an integer identifier to obtain an IPv6 seed address data vector, and then multiple IPv6 seed address data sets are obtained.

[0019] Furthermore, the neural network model in step 2 includes an embedding layer, N decoder layers, a linear output layer and a normalization layer. , the embedding layer is used to convert the IPv6 seed address data vector in the IPv6 seed address dataset into an IPv6 seed address two-dimensional tensor;

[0020] The decoder layer is used to extract hierarchical features of the output sequence of the embedding layer during training and to autoregressively (bit by bit) generate the target address sequence during inference;

[0021] The linear output layer is used to map the high-dimensional vector output by the decoder to a dimension with the same size as the IPv6 address encoding dictionary;

[0022] The normalization layer is used to convert the linear layer output into the probability of occurrence of each character in the IPv6 address encoding dictionary, and then randomly select a character as the output from the top_k characters with the highest probability using the probability as the weight. .

[0023] Furthermore, the decoder layer in step 2 includes two identical masked multi-head attention layers and a feedforward network layer. The layers are connected with residual connections and layer normalization is performed, and the activation function of each layer is Gelu.

[0024] The input and output dimensions of the masked multi-head attention layer are both the model width, and the input and output dimensions of the feedforward network layer are both the model width.

[0025] Furthermore, each bit of the output vector of the linear output layer in step 2 corresponds to a character in the IPv6 address encoding dictionary; during model training, the output of the linear output layer is directly used to calculate the cross entropy loss with the target data to adjust the model parameters; in the inference stage, the output of the linear output layer needs to be processed in combination with the normalization layer to generate a candidate IPv6 address.

[0026] Furthermore, in the inference phase, the normalization layer described in step 2 converts the output vector of the linear output layer into a probability distribution corresponding to the probability of occurrence of each character in the IPv6 address encoding dictionary. The output probability distribution is controlled by the temperature parameter T, and one of the top_k characters with the largest probability is randomly selected as the generated character using the probability as the weight to generate a candidate IPv6 address.

[0027] Furthermore, the use of multiple IPv6 seed address data sets to train the neural network model in step 2 includes: adopting a self-supervised training mode to train multiple IPv6 seed address data sets separately; there is no need to manually label the data, and the target data is generated by the IPv6 seed address data itself; the model loss function selects cross entropy loss, the optimizer selects Adam, and the learning rate is adaptively adjusted, and the first-order moment and second-order moment of the gradient are used to update the parameters.

[0028] Furthermore, the step 2 of using the trained neural network model to generate a candidate IPv6 address includes:

[0029] Calculate the IPv6 address bit by bit to generate the initial candidate IPv6 address;

[0030] De-duplicate the initial candidate IPv6 addresses to obtain the final candidate IPv6 addresses.

[0031] Furthermore, in step 3, before performing active address detection on the generated candidate IPv6 addresses, the final candidate IPv6 addresses are removed from the addresses that are duplicated with the IPv6 seed addresses and the alias prefix addresses.

[0032] Furthermore, the model width, the number of decoder layers N, and the internal dimension d of the feedforward network of the decoder layer in the neural network model in step 2 are ff, the number of attention taps h, the dropout rate, the normalization layer temperature parameter and top_k are all hyperparameters and are adjusted according to the actual situation of the training dataset.

[0033] Beneficial effects:

[0034] This application proposes an efficient neural network-based IPv6 address detection method. The neural network model treats IPv6 addresses as text and uses seed addresses for self-supervision. Training on seed sets of 10K, 100K, and 1M, respectively, generates candidate addresses ranging from 1M to 10M. The detection hit rate ranges from 56% to 74%, 1.1 to 292 times higher than other current methods. In the future, seed addresses can be manually labeled with address patterns or domains, allowing the network model to generate desired address patterns or IPv6 addresses in certain domains, further contributing to network security research. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more apparent.

[0036] Figure 1 A flowchart of an efficient IPv6 address detection method based on a neural network model is provided in an embodiment of the present application.

[0037] Figure 2 A schematic diagram of the neural network model structure in an efficient IPv6 address detection method based on a neural network model provided in an embodiment of the present application.

[0038] Figure 3 A schematic diagram of the internal structure of the decoder layer of a neural network model in an efficient IPv6 address detection method based on a neural network model provided in an embodiment of the present application.

[0039] Figure 4 A curve showing changes in the activity rate of IPv6 addresses over time in a dataset used in an experimental evaluation of an efficient IPv6 address detection method based on a neural network model provided in an embodiment of the present application.

[0040] Figure 5 A histogram of the proportion of the top 10 autonomous systems in the data during experimental evaluation of an efficient IPv6 address detection method based on a neural network model provided in an embodiment of the present application.

[0041] Figure 6 A histogram of the proportion of the top 10 network prefixes in the data during experimental evaluation of an efficient IPv6 address detection method based on a neural network model provided in an embodiment of the present application.

[0042] Figure 7 An efficient IPv6 address detection method based on a neural network model provided in an embodiment of the present application uses seed address data sets S1 to S3 to generate hit rates of candidate addresses when using different hierarchical models during experimental evaluation.

[0043] Figure 8 An efficient IPv6 address detection method based on a neural network model provided in an embodiment of the present application uses seed address data sets S1 to S3 to generate hit rates of candidate addresses at different model dimensions during experimental evaluation.

[0044] Figure 9 An efficient IPv6 address detection method based on a neural network model provided in an embodiment of the present application uses seed address data sets S1 to S3 to generate the hit rate of candidate addresses when the hidden layer dimensions of the model feedforward network are different during experimental evaluation.

[0045] Figure 10 An efficient IPv6 address detection method based on a neural network model provided in an embodiment of the present application uses seed address data sets S1 to S3 to generate the hit rate of candidate addresses at different numbers of model attention taps during experimental evaluation.

[0046] Figure 11 An efficient IPv6 address detection method based on a neural network model provided in an embodiment of the present application uses a seed address data set S1 to S3 to generate a hit rate of candidate addresses in a model with different discard rates during experimental evaluation.

[0047] Figure 12 An efficient IPv6 address detection method based on a neural network model provided in an embodiment of the present application uses seed address data sets S1 to S3 to generate hit rates of candidate addresses at different temperatures during experimental evaluation.

[0048] Figure 13 An efficient IPv6 address detection method based on a neural network model provided in an embodiment of the present application uses seed address data sets S1 to S3 to generate the hit rate of candidate addresses at different top_k during experimental evaluation.

[0049] Figure 14 A schematic diagram comparing the hit rates of an efficient IPv6 address detection method based on a neural network model provided in an embodiment of the present application and an existing method for generating candidate addresses using seed address data sets S1 to S3.

[0050] Figure 15 A pseudocode diagram of generating candidate IPv6 addresses in an efficient IPv6 address detection method based on a neural network model provided in an embodiment of the present application. DETAILED DESCRIPTION

[0051] The embodiments of the present invention will be described below with reference to the accompanying drawings.

[0052] IPv6 addresses are 128 bits long and are typically represented using 32-bit hexadecimal characters separated by colons. The type of IPv6 address is identified by the high-order bit of the address. These include: unassigned address :: / 128, loopback address ::1 / 128, multicast address FF00:: / 8, link-local address FE80:: / 10, and global unicast address 2000:: / 3. Current IPv6 address detection techniques are primarily focused on global unicast addresses.

[0053] IPv6 addresses are allocated downwards by the Internet Assigned Numbers Authority (IANA). IPv6 addresses assigned to users by carriers typically consist of a / 64-bit segment. The lower 64 bits are called the Interface IDentifier (IID), which is assigned by the user. IIDs can be randomly generated, encoded in EUI-64, or manually configured. Manual configuration includes using low-byte addresses, embedding IPv4 addresses, embedding service ports, and embedding characters. Analysis of IPv6 address encoding can significantly reduce the address scanning space, providing a method for IPv6 address scanning. However, this method may not work for randomly generated or complex manually configured addresses.

[0054] The present application embodiment discloses an efficient IPv6 address detection method based on a neural network model, such as Figure 1 As shown, the following steps are included:

[0055] Step 1: Collect IPv6 seed addresses and pre-process the IPv6 seed addresses to obtain multiple IPv6 seed address data sets, specifically including:

[0056] Perform activity detection on the collected IPv6 seed addresses to obtain responsive active IPv6 seed addresses;

[0057] Downsampling the active IPv6 seed addresses to obtain multiple IPv6 seed address sets with different sampling numbers;

[0058] The IPv6 seed addresses in each IPv6 seed address set are expanded, the colon is removed, a start identifier and an end identifier are added to the beginning and end of the IPv6 seed address respectively to obtain a character string, each character of the character string is converted into an integer identifier to obtain an IPv6 seed address data vector, and then multiple IPv6 seed address data sets are obtained.

[0059] Step 2: constructing a neural network model and training the neural network model using the multiple IPv6 seed address data sets; generating candidate IPv6 addresses using the trained neural network model;

[0060] The neural network model includes an embedding layer (Embedding), N decoder layers, a linear output layer (Linear) and a normalization layer (Softmax). , the model structure is as follows Figure 2 As shown, the embedding layer is used to convert the IPv6 seed address data vector in the IPv6 seed address data set into an IPv6 seed address two-dimensional tensor;

[0061] The decoder layer is used to extract hierarchical features of the output sequence of the embedding layer during training and to autoregressively generate the target address sequence during inference;

[0062] The linear output layer is used to map the high-dimensional vector output by the decoder to a dimension with the same size as the IPv6 address encoding dictionary;

[0063] The normalization layer is used to convert the linear layer output into the probability of occurrence of each character in the IPv6 address encoding dictionary, and then randomly select a character as the output from the top_k characters with the highest probability using the probability as the weight. .

[0064] Compared with the classic Transformer, the neural network model in this embodiment has the following optimizations:

[0065] (1) IPv6 address generation is a text generation task, and the Transformer encoder has little effect. This embodiment removes the encoder, reduces the number of model parameters, improves the efficiency of the algorithm, and achieves better IPv6 address generation. model =512, the internal dimension of the feedforward network of the decoder layer is d ff =2048 as an example, the number of model parameters is reduced from 44M to 25M.

[0066] (2) An IPv6 address can be considered as 32 characters, and the text length is relatively short. Therefore, this embodiment removes the positional encoding, further optimizes the model, and shortens the model training and generation time.

[0067] According to the RFC5952 standard, IPv6 addresses are represented using colons, and consecutive zeros in the middle can be compressed. Therefore, IPv6 addresses need to be encoded before they can be input into the neural network model. Taking the OpenDNS IPv6 address 2620:0:ccc::2 as an example, it is expanded to 2620:0000:0ccc:0000:0000:0000:0000:0002, removing the colon and adding the start mark. <bos>and end marker <eos>Later became <bos>262000000ccc00000000000000000002 <eos>, convert these strings into integer identifiers (ID) and transform them into a vector [16, 2, 6, 2, 0,0, 0, 0, 0, 0, 12, 12, 12, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 2, 17], where <bos> , <eos>They are represented by 16 and 17 respectively.

[0068] The size of the embedding layer dictionary is 18, i.e. characters 0-9, af and <bos>and <eos>, the embedding dimension needs to be the same as the model width d model To keep it consistent, the one-dimensional vector of length (34,) is converted to a vector of size (34, d model ) can be input into the model for the next step of processing.

[0069] The decoder layer is as follows Figure 3 As shown in the figure, each decoder component consists of two identical masked multi-head attention layers and a feed-forward network layer. The layers are connected via residual connections (Add) and normalized (Norm). The activation function of each layer is Gelu, which can improve model performance.

[0070] The input and output dimensions of the multi-head attention layer are both d model The feedforward network layer is actually two linear layers in series, and the input and output dimensions of the two linear layers are (d model , d ff ) and (d ff , d model ), so the input and output dimensions of the feedforward network are still d model .

[0071] In this embodiment, the width of the entire neural network model is d model The IPv6 address encoding dictionary is only 18 characters long, so a linear output layer is needed for conversion, with each bit of the output corresponding to each character in the dictionary. During model training, the output of the linear output layer is directly compared with the target data to calculate the cross-entropy loss to adjust the model parameters. During inference, the output of the linear output layer is further processed with the Softmax layer to generate candidate addresses.

[0072] In the inference phase, the normalization layer converts the output vector of the linear output layer into a probability distribution, corresponding to the probability of each character in the dictionary, and controls the output probability distribution through the temperature parameter T, and then A random selection with probability as weight is used as the generating character to generate the candidate IPv6 address. . The larger the value, the richer the target addresses generated; The smaller the value, the more repetitive the generated target address will be. When it is a greedy algorithm, there is only one target address generated.

[0073] The use of multiple IPv6 seed address data sets to train the neural network model in step 2 includes: using a self-supervised training mode to train multiple IPv6 seed address data sets respectively; without manually labeling the data, the seed IPv6 address data generates the target data by itself. For example, for an IPv6 address data 240e:06a0:0010:012e:0000:0001:0000:0001, after processing, the training data input to the model is <bos>240e06a00010012e0000000100000001, while the automatically generated target data is 240e06a00010012e0000000100000001 <eos>The model loss function selects cross entropy loss, the optimizer selects Adam, the learning rate is adjusted adaptively, and the first-order moment (mean) and second-order moment (variance) of the gradient are used to update the parameters.

[0074] When training the model, the cross-entropy loss between the output of the model's linear output layer and the target data is calculated, and this error is used to adjust the model parameters. After multiple rounds of training iterations, the neural network is finally trained into a model that can generate IPv6 addresses.

[0075] Generating candidate IPv6 addresses using the trained neural network model described in step 2 includes:

[0076] Calculate the IPv6 address bit by bit to generate the initial candidate IPv6 address;

[0077] De-duplicate the initial candidate IPv6 addresses to obtain the final candidate IPv6 addresses.

[0078] During model training, an IPv6 address is calculated once. However, IPv6 address generation is different. It is generated bit by bit in an autoregressive manner. That is, to generate an IPv6 address, the model needs to calculate 31 times.

[0079] According to RFC 4291, the IPv6 global unicast address range is 2000:: / 3, that is, all IPv6 global unicast addresses begin with 2 or 3. However, all the IPv6 addresses actually collected begin with 2. Therefore, when generating the IPv6 predicted address, this embodiment generates subsequent characters one by one starting with 2. The steps are as follows:

[0080] (1) <bos>2 Input to the model, assuming that the softmax output of this step model is x0 (x represents any character), intercept the last bit of the output and splice it back to the input to get <bos>20;

[0081] (2) <bos>20 is input to the model. Assuming that the softmax output of this step model is xx0, the last digit of the output is intercepted and then spliced ​​to the input to obtain <bos>200;

[0082] (3) Repeat the above steps until 31 characters are generated. Cut the last bit of the softmax output and splice it into the input. Assume that the result is <bos>20014c480200e30000000000000000001, remove the start mark <bos>, and add a colon separator to generate the IPv6 address 2001:4c48:0200:e300:0000:0000:0000:0001.

[0083] The generation of candidate addresses has a certain degree of randomness, so the generated addresses may be repeated. In order to remove duplicates from the generated target addresses, this embodiment uses an ordered set data type OrderedSet, which can not only remove duplicates from the generated addresses, but also maintain the order of the generated addresses, making it easier to verify the indicator differences of the addresses generated successively. The pseudo code for generating candidate addresses is as follows: Figure 15 shown.

[0084] Step 3: Perform active address detection on the generated candidate IPv6 addresses to obtain an active target IPv6 address.

[0085] In this embodiment, due to the characteristics of its generative neural network model, the generated IPv6 addresses must contain addresses that are repeated with the seeds and possible aliased prefixes. Therefore, before performing activity detection, it is necessary to first remove the addresses that are repeated with the seeds, and then use the alias prefixes collected by existing technology to match and remove the alias prefix addresses.

[0086] When detecting active addresses, most IPv6 addresses that respond to protocols like TCP and UDP also respond to ICMPv6. Therefore, to reduce the number of probe packets and the impact of address scanning on the network, you can use ICMPv6 to detect address activity. Use the IPv6-enabled zmap tool for scanning. Control the probe packet rate to 20k pps, with an average transmission rate of approximately 15Mbps.

[0087] The experimental environment of this embodiment is a CentOS 7.9 server with two Intel(R) Xeon(R) Silver 4210 CPUs (2.20GHz), 128GB of memory, and one NVIDIA A10 GPU installed.

[0088] The IPv6 seed addresses used in the experiment were collected from Gasser, a dataset of active addresses after removing alias prefixes. This dataset, released on February 22, 2025, contains 43.28M addresses. Because IP address activity varies across time and space, probing the same IP address at different times and addresses can yield completely different results. Therefore, before the experiment, we performed an activity probe on the seed addresses, retaining only those active IPv6 seed addresses that responded.

[0089] After ICMPv6 protocol data detection, this dataset has a total of 18.62M active IPv6 addresses. Long-term ICMPv6 protocol detection of the dataset found that the active rate (hit rate) of seed addresses gradually decreases over time. Figure 4 shown.

[0090] The PyASN tool was also used to analyze the proportion of network prefixes and autonomous systems in the data. Among the 18.62M active IPv6 addresses, there are 61.41K network prefixes (BGP prefixes) and 20.29K autonomous systems (AS). The address distribution is extremely uneven, with the top 4 ASes accounting for more than 50% of the addresses. The proportion of network prefixes is relatively even, with the top 12 prefixes accounting for more than 50% of the addresses. Figure 5 and Figure 6 The proportions of the top 10 ASes and BPG prefixes are shown respectively.

[0091] To compare the uniformity of the distribution of network prefixes and autonomous systems in IPv6 address sets of different sizes, this embodiment proposes a calculation formula for the uniformity U similar to the entropy value:

[0092]

[0093] Where n represents the total number of addresses in the dataset rather than the number of network prefixes (or autonomous systems), and p i The uniformity value (U) represents the proportion of addresses in each network prefix (or autonomous system). It ranges from 0 to 1, with larger values ​​indicating more uniform distribution. U = 0 is the minimum when the address set contains only one network prefix (or autonomous system); U = 1 is the maximum when the number of addresses in the data set is equal to the number of network prefixes (or autonomous systems).

[0094] According to the above formula, the network prefix uniformity U of the active addresses in the dataset is prefix =0.30, autonomous system uniformity U asn =0.23, which shows that the uniformity of network prefixes in the dataset is slightly higher than that of autonomous systems.

[0095] To analyze the impact of different numbers of seed addresses on the target generation algorithm, the experiment downsampled the dataset to collect three address sets of 10K, 100K, and 1M, namely S1, S2, and S3. The AS and prefix uniformity shown in Table 1 shows that as the number of samples increases, the distribution of AS and prefix approaches that of the dataset.

[0096] Table 1 Seed set characteristics

[0097]

[0098] In the neural network model of this embodiment, the model width d model , number of decoder layers N, internal dimension of feedforward network d ff , number of attention taps h, discard rate P dropt , Softmax temperature T and top_k are all hyperparameters and need to be adjusted according to the actual situation of the training data set. The default hyperparameters N = 6, d model = 512, d ff = 2048, h= 6, P drop = 0.1, T = 1, top_k = 5.

[0099] During the neural network model training process of this embodiment, the stochastic gradient descent (SGD) algorithm is used to update parameters. The learning rate and batch size are two important parameters that affect model convergence. In this experiment, the training parameters batch_size = 64 and learning_rate = 5e-5 are selected.

[0100] In the field of deep learning, maintaining the same total number of gradient updates (i.e., num_updates = epochs × (training_set_size / batch_size)) is crucial to maintaining equivalent training results. Because the data volumes of address sets S1, S2, and S3 differ, the number of training epochs (epochs) should be increased for smaller datasets. A simple experiment selected 200, 160, and 30 epochs for training address sets S1, S2, and S3, respectively. Ten intermediate results were retained during training, and 100,000 candidate addresses were generated each time to verify the model's effectiveness, primarily evaluating the hit rate metric. After removing duplicate addresses and alias addresses from the generated candidate address set C, activity detection was performed to obtain the active target address set T. The hit rate is defined as the ratio of the number of active addresses to the number of candidate addresses, as shown in the following formula:

[0101] .

[0102] Obviously, the hit rate is affected by the address duplication rate and the proportion of alias addresses. A higher address duplication rate and the proportion of alias addresses will lead to a lower hit rate. A higher hit rate indicates a higher target address generation effect.

[0103] 1. Experiment on the number of decoder layers N in the neural network model of this embodiment

[0104] In order to verify how many layers of models can achieve better results, we set N = [1, 3, 6, 9, 12] and conduct experiments on three seed address datasets S1, S2, and S3 respectively. The hit rates of candidate addresses generated by different models are compared. The results are as follows: Figure 7 As shown in the figure, the horizontal axis represents training epochs, and the vertical axis represents hit rate. The above experiments show that a model with one decoder layer is too simple and insufficient to learn the patterns of IPv6 addresses in the dataset, resulting in an underfitting state. As the number of decoder layers in the model increases, the model becomes increasingly complex, requiring a corresponding amount of data. For dataset S1, when epochs exceed 100, the model overfits, resulting in a lower hit rate and an increased repetition rate during address generation, which in turn increases the time it takes to generate candidate addresses. This is particularly true for the 9- and 12-layer models, where address generation is interrupted in the late stages of training due to excessively long training times. Experimental data shows that a model with 6 decoder layers is the most suitable for datasets S1, S2, and S3.

[0105] 2. Model width d of the neural network model in this embodiment model experiment

[0106] In this experiment, we further verified the impact of model dimension on model performance and set d model = [128, 256, 512, 1024, 2048] were tested to compare the hit rates of candidate addresses generated by different models. The results are as follows Figure 8 As shown, for the address sets S1 and S2 with relatively small amounts of data, an excessively large d model Will cause the model to overfit. For these two data sets, choose d model = 256; for S3 with large data volume, select d model = 512 hit rate is higher.

[0107] 3. The internal dimension d of the feedforward network of the decoder layer in the neural network model of this embodiment ff experiment

[0108] The hidden layer dimension d of the feedforward network of the test model in this experiment ff Impact on model performance, for S1, S2, S3 datasets, setting the model width d model The parameters are 256, 256, and 512 respectively. Find the best d in [512, 1024, 2048, 4096, 8192]. ff Parameters. The experimental results are as follows Figure 9 As shown, different d ff Parameters have little effect on the performance of the model. For the S1, S2, and S3 datasets, select the d ff The parameters are 512, 512, 2048.

[0109] 4. Experiment on the number of attention taps h

[0110] The number of model attention taps does not change the total number of model parameters, but the way of calculating attention is different. The more taps there are, the more complex the model calculation will be. Therefore, it will also have a certain impact on the performance of the model. model and d ff The parameter further verifies the impact of the number of attention taps on the model performance. Experiments are performed with h = [2, 4, 8, 16, 32] respectively. The results are as follows: Figure 10 As shown in the figure, when the number of seed addresses is small, the difference in the number of attention taps has little effect on model performance. As the amount of training data increases, the impact of the number of attention taps on model performance becomes more and more obvious. For the S1 dataset, h = 4 has the best hit rate, while for the S2 and S3 datasets, h = 8 has the most stable hit rate.

[0111] 5. Discard rate P drop experiment

[0112] In the process of network model training, Dropout is a commonly used regularization technique, which is mainly used to prevent overfitting of neural networks and improve the generalization ability of the model. In order to verify how large the dropout probability is more suitable for the three seed sets S1, S2, and S3, set P drop = [0, 0.05, 0.1, 0.15, 0.2, 0.25, 0.3] were tested on three data sets respectively, and other hyperparameters were taken as the optimal values ​​according to the previous experiments. The experimental results are as follows Figure 11 As shown in Figure 2, similarly, for seed sets with a small amount of data, the size of the dropout rate has little effect on the model; as the amount of data increases, the impact on model performance becomes more obvious. For the three data sets, all P drop = 0.05 is optimal.

[0113] 6. Normalized layer temperature T experiment

[0114] The Softmax Temperature is a hyperparameter used to adjust the smoothness of the probability distribution. It is mainly used to control the "confidence" or "diversity" of the model output. For high temperature parameters, the difference in probability between different categories decreases, and the model is less certain about the prediction results, that is, the output diversity increases; for low temperature parameters, the difference in probability between different categories increases, the probability distribution becomes more "concentrated", and the model is more certain about the prediction results. To verify the impact of temperature on model performance, the optimal hyperparameters from previous experiments were set to T = [0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0] and tested on three datasets, generating 10M candidate addresses respectively. The experimental results are as follows: Figure 12 As shown in Table 2, the horizontal axis represents budget (number of generated candidate addresses), in M, and the vertical axis represents hit rate. It can be seen that the normalization layer temperature has a crucial impact on model performance. When the number of generated addresses is small, the lower T, the higher the hit rate of generated candidate addresses. However, when the number of generated addresses is large, the hit rate drops rapidly at low T. Furthermore, since low T also reduces diversity, the repetition rate of generated addresses increases, which means that the time to generate unique addresses increases (see Table 2). For a budget of 10M, T = 0.5 is a good choice for balancing hit rate and generation time for the three address sets S1, S2, and S3.

[0115] Table 2 Time to generate 10M candidate addresses at different temperatures (hours: minutes: seconds)

[0116]

[0117] 7. Parameter top_k experiment

[0118] top_k is usually used together with the normalization layer temperature to control the diversity of the model output. The larger the top_k, the more tokens are likely to be selected, the greater the randomness of the generated addresses, and the richer the address diversity. Conversely, the smaller the top_k, the fewer tokens are selected, the more certain the generated addresses are, the higher the address repetition rate is, and the time to generate a specified number of non-repeating addresses will also increase. This experiment uses the optimal model parameters trained by the three seed sets S1, S2, and S3 above, and generates 10M addresses with top_k = [7, 9, 11, 13, 15, 16] respectively. The results are as follows: Figure 13 As shown in the figure, the smaller the top_k is, the higher the hit rate is. Experiments show that top_k = 16 is the best for the three seed sets S1, S2, and S3, and has the highest average hit rate.

[0119] To verify the performance of the method proposed in this example, we used the optimal parameters obtained in the above experiments (see Table 3) to compare with five open-source methods: Entropy / IP, 6GCVAE, 6Graph, 6Sense, and Addriner-S. 6Graph only generates target addresses with a Hamming distance of 1 from the seed address, resulting in a very small number of generated target addresses. This experiment improved 6Graph by prioritizing target addresses with a Hamming distance of 1. If the number of addresses is insufficient, it then generates target addresses with Hamming distances of 2, 3, and so on. 6Sense, likely due to different operating environments, did not achieve the performance described in the paper. Addriner-S is an address generation method that uses a sufficient number of seed addresses. 6VecLM and 6GAN were not included in the comparison due to their long runtimes.

[0120] Table 3 Optimal hyperparameters of the model proposed in this example

[0121]

[0122] Figure 14 The hit rates of candidate addresses generated by different methods on address sets S1-S3, after removing duplicate seed addresses and alias prefix addresses, are obtained through ICMPv6 probing. It is clear that the hit rates of most methods increase with the number of seed addresses, while they decrease with the number of generated candidate addresses. Among the six methods, the hit rate of the method in this embodiment is consistently higher than the others, with hit rates ranging from 56% to 74% for candidate addresses generated on the three seed sets S1-S3.

[0123] This embodiment proposes an effective IPv6 address detection method based on a neural network model. This method treats IPv6 addresses as text and uses seed addresses to train the network through self-supervision. This method achieves a higher hit rate than existing methods. In the future, seed addresses can be manually labeled with their address patterns or domains, allowing the network model to generate desired address patterns or IPv6 addresses in certain domains, further serving network security research.

[0124] In a specific implementation, the present application provides a computer storage medium and a corresponding data processing unit, wherein the computer storage medium is capable of storing a computer program that, when executed by the data processing unit, executes the invention disclosure of an efficient IPv6 address detection method based on a neural network model provided by the present invention, as well as some or all of the steps in each embodiment. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0125] Those skilled in the art will clearly understand that the technical solutions in the embodiments of the present invention can be implemented using computer programs and their corresponding general-purpose hardware platforms. Based on this understanding, the technical solutions in the embodiments of the present invention, or the portion that contributes to the prior art, can be embodied in the form of a computer program, i.e., a software product. This computer program software product can be stored in a storage medium and includes a number of instructions for enabling a device including a data processing unit (such as a personal computer, server, single-chip microcomputer, MUU, or network device) to execute the methods described in various embodiments of the present invention or certain portions of these embodiments.

[0126] The present invention provides an efficient IPv6 address detection method based on a neural network model. There are many methods and approaches to implement this technical solution. The above is only a preferred embodiment of the present invention. It should be noted that those skilled in the art may make various improvements and modifications without departing from the principles of the present invention, and such improvements and modifications should also be considered within the scope of protection of the present invention. Any components not specified in this embodiment may be implemented using existing technologies.< / bos> < / bos> < / bos> < / bos> < / bos> < / bos> < / eos> < / bos> < / eos> < / bos> < / eos> < / bos> < / eos> < / bos> < / eos> < / bos>

Claims

1. An efficient IPv6 address detection method based on a neural network model, characterized in that: The following steps are involved: Step 1: Collect IPv6 seed addresses, pre-process the IPv6 seed addresses, and obtain multiple IPv6 seed address data sets; Step 2: constructing a neural network model and training the neural network model using the multiple IPv6 seed address data sets; generating candidate IPv6 addresses using the trained neural network model; Step 3: Perform active address detection on the generated candidate IPv6 addresses to obtain an active target IPv6 address.

2. The method of claim 1, wherein the method comprises: Step 1 includes: Perform activity detection on the collected IPv6 seed addresses to obtain responsive active IPv6 seed addresses; Downsampling the active IPv6 seed addresses to obtain multiple IPv6 seed address sets with different sampling numbers; The IPv6 seed addresses in each IPv6 seed address set are expanded, the colon is removed, a start identifier and an end identifier are added to the beginning and end of the IPv6 seed address respectively to obtain a character string, each character of the character string is converted into an integer identifier to obtain an IPv6 seed address data vector, and then multiple IPv6 seed address data sets are obtained.

3. The efficient IPv6 address detection method based on a neural network model according to claim 2, characterized in that: The neural network model described in step 2 includes an embedding layer, N decoder layers, a linear output layer, and a normalization layer. , the embedding layer is used to convert the IPv6 seed address data vector in the IPv6 seed address dataset into an IPv6 seed address two-dimensional tensor; The decoder layer is used to extract hierarchical features of the output sequence of the embedding layer during training and to autoregressively generate the target address sequence during inference; The linear output layer is used to map the high-dimensional vector output by the decoder to a dimension with the same size as the IPv6 address encoding dictionary; The normalization layer is used to convert the linear layer output into the probability of occurrence of each character in the IPv6 address encoding dictionary, and then randomly select a character as the output from the top_k characters with the highest probability using the probability as the weight. .

4. The efficient IPv6 address detection method based on a neural network model according to claim 3, characterized in that: The decoder layer in step 2 consists of two identical masked multi-head attention layers and a feedforward network layer. The layers are connected with residual connections and layer normalization, and the activation function of each layer is Gelu. The input and output dimensions of the masked multi-head attention layer are both the model width, and the input and output dimensions of the feedforward network layer are both the model width.

5. The efficient IPv6 address detection method based on a neural network model according to claim 4, characterized in that: Each bit of the output vector of the linear output layer in step 2 corresponds to a character in the IPv6 address encoding dictionary. During model training, the output of the linear output layer is directly used to calculate the cross entropy loss with the target data to adjust the model parameters. During the inference phase, the output of the linear output layer is also processed in conjunction with the normalization layer to generate a candidate IPv6 address.

6. The efficient IPv6 address detection method based on a neural network model according to claim 5, characterized in that: During the inference phase, the normalization layer described in step 2 converts the output vector of the linear output layer into a probability distribution corresponding to the probability of occurrence of each character in the IPv6 address encoding dictionary. The output probability distribution is controlled by the temperature parameter T, and one of the top_k characters with the largest probability is randomly selected as the generated character using the probability as the weight to generate the candidate IPv6 address.

7. The efficient IPv6 address detection method based on a neural network model according to claim 6, characterized in that: In step 2, the neural network model is trained using the multiple IPv6 seed address data sets, including: adopting a self-supervised training mode to train the multiple IPv6 seed address data sets separately; there is no need to manually label the data, and the target data is generated by the IPv6 seed address data itself; the model loss function selects the cross entropy loss, the optimizer selects Adam, and the learning rate is adaptively adjusted, and the first-order moment and second-order moment of the gradient are used to update the parameters.

8. The efficient IPv6 address detection method based on a neural network model according to claim 7, characterized in that: Generating candidate IPv6 addresses using the trained neural network model described in step 2 includes: Calculate the IPv6 address bit by bit to generate the initial candidate IPv6 address; De-duplicate the initial candidate IPv6 addresses to obtain the final candidate IPv6 addresses.

9. The efficient IPv6 address detection method based on a neural network model according to claim 8, characterized in that: Step 3: Before performing active address detection on the generated candidate IPv6 addresses, remove the addresses that are duplicated with the IPv6 seed addresses and remove the alias prefix addresses from the final candidate IPv6 addresses.

10. The efficient IPv6 address detection method based on a neural network model according to claim 9, characterized in that: The model width, number of decoder layers N, and internal dimension d of the feedforward network of the decoder layer in the neural network model described in step 2 ff , the number of attention taps h, the dropout rate, the normalization layer temperature parameter and top_k are all hyperparameters and are adjusted according to the actual situation of the training dataset.

Citation Information

Patent Citations

  • IPv6 address generation model creation method and device and address generation method

    CN110809066A

  • Active IPV6 address prediction method based on deep learning

    CN115422914A

  • Seedless area address detection method and device

    CN118250258A

  • IPv6 address prediction method and system based on double space trees

    CN119603043A

  • IPv6 network space active address and port detection method and device

    CN119629080A

Cited By

  • Large-scale IPv6 network node detection method based on self-learning algorithm

    CN121284005A

  • A large-scale ipv6 network node detection method based on self-learning algorithm

    CN121284005B

  • IPv6 asset efficient discovery method and system with unified address and port mapping

    CN121567675A