Remote autonomous domain alias analysis method and system based on IPID sequence reconstruction
Through the IPID sequence reconstruction method, the problem that the prior art cannot effectively handle the IPID sequence of multi-core routers is solved, and alias analysis of single-core and multi-core routers is realized, which improves accuracy and efficiency.
Patent Information
- Application Number
- CN202510017642.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-06
AI Technical Summary
The existing IPID-based alias resolution method cannot effectively process the IPID sequence of multi-core routers, and there are problems with high false alarm rates and low efficiency.
A method of alias in a remote autonomous domain based on IPID sequence reconstruction is proposed. By obtaining the target IP address of the backbone router in the target autonomous domain, probing it to obtain the detection feature set, data preprocessing is performed to obtain the feature vector, and clustering the IP addresses according to the feature vector and the preset clustering rules to realize the identification of alias.
It realizes alias analysis of single-core routers and multi-core routers while ensuring accuracy and efficiency, which significantly improves the coverage and efficiency of alias analysis.
Smart Images

Figure CN119946016A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of alias resolution, and particularly to an alias resolution method and system within a remote autonomous domain based on IPID sequence reconstruction. Background Art
[0002] There are many implementation schemes for alias resolution around IPID. The core principle of this technology is that different interfaces of the same router usually share an IPID counter. Therefore, if two IP addresses show a consistent IPID growth pattern, they are very likely to belong to different interfaces of the same router.
[0003] This technology was initially implemented in Ally. To check whether two addresses A and B are aliases of the same router, Ally first sends a probe to A, waits for 1 ms, then sends a probe to B, and records the IPIDs of the received response packets as A1 and B1 respectively. Ally determines whether A1 - 10 < B1 < A1 + 100 is satisfied. If it is satisfied, Ally waits for 400 ms, then sends a probe to B, waits for 1 ms, and then sends a probe to A to get A2 and B2. If the second set of probes satisfies a similar condition, Ally regards A and B as aliases.
[0004] RadarGun improved the method of Ally. It improves scalability by probing a whole set of addresses in parallel instead of a series of pairs of addresses. If two addresses are aliases of each other, the IPID values are close when there is time overlap. RadarGun performs a distance test on the time series of a series of addresses. If their distance is less than a certain threshold, they are considered aliases.
[0005] Keys et al. proposed MIDAR, which is also considered the most classic alias resolution method so far. The focus of MIDAR is the monotonic boundary test MBT. MIDAR constructs a time series by sampling IPIDs, and then separately checks whether each element of the time series A and B satisfies the monotonicity of the sample, so as to prove whether the whole time series meets the monotonicity requirement. The monotonic boundary test determines whether two IPs are aliases by alternately sending probe packets to two IP addresses and detecting whether the IPID sequence of the response probe packets forms a monotonically increasing sequence. To achieve the scalability of probing, MIDAR first sorts the sample time series in descending order of IPID growth rate, and then realizes multi-round incremental probing of the target list through a sliding window. MIDAR performs a complete alias resolution through an estimation phase, a discovery phase, an elimination phase, and a confirmation phase.
[0006] Ally has a high false alarm rate because it only relies on 4 samples; RadarGun uses a manually set threshold to judge aliases, so the growth rate for different IPIDs is not universal, and there is a high false alarm rate on large-scale data sets; MIDAR, as the most classic alias resolution method, has a high accuracy rate. However, MIDAR's monotonic boundary test cannot be applied to multi-core routers, which limits its coverage. MIDAR uses a sliding window to implement multiple rounds of incremental detection of the target IP list, which requires a long detection time under a large-scale address set, limiting its efficiency. Existing technologies only perform alias resolution for a given target address set, and there is no remote autonomous domain alias resolution system that performs alias resolution for routers within a target autonomous system. Summary of the invention
[0007] The present invention aims to solve one of the technical problems in the related art at least to a certain extent.
[0008] In order to obtain accurate Internet router-level topology, alias resolution is crucial. Alias resolution is to cluster IPs and identify which IPs belong to the same device. With the expansion of the scale of the Internet and the iterative update of network infrastructure, multi-core routers account for an increasingly larger proportion of the Internet. Existing methods for alias resolution based on IPID are usually judged based on the monotonicity of IPID, but the IPID sequence of a multi-core router is a mixture of multiple monotonic IPID sequences, and existing methods cannot perform alias resolution of multi-core routers. In addition, existing measurement-based methods consume more time, and the efficiency of alias resolution needs to be improved. To this end, the present invention proposes a remote autonomous domain alias resolution method based on IPID sequence reconstruction, which can achieve alias resolution of single-core routers and multi-core routers while ensuring accuracy and efficiency.
[0009] Another object of the present invention is to propose a remote autonomous domain alias resolution system based on IPID sequence reconstruction.
[0010] To achieve the above object, the present invention proposes, on one hand, a remote autonomous domain alias resolution method based on IPID sequence reconstruction, comprising:
[0011] Obtain the target IP address of the backbone router in the target autonomous domain;
[0012] Perform detection based on the acquired target IP address to obtain a detection feature set of the target IP;
[0013] Performing data preprocessing on the detection feature set to obtain a feature vector;
[0014] The IP addresses in the same batch are clustered according to the feature vector and a preset clustering rule to obtain an alias cluster set.
[0015] The remote autonomous domain alias resolution method based on IPID sequence reconstruction according to the embodiment of the present invention may also have the following additional technical features:
[0016] In one embodiment of the present invention, obtaining a target IP address of a backbone router in a target autonomous domain includes:
[0017] Obtain all IPv4 address spaces under the target autonomous system to obtain an IPv4 address block;
[0018] Performing Traceroute detection on the router IP under the / 28 subnet of the IPv4 address block to obtain traceroute detection results; wherein the traceroute detection results include an IP-level topology constructed based on all intermediate router IPs in the path from the observation point to the target network;
[0019] Extract all IP addresses in the path from the traceroute detection results to filter out non-router IPs;
[0020] Based on the results of filtering non-router IPs and using the characteristics of back-to-back connections of routers, the IP interfaces of adjacent routers located in the same / 30 subnet are found to infer potential IP addresses.
[0021] In one embodiment of the present invention, detection is performed based on the acquired target IP address to obtain a detection feature set of the target IP, including:
[0022] Sort the target IP addresses according to the TTL labels to divide the IPs into multiple batches;
[0023] For IPs in the same batch, start sending probe packets from the same time point, send multiple probe packets at the same sending interval, and record the time sequence of sending probe packets;
[0024] Collect the TTL and IPID sequence of the response probe packet and the time sequence of the received probe packet.
[0025] In one embodiment of the present invention, data preprocessing is performed on the detection feature set to obtain a feature vector, including:
[0026] The timestamp sequence when the target IP generates the response probe packet is estimated through the time sequence of sending the probe packet and the time sequence of receiving the probe packet;
[0027] Using the estimated timestamp sequence and IPID sequence, calculate the IPID growth rate at each time point to obtain the IPID growth rate sequence;
[0028] The highest 5% and the lowest 5% of the values in the IPID growth rate sequence are deleted to obtain the average growth rate and the growth rate variance.
[0029] In one embodiment of the present invention, clustering IP addresses in the same batch according to the feature vector and a preset clustering rule to obtain an alias cluster set includes:
[0030] For each pair of IPs in the same batch, the similarity is defined as the ratio of the smaller value of two non-negative values divided by the larger value, and the proximity of the first IP ID of similar IPs is determined;
[0031] Reconstruct the IPID sequence of the IP pairs determined by similarity and proximity, bind the elements in each IPID sequence and its corresponding timestamp into a tuple, merge and sort the tuples of the two sequences by timestamp, obtain the merged IPID sequence, and count the number of negative increments of the merged IPID sequence; if the number of negative increments of the merged sequence does not exceed the number of negative increments of the single sequence before the merge, the IP pair is determined to be an alias;
[0032] Iterate multiple rounds of scanning. The target IP for the next batch of scanning is a representative IP randomly selected from each cluster and divided into different batches for scanning according to TTL sorting. The clustering results of the previous round are merged with the clustering results of this round. If the difference between the number of clusters after the merger and the number of clusters in the previous round does not exceed the preset threshold, the alias resolution process is terminated; otherwise, the iterative process continues until the termination condition is met.
[0033] To achieve the above object, the present invention further proposes a remote autonomous domain alias resolution system based on IPID sequence reconstruction, comprising:
[0034] An address acquisition module is used to obtain a target IP address of a backbone router in a target autonomous domain;
[0035] An active detection module is used to detect based on the acquired target IP address to obtain a detection feature set of the target IP;
[0036] A preprocessing module, used for performing data preprocessing on the detection feature set to obtain a feature vector;
[0037] The clustering module is used to cluster the IP addresses in the same batch according to the feature vector and a preset clustering rule to obtain an alias cluster set.
[0038] The alias resolution method and system in a remote autonomous domain based on IPID sequence reconstruction of the embodiment of the present invention can accurately capture the IPID change pattern of the same device in different time windows through the IPID sequence reconstruction method. First, a preliminary division is made based on similarity and proximity, and then an accurate judgment is made in combination with IPID sequence reconstruction. Through iterative scanning and multiple rounds of clustering, the system can complete the alias resolution of a large-scale address set in a relatively short time. In each iteration, the system only further detects the representative IP in each cluster, which greatly reduces unnecessary repeated detection and improves overall efficiency.
[0039] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0041] Figure 1 is a flow chart of a method for resolving aliases in a remote autonomous domain based on IPID sequence reconstruction according to an embodiment of the present invention;
[0042] Figure 2 is a schematic diagram of IPID sequence reconstruction according to an embodiment of the present invention;
[0043] Figure 3 is an architectural diagram of alias resolution within a remote autonomous domain based on IPID sequence reconstruction according to an embodiment of the present invention;
[0044] Figure 4 It is a structural diagram of an alias resolution system within a remote autonomous domain based on IPID sequence reconstruction according to an embodiment of the present invention. DETAILED DESCRIPTION
[0045] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0046] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0047] The following describes a method and system for alias resolution within a remote autonomous domain based on IPID sequence reconstruction according to an embodiment of the present invention with reference to the accompanying drawings.
[0048] Figure 1 is a flow chart of a method for resolving aliases in a remote autonomous domain based on IPID sequence reconstruction according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0049] S1, obtain the target IP address of the backbone router in the target autonomous domain.
[0050] In one embodiment of the present invention, all IPv4 address spaces under the target autonomous system are obtained to obtain an IPv4 address block; a traceroute detection is performed on the router IP under the / 28 subnet of the IPv4 address block to obtain a traceroute detection result; wherein the traceroute detection result includes an IP-level topology constructed based on all intermediate router IPs in the path from the observation point to the target network; all IP addresses in the path are extracted from the traceroute detection result to obtain filtered non-router IPs; based on the result of filtering non-router IPs and utilizing the characteristics of back-to-back connection of routers, the IP interfaces of adjacent routers located in the same / 30 subnet are found to infer potential IP addresses.
[0051] Specifically, for the target autonomous system (AS), the present invention can obtain the IPv4 address space under it through Whois. For each IPv4 address block, traceroute detection is performed on the router IP under its / 28 subnet to obtain a comprehensive IP-level topology. For the traceroute results, the IPs in the path are extracted, and the end IPs are ignored, so as to obtain filtered non-router IPs. For these IPs, check whether they are in the address space of the target AS. Since some ASs have only one detection entrance, no matter how many observation points there are, the IP path obtained by traceroute is only the result of one direction. Based on the characteristics of back-to-back connections between adjacent routers, for adjacent routers, their interface IPs belong to the same / 30 subnet. Therefore, the present invention can expand the IP location obtained previously, find potential IP addresses based on the characteristics of back-to-back connections between routers, thereby improving the coverage of addresses.
[0052] S2, performing detection based on the acquired target IP address to obtain a detection feature set of the target IP.
[0053] In one embodiment of the present invention, the target IP addresses are sorted according to the TTL labels to divide the IPs into multiple batches; for the IPs in the same batch, probe packets are sent starting from the same time point, and multiple probe packets are sent at the same sending interval, and the time series of sending the probe packets is recorded; the TTL and IPID sequence of the response probe packets and the time series of the received probe packets are collected.
[0054] Specifically, the present invention sorts the target IPs according to the TTL labels, and divides the IPs into multiple batches, with the same number of IPs in each batch. Serial scanning is performed between batches, and parallel scanning is performed within batches. For the IPs in the same batch, detection packets are sent starting from the same time point, and multiple detection packets are sent at the same sending interval, and the time series of the sent data packets is recorded. The TTL of the response data packets, the IPID sequence, and the time series of the received data packets are collected. Note that due to the lack of strict time control, there is still a gap when the packet sending time in the same batch is accurate to the ms level, which also provides a basis for the subsequent sequence reconstruction of the present invention. After the parallel scanning of the IPs in one batch is completed, all IP addresses in the next batch are scanned in parallel until all IP addresses in the target address set are scanned.
[0055] S3, performing data preprocessing on the detection feature set to obtain a feature vector.
[0056] In one embodiment of the present invention, the timestamp sequence when the target IP generates a response number probe packet is estimated by using the time sequence of sending the probe packet and the time sequence of receiving the probe packet; the IPID growth rate at each time point is calculated using the estimated timestamp sequence and IPID sequence to obtain the IPID growth rate sequence; the highest 5% and the lowest 5% values in the IPID growth rate sequence are deleted to obtain the average growth rate and the growth rate variance.
[0057] Specifically, the present invention preprocesses the data obtained by S2 to make it a feature basis that can be used to determine the alias. IPID sequences with negative increments greater than 30% are excluded to exclude random IPID counters. Through the time sequence of sending data packets and the time sequence of receiving data packets, the present invention estimates the timestamp sequence when the target IP generates a response data packet. Through the timestamp sequence and the IPID sequence, the present invention can obtain the IPID growth rate sequence. In order to exclude the situation of loopback, core replacement, and sudden increase that produces mutations in the IPID sequence and triggers the extreme values of the IPID growth rate sequence, the present invention deletes the high 5% and low 5% values of the IPID growth rate sequence, and obtains the average growth rate and growth rate variance from the IPID growth rate sequence. The average growth rate can measure the growth rate of the IPID within the measured time window, and the growth rate variance can describe the stability of the IPID growth.
[0058] S4, clustering the IP addresses in the same batch according to the feature vector and the preset clustering rule to obtain an alias cluster set.
[0059] In one embodiment of the present invention, for each pair of IPs in the same batch, similarity is defined as the ratio of the smaller of two non-negative values divided by the larger value, and the first IPID of similar IPs is judged for proximity; the IP pairs that pass the similarity and proximity judgments are reconstructed for IPID sequences, the elements in each IPID sequence and their corresponding timestamps are bound into tuples, the tuples of the two sequences are merged and sorted by timestamps to obtain a merged IPID sequence, and the number of negative increments of the merged IPID sequence is counted; if the number of negative increments of the merged sequence does not exceed the number of negative increments of the single sequence before the merge, the IP pair is determined to be an alias; multiple rounds of scanning are performed iteratively, and the target IP for the next batch of scanning is a representative IP randomly selected from each cluster, and is divided into different batches for scanning according to TTL sorting, and the clustering results of the previous round are merged with the clustering results of this round. If the difference between the number of clusters after the merge and the number of clusters in the previous round does not exceed a preset threshold, the process of alias resolution is terminated; otherwise, the iterative process is continued until the termination condition is met.
[0060] Specifically, the goal of this step is to cluster the IPs in the same batch based on the given clustering rules according to the preprocessed data, that is, to determine whether they belong to aliases. Because only IPs belonging to the same batch are measured in the same time window, if they belong to the same device, their average growth rate and growth rate variance are similar.
[0061] The present invention defines similarity as the ratio of the smaller of two non-negative values divided by the larger value, and the similarity value is [0,1]. The greater the similarity, the closer the two IPs are in the law of IPID change. Judging by the ratio instead of a given threshold can reduce misjudgment or missed judgment caused by the difference in growth rate of different devices. If the average growth rate similarity of a pair of IPs is higher than 0.95 and the similarity of the growth rate variance is higher than 0.8, it means that the IPID change law of the two IPs is similar, and they may be aliases.
[0062] Afterwards, the present invention conducts proximity determination to the first IPID of the pair of IPs, ensuring that not only are their IPID variation rules similar, but their IPIDs also have proximity. If the first IPID difference of the pair of IPs is less than the smaller value of the difference of the first two IPIDs of the two IPs, then it is considered that the pair of IPs meets the proximity criterion. This is done based on the timestamp that generates the first IPID being earlier than the timestamp that generates the second IPID. Passing the similarity criterion and the proximity criterion only indicates that the pair of IPs is likely to be an alias pair, but for large-scale address sets, it is easy to cause misjudgment based on these two bases alone.
[0063] Therefore, the present invention uses the IPID sequence reconstruction method to make a more accurate judgment. If a pair of IPs belong to the same device, then within the same time window, the changes in their IPIDs are synchronized, which is reflected in the IPID sequence as the same number of negative increments for the two, denoted as nn. If this is met, the sequence of the pair of IPs is reconstructed. Each element in the IPID sequence and the element in its corresponding timestamp sequence are bound into a tuple, and the tuples of the two sequences are merged and sorted by timestamp to obtain a merged IPID sequence. The number of negative increments of the merged IPID sequence is counted, denoted as nn'. If nn'≤nn is satisfied, the pair of IPs is considered to be an alias. The principle of this process is as follows: Figure 2 As shown. The sequence length = 5 is an example for easy understanding. In practice, a longer sequence can be selected to ensure accuracy. In this way, the loopback and core switching phenomena for different IPID counter upper limits can be solved by only one judgment formula, which is simple and universal. For each pair of IPs in the same batch, the similarity judgment, proximity judgment and sequence reconstruction mentioned above are performed, and a cluster set belonging to the same batch is obtained.
[0064] In order to merge the results of different batches, the present invention iteratively performs multiple rounds of scanning, which can make the detection time exponential converge. The target IP for the next batch of scans is a representative IP randomly selected from each cluster. These representative IPs are sorted according to TTL, divided into different batches for scanning, and the above process is repeated. The clustering results of the previous round are merged with the clustering results of this round. If the difference between the number of clusters after the merger and the number of clusters in the previous round does not exceed a given threshold, it can be considered that the alias resolution process is terminated. Otherwise, continue the above iterative process until the termination condition is met.
[0065] The overall architecture of the alias resolution method in a remote autonomous domain based on IPID sequence reconstruction of the present invention is as follows: Figure 3 As shown, the present invention can achieve alias resolution of single-core routers and multi-core routers while ensuring accuracy and efficiency.
[0066] In summary, the beneficial effects of the present invention are as follows: the method based on IPID sequence reconstruction can make this solution easier to deploy and discover more alias IPs than the classic MIDAR; through experiments on public data sets, comparative experiments on campus networks and test experiments on real ASs, it is proved that the solution of the present invention can achieve higher coverage than the classic solution MIDAR while achieving high accuracy, and the time required is shorter.
[0067] According to the alias resolution method in a remote autonomous domain based on IPID sequence reconstruction according to an embodiment of the present invention, alias resolution of single-core routers and multi-core routers can be realized while ensuring accuracy and efficiency, and the time exponential convergence of alias resolution based on batched multi-round scanning can improve efficiency; and the judgment based on similarity and proximity can filter non-alias pairs in advance to improve efficiency.
[0068] In order to implement the above embodiment, Figure 4 As shown, this embodiment also provides a remote autonomous domain alias resolution system 10 based on IPID sequence reconstruction, including:
[0069] The address acquisition module 100 is used to acquire the target IP address of the backbone router in the target autonomous domain;
[0070] An active detection module 200 is used to perform detection based on the acquired target IP address to obtain a detection feature set of the target IP;
[0071] A preprocessing module 300, configured to perform data preprocessing on the detection feature set to obtain a feature vector;
[0072] The clustering module 400 is used to cluster the IP addresses in the same batch according to the feature vector and a preset clustering rule to obtain an alias cluster set.
[0073] Furthermore, the address acquisition module 100 is also used for:
[0074] Obtain all IPv4 address spaces under the target autonomous system to obtain an IPv4 address block;
[0075] Performing Traceroute detection on the router IP under the / 28 subnet of the IPv4 address block to obtain traceroute detection results; wherein the traceroute detection results include an IP-level topology constructed based on all intermediate router IPs in the path from the observation point to the target network;
[0076] Extract all IP addresses in the path from the traceroute detection results to filter out non-router IPs;
[0077] Based on the results of filtering non-router IPs and using the characteristics of back-to-back connections of routers, the IP interfaces of adjacent routers located in the same / 30 subnet are found to infer potential IP addresses.
[0078] Furthermore, the active detection module 200 is also used for:
[0079] Sort the target IP addresses according to the TTL labels to divide the IPs into multiple batches;
[0080] For IPs in the same batch, start sending probe packets from the same time point, send multiple probe packets at the same sending interval, and record the time sequence of sending probe packets;
[0081] Collect the TTL and IPID sequence of the response probe packet and the time sequence of the received probe packet.
[0082] Furthermore, the preprocessing module 300 is also used for:
[0083] The timestamp sequence when the target IP generates the response probe packet is estimated through the time sequence of sending the probe packet and the time sequence of receiving the probe packet;
[0084] Using the estimated timestamp sequence and IPID sequence, the IPID growth rate at each time point is calculated to obtain the IPID growth rate sequence;
[0085] The highest 5% and the lowest 5% of the values in the IPID growth rate sequence are deleted to obtain the average growth rate and the growth rate variance.
[0086] Furthermore, the clustering module 400 is also used for:
[0087] For each pair of IPs in the same batch, the similarity is defined as the ratio of the smaller value of two non-negative values divided by the larger value, and the proximity of the first IP ID of similar IPs is determined;
[0088] Reconstruct the IPID sequence of the IP pairs determined by similarity and proximity, bind the elements in each IPID sequence and its corresponding timestamp into a tuple, merge and sort the tuples of the two sequences by timestamp, obtain the merged IPID sequence, and count the number of negative increments of the merged IPID sequence; if the number of negative increments of the merged sequence does not exceed the number of negative increments of the single sequence before the merge, the IP pair is determined to be an alias;
[0089] Iterate multiple rounds of scanning. The target IP for the next batch of scanning is a representative IP randomly selected from each cluster and divided into different batches for scanning according to TTL sorting. The clustering results of the previous round are merged with the clustering results of this round. If the difference between the number of clusters after the merger and the number of clusters in the previous round does not exceed the preset threshold, the alias resolution process is terminated; otherwise, the iterative process continues until the termination condition is met.
[0090] According to the alias resolution system in a remote autonomous domain based on IPID sequence reconstruction according to the embodiment of the present invention, it is possible to accurately capture the IPID change rule of the same device in different time windows. The present invention can solve the judgment of single-core routers and multi-core routers at the same time. By analyzing the characteristics of multi-core routers, it is found that the IPID change rule is the superposition of multiple monotonically increasing IPID sequences, mainly based on the preliminary division based on similarity and proximity, and then combined with IPID sequence reconstruction for accurate judgment. The present invention utilizes a dual judgment mechanism of similarity and proximity, and combines IPID sequence reconstruction to significantly reduce the problem of misjudgment or missed judgment caused by the difference in growth rates of different devices; it can complete the alias resolution of large-scale address sets in a relatively short time. Taking into account the effective use of resources, especially when processing large-scale data, scanning is performed in a serial manner between batches and in parallel within batches, which makes full use of computing resources and shortens processing time.
[0091] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0092] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
Claims
1. A remote autonomous domain alias resolution method based on IPID sequence reconstruction, characterized in that: include: Obtain the target IP address of the backbone router in the target autonomous domain; Perform detection based on the acquired target IP address to obtain a detection feature set of the target IP; Performing data preprocessing on the detection feature set to obtain a feature vector; The IP addresses in the same batch are clustered according to the feature vector and a preset clustering rule to obtain an alias cluster set.
2. The method according to claim 1, characterized in that Obtain the target IP address of the backbone router in the target autonomous domain, including: Obtain all IPv4 address spaces under the target autonomous system to obtain an IPv4 address block; Performing Traceroute detection on the router IP under the / 28 subnet of the IPv4 address block to obtain traceroute detection results; wherein the traceroute detection results include an IP-level topology constructed based on all intermediate router IPs in the path from the observation point to the target network; Extract all IP addresses in the path from the traceroute detection results to filter out non-router IPs; Based on the results of filtering non-router IPs and taking advantage of the back-to-back connection characteristics of routers, the IP interfaces of adjacent routers located in the same / 30 subnet are found to infer potential IP addresses.
3. The method according to claim 2, characterized in that Based on the acquired target IP address, detection is performed to obtain the detection feature set of the target IP, including: Sort the target IP addresses according to the TTL labels to divide the IPs into multiple batches; For IPs in the same batch, start sending probe packets from the same time point, send multiple probe packets at the same sending interval, and record the time sequence of sending probe packets; Collect the TTL and IPID sequence of the response probe packet and the time sequence of the received probe packet.
4. The method according to claim 3, characterized in that Performing data preprocessing on the detection feature set to obtain a feature vector includes: The timestamp sequence when the target IP generates the response probe packet is estimated through the time sequence of sending the probe packet and the time sequence of receiving the probe packet; Using the estimated timestamp sequence and IPID sequence, calculate the IPID growth rate at each time point to obtain the IPID growth rate sequence; The highest 5% and the lowest 5% of the values in the IPID growth rate sequence are deleted to obtain the average growth rate and the growth rate variance.
5. The method according to claim 4, characterized in that The IP addresses in the same batch are clustered according to the feature vector and a preset clustering rule to obtain an alias cluster set, including: For each pair of IPs in the same batch, the similarity is defined as the ratio of the smaller value of two non-negative values divided by the larger value, and the proximity of the first IP ID of similar IPs is determined; Reconstruct the IPID sequence of the IP pairs determined by similarity and proximity, bind the elements in each IPID sequence and its corresponding timestamp into a tuple, merge and sort the tuples of the two sequences by timestamp, obtain the merged IPID sequence, and count the number of negative increments of the merged IPID sequence; if the number of negative increments of the merged sequence does not exceed the number of negative increments of the single sequence before the merge, the IP pair is determined to be an alias; Iterate multiple rounds of scanning. The target IP for the next batch of scanning is a representative IP randomly selected from each cluster and divided into different batches for scanning according to TTL sorting. The clustering results of the previous round are merged with the clustering results of this round. If the difference between the number of clusters after the merger and the number of clusters in the previous round does not exceed the preset threshold, the alias resolution process is terminated; otherwise, the iterative process continues until the termination condition is met.
6. A remote autonomous domain alias resolution system based on IPID sequence reconstruction, characterized in that: include: An address acquisition module is used to obtain a target IP address of a backbone router in a target autonomous domain; An active detection module is used to detect based on the acquired target IP address to obtain a detection feature set of the target IP; A preprocessing module, used for performing data preprocessing on the detection feature set to obtain a feature vector; The clustering module is used to cluster the IP addresses in the same batch according to the feature vector and a preset clustering rule to obtain an alias cluster set.
7. The system according to claim 6, characterized in that The address acquisition module is also used to: Obtain all IPv4 address spaces under the target autonomous system to obtain an IPv4 address block; Performing Traceroute detection on the router IP under the / 28 subnet of the IPv4 address block to obtain traceroute detection results; wherein the traceroute detection results include an IP-level topology constructed based on all intermediate router IPs in the path from the observation point to the target network; Extract all IP addresses in the path from the traceroute detection results to filter out non-router IPs; Based on the results of filtering non-router IPs and taking advantage of the back-to-back connection characteristics of routers, the IP interfaces of adjacent routers located in the same / 30 subnet are found to infer potential IP addresses.
8. The system according to claim 7, characterized in that The active detection module is also used to: Sort the target IP addresses according to the TTL labels to divide the IPs into multiple batches; For IPs in the same batch, start sending probe packets from the same time point, send multiple probe packets at the same sending interval, and record the time sequence of sending probe packets; Collect the TTL and IPID sequence of the response probe packet and the time sequence of the received probe packet.
9. The system according to claim 8, characterized in that The preprocessing module is also used to: The timestamp sequence when the target IP generates the response probe packet is estimated through the time sequence of sending the probe packet and the time sequence of receiving the probe packet; Using the estimated timestamp sequence and IPID sequence, calculate the IPID growth rate at each time point to obtain the IPID growth rate sequence; The highest 5% and the lowest 5% of the values in the IPID growth rate sequence are deleted to obtain the average growth rate and the growth rate variance.
10. The system according to claim 9, characterized in that Clustering modules are also used to: For each pair of IPs in the same batch, the similarity is defined as the ratio of the smaller value of two non-negative values divided by the larger value, and the proximity of the first IP ID of similar IPs is determined; Reconstruct the IPID sequence of the IP pairs determined by similarity and proximity, bind the elements in each IPID sequence and its corresponding timestamp into a tuple, merge and sort the tuples of the two sequences by timestamp, obtain the merged IPID sequence, and count the number of negative increments of the merged IPID sequence; if the number of negative increments of the merged sequence does not exceed the number of negative increments of the single sequence before the merge, the IP pair is determined to be an alias; Iterate multiple rounds of scanning. The target IP for the next batch of scanning is a representative IP randomly selected from each cluster. The IPs are divided into different batches for scanning according to TTL sorting. The clustering results of the previous round are merged with the clustering results of this round. If the difference between the number of clusters after the merger and the number of clusters in the previous round does not exceed the preset threshold, the alias resolution process is terminated. Otherwise, the iteration process continues until the termination condition is met.
Citation Information
Cited By
Router alias identification system
CN120750909A