A method for deduplication and aggregation of IP addresses
By using the ordered sequence of IP address segments to represent IP collection data, and perform binary search and aggregation operations, the problem of inefficient IP address dereplication and aggregation in the prior art is solved, and an efficient IP address dereplication and aggregation effect is achieved.
Patent Information
- Application Number
- CN202111534121.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-12-15
AI Technical Summary
In the prior art, when performing IP address dere-reconstruction, the use of subnet masks has problems such as inefficiency and unintuitiveness, especially when the IP address is missing, a large amount of fragmentation will be generated.
The ordered sequence of IP address segments is used to represent the IP set data. Each IP address segment stores the number of starting IP and continuous IP addresses. The position to be inserted is determined through a binary search algorithm, and the related union aggregation action is performed to realize the dereplication and aggregation of IP addresses.
This method can effectively improve the efficiency of IP address derecoration and aggregation, especially when IP addresses reach the order of tens of millions, and still maintain good performance and will not occupy too much memory space.
Smart Images

Figure CN114265867B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information technology, and in particular to an IP address deduplication aggregation method. Background Art
[0002] IP addresses (the IP addresses in this article only refer to IPv4 addresses) are composed of 32-bit binary numbers. Generally, IP addresses have the following expressions, such as single IP: 192.168.1.1, IP range: 192.168.1.1-255, IP range: 200.1.1.1-220.2.2.2, and mask form: 192.168.1.0 / 24. In applications such as network mapping, it is necessary to quickly deduplicate and aggregate IP address data sets composed of the above expressions. This operation requires attention to controlling time complexity and space complexity.
[0003] At present, IP deduplication aggregation generally adopts the subnet mask method. Aggregating IP in the form of subnet mask requires that all IP addresses in the IP segment must have the same prefix, and IP addresses cannot be missing. First of all, this method is not conducive to improving the aggregation effect. When a large network segment has individual IP missing, IP aggregation using subnet mask will generate a large number of fragments. Secondly, this expression is not intuitive and must be converted into binary data to better analyze related problems. Summary of the invention
[0004] In view of the defects in the prior art, the present invention provides an IP address deduplication aggregation method, which uses an ordered sequence of IP address segments to represent IP set data, and each IP address segment needs to store the starting IP and the number of consecutive IP addresses. Each IP address data to be processed is converted into an IP address segment, and a suitable position is selected in the ordered sequence of IP address segments to insert and perform related deletion and update actions to achieve deduplication aggregation.
[0005] An IP address deduplication aggregation method comprises the following steps:
[0006] S1: Obtain the IP address to be aggregated, convert the IP address to be aggregated into an address segment format, obtain the IP address segment to be aggregated, and execute step S2;
[0007] S2: Use a binary search algorithm to determine the position where the IP address segment to be aggregated is to be inserted in the IP address segment data container, and execute step S3;
[0008] S3: Search backward according to the position to be inserted, determine whether to perform a union aggregation according to the search result, and execute step S4;
[0009] S4: Search forward according to the position to be inserted, determine whether to perform secondary union aggregation according to the search result, and execute step S5;
[0010] S5: Determine whether to insert the IP address to be aggregated into the position to be inserted according to the first union aggregation result and the second aggregation result, and output the aggregated IP address segment data container.
[0011] Preferably, in step S2, the IP address segment data container includes two arrays, defined as As and Ac, the As array is used to store the starting IP address, and the Ac array is used to store the number of consecutive IP addresses including the starting IP address.
[0012] Preferably, in step S2, the IP address segment to be aggregated includes a starting IP address and the number of consecutive IP addresses including the starting IP address, which are expressed as Is and Ic.
[0013] Preferably, in step S2, the value of the next IP address adjacent to the IP address segment to be aggregated is determined according to the sum of the starting IP address and the number of consecutive IP addresses, which is denoted as Ive, wherein Is and Ive are numerical data.
[0014] Preferably, in step S3, the method of using a binary search algorithm to confirm the position to be inserted of the IP address segment to be aggregated in the IP address segment data container is: using a binary search algorithm in the As array according to the size of the Is value to determine the position i of the Is value in the As array.
[0015] Preferably, in step S4, the method of searching backward according to the position to be inserted and determining whether to perform a union aggregation according to the search result is:
[0016] Start searching from position i of array As and search backward until you find a number greater than Ive or reach the end of the array, and record the number of elements covered, which is recorded as C;
[0017] If C>0, it means that the IP address segment to be aggregated and the existing IP address segment can form a union of continuous IP address blocks. A union aggregation is performed on this continuous IP address block starting with Is, and the value of the i position in the array As is changed to Is, and the number of continuous IP addresses in the corresponding array Ac is changed to the size of the union of continuous IP address blocks.
[0018] Preferably, in step S5, the method of searching forward according to the position to be inserted and determining whether to perform secondary union aggregation according to the search result is: take the previous position of position i, recorded as ib, if there is data at position ib, determine whether the address segment represented by position ib contains or is adjacent to Is, if it contains or is adjacent to Is, perform secondary union aggregation on this overlapping area, the value of position ib in array As remains unchanged, and the number of consecutive IP addresses in the corresponding array Ac is changed to the size of the union of consecutive IP addresses.
[0019] Preferably, in step S6, it is determined whether to insert the IP address to be aggregated at the position to be inserted according to the result of the first union aggregation and the second union aggregation: if the first union aggregation and the second union aggregation are not performed, the IP address segment to be aggregated is inserted at position i in the array As, and the aggregated IP address segment data container is output; if the first aggregation or the second aggregation is performed, the aggregated IP address segment data container is directly output.
[0020] The beneficial effects of the present invention are as follows: the present invention uses an ordered sequence of IP address segments to represent IP set data, each IP address segment needs to store the starting IP address and the number of consecutive IP addresses including the starting IP address, converts the IP addresses to be aggregated into IP address segments to be aggregated, selects a suitable position in the ordered sequence of IP address segments to insert and execute related union aggregation actions, so as to achieve deduplication aggregation of IP addresses. Since the multiple public IP addresses provided by the operator are generally a continuous address segment, this method can achieve better results. When the IP address reaches the tens of millions level, the method can still maintain good performance and will not take up too much memory space. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for the specific embodiments or the description of the prior art. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn according to the actual scale.
[0022] Figure 1 This is a flow chart of Embodiment 1 of the present invention;
[0023] Figure 2 This is a flow chart of Example 2 of the present invention. DETAILED DESCRIPTION
[0024] The following embodiments of the technical solution of the present invention are described in detail in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and are therefore only used as examples, and cannot be used to limit the protection scope of the present invention.
[0025] It should be noted that, unless otherwise specified, the technical terms or scientific terms used in this application should have the common meanings understood by those skilled in the art to which the present invention belongs.
[0026] Example 1
[0027] refer to Figure 1 , an IP address deduplication aggregation method provided by an embodiment of the present invention comprises the following steps:
[0028] S1: Obtain the IP address to be aggregated, convert the IP address to be aggregated into an address segment format, obtain the IP address segment to be aggregated, and execute step S2;
[0029] S2: Use a binary search algorithm to determine the position where the IP address segment to be aggregated is to be inserted in the IP address segment data container, and execute step S3;
[0030] S3: Search backward according to the position to be inserted, determine whether to perform a union aggregation according to the search result, and execute step S4;
[0031] S4: Search forward according to the position to be inserted, determine whether to perform secondary union aggregation according to the search result, and execute step S5;
[0032] S5: Determine whether to insert the IP address to be aggregated into the position to be inserted according to the first union aggregation result and the second aggregation result, and output the aggregated IP address segment data container.
[0033] To summarize, the present invention uses an ordered sequence of IP address segments to represent IP set data. Each IP address segment needs to store the starting IP address and the number of consecutive IP addresses including the starting IP address. The IP addresses to be aggregated are converted into IP address segments to be aggregated. A suitable position is selected in the ordered sequence of IP address segments to insert and execute related union aggregation actions to achieve deduplication aggregation of IP addresses.
[0034] Example 2
[0035] refer to Figure 2 , an IP address deduplication aggregation method provided by an embodiment of the present invention comprises the following steps:
[0036] Step 1: Initialize two empty arrays, defined as As and Ac. The As array is used to store the starting IP address, which is stored as an unsigned long 32-bit numeric data. The Ac array is used to store the number of consecutive IP addresses including the starting IP address. The elements in the two arrays establish a one-to-one correspondence based on the subscripts to represent an IP address segment data. The array in As is always kept in order.
[0037] Step 2: Get an IP address to be aggregated in the data set to be processed, and convert the IP address into an IP address segment format, which is the starting IP address and the number of consecutive IP addresses including the starting IP address, expressed as Is, Ic. The current starting IP address + the number of consecutive IP addresses can be used to obtain the value of the next adjacent IP address outside the current IP address segment, recorded as Ive, where Is and Ive are numerical data;
[0038] Step 3: Use the binary search algorithm to determine the position i where it needs to be inserted in the As array according to the value of Is. If they are equal, take the position of the equal value, and the value corresponding to the i position is recorded as I;
[0039] Step 4: Search backward from position i of array As until a number greater than Ive is found or the end of the array is reached, and the number of elements covered is recorded, which is recorded as C; if C>0, it means that the IP segment data to be inserted and the existing IP segment data can form a union of continuous IP address blocks, and the continuous IP address block starting with Is is replaced by union aggregation, the value of position i in array As is changed to Is, and the number of continuous IP addresses in the corresponding array Ac is changed to the size of the union of continuous IP address blocks;
[0040] Step 5: Get the previous position of position i, that is, ib. If there is data at position ib, determine whether the address segment represented by position ib contains or is adjacent to Is. If so, perform aggregation replacement on this overlapping area. The value of position ib in array As remains unchanged, and the number of consecutive IP addresses in the corresponding array Ac is changed to the size of the union of consecutive IP address blocks.
[0041] Step 6: If no aggregation replacement action has been performed after the above steps, it means that the IP address segment to be aggregated does not overlap with the existing IP segment array in the data container, and the IP address segment to be aggregated is inserted at position i.
[0042] Step 7: Determine whether there is an IP address to be aggregated in the data set to be processed. If so, execute step 2. If not, output the aggregated data container.
[0043] In summary, the purpose of the present invention is to provide a new deduplication aggregation algorithm, which uses an ordered sequence of IP address segments to represent IP set data. Each IP address segment needs to store the starting IP and the number of consecutive IP addresses. Each IP address data to be processed is converted into an IP address segment, and a suitable position is selected in the ordered sequence of IP address segments to insert and perform related deletion and update actions to achieve deduplication aggregation. Since multiple public IP addresses provided by operators are generally a continuous address segment, this method can achieve better results. When the number of IP addresses reaches tens of millions, this method can still maintain good performance and will not take up too much memory space.
[0044] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. These modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be included in the scope of the claims and specification of the present invention.
Claims
1. A method for deduplication and aggregation of IP addresses, characterized in that: The following steps are involved: S1: Obtain the IP address to be aggregated, convert the IP address to be aggregated into an address segment format, obtain the IP address segment to be aggregated, and execute step S2; S2: Use a binary search algorithm to determine the position where the IP address segment to be aggregated is to be inserted in the IP address segment data container, and execute step S3; S3: Search backward according to the position to be inserted, determine whether to perform a union aggregation according to the search result, and execute step S4; S4: Search forward according to the position to be inserted, determine whether to perform secondary union aggregation according to the search result, and execute step S5; S5: judging whether to insert the IP address to be aggregated into the position to be inserted according to the first union aggregation result and the second union aggregation result, and outputting the aggregated IP address segment data container; In step S2, the IP address segment data container includes two arrays, defined as As and Ac, the As array is used to store the starting IP address, and the Ac array is used to store the number of consecutive IP addresses including the starting IP address; In step S1, the IP address segment to be aggregated includes a starting IP address and the number of consecutive IP addresses including the starting IP address, which are expressed as Is and Ic.
2. The IP address deduplication aggregation method according to claim 1, characterized in that: In step S1, the value of the next IP address adjacent to the IP address segment to be aggregated is determined according to the sum of the starting IP address and the number of consecutive IP addresses, which is recorded as Ive, where Is and Ive are numerical data.
3. The IP address deduplication aggregation method according to claim 2, characterized in that: In step S2, the method of using a binary search algorithm to confirm the position to be inserted of the IP address segment to be aggregated in the IP address segment data container is: using a binary search algorithm in the As array according to the size of the Is value to determine the position i to be inserted of the Is value in the As array.
4. The IP address deduplication aggregation method according to claim 3, characterized in that: In step S3, the method of searching backward according to the position to be inserted and determining whether to perform a union aggregation according to the search result is: Start searching from position i of array As and search backward until you find a number greater than Ive or reach the end of the array, and record the number of elements covered, which is recorded as C; If C>0, it means that the IP address segment to be aggregated and the existing IP address segment form the union of a continuous IP address block. A union aggregation is performed on this continuous IP address block starting with Is, and the value of the i position in the array As is changed to Is, and the number of continuous IP addresses in the corresponding array Ac is changed to the size of the union of the continuous IP address blocks.
5. The IP address deduplication aggregation method according to claim 4, characterized in that: In step S4, forward search is performed according to the position to be inserted, and the method for determining whether to perform secondary union aggregation is as follows: take the previous position of position i, recorded as ib. If there is data at position ib, determine whether the address segment represented by position ib contains or is adjacent to Is. If it contains or is adjacent to Is, secondary union aggregation is performed on this overlapping area. The value of position ib in array As remains unchanged, and the number of consecutive IP addresses in the corresponding array Ac is changed to the size of the union of consecutive IP addresses.
6. The IP address deduplication aggregation method according to claim 5, characterized in that: In step S5, it is determined whether to insert the IP address to be aggregated at the position to be inserted based on the results of the first union aggregation and the second union aggregation: if the first union aggregation and the second union aggregation are not performed, the IP address segment to be aggregated is inserted at position i in the array As, and the aggregated IP address segment data container is output; if the first union aggregation or the second union aggregation is performed, the aggregated IP address segment data container is directly output.
Citation Information
Patent Citations
IP address list storage and query method applied to DNS query
CN105635343A