Method and device for collecting large-scale IP address data

By converting IP addresses into four-dimensional sparse matrices and using TLMB and SSMB mapping mechanisms, the problem of excessive computing costs and memory usage in large-scale IP address data statistics is solved, and the rapid and effective determination of frequent IP addresses is achieved, which is suitable for network traffic management.

CN120281673APending Publication Date: 2025-07-08FOSHAN POLYTECHNIC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510296301.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

When collecting large-scale IP address data, the computing cost and memory usage are too high, making it difficult to effectively determine the frequently occurring IP addresses.

Method used

Two relationship mapping mechanisms are adopted between memory blocks and IP addresses, namely two-layer memory block architecture (TLMB) and one-layer memory architecture (SSMB). By converting IP addresses into four-dimensional sparse matrices, the mapping mechanism is used to store and count the number of IP addresses, and the frequency IP addresses are determined in combination with the minimum heap algorithm.

Benefits of technology

The balance between computational cost and memory usage is achieved, and the first k IP addresses that appear most frequently in IP addresses is achieved quickly and effectively, which is suitable for user behavior analysis in network traffic management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281673A_ABST
    Figure CN120281673A_ABST
Patent Text Reader

Abstract

The invention relates to the field of network data processing, discloses a method and a device for collecting large-scale IP address data, provides two efficient large-scale IP address statistical algorithms, and aims to balance time efficiency and memory consumption. According to the method, the sparsity characteristic of IP address statistics is fully considered, and optimization is achieved by dynamically adjusting the mapping relation between the layered memory blocks. In one method, a double-layer structure is adopted, and each layer is composed of a plurality of memory blocks with a fixed number. Each memory block comprises 256 elements, the size of a single element is 8 bytes, and the memory block is adaptive to a 64-bit system. Compared with a built-in hash table, the method has the advantages that hash conflicts are completely avoided, and meanwhile, the linear time complexity based on the hash method is kept. Further, a parallel optimization scheme is proposed to accelerate data statistics. Experimental results show that the time efficiency and the space efficiency of the method on synthesis and real data sets are obviously superior to those of a baseline algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network data processing, and in particular, to a method and apparatus for collecting large-scale IP address data. Background Art

[0002] With the rapid development of emerging network services such as video streams, instant messaging, and online payment services, network traffic has increased significantly. The evaluation of user behavior characteristics in network traffic management is crucial. User behavior characteristics are usually extracted from IP data packets containing IP addresses, and there is a close correspondence between user behavior characteristics containing a large amount of information and IP addresses frequently appearing in the data packets. Therefore, how to effectively obtain large-scale IP address data statistics within a few minutes is a challenging network traffic measurement problem. Obtaining large-scale IP address data statistics usually includes two tasks: calculating the number of occurrences of each IP address and sorting the results in a specific order.

[0003] In the prior art, many IP address data statistics methods have been studied; for example, a classic divide-and-conquer strategy is to first divide the IP addresses into multiple subsets; then, each subset is calculated separately using a statistical collection method; finally, the results of multiple subsets are merged and sorted. Sorting algorithms such as bubble sort, insertion sort, merge sort, selection sort, or quick sort, etc. However, in the sorting process, reducing the computational cost is a difficult problem. For example, the average complexity and worst-case complexity of bubble, insertion, and selection sort algorithms are O(n 2 ), where n represents the number of unsorted records. This indicates that the merging and sorting steps of multiple subsets occupy a large amount of memory, and when collecting statistical data of millions or tens of millions of IP addresses, the computational cost of executing this step is very high. Therefore, the statistical collection algorithm using the divide-and-conquer strategy still faces some challenges due to the rapid growth of large-scale records, such as limited memory and computational cost limitations.

[0004] Hash tables are an effective method for collecting IP address statistics. It uses a hash function to calculate the hash code of a bucket array and obtain the statistical results. The hash function assigns each key to a unique bucket for each IP address. The hash function can generate the same hash code for multiple IP addresses. With the increase in the amount of big data generated, millions or tens of millions of records are ubiquitous in network traffic. Therefore, this method may cause multiple hash collisions, especially for a large number of IP addresses. Although many strategies can be adopted to avoid collisions, such as linear probing, quadratic probing, and double hashing, they require additional storage space and computation.

[0005] Therefore, it is of great practical significance to develop a data statistics method that can balance the computational cost and memory usage, and then quickly and effectively determine the top k IP addresses that appear most frequently in the IP address set. Summary of the Invention

[0006] The purpose of the present invention is to provide a large-scale IP address data statistics method that can balance the computational cost and memory usage, and then quickly and effectively determine the top k IP addresses that appear most frequently in the IP address set.

[0007] To achieve the above objectives, the present invention adopts the following technical solutions.

[0008] A method for collecting large-scale IP address data, which converts the IP address set into a four-dimensional sparse matrix with the numerical values of the four parts of the IP address as matrix elements, and maps the four-dimensional sparse matrix to a memory block by using any one of the following two mapping mechanisms to achieve the storage of the IP address set.

[0009] 1) Adopt a two-layer memory block architecture. The first layer is pre-allocated continuous memory space with a dimension of 256 3 pointer slots, and the total space is 256³×8B = 128MB; the second layer is a counter block with 256×8B effective slots.

[0010] The numerical values of the first three parts of the IP address are mapped to the corresponding positions of the pointer slots in the pre-allocated continuous memory space of the first layer, and the statistical information of the corresponding IP address is stored in each effective slot in the second layer.

[0011] 2) Adopt a one-layer memory architecture, which is pre-allocated shared continuous memory space with a dimension of 256³ pointer slots, and the total space is 256³×8B = 128MB.

[0012] The IP address set is divided into several subsets according to the value of the first part of each IP address; for each subset, the numerical values of the last three parts of the IP address are mapped to the corresponding positions of the pointer slots in the pre-allocated shared continuous memory space, which is always initialized at the beginning of the relationship mapping, and the statistical information of the IP addresses in this subset is stored by using the pointer slots in the pre-allocated shared continuous memory space.

[0013] More preferably, in the above mapping mechanism 1), the address processing process is as follows: 11) Address resolution: Convert the IP address a.b.c.d into four integer values in the range of [0, 255]; 12) Index calculation: According to p = a×256 2Determine the first-level index position P as p = b × 256 + c, where 0 ≤ p ≤ 256³ - 1; 13) Memory operation: Access the position p in the first layer. If p is empty, allocate a new counter block as the second layer, store the starting address of the second layer in the first layer at position p, and perform an atomic increment of 1 on the position d in the second layer.

[0014] More preferably, in the above mapping mechanism 2), the address processing flow is as follows: 21) Address resolution: Convert the IP address a.b.c.d into four integer values within the range of [0, 255]; 22) Index calculation: Determine the first-level index position according to p = b × 256 2 + c × 256 + d, where 0 ≤ p ≤ 256³ - 1; 23) Memory operation: Divide the set of IP addresses into q parts, set all values in the pre-allocated shared continuous memory space to 0, calculate the index p for the IP addresses belonging to the same part, and increment the position p in the pre-allocated shared continuous memory space by 1.

[0015] More preferably, in the above mapping mechanism 1), traverse the second layer to construct a minimum heap of size k for all non-zero elements, traverse the nodes in the heap to obtain k IP addresses and their occurrence times, and thus determine the top k IP addresses that appear most frequently in the IP address set; k is a natural number.

[0016] More preferably, in the above mapping mechanism 2), traverse the pre-allocated shared continuous memory space to construct a minimum heap of size k for all non-zero elements, traverse the nodes in the heap to obtain k IP addresses and their occurrence times, and thus determine the top k IP addresses that appear most frequently in the IP address set; k is a natural number.

[0017] More preferably, in the above mapping mechanism 1), divide the task of collecting IP address statistics information into multiple subtasks, each subtask is executed by a computer, and the second layers of multiple computers are merged into a complete second layer for constructing a minimum heap of size k.

[0018] More preferably, in the above mapping mechanism 2), the statistical collection task for each subset is executed by a computer, and a minimum heap of size k is shared among these computers.

[0019] On the other hand, the present invention also provides a device for collecting large-scale IP address data, which is used to implement a method for collecting large-scale IP address data as described above.

[0020] The present invention adopts the above solution and has at least the following beneficial effects.

[0021] The statistical collection of large-scale IP address data is one of the most basic problems in network traffic measurement. To solve this problem, this paper proposes two different relationship mapping mechanisms between memory blocks and IP addresses. These two mapping mechanisms strike a balance between computational cost and memory usage, and can be used to search for frequently appearing IP addresses in practical applications. A large number of experimental results verify the effectiveness of the proposed method.

[0022] In the present invention, two effective algorithms are proposed to collect statistics of large-scale IP address data, and frequently occurring IP addresses can be obtained from the statistics, which can be regarded as a preprocessing step for user behavior analysis in network traffic management. Due to the increase and acceleration of network traffic, it becomes expensive and impractical to process all IP addresses containing IP packets. By making full use of the continuous characteristics of memory addresses and the fixed range of each part of IP addresses, two relationship mapping mechanisms between memory blocks and IP addresses are designed for four-dimensional sparse matrices. The sparse matrix stores the number of occurrences of a single IP address, where the positions of rows and columns are used to represent the mapping relationship between memory blocks and IP addresses. Specifically, the present invention constructs a two-layer memory block (TLMB) to implement the first mapping mechanism of IP addresses. In addition, the present invention adopts a single shared memory block (SSMB) for all IP addresses to implement other mapping mechanisms of IP addresses. The mechanism of the mapping relationship effectively deletes information about trivial user behavior that is irrelevant to statistical analysis. The proposed method can be extended to the corresponding parallel version of a specific hardware architecture. Extensive experiments on multiple synthetic data sets demonstrate the effectiveness of the method. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 The figure shows the first mapping relationship between the memory block and the IP address in the present invention.

[0024] Figure 2 The figure shows the second mapping relationship between the memory block and the IP address in the present invention. DETAILED DESCRIPTION

[0025] The following is a further description of the specific implementation of the present invention in conjunction with the drawings of the specification, so that the technical solution and its beneficial effects of the present invention are clearer and more explicit. The following description of the embodiments with reference to the drawings is exemplary and intended to explain the present invention, but cannot be understood as limiting the present invention.

[0026] Additional aspects and advantages of the present invention will become apparent from the following description or may be learned by practice of the present invention.

[0027] The IP flow data FD is a sequence of IP records, i.e., FD = {(x1, p1), ..., (xn, pn)}, and , where each pair of elements (xi, pi) (i ∈ [1, n]) consists of an IP address xi and a corresponding set of user behavior attributes pi. Given a finite set of IP addresses X = {x1, x2,..., xn} ∈ R m×n , the purpose of the IP address statistics task is to effectively determine the top k most frequently occurring IP addresses in X, where m is the dimension of a single IP address and k ≪ n.

[0028] A standard IP address consists of four decimal digits from 0 to 255, separated by dot symbols. A single IP address is logically divided into four parts, and each part of the IP address has a numerical value. Therefore, a four-dimensional array can be created to count IP addresses, where the length of each dimension in the array is 256, and each element in the array can store the occurrence count of the IP address according to the mapping relationship between the index of each dimension of the array and the corresponding integer value. A single IP address in the host log usually only accounts for a small part of all IP addresses, and the array can be considered sparse. Therefore, a four-dimensional sparse matrix can be designed to store the quantity of a single IP address by making full use of the continuous characteristics of the array address and the fixed range of a single part.

[0029] The position of the array can be indexed by integer values. Two efficient methods are proposed below to collect statistical information of large-scale IP address data, and each method includes a mapping mechanism between a memory block and the IP addresses in the four-dimensional sparse matrix.

[0030] Combined Figure 1 As shown, a method for collecting large-scale IP address data (hereinafter referred to as TLMB for convenience of description) has a storage strategy: adopting a two-layer memory block architecture, the first layer is pre-allocated continuous memory space with a dimension of 256 3 pointer slots, and the total space is 256³ × 8B = 128MB; the second layer is a counter block with 256 × 8B effective slots; the numerical values of the first three parts of the IP address are mapped to the corresponding positions of the pointer slots in the pre-allocated continuous memory space of the first layer, and each effective slot in the second layer stores the corresponding IP address statistical information.

[0031] Specifically, the address processing flow is as follows: 11) Address resolution: Convert the IP address a.b.c.d into four integer values within the range of [0, 255]; 12) Index calculation: Determine the first-level index position P according to p = a × 256 2 + b × 256 + c, where 0 ≤ p ≤ 256³ - 1; 13) Memory operation: Access the position p in the first layer. If p is empty, allocate a new counter block as the second layer, store the starting address of the second layer into the first layer at position p, and perform an atomic increment of 1 on the position d in the second layer.

[0032] Taking the set of IP addresses 192.168.1.2, 192.168.1.3, and 192.168.2.3 as an example, the address processing flow is as follows:

[0033] The first layer layer1: Initialize an array of length 256 3 Each element in the array is a pointer to an address and is initially empty.

[0034] Process the first 192.168.1.2.

[0035] Index p: p = 192 × 256 2 + 168 × 256 + 1.

[0036] layer1[p] is empty. Create an array of length 256 as the second layer layer2_0. Each element in the array is an integer. Store the starting address of this array in layer1[p].

[0037] Increment the value at the position layer2_0[2] by 1.

[0038] Process the second 192.168.1.3.

[0039] Index p: p = 192 × 256 2 + 168 × 256 + 1.

[0040] layer1[p] is not empty. Access the address of layer1[p] to get layer2_0. Increment the value at the position layer2_0[3] by 1.

[0041] Process the third 192.168.2.3.

[0042] Index p: p = 192 × 256 2 + 168 × 256 + 2.

[0043] layer1[p] is empty. Create an array of length 256 as the second layer layer2_1. Each element in the array is an integer. Store the starting address of this array in layer1[p].

[0044] Increment the value at the position layer2_1[2] by 1.

[0045] Traverse all the addresses of layer2 and construct a min-heap of size k from all non-zero elements.

[0046] Traverse the nodes in the heap to obtain k IP addresses and their occurrence counts.

[0047] Combine Figure 2As shown, a method for collecting large-scale IP address data (for convenience of description, hereinafter referred to as SSMB) has a storage strategy as follows: It adopts a one-layer memory architecture, which is a pre-allocated shared continuous memory space with a dimension of 256³ pointer slots and a total space of 256³×8B = 128MB.

[0048] Divide the IP address set into several subsets according to the value of the first part of each IP address; for each subset, map the numerical values of the last three parts of the IP address to the corresponding positions of the pointer slots in the pre-allocated shared continuous memory space, always initialize at the beginning of the relationship mapping, and use the pointer slots in the pre-allocated shared continuous memory space to store the IP address statistical information in this subset.

[0049] In the above mapping mechanism 2), the address processing flow is as follows: 21) Address resolution: Convert the IP address a.b.c.d into four integer values within the range of [0, 255]; 22) Index calculation: Determine the first-level index position according to p = b×256 2 +c×256 + d, where 0 ≤ p ≤ 256³ - 1; 23) Memory operation: Divide the IP address set into q parts, set all values in the pre-allocated shared continuous memory space to 0, calculate the index p for the IP addresses belonging to the same part, and increment the value at position p in the pre-allocated shared continuous memory space by 1.

[0050] Taking the IP address set consisting of 192.168.1.2, 193.168.1.2, and 192.168.2.3 as an example, the address processing flow is as follows:

[0051] The first layer layer1: Initialize an array with a length of 256 3 and each element in the array is an integer.

[0052] Divide the IP addresses into q parts according to the first part. The IP addresses belonging to the same part are processed together.

[0053] Set all values in layer1 to 0, and process 192.168.1.2 and 192.168.2.3.

[0054] Index p: p = 168×256 2 +1×256 + 2.

[0055] Increment the value at position Layer1[p] by 1.

[0056] Index p: p = 168×256 2 +2×256 + 3.

[0057] Increment the value at position Layer1[p] by 1.

[0058] Traverse all the addresses of layer1 and construct a minimum heap of size k from all non-zero elements.

[0059] Set all the values in layer1 to 0 and process 193.168.1.2.

[0060] Index p: p = 168 × 256 2 + 1 × 256 + 2.

[0061] Increment the value at position Layer1[p] by 1.

[0062] Traverse all the addresses of layer1 and construct a minimum heap of size k from all non-zero elements.

[0063] Traverse the nodes in the heap to obtain k IP addresses and their occurrence counts, and thus obtain the top k most frequently occurring IP addresses in the heap.

[0064] Memory usage and complexity analysis.

[0065] In the first method, assume that the number of the first three different parts of the IP address is s. Then, the counter block slots in the second layer are linearly proportional to s. In addition, the memory size of the minimum heap is (k + 8) bytes, where each tree node contains two attributes, namely the IP address and the occurrence count. The total memory size of the proposed method is approximately the sum of three parts, namely 128MB, sKB, and (k + 8) bytes. In the first method, the computational complexity of the two layers for calculating IP address statistics is O(n), where n is the number of IP addresses. In addition, the computational complexity of constructing a minimum heap of size k is O(k log k), where k is the number of tree nodes in the heap. Therefore, the overall computational complexity of the proposed algorithm is O(k log k + n).

[0066] In the second method, assume that the number of different parts of the IP address is q. Then, the computational complexity of the IP address mapping ping mechanism is O(qn), where n is the number of IP addresses. Similarly, for each subset, the computational complexity of constructing a minimum heap of size k is O(k log k). Therefore, the overall computational complexity of the proposed algorithm is O(q(k log k + n)).

[0067] Parallel computing optimization.

[0068] To improve the computational efficiency and reduce memory occupancy, parallel computing can be performed on multiple computers. In the above mapping mechanism 1), the task of collecting IP address statistical information is divided into multiple subtasks, each subtask is executed by one computer, and the second layers of multiple computers are merged into a complete second layer for constructing a minimum heap of size k. In the above mapping mechanism 2), the statistical collection task for each subset is executed by one computer, and the minimum heap of size k is shared among these computers.

[0069] To better demonstrate the progressiveness of the technology, the present invention also conducted verification experiments.

[0070] 1. Experimental setup.

[0071] Test the performance on three synthetic datasets and two real-world datasets. The three synthetic datasets contain 5 million, 10 million, and 50 million randomly generated IP records. Each individual IP address contains one or more IP records, and the average number of IP records for each individual IP address is 100. The parameter k represents the number of frequently occurring IP addresses. Two actual network traffic datasets provided by the Cooperative Association for Internet Data Analysis (CAIDA) are used as real-world datasets. These two datasets are collected from various parts of the Internet and are widely used in network and traffic analysis research. The parameters of the three synthetic datasets and the two real-world datasets are shown in Table 1.

[0072] Table 1. Dataset statistics Data number IP record Separate IP address Size Type 1 5,000,000 50,000 77.5MB Synthetic data 2 10,000,000 100,000 155MB Synthetic data 3 50,000,000 500,000 775MB Synthetic data 4 1,114,633 107,988 14.4MB Real data 5 1,430,258 133,116 18.5MB Real data 。

[0073] Hash mapping. Each IP record is mapped to a statistical result entry using a hash table. Next, these statistical results are used to construct a minimum heap of size k.

[0074] IP mapping. All IP records are divided into q subsets according to the first part of each IP address. Then, the statistical information of the IP records in each subset is mapped to an array, and the memory of this array is pre-allocated on the computer according to the last three parts of each IP address. The top k most frequent IP addresses are selected from each subset. Next, a minimum heap of size k is constructed using q×k IP addresses.

[0075] The size ranges from small (e.g., 5 million) to large (e.g., 50 million) to test scalability. Two metrics are used to evaluate the sorting performance, namely the computational cost and the memory usage. All experiments are implemented in C language on the Windows platform, which has an Intel i7-9700k CPU and 32GB of RAM.

[0076] 2. Experimental results.

[0077] Set the experimental evaluation parameter k of the synthetic data to 10 or 100, and each experiment is repeated 10 times. The average computational cost and standard deviation are listed in Table 2, and the average memory usage and standard deviation are listed in Table 3.

[0078] Table 2. Computational Costs of Different Methods on Three Synthetic Datasets Data k Hash mapping IP mapping TLMB SSMB 1 10100 1,458.14(2.32)1,458.56(2.76) 15.18(0.07)15.21(0.04) 2.51(0.01)2.52(0.01) 18.43(0.05)18.43(0.06) 2 10100 2,927.23(11.64)2,934.65(28.93) 17.41(0.11)17.45(0.05) 4.92(0.01)24.95(0.01) 27.59(0.05)27.65(0.06) 3 10100 14,547.09(13.95)14,566.36(26.78) 35.33(0.09)35.39(0.05) 24.22(0.07)24.24(0.04) 101.01(0.17)101.20(0.2) 。

[0079] Table 3. Memory Usages of Different Methods on Three Synthetic Datasets Data k Hash mapping IP mapping TLMB SSMB 1 10100 43.63(0.12)43.73(0.06) 16,332.81(0.07)16,332.77(0.05) 189.61(0.16)189.64(0.07) 139.77(0.13)139.79(0.07) 2 10100 74.8(0.12)74.83(0.05) 16,332.75(0.12)16,332.77(0.07) 239.11(0.1)239.06(0.15) 139.74(0.13)139.7(0.12) 3 10100 324.45(0.99)324.4(1.02) 16,332.59(0.06)16,332.6(0.16) 628.33(0.07)628.27(0.01) 139.6(0.15)139.71(0.16) 。

[0080] As can be seen from Table 2 and Table 3, TLMB is always superior to all other methods in terms of computational cost. For example, when k = 10 and k = 100, the computational costs of TLMB are 2.51 s and 2.52 s respectively. When the number of IP addresses increases from 5 million to 50 million, with k = 10, the computational cost gaps between TLMB and IP mapping are 12.6 s and 11.11 s respectively. The same advantage is also observed when k = 100. In addition, in terms of memory usage, SSMB has competitive results compared with the comparison methods. Even when the number of IP addresses changes, the memory usage of SSMB is always about 139 MB. In contrast, for different numbers of IP addresses, the computational cost of SSMB is significantly better than that of hash mapping. For all numbers of IP addresses, IP mapping obtains the lowest computational cost. However, the memory usage of IP mapping is large.

[0081] Set the experimental evaluation parameter k of the real data to 10 or 100, and each experiment is repeated 10 times. The average computational cost is listed in Table 4, and the average memory usage is listed in Table 5.

[0082] Table 4. Computational Costs of Different Methods on Two Real Datasets Data k Hash mapping IP mapping TLMB SSMB 4 10100 1675.31673.9 6.736.90 0.610.62 7.007.09 5 10100 1996.81989.9 7.337.47 0.780.80 6.887.02 。

[0083] Table 5. Memory Usages of Different Methods on Two Real Datasets Data k Hash mapping IP mapping TLMB SSMB 4 10100 238.9238.8 16,392.116,392.1 249.8249.8 183.3183.4 5 10100 266.7266.7 16,391.916,392.3 281.4281.4 183.4183.4 。

[0084] As can be seen from Tables 4 and 5, in terms of frame number data, TLMB always produces a lower computational cost than other methods. SSMB and IP mapping show comparable computational costs. However, SSMB significantly reduces the memory requirements relative to IP mapping. In addition, hash mapping produces a higher computational cost than competing methods because it uses hash calculations, which makes it very time-consuming in practice. As the number of IP addresses increases on two real-world datasets, SSMB maintains relatively stable memory usage. This also highlights the effectiveness of TLMB and SSMB.

[0085] Ablation study.

[0086] To study the impact of memory blocks in the TLMB and SSMB methods, ablation experiments are conducted below on three synthetic datasets. Specifically, two specific cases are examined in the experiment. Hash tables are used to replace the second part (memory blocks) of TLMB and SSMB respectively. The main purpose of the ablation experiment is to demonstrate the importance of memory blocks in collecting statistical information of large-scale IP address data. The corresponding TLMB and SSMB invariants in these two cases are called TLMBhash and SSMBhash respectively. Tables 7 and 8 show the results of the ablation study on computational cost and memory usage.

[0087] Table 7. Ablation of computational cost occurring on three synthetic datasets Data k <![CDATA[TLMB 哈希 > <![CDATA[SSMB 哈希 > TLMB SSMB 1 10100 37.9537.9 37.1637.34 2.512.52 18.4318.43 2 10100 75.6875.19 74.5174.05 4.924.95 27.5927.65 3 10100 384.72384.23 374.49375.9 24.2224.24 101.01101.20 。

[0088] Table 8. Ablation of memory usage required on three synthetic datasets Data k <![CDATA[TLMB 哈希 > <![CDATA[SSMB 哈希 > TLMB SSMB 1 10100 3,397.33,397.5 3,397.23,397.2 189.61189.64 139.77139.79 2 10100 6,654.46,654.4 6,654.36,654.4 239.11239.06 139.74139.7 3 10100 25,665.625,008.6 27,367.627,517.5 628.33628.27 139.6139.71 。

[0089] As can be seen from Tables 7 and 8, TLMB hash compared with SSMB on the first two synthetic datasets hash shows similar computational costs and memory usage; in addition, the computational cost of TLMB hash is slightly higher than that of SSMB hash , while the memory usage rate is slightly lower. TLMB hash and SSMB hash achieve acceptable computational costs. However, compared with TLMB hash and SSMB hash , TLMB and SSMB achieve more superior performance in terms of computational cost and memory usage. These results indicate that integrating the hash table scheme into TLMB and SSMB is both time-consuming and memory-consuming. Therefore, the results of the ablation study prove the effectiveness of the storage blocks in the proposed TLMB and SSMB methods.

[0090] Note that the logics and / or steps represented or otherwise described in the present invention, for example, can be considered as a defined sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus or device (such as a computer-based system, a system including a processor or other systems that can fetch instructions from the instruction execution system, apparatus or device and execute the instructions), or used in combination with these instruction execution systems, apparatus or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by or in connection with an instruction execution system, apparatus or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation or other appropriate processing as necessary, and then stored in a computer memory.

[0091] It should be understood that those of ordinary skill in the art can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0092] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for collecting large-scale IP address data, characterized in that, Convert the IP address set into a four-dimensional sparse matrix with the numerical values of the four parts of the IP address as matrix elements, and map the four-dimensional sparse matrix to a memory block by using any one of the following two mapping mechanisms to realize the storage of the IP address set; 1) Adopt a two-layer memory block architecture. The first layer is pre-allocated continuous memory space with a dimension of 256 3 pointer slots, and the total space is 256³ × 8B = 128MB; The second layer is a counter block with 256×8B of valid slots; The numerical values of the first three parts of the IP address are mapped to the corresponding positions of the pointer slots in the pre-allocated continuous memory space in the first layer, and the statistical information of the corresponding IP address is stored in each valid slot in the second layer; 2) Adopt a one-layer memory architecture, which is a pre-allocated shared continuous memory space with a dimension of 256³ pointer slots and a total space of 256³×8B = 128MB; Divide the IP address set into several subsets according to the value of the first part of each IP address; for each subset, map the numerical values of the last three parts of the IP address to the corresponding positions of the pointer slots in the pre-allocated shared continuous memory space, always initialize at the beginning of the relationship mapping, and use the pointer slots in the pre-allocated shared continuous memory space to store the IP address statistical information in this subset.

2. The method for collecting large-scale IP address data according to claim 1, wherein In the above mapping mechanism 1), the address processing process is as follows: 11) Address resolution: Convert the IP address a.b.c.d into four integer values within the range of [0,255]; 12) Index calculation: Calculate the first-level index position P according to p = a × 256 2 + b × 256 + c, where 0 ≤ p ≤ 256³ - 1; 13) Memory operation: Access the position p in the first layer. If p is empty, allocate a new counter block as the second layer, store the starting address of the second layer in the first layer p, and perform an atomic increment of 1 on the position d in the second layer.

3. A method for collecting large-scale IP address data according to claim 1, characterized in that In the above mapping mechanism 2), the address processing process is as follows: 21) Address resolution: Convert the IP address a.b.c.d into four integer values within the range of [0,255]; 22) Index calculation: According to p = b×256 2 + c×256 + d to determine the position of the first-level index, where 0 ≤ p ≤ 256³ - 1; 23) Memory operation: Divide the IP address set into q parts, set all values in the pre-allocated shared continuous memory space to 0, calculate the index p of the IP addresses belonging to the same part, and increment the position p in the pre-allocated shared continuous memory space by 1.

4. A method for collecting large-scale IP address data according to claim 1, characterized in that In the above mapping mechanism 1), traverse the second layer to construct a minimum heap of size k from all non-zero elements, traverse the nodes in the heap to obtain k IP addresses and their occurrence times, and then determine the top k IP addresses that appear most frequently in the IP address set; k is a natural number.

5. A method for collecting large-scale IP address data according to claim 1, characterized in that, In the above mapping mechanism 2), traverse the pre-allocated shared continuous memory space to construct a minimum heap of size k from all non-zero elements, traverse the nodes in the heap to obtain k IP addresses and their occurrence times, and then determine the top k IP addresses that appear most frequently in the IP address set; k is a natural number.

6. A method for collecting large-scale IP address data according to claim 4, characterized in that, In the above mapping mechanism 1), divide the task of collecting IP address statistical information into multiple subtasks, each subtask is executed by one computer, and the second layers of multiple computers are merged into a complete second layer for constructing a minimum heap of size k.

7. A method for collecting large-scale IP address data according to claim 5, characterized in that, In the above mapping mechanism 2), the statistical collection task of each subset is executed by one computer, and a minimum heap of size k is shared among these computers.

8. An apparatus for collecting large-scale IP address data, characterized in that, Used to implement a method for collecting large-scale IP address data as described in any one of claims 1-7.