Data packet indexing and retrieval method based on long-short characteristics of attribute values of network data packet
By classifying the attribute values of network data packets in length and selecting suitable indexing algorithms, the problems of insufficient efficiency and storage space in the existing technology are solved, fast indexing and efficient retrieval are achieved, and the needs of different network traffic data processing scenarios are adapted.
Patent Information
- Application Number
- PCT/CN2023/143124
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-03
AI Technical Summary
Existing network traffic packet indexing technology cannot balance efficiency, accuracy and flexibility, especially when large-scale data processing requires a long time and large storage space, affecting processing efficiency and real-time.
The attribute values of network data packets are divided into long attributes and short attributes, and different indexing algorithms are used for processing. The long attributes are indexed using hash-Trie, and the short attributes are direct addressing hash method. When constructing indexes, the attribute length is considered to optimize efficiency and storage space.
It realizes rapid index creation, reduces storage space usage, improves retrieval efficiency and flexibility, and adapts to the needs of different application scenarios.
Smart Images

Figure CN2023143124_03072025_PF_FP_ABST
Abstract
Description
Data packet indexing and retrieval method based on the length characteristics of network data packet attribute values Technical Field
[0001] The present invention relates to the technical field of network flow data processing, and in particular to a data packet indexing and retrieval method based on the length characteristics of network data packet attribute values. Background Art
[0002] With the rapid development of network technology, the number and complexity of network traffic packets are constantly increasing. In this context, traditional single indexing algorithms are no longer able to meet the requirements for efficient, accurate, and flexible indexing and retrieval of network traffic packets. Although existing network traffic packet indexing technologies have matured, they primarily focus on a single indexing algorithm and fail to fully consider the impact of varying length attributes on indexing efficiency. Therefore, when processing network traffic packets, a single indexing algorithm often fails to achieve optimal performance. Existing technologies also suffer from the high time and space overhead of index construction. Especially when processing large-scale network traffic packets, existing indexing algorithms may require a long time and a large amount of storage space to build the index, which directly impacts the efficiency and real-time performance of network traffic packet processing. This means that the application of these technologies is limited in scenarios where large amounts of data are processed or indexing efficiency is optimized.
[0003] Document 1: Chinese invention patent CN202110457333.7 discloses a real-time indexing method and system for network traffic. Although the index establishment is also segmented according to the attribute value, its function is to highlight all attributes using the hash-Trie shortest non-overlapping prefix algorithm and index compression algorithm; this inverted index method has an inverted list with a non-fixed record number table length. When inserting and modifying elements, it is necessary to reallocate the occupied memory and move the element memory, which causes chain update problems, excessive memory consumption, and retrieval delays.
[0004] Summary of the Invention
[0005] The purpose of the present invention is to provide a data packet indexing and retrieval method based on the length characteristics of network data packet attribute values, which divides the attribute values in the network traffic data packet to be indexed into long attributes and short attributes, and uses different indexing algorithms to index the attribute values to improve the efficiency of network traffic data packet indexing and retrieval.
[0006] The technical solutions for achieving the purpose of the present invention are:
[0007] A data packet indexing and retrieval method based on the length characteristics of network data packet attribute values, the method comprising:
[0008] Get the network data packet, decode the network data packet to get the attribute value;
[0009] Classify attribute values by length, build an index for each category and save it;
[0010] Based on the index, a search request is initiated, the corresponding network data packet is obtained, and the search is completed.
[0011] Furthermore, the attribute value includes a source IPv6 address, a destination IPv6 address, a source IPv4 address, a destination IPv4 address, a timestamp, a source port, a destination port, and a protocol number.
[0012] Furthermore, attribute values are divided into two categories according to their length: long attributes and short attributes; a long attribute is an attribute value whose length is greater than or equal to a set number of bytes, and a short attribute is an attribute value whose length is less than a set number of bytes.
[0013] Furthermore, for long attributes, the hash-Trie index algorithm is used to build the index.
[0014] Furthermore, for long attributes, the hash-Trie index algorithm is used to build the index. The specific construction process is as follows:
[0015] Split long attributes into network prefix and host part, and segment the network prefix;
[0016] Hash each network prefix and use the hash function to map the hash values of all segmented network prefixes into a hash table.
[0017] Construct a Tire tree where each node represents a network prefix or host part, and each node contains an array of pointers to its child nodes.
[0018] The hash values of all segmented network prefixes are inserted into the nodes of the Tire tree one by one, and each node records the position offset of the network data packet corresponding to the inserted segmented network prefix.
[0019] Furthermore, for short attributes, direct addressing hashing is used to build indexes.
[0020] Furthermore, for short attributes, a direct addressing hash method is used to construct an index. The specific construction process is as follows:
[0021] The hash value of the short attribute is used as the index of the hash table and directly inserted into the hash table; the offset linked list corresponding to the short attribute is obtained, and the offset of the data packet containing the short attribute is inserted into the offset linked list to complete the creation of a short attribute index;
[0022] The index structure consists of two parts: a hash table and an offset list. The hash table is used to store pointers to the offset list, and the offset list is used to record the offsets of packets containing short attributes.
[0023] The size of the hash table is determined by the range of values of the short attribute length.
[0024] Furthermore, a search request is initiated, including:
[0025] Input the search data and calculate the attribute value of the search data;
[0026] Match the attribute value of the retrieved data with the corresponding long attribute index data structure or short attribute index data structure to obtain the index position corresponding to the retrieved data, and extract the network data packet at the index position.
[0027] A computer system comprising a memory and a processor; wherein:
[0028] Memory, used to store programs;
[0029] The processor is used to execute the program to implement various steps of the data packet indexing and retrieval method based on the length characteristics of the network data packet attribute value.
[0030] A readable storage medium stores a computer program, which, when executed by a processor, implements various steps of a data packet indexing and retrieval method based on the length characteristics of network data packet attribute values.
[0031] Compared with the prior art, the present invention has the following significant advantages:
[0032] (1) Fast index creation: Select different indexing algorithms based on attribute length to achieve fast index creation and improve index building efficiency. Since direct addressing hashing is faster than hash-Trie indexing algorithms, it is beneficial for faster index creation for short attributes.
[0033] (2) The index occupies less space: Selecting an adaptive indexing algorithm and considering the impact of attribute length on indexing efficiency can effectively reduce the space occupied by the index; among them, the hash table and offset linked list in the short attribute index data structure are more compact than the Trie tree in the long attribute index data structure, which is more conducive to saving storage resources.
[0034] (3) Flexible retrieval with high efficiency: By analyzing and dividing the length and short features of attribute values, and according to the adapted indexing algorithm, the attribute values to be indexed are flexibly selected to achieve the effect of quickly retrieving accurate data packets; at the same time, it meets the retrieval needs in different application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] FIG1 is a flow chart of a data packet indexing and retrieval method based on the length characteristics of network data packet attribute values according to the present invention.
[0036] FIG2 is a schematic diagram of a short index data structure in one embodiment of the present invention.
[0037] FIG3 is a schematic diagram of deriving short attributes of a short index data structure in one embodiment of the present invention. DETAILED DESCRIPTION
[0038] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0039] As shown in FIG1 , a data packet indexing and retrieval method based on the length characteristics of network data packet attribute values includes:
[0040] Get the network data packet, decode the network data packet to get the attribute value;
[0041] Classify attribute values by length, build an index for each category and save it;
[0042] Based on the index, a search request is initiated, the corresponding network data packet is obtained, and the search is completed.
[0043] Specifically, the attribute values include source IPv6 address, destination IPv6 address, source IPv4 address, destination IPv4 address, timestamp, source port, destination port, and protocol number.
[0044] Specifically, attribute values are divided into two categories according to their length: long attributes and short attributes; a long attribute is an attribute value whose length is greater than or equal to a set number of bytes, and a short attribute is an attribute value whose length is less than a set number of bytes.
[0045] Specifically, for long attributes, the hash-Trie index algorithm is used to build the index.
[0046] Specifically, for long attributes, the hash-Trie index algorithm is used to build the index. The specific construction process is as follows:
[0047] Split long attributes into network prefix and host part, and segment the network prefix;
[0048] Hash each network prefix and use the hash function to map the hash values of all segmented network prefixes into a hash table.
[0049] Construct a Tire tree where each node represents a network prefix or host part, and each node contains an array of pointers to its child nodes.
[0050] The hash values of all segmented network prefixes are inserted into the nodes of the Tire tree one by one, and each node records the position offset of the network data packet corresponding to the inserted segmented network prefix.
[0051] Specifically, for short attributes, a direct addressing hash method is used to build the index.
[0052] Specifically, for short attributes, the direct addressing hash method is used to build the index. The specific construction process is as follows:
[0053] The hash value of the short attribute is used as the index of the hash table and directly inserted into the hash table; the offset linked list corresponding to the short attribute is obtained, and the offset of the data packet containing the short attribute is inserted into the offset linked list to complete the creation of a short attribute index;
[0054] The index structure consists of two parts: a hash table and an offset list. The hash table is used to store pointers to the offset list, and the offset list is used to record the offsets of packets containing short attributes.
[0055] The size of the hash table is determined by the range of values of the short attribute length.
[0056] Specifically, initiating a search request includes:
[0057] Input the search data and calculate the attribute value of the search data;
[0058] Match the attribute value of the retrieved data with the corresponding long attribute index data structure or short attribute index data structure to obtain the index position corresponding to the retrieved data, and extract the network data packet at the index position.
[0059] A computer system comprising a memory and a processor; wherein:
[0060] Memory, used to store programs;
[0061] The processor is used to execute the program to implement various steps of the data packet indexing and retrieval method based on the length characteristics of the network data packet attribute value.
[0062] A readable storage medium stores a computer program, which, when executed by a processor, implements various steps of a data packet indexing and retrieval method based on the length characteristics of network data packet attribute values.
[0063] The following describes in detail the operation process of the data packet indexing and retrieval method based on the length characteristics of network data packet attribute values in conjunction with the actual application scenarios of the present invention.
[0064] S1. Index generation process
[0065] (1) Receive data packets: Receive network traffic data packets through a physical layer receiver (such as a network card).
[0066] (2) Decoding data packets: Decode the received data packets and obtain their attribute values, including source IPv6 address, destination IPv6 address, source IPv4 address, destination IPv4 address, timestamp, source port, destination port, and protocol number.
[0067] (3) Classify and process attribute values: The attribute values in the network traffic data packets to be indexed are divided into long attributes and short attributes. Long attributes include the source IPv6 address, destination IPv6 address, source IPv4 address, destination IPv4 address, and timestamp. The length of these attributes is greater than or equal to 4 bytes. Short attributes include the source port, destination port, and protocol number. The length of these attributes is less than 4 bytes.
[0068] (4) Index long attributes: Use hash-Trie index algorithm for processing.
[0069] First, split the long attribute into a network prefix and a host portion. The network prefix is further divided into several segments. A hash function is selected to map the values of the segments into a hash table, preferably MD5, SHA-1, or other commonly used hash functions. For example, the IPv4 address 192.168.0.1 is converted to 11000000 10101000 00000000 00000001. Based on the length of the network prefix, the corresponding number of bits are truncated from the left side of the IPv4 address as the network prefix. Assuming the network prefix length is 24, 24 bits are truncated from the left side as the network prefix: 11000000 10101000 00000000; the remaining bits are used as the host portion, which is 00000001.
[0070] For example, when converting the IPv6 address 2001:db8::1 to 2001:db8::1, the network prefix is truncated by a corresponding number of bits from the left side of the IPv6 address, depending on the network prefix length. Assuming the network prefix length is 48, 48 bits are truncated from the left side to form the network prefix, that is, 2001:db8::, and the host portion is ::1.
[0071] Then, a trie tree is constructed based on the output of the hash function. Each node represents a prefix or host part. Each node contains an array of pointers to its child nodes to support more precise matching.
[0072] Finally, insert the data, inserting the hash value of each segment's network prefix into the Trie tree. For each inserted node, the location offset of the corresponding network packet under that node can be recorded. This allows for quick retrieval of the packet corresponding to the target attribute value.
[0073] It should be noted that the network prefix is further segmented to account for the long length of IP addresses. Dividing the network prefix into smaller segments effectively reduces the number of child nodes per node, enabling faster retrieval of data packets corresponding to the target attribute value. In a Trie tree, each node represents a prefix or host portion, and each node contains an array of pointers to its child nodes. When querying the target attribute value, starting from the root node, the tree structure is traversed downwards, level by level, until the corresponding node is found. Dividing the network prefix into smaller segments effectively reduces the number of levels in the tree structure to be traversed, thereby improving query efficiency. The smaller the segments, the more accurately the network prefix and host portion can be matched. This allows for faster retrieval of matching nodes and precise determination of the packet's location offset when querying the target attribute value.
[0074] (5) Indexing short attributes: For short attributes, based on the relatively short width of the short attribute domain, the complete attribute value is used to create an index, and the direct addressing hash method is used for processing, which can not only ensure the insertion speed, but also keep the space overhead within a reasonable range. The size of the hash table is determined according to the value range of the short attribute length, and a hash function is used to map the attribute value to the corresponding position in the hash table, so that the data packet corresponding to the target attribute value can be quickly found. When a short attribute (such as port, protocol, etc.) is obtained and an index is created for it, its attribute value is used as the index of the hash table and directly inserted into the hash table. The offset linked list corresponding to the attribute value is obtained, and then the offset of the current data packet is inserted into the offset linked list to complete the creation of a short attribute index.
[0075] S2. Build index data structure
[0076] (1) Short attribute index data structure
[0077] The short attribute index data structure includes a hash table and an offset linked list; among them:
[0078] The hash table is a pointer array. The size of the array is determined by the range of the length of the indexed attribute. For example, if the port number is 2 bytes, the corresponding array size is 65536, and if the protocol number is one byte, the corresponding array size is 256.
[0079] The offset linked list is used to record the offsets of packets containing a specific attribute value. An offset field is set at the beginning of each packet. When indexing a short attribute of a packet (such as the source port), the short attribute value is used as an index into a hash table. In the hash table, each position stores a pointer to the offset linked list. For each attribute value, the packet corresponding to the offset recorded in the offset linked list is the packet containing that attribute value. Based on the target attribute value, the corresponding packet is quickly found, the offset is accessed, and the actual content of the packet is obtained.
[0080] Taking the destination port as an example, suppose a traffic set contains six packets. The two packets with destination port number 21 have offsets 0 and 1330 in the file; the two packets with destination port number 80 have offsets 450 and 2000 in the file; and the two packets with destination port number 445 have offsets 1024 and 2880 in the file. For the received packet with destination port number 21, we first use direct hashing to find the corresponding location of 21. If it is empty, we create an offset linked list pointer and write a node with offset 0 to the offset linked list. Otherwise, we find the corresponding offset linked list and create a new offset node to be attached after the offset node. The remaining five packets are processed similarly, resulting in a hash table with buckets 21, 80, and 445 storing pointers to three offset linked lists, respectively. Bucket 21 points to linked lists with offsets 0 and 1330, bucket 80 points to linked lists with offsets 512 and 2000, and bucket 445 points to linked lists with offsets 1024 and 2880. The corresponding index structure is shown in Figure 2.
[0081] (2) Long attribute index data structure
[0082] The long attribute index data structure also consists of a hash table and an offset linked list, but its specific construction process and storage method are different from those of the short attribute index data structure.
[0083] A hash table acts as an array, mapping each segment value of a long attribute to a corresponding Trie tree node. Each array element is a pointer to a Trie tree node. A Trie tree is a multi-branch tree structure used to efficiently store and look up hash values for long attributes. Each node contains a hash value and an array of pointers to its child nodes. The root node of the Trie tree is the starting point for all long attributes.
[0084] The offset linked list records the offset information of all data packets under the node in the storage space. For each Trie tree node, a pointer to the offset linked list is stored. By using the hash-Trie index algorithm, the corresponding data packet can be efficiently found according to long attributes (such as IP address), and fast retrieval and update operations are supported.
[0085] S3, export and save index data
[0086] Because storage devices have limited memory space, when the amount of data in memory reaches a certain level, it is exported to the hard disk for permanent storage. When exporting indexes to the hard disk, it is important to ensure that their logical structure remains consistent with that in memory. For example, the direct hashing algorithm for short indexes is shown in Figure 3.
[0087] S4. Retrieve data
[0088] When a search request is initiated, the attribute value of the search data is searched as needed, and a quick search is performed using the corresponding index data structure. Because the present invention utilizes a method that selects different indexing methods based on attribute length, data packets corresponding to the target attribute value can be found more efficiently, improving search efficiency. This flexible indexing method can play an important role in various network traffic data processing scenarios and has high practicality and value.
[0089] It should be noted that the order of the embodiments of the present application described above is for descriptive purposes only and does not represent the superiority or inferiority of the embodiments. The above description is of specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0090] The various embodiments in this application are described in a progressive manner. Similar portions between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the device, equipment, and storage medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simple. For relevant portions, refer to the descriptions of the method embodiments.
[0091] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0092] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A method for packet indexing and retrieval based on the length characteristics of network packet attribute values, characterized in that: The method includes: Obtain network data packets and decode the network data packets to obtain attribute values; Classify the attribute values according to their lengths, construct an index for each class respectively, and save it; Based on the index, initiate a retrieval request to obtain the corresponding network data packets, and the retrieval is completed.
2. The data packet indexing and retrieval method based on the length characteristics of network data packet attribute values according to claim 1, wherein: The attribute values include source IPv6 address, destination IPv6 address, source IPv4 address, destination IPv4 address, timestamp, source port, destination port, and protocol number.
3. The data packet indexing and retrieval method based on the length characteristics of network data packet attribute values according to claim 2, characterized in that: The attribute values are classified into two categories according to their lengths: long attributes and short attributes; where long attributes are attribute values with a length greater than or equal to the set number of bytes, and short attributes are attribute values with a length less than the set number of bytes.
4. The data packet indexing and retrieval method based on the length characteristics of network data packet attribute values according to claim 3, wherein: For long attributes, use the hash-Trie index algorithm to construct an index.
5. The data packet indexing and retrieval method based on the length characteristics of network data packet attribute values according to claim 4, wherein: For long attributes, the specific construction process of using the hash-Trie index algorithm is as follows: Divide the long attribute into a network prefix and a host part, and segment the network prefix; Perform hash processing on each segment of the network prefix, and use a hash function to map the hash values of all segmented network prefixes obtained to the hash table; Construct a Tire tree, where each node in the Tire tree represents a segment of the network prefix or the host part, and each node contains an array of pointers pointing to its child nodes; Insert the hash values of all segmented network prefixes into the nodes of the Tire tree one by one, and each node records the position offset of the network data packet corresponding to the inserted segmented network prefix.
6. The data packet indexing and retrieval method based on the length characteristics of network data packet attribute values according to claim 4, wherein: For short attributes, use the direct addressing hashing method to construct an index.
7. The data packet indexing and retrieval method based on the length characteristics of network data packet attribute values according to claim 6, characterized in that: For short attributes, the specific construction process of using the direct addressing hashing method is as follows: Use the hash value of the short attribute as the index of the hash table and directly insert it into the hash table; obtain the offset linked list corresponding to the short attribute, and insert the offset of the data packet containing the short attribute into the offset linked list to complete the creation of a short attribute index once; The structure of the index includes two parts: a hash table and an offset linked list, where the hash table is used to store the pointers of the offset linked list, and the offset linked list is used to record the offsets of the data packets containing the short attribute; The size of the hash table is determined by the value range of the short attribute length.
8. The method for packet indexing and retrieval based on the length characteristics of network packet attribute values according to claim 5 or 7, characterized in that: The initiation of the retrieval request includes: Input retrieval data and calculate the attribute values of the retrieval data; Match the attribute values of the retrieval data with the corresponding long attribute index data structure or short attribute index data structure to obtain the index position corresponding to the retrieval data, and extract the network data packet at the index position.
9. A computer system, characterized in that, Includes: A memory and a processor; The memory is used to store programs; The processor is used to execute the program to implement each step of the data packet index and retrieval method based on the length characteristics of network data packet attribute values as described in any one of claims 1 to 8.
10. A readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements each step of the data packet index and retrieval method based on the length characteristics of network data packet attribute values as described in any one of claims 1 to 8.
Citation Information
Patent Citations
An efficient RDF data storage and query system
CN109684325A
Network flow real-time indexing method and system
CN113139100A
Data packet indexing and retrieval method based on long and short characteristics of network data packet attribute values
CN117435912A
Enhanced prefix matching
US10230639B1
Cited By
Non-motor vehicle violation data processing method and system
CN121166677A
High-frequency efficient data addressing query method based on shared memory
CN122364109A