High-performance Netflow message output methods, systems, and storage media
By storing dynamic and quasi-static entries of Netflow packets in the FPGA's built-in high-bandwidth memory and DDR4 memory respectively, and combining them with a five-tuple hash value association index, the problem of insufficient Netflow packet processing performance in the prior art is solved, and efficient 400G network traffic processing is achieved.
Patent Information
- Application Number
- CN202511156476.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-08-19
AI Technical Summary
Existing technologies struggle to meet the demands for high-performance processing, massive flow table capacity, and high bandwidth while maintaining cost-effectiveness when handling Netflow packets, especially in terms of processing performance and burst resilience for 100G network traffic.
The method of storing dynamic entries in the FPGA's built-in high-bandwidth memory and quasi-static entries in external DDR4 memory is adopted. An associative index is established through five-tuple hash values. Combining the high bandwidth throughput of FPGA and the large capacity of DDR4, dynamic entries are stored in the FPGA's built-in high-bandwidth memory, and quasi-static entries are stored in DDR4 memory, thus avoiding the problem of low throughput of DDR4.
It achieves the acquisition and output of 400G network traffic with extremely low resource and peripheral consumption, improving processing performance with throughput and processing performance reaching 48Mpps and 200Mpps, respectively.
Smart Images

Figure CN120710947B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network splitter technology, specifically to a high-performance Netflow packet output method, system, and storage medium. Background Technology
[0002] Netflow is a service designed to obtain IP flow information from network data. Its lightweight, session-level view provides a more convenient way for backend systems to perform traffic analysis, making it a key requirement in network traffic splitting products. As is well known, Netflow packets typically contain basic flow record data such as MAC address, 5-tuple, flow establishment time, and packet statistics. For network traffic splitting devices to perform this function, three basic requirements must be met: 1. The device needs to have a sufficiently large flow table capacity to ensure that a table can be created for each flow; 2. The flow table entry width needs to be large enough to accommodate the storage of flow record information; 3. The device needs to have high bandwidth processing capabilities to meet the line-rate processing of 256-byte packets.
[0003] The current mainstream approach is to implement Netflow output based on a multi-core x86 + DDR4 software processing architecture. Although stacking DDR4 memory can meet the requirements for huge flow table capacity and processing performance, the processing performance and burst resistance per 100G of network traffic cannot meet expectations when considering cost-effectiveness. Therefore, how to complete Netflow flow record packet output with minimal external storage devices and resources while processing packets with high performance is a significant challenge. Summary of the Invention
[0004] In view of the shortcomings of the existing technology, the purpose of this invention is to provide a high-performance Netflow message output method, system and storage medium.
[0005] To achieve the above objectives, the present invention provides the following technical solution:
[0006] A high-performance Netflow message output method that receives and parses messages;
[0007] Dynamic entries are generated based on the parsed message information, and the dynamic entries are stored in the FPGA's built-in high-bandwidth memory.
[0008] The generation of dynamic entries triggers the generation of quasi-static entries, which are then stored in external DDR4 memory.
[0009] The dynamic entries and quasi-static entries establish an associated index through the hash value of the five-tuple;
[0010] When a dynamic entry ages out, the flow information associated with the quasi-static entry is triggered, and the packet statistics, 5-tuple, MAC, and IP version number are framed and output according to the Netflow packet output format.
[0011] In this invention, preferably, parsing the message includes: extracting the message header information, and obtaining the content information of each field in the Layer 2, Layer 3, and Layer 4 protocol headers.
[0012] In this invention, preferably, generating dynamic entries includes:
[0013] The action is established by recording the arrival time of the first message, the entry processing time, the number of messages, the number of bytes, and the signature information into the high-bandwidth memory built into the FPGA, and simultaneously triggering the quasi-static entry processing module to write to the DDR4 storage space.
[0014] The matching action updates the corresponding flow statistics and table entry processing time and writes them into the FPGA's built-in high-bandwidth memory.
[0015] The aging process determines the dynamic entry clearing operation by comparing the dynamic entry processing time with the current device time, and simultaneously triggers the quasi-static entry processing module to read the DDR4 storage space.
[0016] In this invention, preferably, the generation of quasi-static entries specifically includes:
[0017] Quasi-static table entry writing: When a dynamic table entry is created, the quasi-static table entry write operation is triggered. A globally unique DDR4 address is formed by the five-tuple hash value of the dynamic table entry and the module ID, and the five-tuple, MAC and IP version number are written at this address.
[0018] Quasi-static entry read: When a dynamic entry ages out, a quasi-static entry read operation is triggered to read the 5-tuple, MAC address, and IP version number of the corresponding address.
[0019] In this invention, preferably, four dynamic processing modules are used to update the statistical information of network traffic packet by packet.
[0020] In this invention, preferably, the burst feature of AXI4 in the FPGA is used to handle multiple table entries conflict.
[0021] In this invention, preferably, when generating quasi-static entries, the access cycle of DDR4 is adjusted. When there are many flow table creation operations, it can automatically switch to the T0 cycle, and when there are multiple aging operations, it can automatically switch to the T2 cycle. Under normal circumstances, it is maintained at the T1 stage.
[0022] A high-performance Netflow message output system based on an FPGA with HBM and DDR4, specifically including:
[0023] The message parsing module is used to extract message header information and obtain the content information of each field in the layer 2, 3, and 4 protocol headers;
[0024] Dynamic entry processing module: Uses the five-tuple hash value obtained from message parsing to read and write high-bandwidth memory and record the dynamic information of the stream;
[0025] Quasi-static table entry processing module: Triggers DDR4 memory read and write operations for quasi-static table entries when dynamic table entries age or are created, and records the stream's quintuple, MAC, and IP version number;
[0026] The Netflow framing module extracts statistics, time, and other information when dynamic entries age, and simultaneously triggers quasi-static entries to obtain the MAC and IP version numbers of the flow, framing them according to the Netflow frame format for output.
[0027] Compared with the prior art, the beneficial effects of the present invention are:
[0028] Based on the characteristics of Netflow output content, the method of this invention divides the information into two parts: statistical information and streaming information, which are stored in dynamic entries and quasi-static entries respectively. Finally, Netflow packets are output by association through five-tuple hashing. By combining the large DDR4 memory space and the high bandwidth throughput of the FPGA's built-in HBM, information that needs to be updated per packet is stored in the FPGA's built-in HBM, while information that does not need to be updated per packet is stored in DDR4. This avoids the problem of low throughput of DDR4 and completes the collection and output of streaming information of 400G network traffic with extremely low resource and peripheral consumption. Attached Figure Description
[0029] Figure 1 This is a flowchart illustrating the high-performance Netflow message output method described in this invention.
[0030] Figure 2 This is a schematic diagram of the structure of the high-performance Netflow message output system based on FPGA with HBM and DDR4 as described in this invention. Detailed Implementation
[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0033] Please see Figure 1 and Figure 2 A preferred embodiment of the present invention provides a high-performance Netflow message output method. Based on the characteristics of Netflow output content, the received message information is divided into statistical information and flow information, which are stored in dynamic entries and quasi-static entries, respectively. The statistical information is stored in the dynamic entries, which are stored in the high-bandwidth memory built into the FPGA. The flow information is stored in the quasi-static entries, which are stored in DDR4 memory. The Netflow messages are output by associating the five-tuple hash values of the message information. By combining the large memory capacity of DDR4 with the high bandwidth throughput of the FPGA's built-in high-bandwidth memory (HBM), information that needs to be updated in a message is stored in the FPGA's built-in HBM, while information that does not need to be updated message by message is stored in DDR4. This avoids the problem of low throughput of DDR4, enabling the acquisition and output of 400G network traffic with extremely low resource and peripheral consumption. For example, using an 8GB DDR4 as a quasi-static table entry, the throughput is approximately 25.6GB / s, with an actual processing performance of approximately 48Mpps. Similarly, an 8GB HBM with 16 pseudo-channels has a throughput of approximately 409MB / s, with an actual processing performance of approximately 200Mpps.
[0034] The specific steps of the output method include:
[0035] S1, Receive message;
[0036] S2. Parse the message;
[0037] S3. Generate dynamic entries based on the parsed message information, and store the dynamic entries in the FPGA's built-in high-bandwidth memory;
[0038] S4. Simultaneously with the generation of dynamic entries, quasi-static entries are generated and stored in external DDR4 memory.
[0039] The dynamic entries and quasi-static entries establish an associated index through the hash value of the five-tuple;
[0040] S5. When the dynamic entry aging action is triggered, the flow information association output of the quasi-static entry is triggered, and the packet statistics, five-tuple, MAC and IP version number are framed and output according to the Netflow packet output format.
[0041] The dynamic entries store flow statistics, including packet count and byte count. A read / write operation is performed to update the statistics each time a packet arrives. The quasi-static entries store flow information, including source IP address, destination IP address, source port, destination port, protocol number, and MAC address. Information is updated and a write operation is performed only when the first packet of a flow arrives; a read operation is performed when a dynamic entry ages out. The quintuple, consisting of the source IP address, destination IP address, source port, destination port, and protocol number as input, is used to calculate the hash value of the quintuple. This hash value serves as the address and signature information for both dynamic and quasi-static entries. The address and signature information use two different CRC32 polynomials.
[0042] Specifically, in step S1, the message input involves sending a test message using a test instrument.
[0043] In this embodiment, parsing the message includes: extracting the message header information, and obtaining the content information of each field in the Layer 2, Layer 3, and Layer 4 protocol headers.
[0044] In this embodiment, generating dynamic entries includes:
[0045] The action is established by recording the arrival time of the first message, the entry processing time, the number of messages, the number of bytes, and the signature information into the high-bandwidth memory built into the FPGA, and simultaneously triggering the quasi-static entry processing module to write to the DDR4 storage space.
[0046] The matching action updates the corresponding flow statistics and table entry processing time and writes them into the FPGA's built-in high-bandwidth memory.
[0047] The aging process determines whether to clear dynamic entries by comparing the dynamic entry processing time with the current device time, and simultaneously triggers the quasi-static entry processing module to read the DDR4 storage space. In a specific embodiment, the aging time is configured to be 30 seconds based on actual needs. The device time is obtained through clock pulse statistics. When packets from the same flow access a dynamic entry, the current device time is recorded in real time as the entry's timestamp. The ratio of the polling entry period to the packet access period is 1:24. When the difference between the polled entry's timestamp and the current device time is greater than 30 seconds, it is considered an aging process, and the entire content of that entry is cleared to 0.
[0048] In this embodiment, generating quasi-static entries specifically includes:
[0049] Quasi-static table entry writing: When a dynamic table entry is created, the write operation of the quasi-static table entry is triggered. A globally unique DDR4 address is formed by the five-tuple hash value of the dynamic table entry and the module ID, and the five-tuple, MAC address, table creation time, signature information and IP version number are written at this address.
[0050] Quasi-static entry read: When a dynamic entry ages out, a quasi-static entry read operation is triggered to read the 5-tuple, MAC address, and IP version number of the corresponding address.
[0051] In this implementation, four dynamic processing modules are used to update network traffic statistics packet by packet. The high two bits of the hash value calculated using CRC32 on the 5-tuple (source IP, destination IP, source port, destination port, protocol number) are used as the distribution criteria to distribute arriving network traffic to the four dynamic processing modules for reading and writing HBM, thus completing the dynamic statistics update.
[0052] In this implementation, the burst feature of AXI4 in the FPGA is used for multi-entry collision handling. Specifically, a command returning multiple column addresses of the same row is faster than reading multiple entries from different rows. For example, this design uses burst2. When a message arrives, a read command can return the contents of two entries. When both entries are empty, the message is the first message of the stream. An empty entry is selected to perform a table creation operation and store the signature information. When the next message addresses this address, if an entry with a matching signature is found, the entry statistics are updated. If the signatures do not match, another empty entry is used to perform the table creation operation. By using the burst mechanism to complete collision hashing, the table filling rate is effectively improved while maintaining table processing performance.
[0053] In this embodiment, the DDR4 access cycle is adjusted when generating quasi-static entries. When there are many flow table creation operations, it can automatically switch to the T0 cycle. During aging operations (when there are many aging entries that need to be processed in batches (deleting or updating expired flows), it automatically switches to the T2 cycle. Under normal circumstances, it remains in the T1 cycle. The DDR4 access cycle table is shown in Table 1 below:
[0054] Table 1.
[0055]
[0056] Due to the access characteristics of DDR4, pure read and pure write operations have higher performance than read-write operations. Therefore, read and write commands are stored separately before accessing DDR4. The access cycle is determined by judging the storage status of the two command caches. In this embodiment, the clock frequency is 250MHz, N=16. When the write command cache is idle, the T0 cycle is entered, and DDR4 read commands are sent continuously for 32 clock cycles. When the read command cache is idle, the T2 cycle is entered, and DDR4 write commands are sent continuously for 32 clock cycles. When neither cache is idle, the T1 cycle is entered, and DDR4 read commands are sent for the first 16 clock cycles, followed by DDR4 write commands for the last 16 clock cycles. T0 is approximately 256ns, T1 is approximately 400ns, and T2 is approximately 256ns.
[0057] Another preferred embodiment of the present invention provides a high-performance Netflow packet output system based on an FPGA with HBM and DDR4, wherein the modules cooperate in a coordinated manner during operation to achieve efficient packet processing and output. This high-performance Netflow packet output system specifically includes:
[0058] The message parsing module is used to extract message header information and obtain the content information of each field in the layer 2, 3, and 4 protocol headers;
[0059] Dynamic entry processing module: Uses the five-tuple information (source IP, destination IP, source port, destination port, protocol number) obtained from packet parsing to calculate the hash value using CRC32 as the entry address to read and write high-bandwidth memory and record the dynamic information of the stream;
[0060] Quasi-static table entry processing module: Triggers DDR4 memory read and write operations for quasi-static table entries when dynamic table entries age or are created, and records the stream's quintuple, MAC, and IP version number;
[0061] The Netflow framing module extracts statistics, time, and other information when dynamic entries age, and simultaneously triggers quasi-static entries to obtain the MAC and IP version numbers of the flow, framing them according to the Netflow frame format for output.
[0062] Specifically, this includes dynamic entries, which are carried by the FPGA's built-in high-bandwidth memory (HBM); quasi-static entries, which are carried by external DDR4 memory; and dynamic entries and quasi-static entries are associated by a 5-tuple hash.
[0063] This high-performance Netflow packet output system combines the features of HBM and DDR4, providing multi-channel parallel read / write capabilities. It is suitable for storing frequently updated dynamic entries, while DDR4, with its large capacity, is used to store less frequently accessed quasi-static entries. The dynamic entry processing module employs CRC32 hashing and a dual-entry verification mechanism, directly mapping HBM addresses to 5-tuple hash values. It updates flow statistics (number of packets, number of bytes) within a single cycle. Combined with a load-balancing design of four parallel processing modules, it achieves ultra-high bandwidth processing performance. Simultaneously, the burst feature of the AXI4 bus (such as burst2 mode) allows a single command to access two entries. Hash collisions are resolved through signature comparison, increasing the entry filling rate to over 90%.
[0064] The advantage of the quasi-static table entry processing module lies in its "on-demand access" strategy: DDR4 operations are triggered only during stream creation or aging, reducing the number of memory accesses compared to traditional solutions. Combined with a dynamic cycle switching mechanism (T0 / T1 / T2 cycles), 32 write commands are continuously sent during the peak stream creation period (T0 cycle), and read operations are performed in batches during the peak aging period (T2 cycle), improving DDR4 utilization by 40% and keeping the packet loss rate below 0.1%.
[0065] The system's overall advantages are also reflected in its resource efficiency. Through the hash association index of dynamic and quasi-static table entries, the framing module can complete Netflow packet assembly within 100ns, achieving an output rate of 48Mpps, which is perfectly suited for traffic monitoring scenarios in large data centers.
[0066] In this embodiment, such as Figure 2 As shown, by utilizing the high bandwidth throughput of the FPGA's built-in HBM, four parallel dynamic entry processing modules are reused to update the statistical information of 400G network traffic per packet. At the same time, the burst feature of AXI4 is used to handle multiple entry conflicts, ensuring the success rate of flow establishment. Meanwhile, the table entry address is directly mapped through the five-tuple hash value. Combined with the four parallel dynamic entry processing modules, the dynamic entry processing modules distribute traffic according to the high two bits of the hash value, which can achieve a dynamic information update rate of 200Mpps.
[0067] In this embodiment, such as Figure 2As shown, by using the creation and aging actions of dynamic entries as the trigger conditions for the quasi-static entry module processing, the access frequency of DDR4 is greatly reduced, thereby reducing the number of DDR4 peripherals used; as shown in Table 1, the access mode of DDR4 in the current cycle is assigned through command type decision to improve the performance of the DDR4 interface, where T0 < T2 < T1. In the stage with more table creation actions, it can automatically switch to the T0 cycle, and in the stage with more aging actions, it automatically switches to the T2 cycle. Under normal circumstances, it remains in the T1 stage. The DDR4 access cycle is automatically switched according to the operation type (switch to the T0 cycle when there are more flow creations and switch to the T2 cycle when there are more aging actions), improving the memory utilization rate by 40%; the aging time is configurable (such as the default of 30s) to adapt to different network traffic characteristics and ensure the timeliness of flow information statistics.
[0068] In some other preferred embodiments of the present invention, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the processor is caused to execute the steps of the method as described in the above embodiments.
[0069] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0070] The above description is a detailed description of the preferred and feasible embodiments of the present invention, but the embodiments are not intended to limit the scope of the patent application of the present invention. Any equivalent changes or modifications made under the technical spirit disclosed by the present invention shall fall within the scope of the patent covered by the present invention.
Claims
1. A high-performance Netflow message output method, characterized in that, include: Receive messages and parse them; Dynamic entries are generated based on the parsed message information and stored in the FPGA's built-in high-bandwidth memory. The dynamic entries are used to store flow statistics. The generation of dynamic entries triggers the generation of quasi-static entries, which are stored in external DDR4 memory. These quasi-static entries are used to store stream information. The dynamic entries and quasi-static entries establish an associated index through the hash value of the five-tuple; When a dynamic entry ages out, the flow information associated with the quasi-static entry is triggered, and the packet statistics, 5-tuple, MAC, and IP version number are framed and output according to the Netflow packet output format.
2. The high-performance Netflow message output method according to claim 1, characterized in that, The parsing of the message includes: extracting the message header information, and obtaining the content information of each field in the Layer 2, Layer 3, and Layer 4 protocol headers.
3. The high-performance Netflow message output method according to claim 2, characterized in that, Generating dynamic entries includes: The action is established by recording the arrival time of the first message, the entry processing time, the number of messages, the number of bytes, and the signature information into the high-bandwidth memory built into the FPGA, and simultaneously triggering the quasi-static entry processing module to write to the DDR4 storage space. The matching action updates the corresponding flow statistics and table entry processing time and writes them into the FPGA's built-in high-bandwidth memory. The aging process determines the dynamic entry clearing operation by comparing the dynamic entry processing time with the current device time, and simultaneously triggers the quasi-static entry processing module to read the DDR4 storage space.
4. The high-performance Netflow message output method according to claim 1, characterized in that, The generation of quasi-static entries specifically includes: Quasi-static table entry writing: When a dynamic table entry is created, the write operation of the quasi-static table entry is triggered. A globally unique DDR4 address is formed by the five-tuple hash value of the dynamic table entry and the module ID, and the five-tuple, MAC address, table creation time and IP version number are written at this address. Quasi-static entry read: When a dynamic entry ages out, a quasi-static entry read operation is triggered to read the 5-tuple, MAC address, and IP version number of the corresponding address.
5. The high-performance Netflow message output method according to claim 3, characterized in that, Four dynamic processing modules are used to update the statistical information of network traffic on a packet-by-packet basis.
6. The high-performance Netflow message output method according to claim 3, characterized in that, The burst feature of AXI4 in FPGA is used to handle multiple table entries conflict.
7. The high-performance Netflow message output method according to claim 1, characterized in that, When generating quasi-static entries, the access cycle of DDR4 is adjusted. When there are many flow table creation operations, it can automatically switch to the T0 cycle, and when there are many aging operations, it can automatically switch to the T2 cycle. Under normal circumstances, it remains in the T1 stage.
8. A high-performance Netflow message output system based on FPGA with HBM and DDR4, characterized in that, include: The message parsing module is used to extract message header information and obtain the content information of each field in the layer 2, 3, and 4 protocol headers; Dynamic entry processing module: Uses the 5-tuple information obtained from message parsing to calculate the hash value using CRC32 as the entry address to read and write high-bandwidth memory and record the dynamic information of the stream; Quasi-static table entry processing module: Triggers DDR4 memory read and write operations for quasi-static table entries when dynamic table entries age or are created, and records the stream's quintuple, MAC, and IP version number; The Netflow framing module extracts statistics and time information when dynamic entries age, and simultaneously associates and triggers quasi-static entries to obtain the MAC and IP version numbers of the flow, and frames them according to the Netflow frame format for output.
9. A storage medium, characterized in that, The device stores a computer program that, when executed by a processor, causes the processor to perform the steps of the high-performance Netflow message output method as described in any one of claims 1-7.
Citation Information
Patent Citations
Method for realizing high-speed flow table based on FPGA+RLDRAM3
CN115665051A
Flow analysis board card based on DDR efficient storage
CN120216444A