HTTP text restoration method based on DPDK
By using a multi-process architecture based on DPDK and a file-to-disk approach to process TCP fragmented data, the memory limitations and coupling issues in existing TCP fragmented data processing technologies are resolved, achieving efficient and reliable HTTP text restoration and improving the system's debuggability and scalability.
Patent Information
- Application Number
- CN202511441181.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-10-10
AI Technical Summary
Existing technologies struggle to reliably handle complex TCP fragmented data in high-concurrency, bypass monitoring, or low-latency scenarios. They suffer from issues such as memory limitations, unrecoverable failures, and high thread coupling, failing to meet the streaming restoration processing requirements of complex HTTP application layers.
It adopts a multi-process architecture based on DPDK, running process, reassembly, sink and monitor processes respectively. Data is transferred asynchronously through DPDK ring, generating a unique identifier fileID. The reassembled content is stored by file persistence to disk, supporting fault-tolerant reassembly with out-of-order delivery, packet loss and retransmission. The atomic renaming mechanism ensures clear state, decouples processes, and improves the consistency and concurrency security of the processing flow.
It achieves efficient text restoration in complex TCP scenarios, has strong fault tolerance and replay capabilities, reduces the system crash domain, improves system debuggability and scalability, and ensures that downstream modules can easily consume data.
Smart Images

Figure CN120915769A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer networks, and particularly relates to a HTTP text restoration method based on DPDK. BACKGROUND
[0002] HTTP (Hypertext Transfer Protocol) is the most widely used application layer protocol on the Internet, supporting the communication of most Web services, API interfaces, cloud storage, mobile Apps and other systems. In addition, with the popularity of HTTPS (HTTP over TLS), although the number of clear HTTP has decreased, in many border detection devices, further data restoration and behavior analysis of HTTP content after TLS decryption are still required. In a real network environment, due to MTU limitations, unstable links, and large application layer packets, TCP packets are often split into multiple fragments for transmission, forming so-called TCP fragmentation, out-of-order, and retransmission packets. Therefore, it is difficult to accurately obtain complete HTTP packet content using only packet capture tools or protocol stack logs.
[0003] Traditional methods based on kernel protocol stacks or single-thread user mode splicing cannot stably process complex TCP fragmentation data in high-concurrency, bypass monitoring or low-latency scenarios. Existing technologies mostly use memory-based packet buffering methods for recombination, which have problems such as being susceptible to memory limitations, invalidity being unrecoverable, and high thread coupling degree. In the process of fragmentation recombination and subsequent HTTP content analysis coupling, the maintainability and scalability of the system are severely restricted.
[0004] For the above problems, the existing technologies and comparative patents are analyzed as follows: A kind of IP packet fragmentation recombination method and system based on FPGA in Chinese patent publication No. CN119420707A, proposes a scheme that offloads the IP packet fragmentation recombination process from CPU to FPGA for execution, solves the problem of high CPU load of IP fragmentation packets, and uses TCAM for IP packet fragmentation matching to avoid the hash collision problem caused by the common hash operation-based matching method. However, it has the defects of limited processing logic and poor flexibility, cannot be implemented in a pure software environment, has limited application scope, and the processing still stays at the IP layer, which is not suitable for complex HTTP application layer stream restoration processing. The DPDK packet fragmentation processing method and device of Chinese patent publication No. CN116634044A proposes a DPDK-based packet fragmentation processing method, which judges whether there is packet information in a preset hash table, and if not, stores the fragmented packet in the hash table based on the packet information, and transfers the fragmented packet to the corresponding CPU for processing by extracting the corresponding logical core number. However, the scheme is insufficient in handling abnormal situations such as packet disorder, packet loss, and retransmission, and still has certain coupling, and fails to completely decouple the operations of packet fragmentation and reassembly and the main process, lacking more fine-grained fault tolerance mechanism and robustness. The DPDK-based high-speed network packet processing system and method of Chinese patent publication No. CN116996444A proposes a DPDK-based high-speed network packet processing system and method, which has the core advantage of improving the efficiency of packet processing by bypassing the operating system kernel, and is suitable for high-throughput network scenarios. However, its limitation lies in the lack of in-depth design for the reassembly process of TCP fragments, and the lack of processing of situations such as TCP disorder, retransmission, and packet loss in data packets. This scheme focuses more on system optimization at the framework level, and is insufficient in detail processing such as fragmented packet reassembly and text restoration, and cannot meet the high reliability requirements in complex network environments. SUMMARY
[0005] The present application aims to provide a DPDK-based HTTP text restoration method to solve the problems raised in the background art.
[0006] To solve the above technical problems, the technical solution adopted by the present application is: A DPDK-based HTTP text restoration method, comprising the following steps: Step 1, a multi-process architecture based on DPDK is constructed, and process, reassembly, sink, and monitor processes are respectively run, and each process performs asynchronous data transmission through a DPDK ring; Step 2, TCP fragmented HTTP streams are collected by the process, and connections meeting the conditions are parsed and identified to generate a unique identifier fileID combined with a timestamp and a core ID in quintuple; Step 3, fileID is named as fileID1 and is directed to a temporary path, payload and TCP sequence number are extracted, and are sent to the reassembly process for reassembly processing through the ring; Step 4, the reassembly process calculates the offset position of the payload in the target file according to the TCP sequence number, writes the fileID1 file, and supports fault-tolerant reassembly of disorder, retransmission, and packet loss; Step 5, when the recombination completion condition is met, the file indicated by fileID1 is atomically renamed as fileID2 and moved to the success path, ensuring the accurate migration of the file state; Step 6, the sink process judges whether the file exists according to the fileID2 carried by the analysis result, reads the text content splicing result data, and if the file does not exist, transmits the data to the monitor process through the ring for asynchronous polling detection; Step 7, the monitor process asynchronously polls the existence state of fileID2, notifies the sink process after completing recombination, otherwise marks recombination failure after timeout, the sink process reads the file content according to fileID2, splices into formatted data and sends to the middleware for use by the upper module.
[0007] The further improvement of the technical scheme of the application is that the step 1 specifically comprises: The system environment is initialized by using the DPDK framework, including configuring a memory pool, creating a CPU affinity setting and initializing a ring structure of the DPDK for inter-process communication, and four independent processes are created, which are a process, a reassembly process, a sink process and a monitor process, each process is bound to a different CPU core for running, so as to realize decoupling and parallel processing between processes, wherein the process is responsible for traffic collection and protocol analysis, the reassembly process is responsible for fragmentation recombination and text restoration, the sink process is responsible for formatted message splicing, and the monitor process is responsible for text restoration timeout waiting; According to the designed architecture, the processes are deployed to different CPU cores, the DPDK ring structure is created for each process, and the processes perform asynchronous data transmission through the DPDK ring, forming a collaborative workflow.
[0008] The further improvement of the technical scheme of the application is that the step 2 specifically comprises: The process acquires the TCP fragment HTTP stream in the network through multiple ways including network card capture, pcap file reading or socket communication, and performs protocol analysis on the collected TCP fragments, identifies the messages conforming to the HTTP protocol characteristics, including checking the header information of the TCP message, identifying the HTTP request and response message, etc. In the analysis process, the process identifies each HTTP connection by analyzing the five-tuple of the TCP message, wherein each HTTP connection has a unique five-tuple identifier, and the five-tuple is the source IP address, the destination IP address, the source port, the destination port and the protocol type. For each identified HTTP connection, a unique identifier, i.e., fileID, is generated by the process, which is composed of a five-tuple, a timestamp and a core ID of a CPU core currently running the process, and the generated fileID is recorded and associated with the TCP packet fragments related to the HTTP connection, and is sent to the reassembly process through a DPDK ring structure for further processing.
[0009] The further improvement of the technical scheme of the application is that the fileID includes two states, i.e., fileID1 and fileID2, and the fileID is generated in the form of a UUID based on a TCP five-tuple, a timestamp and a core ID of a process running, so as to ensure that the reassembled files of each session are unique and do not conflict. The fileID1 is used to represent an intermediate file that has not been completed reassembled and is located in a temporary path. The fileID2 represents a file that has been completed reassembled or only partially completed reassembled and is located in a success path, and the fileID1 and the fileID2 are converted in state by renaming, so as to avoid that the intermediate state file is misprocessed.
[0010] The further improvement of the technical scheme of the application is that the step 3 specifically includes: The original packet is captured by DPDK, the TCP header information is parsed, the five-tuple is extracted, and the five-tuple is also parsed if the data comes from a pcap file or a socket, a hash table is used to maintain the HTTP connection state, if the five-tuple already exists, the last active time is updated, if the five-tuple does not exist, a new connection context is created, a first identification timestamp and a current CPU core ID are recorded, and then the fileID1 is generated based on the five-tuple, the timestamp and the core ID. It is checked whether the TCP load contains HTTP features, non-HTTP traffic is filtered, and key fields including a payload, a TCP sequence number and the fileID1 are extracted from the packet, wherein the payload is an HTTP header and text data, the TCP sequence number is used for out-of-order rearrangement, and the fileID1 is used for association of the current connection, and then the payload and the sequence number are written into a temporary file, and metadata is bound in the memory, and if the packet is a fragment, the temporary file is appended in sequence number order. A DPDK Ring message structure is constructed, which includes the fileID1, a temporary file path and a sequence number range, the message structure is sent to the reassembly process through the DPDK Ring, and if the Ring is full, a spin waiting or back pressure mechanism is used to avoid packet loss.
[0011] The further improvement of the technical scheme of the present application is that the step 4 specifically comprises: The reassembly process receives the message structure from the process process in the DPDK ring queue, which contains the key information of fileID1, temporary file path and sequence number range, and parses the message structure, extracts fileID1 and related data, locates to the corresponding temporary file path, and then checks whether the corresponding temporary file exists according to fileID1, and if the file does not exist, a new temporary file is created according to fileID1, and if the file already exists, the file is prepared for writing operation; According to the TCP sequence number, the offset position of the payload in the target file is calculated, wherein the TCP sequence number and the initial sequence number of the current message are obtained from the message structure, the offset is calculated, and the payload is written into the corresponding position of the fileID1 file according to the calculated offset, if it is the first packet, offset=0, the payload is written at the beginning of the file, and if it is a non-first packet, the payload is directly written into the specified position of the file according to the offset.
[0012] The further improvement of the technical scheme of the present application is that the reassembly process calculates the offset offset according to the difference between the TCP sequence number and the initial sequence number, that is, the writing offset position, and directly writes into the fileID1 corresponding file, supporting fault tolerance recombination of out-of-order, retransmission and packet loss; The expression of the offset is: ; In the formula, is the offset, which represents the offset position of the payload of the current message in the target file, and is used to determine which position of the file the data of the current message should be written into, is the TCP sequence number, which represents the TCP sequence number of the current message, is the initial sequence number, which represents the TCP sequence number of the first message of the HTTP connection.
[0013] The further improvement of the technical scheme of the present application is that the step 5 specifically comprises: In the reassembly process, after completing the writing operation of the payload of the current message, whether the recombination completion condition is met is checked by comparing whether the current length of the file is consistent with the expected content-length, if the file length is equal to the content-length, it means that all data has been completely written, and the recombination is completed, if the file length is less than the content-length, it means that there are subsequent fragment messages that have not arrived, and the waiting continues, if the recombination is not completed within the preset timeout time, it is considered that the recombination is completed; When it is confirmed that the reorganization is completed or timeout, an atomic renaming operation is performed to atomically rename the temporary file pointed by fileID1 to fileID2, wherein the renaming operation is completed by using an atomic file operation function provided by an operating system, and after the renaming, the file path is migrated from the temporary path to the success path; After the atomic renaming is completed, the file path is migrated from the temporary path to the success path, and the storage location of the file is updated, so that the file after the reorganization is completed can be correctly recognized and processed by a downstream module, and in the success path, the file is identified by fileID2, which indicates that the file has completed the reorganization and is used for subsequent text restoration and analysis.
[0014] The further improvement of the technical scheme of the application is that the step 6 specifically comprises: The sink process receives the data from the process in the DPDK ring queue, parses the data, extracts the key information including fileID2 and related metadata, and locates the corresponding file path according to fileID2, and prepares for the next operation; According to the fileID2 in the analysis result, it is checked whether the file exists, if the file exists, the file content is read, the reorganized HTTP text data is obtained, the text content is spliced with fileID2 to form formatted data, and the formatted data is sent to the middleware or the upper module, if the file does not exist, it is checked whether the current time and the data recording time exceed a preset waiting time threshold, if the threshold is not exceeded, the data is sent to the monitor process through the DPDK ring to request asynchronous polling detection, if the threshold is exceeded, the reorganization failure is marked, and the failure information is spliced to form formatted data, and the formatted data is sent to the middleware or the upper module; According to the file existence and the reorganization state, the processing result data is processed, if the file exists and the reorganization is successful, the spliced formatted data is sent to the middleware or the upper module to complete the final output of the data, if the file does not exist and the asynchronous polling detection is required, the data is sent to the monitor process, and the subsequent processing result is continued to be waited, if the reorganization fails, the failure information is sent to the middleware or the upper module.
[0015] The further improvement of the technical scheme of the application is that the step 7 specifically comprises: The monitor process receives the data from the sink process in the DPDK ring queue, parses the data to obtain fileID2 and related metadata, locates the corresponding file path according to fileID2, and records the receiving time of the data, and at the same time, an asynchronous polling mechanism is started to periodically check whether the fileID2 file exists, and the polling interval can be adjusted according to system load and performance requirements; In the polling process, the existence state of the fileID2 file is checked, if the file exists, the recombination success is marked, the success state information and the fileID2 are packaged, and are sent back to the sink process through the DPDK ring, the sink process is informed that the recombination has been completed, if the file does not exist, whether the current time and the data receiving time exceed the preset timeout threshold is checked, if not, the polling is continued to wait for the file to appear, if yes, the recombination failure is marked, the failure state information and the fileID2 are packaged, and are sent back to the sink process through the DPDK ring, the sink process is informed that the recombination fails; According to the polling result, the state information is sent back to the sink process, for the state information of the recombination success, after the sink process receives the success state, the file content is read according to the fileID2, the file content and the fileID2 are spliced as formatted data, and are sent to the middleware for use of an upper module, for the state information of the recombination failure, after the sink process receives the failure state, the failure information is spliced as formatted data, and is sent to the middleware for subsequent error processing or retry operation of the upper module.
[0016] Due to the adoption of the above technical scheme, the present application has the following technical progress compared with the prior art: The present application provides a HTTP text restoration method based on DPDK, adopts a DPDK multi-process architecture, each function (such as fragmentation recombination, HTTP text restoration, data processing, etc.) independently runs on different CPU cores, asynchronous communication is realized between processes through a Ring queue, and the method has natural isolation and error tolerance capability, so that the system crash domain can be effectively reduced, and the debuggability and expansibility of the system are improved.
[0017] The present application provides a HTTP text restoration method based on DPDK, the prior art usually stores fragmentation message recombination information and cache in the memory, information loss is prone to occur under abnormal restart, memory leakage or large flow, and the memory consumption is high, which is not conducive to engineering implementation, the present application stores the recombination content in a file writing disk mode, calculates the writing offset according to the TCP sequence number, accurately processes complex TCP scenes such as out-of-order, packet loss and retransmission, and does not affect data recombination in the process of machine restart, so that strong fault tolerance and replay capability are achieved.
[0018] The present application provides a HTTP text restoration method based on DPDK, adopts an atomic renaming mechanism of fileID1 to fileID2, ensures clear state and process decoupling, improves the consistency and concurrent safety of the processing flow, and is more easily consumed by downstream modules.
[0019] The application provides a HTTP text restoring method based on DPDK, unfinished tasks are temporarily transferred to a monitor by using a sink and a monitor process bidirectional asynchronous scheduling mechanism, the sink keeps high throughput, the mainstream process can be greatly avoided from being slowed down by IO, and the overall system processing capacity is ensured. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art based on these drawings.
[0021] Figure 1 The figure is a workflow diagram of the present application. Figure 2 The figure is a method flow diagram of the present application. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical solutions and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0023] Embodiment 1, as shown in the present application provides a HTTP text restoring method based on DPDK, including the following steps: Figure 1 , Figure 2 Step 1, a DPDK-based multi-process architecture is constructed, and process processes, reassembly processes, sink processes and monitor processes are respectively run, each process performs asynchronous data transmission through a DPDK ring, a system environment is initialized by using a DPDK framework, including configuring a memory pool, creating a CPU affinity setting and initializing a ring structure of the DPDK for inter-process communication, and four independent processes are created, which are process processes, reassembly processes, sink processes and monitor processes, each process is bound to a different CPU core to run, so as to realize decoupling and parallel processing between processes, wherein the process processes are responsible for traffic collection and protocol analysis, the reassembly processes are responsible for fragmentation and text restoration, the sink processes are responsible for formatted message assembly, and the monitor processes are responsible for text restoration timeout waiting, according to the designed architecture, each process is deployed to a different CPU core, a DPDK ring structure is created for each process, each process performs asynchronous data transmission through the DPDK ring, and a cooperative workflow is formed, wherein the process processes collect HTTP streams, parse key information and send the information to the reassembly processes through the ring, the reassembly processes complete fragmentation and text restoration, and transmit the results to the sink processes, when the sink processes assemble messages, if it is found that text restoration is not completed, the sink processes transmit data to the monitor processes through the ring for timeout waiting detection, the monitor processes feed back state information to the sink processes according to the detection results, and finally data assembly and output are completed, in the whole process, each process realizes efficient and low-delay data exchange through the DPDK ring, and high throughput and low coupling of the system are ensured; Step 2, the process process collects TCP fragment HTTP stream, parses and identifies the connection meeting the condition, generates a unique identifier fileID combined with the quintuple timestamp and core ID, the process process obtains the TCP fragment HTTP stream in the network through various ways including network card capture, pcap file reading or socket communication, and performs protocol analysis on the collected TCP fragments, identifies the messages meeting the HTTP protocol characteristics, including checking the header information of TCP message, identifying HTTP request and response message, etc., in the analysis process, the process process identifies each HTTP connection by analyzing the quintuple of TCP message, each HTTP connection has its unique quintuple identifier, the quintuple is source IP address, destination IP address, source port, destination port and protocol type, for each identified HTTP connection, a unique identifier fileID is generated by the process process, the fileID is composed of quintuple, timestamp and core ID (coreID) of the process process currently running, wherein the quintuple ensures the uniqueness of the connection at the network layer, the timestamp records the time when the connection is identified, which helps to distinguish different connection instances under the same quintuple in subsequent processing, the core ID ensures that the process process instances running on different CPU cores can generate non-conflicting fileIDs, and then records the generated fileID, and associates the TCP fragment messages (including the payload and TCP sequence number) related to the HTTP connection with the fileID, and sends it to the reassembly process through the DPDK ring structure for further processing; In addition, the fileID includes two states of fileID1 and fileID2, the generation method of the fileID includes the UUID generated based on the TCP quintuple, timestamp and core ID of the process running, which ensures the uniqueness of the reassembly file of each session without conflict, wherein fileID1 is used to represent the intermediate file which has not completed reassembly, located in the temporary path, fileID2 represents the file which has completed reassembly or only partially completed reassembly, located in the success path, fileID1 and fileID2 realize state conversion through renaming, avoiding the intermediate state file being misprocessed; Step 3, name fileID as fileID1 and point to a temporary path, extract payload and TCP sequence number, send to reassembly process through ring for reassembly, capture raw packets through DPDK, parse TCP header information, extract five-tuple, if data comes from pcap file or socket, also need to parse five-tuple, use hash table to maintain HTTP connection state, if five-tuple already exists, update its last active time, if not, create a new connection context, record the first identification timestamp and current CPU core ID, then generate fileID1 based on five-tuple, timestamp and core ID, fileID1 is used to identify a unique connection instance and points to a temporary storage path, subsequent packets will be reassembled according to this path, check if TCP payload contains HTTP features, filter non-HTTP traffic, and extract key fields including payload, TCP sequence number and fileID1 from the packet, where payload is HTTP header and body data, TCP sequence number is used for out-of-order rearrangement, and fileID1 is used to associate the current connection, then write payload and sequence number to a temporary file (path pointed by fileID1), and bind metadata in memory at the same time, if the packet is fragmented, append it to the temporary file in sequence number order, construct a DPDK ring message structure, including fileID1, temporary file path and sequence number range, fileID1 is used to identify the target connection, temporary file path is used to point to the payload storage location, and sequence number range is used for reassembly verification, send the message structure to the reassembly process through the DPDK ring, if the ring is full, use spin waiting or back pressure mechanism to avoid packet loss; Step 4, the reassembly process calculates the offset position of the payload in the target file according to the TCP sequence number, writes the fileID1 file, supports fault-tolerant recombination of out-of-order, retransmission and packet loss, the reassembly process receives a message structure from the process process in the DPDK ring queue, the structure contains the key information of fileID1, temporary file path and sequence number range, parses the message structure, extracts fileID1 and related data, locates to the corresponding temporary file path, and then checks whether the corresponding temporary file exists according to fileID1, if the file does not exist, creates a new temporary file according to fileID1, if the file already exists, prepares to write the file, calculates the offset position of the payload in the target file according to the TCP sequence number, wherein the TCP sequence number and the initial sequence number of the current message are obtained from the message structure, the offset is calculated, and the payload is written to the corresponding position of the fileID1 file according to the calculated offset, if it is the first packet, offset = 0, the payload is written at the beginning of the file, if it is not the first packet, the payload is directly written to the specified position of the file according to the offset, instead of simply appending the content, which can effectively handle the problems of out-of-order, retransmission and packet loss; In addition, the reassembly process calculates the offset offset according to the difference between the TCP sequence number and the initial sequence number, that is, the writing offset position, and directly writes the fileID1 corresponding file, supports fault-tolerant recombination of out-of-order, retransmission and packet loss; Wherein, the expression of the offset is: ; In the formula, is the offset, which represents the offset position of the payload of the current message in the target file, which is used to determine where the data of the current message should be written in the file, is the TCP sequence number, which represents the TCP sequence number of the current message, which is a field extracted from the TCP message header, the TCP sequence number is a mechanism used by TCP protocol to identify the order of data segments, each TCP message segment has a sequence number, which is used to ensure reliable transmission of data, in the recombination process, the current sequence number is used to determine the position of the current message in the entire data stream, is the initial sequence number, which represents the TCP sequence number of the first message of the HTTP connection, which is recorded when the connection is established, the initial sequence number is used as a reference point to calculate the offset of the subsequent messages, by taking the difference between the current sequence number and the initial sequence number as the offset, the position of the current message in the entire file can be determined; Step 5, when the reassembly completion condition is met, the file pointed to by fileID1 is atomically renamed as fileID2 and moved to the success path, ensuring accurate migration of the file state. In the reassembly process, after completing the payload write operation on the current packet, the reassembly completion condition is checked by comparing whether the current file length is consistent with the expected content-length. If the file length is equal to the content-length, it means that all data has been completely written and reassembly is complete. If the file length is less than the content-length, it means that subsequent fragment packets have not arrived and waiting continues. If reassembly is not completed within a preset timeout period, it is considered that reassembly is complete. When reassembly is confirmed to be complete or timeout, the atomic renaming operation is performed to atomically rename the temporary file pointed to by fileID1 as fileID2. The atomic operation ensures that there is no intermediate state during the file state migration process, avoiding other processes from reading incomplete files. The renaming operation is completed through an atomic file operation function provided by the operating system. After renaming, the file path is migrated from the temporary path to the success path. After the atomic renaming is completed, the file path is migrated from the temporary path to the success path, updating the storage location of the file, ensuring that the reassembled file can be correctly recognized and processed by downstream modules. In the success path, the file is identified by fileID2, indicating that the file has completed reassembly and is used for subsequent text restoration and analysis. Step 6, the sink process judges whether the file exists and reads the text content assembly result data according to the fileID2 carried by the parsing result, if the file does not exist, the data is transmitted to the monitor process for asynchronous polling detection through the ring, the sink process receives the data from the process process in the DPDK ring queue, parses the data, extracts the key information including fileID2 and related metadata, and locates to the corresponding file path according to fileID2, and prepares for the next operation, according to fileID2 in the parsing result, checks whether the file exists, if the file exists, reads the file content, obtains the reassembled HTTP text data, assembles the text content and fileID2 into formatted data, and prepares to send to the middleware or upper module, if the file does not exist, checks whether the current time and the data recording time exceed the preset waiting time threshold, if the threshold is not exceeded, the data is sent to the monitor process through the DPDK ring, and asynchronous polling detection is requested, if the threshold is exceeded, the reassembly failure is marked, the failure information is assembled into formatted data, and sent to the middleware or upper module, according to whether the file exists and the reassembly state, the result data is processed, if the file exists and the reassembly is successful, the assembled formatted data is sent to the middleware or upper module, and the final output of the data is completed, if the file does not exist and needs asynchronous polling detection, the data is sent to the monitor process, and the subsequent processing result is continued to wait, if the reassembly fails, the failure information is sent to the middleware or upper module, so as to carry out subsequent error processing or retry operation; Step 7, the monitor process asynchronously polls the existence state of fileID2, notifies the sink process after completing reassembly, otherwise marks reassembly failure after timeout, the sink process reads file content according to fileID2, assembles into formatted data, and sends to middleware for use by upper layer modules, the monitor process receives data from the sink process in the DPDK ring queue, parses the data to obtain fileID2 and related metadata, locates the corresponding file path according to fileID2, and records the reception time of the data, at the same time, an asynchronous polling mechanism is started, and the existence of fileID2 is checked periodically, the polling interval can be adjusted according to system load and performance requirements, in the polling process, the existence state of fileID2 is checked, if the file exists, it is marked as reassembly success, the success state information and fileID2 are packaged, and sent back to the sink process through the DPDK ring, notifying that reassembly is completed, if the file does not exist, it is checked whether the current time exceeds the preset timeout threshold, if not, the polling is continued, if yes, it is marked as reassembly failure, the failure state information and fileID2 are packaged, and sent back to the sink process through the DPDK ring, notifying that reassembly fails, according to the polling result, the state information is sent back to the sink process, for the reassembly success state information, the sink process reads file content according to fileID2 after receiving the success state, assembles it with fileID2 into formatted data, and sends it to middleware for use by upper layer modules, for the reassembly failure state information, the sink process assembles the failure information into formatted data after receiving the failure state, and sends it to middleware for subsequent error handling or retry operation by upper layer modules.
[0024] As shown in Embodiment 2, on the basis of Embodiment 1, the application provides a technical solution: preferably, 1, in the process, complete flow collection and protocol analysis, support pcap / pcapng files and network card flow, socket communication and other ways to obtain data stream, the specific process is as follows: Figure 1 、 Figure 2 As shown in Embodiment 2, on the basis of Embodiment 1, the application provides a technical solution: preferably, 1, in the process, complete flow collection and protocol analysis, support pcap / pcapng files and network card flow, socket communication and other ways to obtain data stream, the specific process is as follows: Process 1.1, when HTTP flow meeting the fragmentation feature is found, a hash node is created or found with five-tuple as key, a file UUID is generated according to timestamp and core id of process running, a file name fileID1:path-temp / UUID is recorded, a copy of message payload is copied, and message tcp sequence number is recorded, and the reassembly ring is thrown to the reassembly process; Flow 1.2, return to the original flow after non-blocking processing, record the file name fileID2:path-success / UUID, and throw the parsed data to the sink ring after processing is completed, which does not affect the performance of the module; Flow 1.3, after receiving the same five-tuple packet again, if it is successful, execute flow 1.1 to directly end, and do not need to continue flow 1.2.
[0025] 2. In the reassembly process, data is collected from the process through the reassembly ring, and the offset is calculated according to the init TcpSeqNum of the five-tuple and the current TcpSeqNum of the packet. The specific flow is as follows: Flow 2.1, offset is 0, indicating that it is the first packet, a new file fileID1 is created and the packet payload is recorded.
[0026] Flow 2.2, non-first packet is recorded in the fileID1 file according to the offset, instead of directly appending the content, which solves the problems of out-of-order, packet loss and retransmission. Flow 2.3, according to whether the length of fileID1 file is equal to the content-length of the packet, it is judged whether the restoration is completed. If it is completed or timed out, fileID1 (path-temp / UUID) is renamed to fileID2 (path-success / UUID), and the file path is moved.
[0027] 3. Sink process, data is received from process and monitor process through sink ring, and data is formatted and assembled according to the specification. If fileID2 field is found in the data, it is considered that the data is a fragmented flow, and it is checked whether the file corresponding to fileID2 exists. The specific flow is as follows: Flow 3.1, if the file exists, it means that the restoration has been completed, the file content is read, the restored text content and fileID2 are assembled together to form formatted data, which is sent to the middleware for use by the upper module.
[0028] Flow 3.2, if the file does not exist, it means that the restoration has not been completed. It is detected whether the actual time and the data recording time exceed the waiting time threshold. If it exceeds, it means that the restoration fails, fileID2 is deleted, and the data is assembled into formatted data and sent to the middleware for use by the upper module.
[0029] Flow 3.3, if the file does not exist but the waiting time threshold has not been exceeded, the data is thrown to the monitor process through the monitor ring, which does not block the sink process.
[0030] 4. The monitor process receives data through the monitor ring and checks if the fileID2 file exists. The specific process is as follows: Process 4.1: If the file exists, mark the data to indicate that fileID2 was successfully restored. The sink process does not need to check again and throws the data back to the sink ring. After receiving the data, the Monitor assembles and formats the data according to the scheme in process 3.1. In process 4.2, if the file does not exist and the waiting time threshold has been exceeded, the data is marked as indicating that fileID2 restoration has failed. The sink process does not need to check further and throws the data back to the sink ring. After receiving the data, the monitor assembles and formats the data according to the scheme in process 3.2. In process 4.3, if the file does not exist and the waiting time threshold has not been exceeded, the data is thrown back into the monitor ring and placed at the end of the ring queue, repeating processes 4.1 and 4.2.
[0031] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A DPDK-based HTTP text restoration method, characterized in that, The method comprises the following steps: Step 1, a DPDK-based multi-process architecture is constructed, and process, reassembly, sink and monitor processes are respectively run, and the processes are connected through a DPDK ring for asynchronous data transmission; Step 2, a TCP fragment HTTP stream is collected by the process, and a connection meeting a condition is parsed and recognized to generate a unique identifier fileID; Step 3, the fileID is named as fileID1 and is pointed to a temporary path, a payload and a TCP sequence number are extracted, and the payload and the TCP sequence number are sent to the reassembly process through the ring for recombination processing; Step 4, the reassembly process calculates an offset position of the payload in a target file according to the TCP sequence number, and writes the fileID1 file; Step 5, when a recombination completion condition is met, a file pointed by the fileID1 is atomically renamed as fileID2 and is moved to a success path; Step 6, the sink process judges whether a file exists according to the fileID2 carried by a parsing result, reads text content splicing result data, and if the file does not exist, transmits the data to the monitor process through the ring for asynchronous polling detection; Step 7, the monitor process asynchronously polls an existence state of the fileID2, notifies the sink process after recombination is completed, otherwise marks recombination failure after a timeout, the sink process reads file content according to the fileID2, splices the file content into formatted data, and sends the formatted data to a middleware for use by an upper module. 2.The HTTP text restoration method based on DPDK of claim 1, wherein: The step 1 specifically comprises: A system environment is initialized by using a DPDK framework, including configuration of a memory pool, creation of a CPU affinity setting, initialization of a ring structure of the DPDK for inter-process communication, and creation of four independent processes, namely, a process, a reassembly process, a sink process and a monitor process, wherein the process is responsible for traffic collection and protocol analysis, the reassembly process is responsible for fragmentation recombination and text restoration, the sink process is responsible for formatted message splicing, and the monitor process is responsible for text restoration timeout waiting; According to the designed architecture, the processes are deployed to different CPU cores, a DPDK ring structure is created for each process, the processes are connected through the DPDK ring for asynchronous data transmission, and a cooperative work flow is formed.
3. The HTTP text restoration method based on DPDK according to claim 1, characterized in that: The step 2 specifically comprises: The process acquires a TCP fragment HTTP stream in a network through multiple ways including a network card capture, a pcap file reading or a socket communication, and performs protocol analysis on the collected TCP fragments to recognize a message meeting a characteristic of an HTTP protocol; In the analysis process, the process process identifies each HTTP connection by analyzing the five-tuple of the TCP message, wherein each HTTP connection has its unique five-tuple identifier, and the five-tuple is the source IP address, the destination IP address, the source port, the destination port and the protocol type; For each identified HTTP connection, a unique identifier, i.e. fileID, is generated by the process process, which is composed of the five-tuple, the timestamp and the CPU core ID currently running by the process process, and then the generated fileID is recorded, and the TCP fragment message related to the HTTP connection is associated with the fileID and sent to the reassembly process for further processing through the DPDK ring structure.
4. The HTTP text restoration method based on DPDK according to claim 3, characterized in that: The fileID includes two states of fileID1 and fileID2, and the generation method of the fileID includes the UUID generated based on the TCP five-tuple, the timestamp and the core ID running by the process. The fileID1 is used to represent the intermediate file which has not completed reassembly, and is located in the temporary path. The fileID2 represents the file which has completed reassembly or only partially completed reassembly, and is located in the success path.
5. The DPDK-based HTTP text restoration method according to claim 4, characterized in that: The step 3 specifically includes: The original packet is captured by DPDK, the TCP header information is analyzed, the five-tuple is extracted, the HTTP connection state is maintained by using a hash table, if the five-tuple already exists, the last active time is updated, if it does not exist, a new connection context is created, the first identification timestamp and the current CPU core ID are recorded, and then the fileID1 is generated based on the five-tuple, the timestamp and the core ID; It is checked whether the TCP load contains the HTTP feature, the non-HTTP flow is filtered, and the key fields including the payload, the TCP sequence number and the fileID1 are extracted from the message, and then the payload and the sequence number are written into the temporary file, and the metadata is bound in the memory, if the message is a fragment, it is appended to the temporary file in sequence according to the sequence number; The DPDK Ring message structure is constructed, including the fileID1, the temporary file path and the sequence number range, and the message structure is sent to the reassembly process through the DPDK Ring.
6. The DPDK-based HTTP text restoration method according to claim 5, characterized in that: The step 4 specifically includes: The reassembly process receives the message structure from the process process in the DPDK ring queue, analyzes the message structure, extracts the fileID1 and the related data, locates to the corresponding temporary file path, and then checks whether the corresponding temporary file exists according to the fileID1, if the file does not exist, a new temporary file is created according to the fileID1, if the file already exists, the file is prepared for writing operation; According to the TCP sequence number, the offset position of the payload in the target file is calculated, wherein the TCP sequence number and the initial sequence number of the current packet are obtained from the message structure, the offset is calculated, and the payload is written into the corresponding position of the fileID1 file according to the calculated offset; if it is the first packet, offset = 0, the payload is written at the beginning of the file, and if it is a non-first packet, the payload is directly written into the specified position of the file according to the offset.
7. The DPDK-based HTTP text restoration method according to claim 6, characterized in that: The reassembly process calculates the offset offset according to the difference between the TCP sequence number and the initial sequence number, that is, the offset position is written, and the file corresponding to fileID1 is directly written; The expression of the offset is: ; In the formula, is an offset, indicating the offset position of the payload of the current packet in the target file, is a TCP sequence number, indicating the TCP sequence number of the current packet, is an initial sequence number, indicating the TCP sequence number of the first packet of the HTTP connection. 8.The HTTP text restoration method based on DPDK of claim 1, wherein: The step 5 specifically includes: In the reassembly process, after the writing of the payload of the current packet is completed, whether the recombination completion condition is met is checked by comparing whether the current length of the file is consistent with the expected content-length; if the file length is equal to the content-length, it is indicated that all data has been completely written, the recombination is completed, if the file length is less than the content-length, it is indicated that there is a subsequent fragment packet that has not arrived, the waiting continues, and if the recombination is not completed within a preset timeout period, it is considered that the recombination is completed; When the recombination is completed or the timeout is confirmed, an atomic renaming operation is performed to atomically rename the temporary file pointed to by fileID1 to fileID2, wherein the renaming operation is completed by using an atomic file operation function provided by the operating system; after the renaming, the file path is migrated from the temporary path to the success path; After the atomic renaming is completed, the file path is migrated from the temporary path to the success path, the storage position of the file is updated, and in the success path, the file is identified by fileID2, indicating that the file has completed the recombination. 9.The HTTP text restoration method based on DPDK of claim 8, wherein: The step 6 specifically includes: The sink process receives the data from the process process from the DPDK ring queue, parses the data, extracts key information including fileID2 and related metadata, and locates to the corresponding file path according to fileID2, and prepares for the next operation; According to the fileID2 in the analysis result, whether the file exists is checked; if the file exists, the file content is read, the recombined HTTP text data is obtained, the text content and fileID2 are spliced into formatted data, and the formatted data is sent to the middleware or the upper module; if the file does not exist, whether the current time and the data recording time exceed a preset waiting time threshold is checked; if the threshold is not exceeded, the data is sent to the monitor process through the DPDK ring to request asynchronous polling detection; if the threshold is exceeded, the recombination failure is marked, the failure information is spliced into formatted data, and the formatted data is sent to the middleware or the upper module; According to the file existence and reorganization state, the processing result data is processed, if the file exists and the reorganization succeeds, the assembled formatted data is sent to the middleware or the upper module, the final output of the data is completed, if the file does not exist and needs to be asynchronously polled and detected, the data is sent to the monitor process, and subsequent processing results are continued to be waited, if the reorganization fails, the failure information is sent to the middleware or the upper module.
10. The HTTP text restoration method based on DPDK according to claim 9, characterized in that: The step 7 specifically comprises: The monitor process receives the data from the sink process in the DPDK ring queue, parses the data to obtain fileID2 and related metadata, positions to the corresponding file path according to fileID2, and records the receiving time of the data, simultaneously, an asynchronous polling mechanism is started, and whether the fileID2 file exists is periodically checked; In the polling process, the existence state of the fileID2 file is checked, if the file exists, the reorganization is marked as successful, the success state information and the fileID2 are packaged, and are sent back to the sink process through the DPDK ring, the reorganization is notified to be completed, if the file does not exist, whether the current time and the data receiving time exceed a preset timeout threshold value is checked, if the timeout is not exceeded, the file is continued to be polled and waited, if the timeout is exceeded, the reorganization is marked as failed, the failure state information and the fileID2 are packaged, and are sent back to the sink process through the DPDK ring, the reorganization is notified to fail; According to the polling result, the state information is sent back to the sink process, for the state information of the reorganization success, after the sink process receives the success state, the file content is read according to the fileID2, the file content is assembled with the fileID2 as formatted data, and is sent to the middleware for the use of the upper module, for the state information of the reorganization failure, after the sink process receives the failure state, the failure information is assembled as formatted data, and is sent to the middleware for the subsequent error processing or retry operation of the upper module.
Citation Information
Patent Citations
DPDK fragmented message processing method and device
CN116634044A
IP message fragmentation and recombination method and system based on FPGA
CN119420707A
Method and device for processing fragment data packets
CN106685862A
TCP / IP (Transmission Control Protocol / Internet Protocol) flow restoration method
CN114629970A
High-speed network packet processing system and method based on DPDK
CN116996444A
Cited By
Timer adjusting method and system based on DPDK dynamic load awareness
CN121957915A
A timer adjustment method and system based on DPDK dynamic load awareness
CN121957915B