A Distributed Data Capture Method Based on DPDK

By adopting a distributed data capture method based on DPDK in the load balancing network, combining the blockchain network and distributed SDN controller, the shortcomings of traditional traffic capture methods in the load balancing scenario are solved, efficient data acquisition and storage are achieved, and system performance and resource utilization are improved.

CN116418700BActive Publication Date: 2025-06-20GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310478318.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-28
Publication Date
2025-06-20
Estimated Expiration
2043-04-28

AI Technical Summary

Technical Problem

Traditional traffic capture methods are not enough to bear huge network traffic in load balancing scenarios and may lead to incomplete data flow.

Method used

The distributed data capture method based on DPDK is adopted to realize data flow information collection, task allocation and data dumping through blockchain network and distributed SDN controller, and the multi-core architecture and memory file system of DPDK are used to improve packet processing efficiency and storage speed.

Benefits of technology

Implement efficient data acquisition in a load-balancing network, improve system throughput and data storage speed, maximize the utilization of network equipment resources, and effectively respond to emergencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116418700B_ABST
    Figure CN116418700B_ABST
Patent Text Reader

Abstract

The present invention discloses a distributed data capture method based on DPDK, including: 1) data stream information collection; 2) task allocation; 3) data dump. This method can perform data collection on a load-balanced network, improve the throughput of the system, and improve the storage speed of data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to computer and information technology, and particularly to a distributed data capture method based on DPDK. Background Art

[0002] Data capture is a process of intercepting and recording data stream packets transmitted in a network. When data packets are transmitted in the network, a capture device captures each data packet and stores the data packet so that a network data analyzer can analyze and audit the traffic data in the network. Traffic capture is very crucial for many fields such as network traffic analysis, network security audit, and network forensics.

[0003] Traditional traffic capture methods are carried out by using the kernel protocol stack of an operating system or modifying the kernel protocol stack of an operating system, such as Tcpdump, Wireshark, etc. Tcpdump is a command-line traffic capture and network monitoring tool based on libpcap, which can help users capture and store the network traffic of the current device. Libpcap runs in the kernel network protocol stack of devices such as the host side or router, and through the filter and bypass mechanism of the kernel network protocol stack, constructs a filter to listen to all data packets flowing through the target network card, copies the listened data packets, filters them according to the rules defined by the user, and then delivers the captured data to the relevant upper-layer applications in the user space. Wireshark is a network protocol analyzer and packet sniffer that can perform real-time capture of data packets and in-depth offline analysis of protocols and packet contents. Shane et al. proposed a C language library named libtrace for network packet capture and processing. Libtrace provides a simple and easy-to-use function interface, which helps to develop more user-friendly, more reliable network trace analysis and monitoring tools. However, this method has the overhead caused by interrupts from the kernel space to the user space and memory copy, resulting in a certain performance waste.

[0004] Some researchers adopt the zero-copy method to reduce the latency caused by data transmission. Luigi et al. proposed a packet processing framework called Netmap. Through means such as memory mapping, Netmap maps the buffer of the captured packets to the user space and implements the main program structure in the user space. By means of memory mapping, memory pre-allocation, etc., Netmap eliminates the overhead of system calls, memory application, memory copy, etc. from the kernel space to the user space, making the framework have the characteristics of zero-copy and greatly improving the packet capture performance; Jiawei et al. proposed an adaptive packet capture scheme based on PF_RING, which can dynamically allocate the buffer according to the amount of data in the network, improving the packet capture performance. When the traffic in the network changes significantly, it can automatically increase or decrease the cache space size in the kernel to ensure that there is no packet loss phenomenon caused by insufficient buffer during the packet capture process, making the occurrence rate of packet loss in this scheme greatly reduced compared with the original PF_RING, and at the same time alleviating the waste of memory resources; Paul et al. proposed a packet capture and storage scheme, which caches the packets in a circular queue in memory and writes the packets to disk when a specified event occurs to improve the writing efficiency; Hyun et al. proposed a malicious packet capture method based on DNS sinkhole to increase the capture ratio of malicious packets; Martino et al. proposed a packet capture and analysis scheme, which combines the Intel Data Plane Development Kit with a traffic analyzer to improve the data processing speed.

[0005] In recent years, some researchers have proposed to use dedicated hardware devices for packet capture. Siyi et al. proposed a network traffic capture and replay solution based on Field-Programmable Gate Array (FPGA), which can ensure high precision and high throughput in the packet capture timestamp. Salvatore et al. proposed a framework for building stateful packet processing functions in hardware, which supports complex network functions and hides the low-level hardware implementation from programmers; Jakub et al. proposed a network traffic capture scheme based on FPGA, which can write packets into the host memory through PCI-E at a transmission rate of 400Gbps; Han et al. proposed a network traffic capture scheme FPC-NM based on FPGA. This scheme is divided into two parts: hardware and software. The hardware part is implemented based on FPGA and preprocesses the packets, such as timestamp, load balancing allocation, TCP segment recombination, etc. The software part further processes different data streams by different CPUs according to the preprocessing results of the hardware part. This scheme combines software and hardware, greatly improving the processing performance.

[0006] However, with the rapid development of network technology, more and more emerging technologies such as cloud computing and big data are becoming increasingly dependent on the network. With the increasing dependence of emerging technologies on the network and the expansion of the application scale, a large amount of data traffic has been generated in the network at the same time. In scenarios such as cloud computing and big data, a single network link is no longer sufficient to carry such a huge amount of traffic. Therefore, network managers usually use flow load balancing to improve the load capacity of the entire network.

[0007] The load balancing association of traffic improves the utilization rate of network resources through methods such as flow redistribution and flow splitting. The flow redistribution technology refers to, through a certain algorithm, dynamically calculating a suitable forwarding path for network flows according to the current network environment, and finally realizing network load balancing and improving the utilization rate of network resources. The flow splitting technology refers to adopting different forwarding paths for traffic of different scales, splitting large-scale data flows into smaller sub-flows, and forwarding the sub-flows along different paths, which improves the reliability and availability of the network.

[0008] In the load balancing scenario, both flow redistribution and flow splitting will change the transmission path of the data flow. The path of the data packets in the same data flow from the source node to the destination node is uncertain. If traffic capture is performed before load balancing traffic splitting, the existing traffic capture methods are not sufficient to handle the huge network traffic; if traffic capture is performed after load balancing traffic splitting, it may cause the situation of incomplete data flow. To sum up, the traditional traffic capture methods are no longer applicable to traffic capture in the load balancing scenario. Summary of the Invention

[0009] The object of the present invention is to provide a distributed data capture method based on DPDK in view of the deficiencies of the prior art. This method can perform data collection on a load balancing network, improve the throughput of the system, and improve the data storage speed.

[0010] The technical solution for achieving the object of the present invention is:

[0011] A distributed data capture method based on DPDK, the method is applicable to a load balancing network, wherein a single node consists of flow rule management, data flow information collection, task allocation, data flow authentication, data packet processing, and network data capture. The upper layer of the method is a blockchain network for communication and a distributed SDN controller based on the blockchain network, and the lower layer is a software router with traffic capture developed based on DPDK. The method includes the following steps:

[0012] 1) Data stream information collection: Network status collection is located on the router node, which periodically collects the status information of the current node and notifies other nodes through the blockchain network. The data stream information collection process collects the data packets passing through the current router during the current time period, classifies and counts the number of forwarded data packets and the total number of forwarded bytes during the current time period according to the data stream characteristics, and organizes them together with the timestamp, the length of the time period, and the system load time series data into a structured message, and broadcasts and publishes it through the blockchain network. First, the data stream forwarding module forwards the data stream according to the matching rules, and records the information of the data stream while forwarding the data stream. The forwarding module obtains the data block storing the network data according to the matching rules and updates the relevant information of the data block. The data collection module regularly obtains the data stream status information dumped by the forwarding module, packs the status information into a specific data structure, and broadcasts it through the blockchain. Specifically, the data stream information collection process is as follows: First, obtain the current time, obtain the storage object storing the information of the current time period from the data stream information, and calculate the duration of the current time period. If the duration of the current time period exceeds the predefined time period length, update the data information of the current time period to the data stream information, generate a data stream status event, broadcast it to all router nodes through the blockchain network, and at the same time generate a storage object that clears the information of the current time period and reset the start time of the time period. Then, accumulate the number of received data packets and the number of received bytes during the current time period according to the data packet information;

[0013] 2) Task allocation: The task allocation algorithm subscribes to the status information of all nodes in the current network through the blockchain, and calculates the task arrangement of each current node according to the network topology information and the current status information of the network, based on the following task allocation algorithm. The task allocation algorithm obtains the network topology and data stream information through the blockchain network. First, the task allocation algorithm traverses the network topology, obtains the initial value of the load capacity of each router node, and stores the initial value of the load capacity in the array R load Secondly, the task allocation algorithm obtains the current network status information through the blockchain network, calculates the processing overhead of each data capture task, and stores the task processing overhead in F cost Then, according to the information collected in the data stream information collection process in step 1, traverse all the data stream information in the network, mark the nodes on the data stream path, indicating that the capture task of the data stream can be executed on the marked nodes, and store the traversal result in the two-dimensional array V. Finally, take the node load capacity R load , the overhead required for the data stream F cost , and the optional matrix V of task allocation as parameters and pass them into the capture task allocation algorithm to obtain the task allocation result F route, the task allocation monitors the data flow information in the network. When the data flow scale changes and exceeds the upper limit of the data that a single node can process, the task allocation algorithm part re-allocates the tasks and obtains a new task allocation scheme. The performance of each router node and the load required to process the data flow are quantified. Assuming that the maximum number of data packets that a single router node route can dump in a unit time is the load capacity of the node, as shown in formula (1), the set of router node load capacity is:

[0014]

[0015] Assume that the number of packets that a data flow needs to process per unit time is the load required by the flow, as shown in formula (2). The set of loads required by the data flow is:

[0016]

[0017] Assume that via i,j Indicates data flow j Passing through router route i , then the symbol V is used to represent the optional matrix of data capture task allocation, as shown in formula (3):

[0018]

[0019] When there is only one data flow in the entire network or the paths of all data flows in the network do not overlap, the load capacity of the nodes through which the data flow passes should be selected. The largest node is used as the execution node of the task. When the transmission path of the data flow in the entire network is the same, the data capture task is captured by any node through which the data flow passes. At this time, a greedy algorithm is used, that is, according to the overhead of the capture task and the load capacity of the node, the node with the largest load capacity is selected in turn to execute the task with the largest overhead, and the load capacity of the node is recalculated until all tasks are assigned. The tasks are evenly assigned to the nodes on the path. The task allocation algorithm gives priority to data flows with relatively few optional task execution nodes to prevent allocation failure problems caused by improper task allocation order. The task allocation process calls tasks with relatively few optional task execution nodes low-heat tasks and tasks with more optional task execution nodes high-heat tasks. The ratio of the maximum available load of the data flow to the data flow overhead indicates the hotness of the data flow, and the hotness is used as the priority for node selection. When the maximum available load of all tasks is the same, the task allocation process degenerates into priority sorting according to overhead:

[0020] Assume that flow is used path Indicates the router path that the data flow passes through, and the hotness and coldness of the data flow flowpop The calculation method is shown in formula (4):

[0021]

[0022] The set of the hot and cold degrees of the data stream is shown in formula (5):

[0023]

[0024] The algorithm calculates the ratio of the maximum available load of all tasks to the data stream overhead, adds all data capture tasks to the priority queue, and sorts them in ascending order according to the ratio of the maximum available load required by the task to the data stream overhead, so that tasks with low heat are preferentially allocated, ensuring that the algorithm will not have the problem of insufficient node margin caused by high-heat tasks piling up on individual nodes, resulting in task allocation failure. Secondly, the optional nodes of each task are sorted according to the load capacity, so that high-performance nodes preferentially select tasks. Then, according to the priority queue, the candidate tasks of all nodes are traversed to select a suitable task arrangement plan. Select the task at the head of the task queue, select the optional nodes of the task according to the priority, and judge whether the node meets the conditions in turn. If the load required by the task is less than or equal to the remaining load capacity of the node, the task is allocated to the current node, and the remaining load of the node is recalculated to end the current loop. Otherwise, continue to traverse the next node. When the task queue is empty, end the loop and return the execution result of the algorithm;

[0025] 3) Data dump: The data dump adopts the DPDK framework and the memory file system. The data dump utilizes the multi-core architecture of DPDK and the characteristics of CPU and thread binding, and allocates a specific logical core for each thread, so that the processing process of each data stream is carried out on a specific core. At the same time, the data dump adopts a memory-hard disk multi-level cache for the collected data files. Workflow: First, when a data packet arrives at the current node, the data stream forwarding module obtains the permission information of the data stream according to the tuple information of the data stream. If the data stream is marked as a captured data stream, the data packet and the data structure storing the data stream are pushed into the queue. The data dump obtains the data packet and the data stream information from the queue, and obtains the file handle of the data dump file of the data stream according to the data stream information, and writes the data packet into the data dump file through the file handle. After writing the file, calculate the size of the data dump file after dumping the current data packet, and judge whether the file size exceeds the upper limit (64M or 128M). If the file exceeds the upper limit, create a new data dump file to replace the current file, reconfigure the data dump object of the data stream, and move the file to the disk for persistent storage.

[0026] This technical solution is used to handle the problem of traffic capture in the load balancing scenario of an industrial control network, capture unauthenticated data streams or data streams specified by network administrators in the network. This technical solution periodically collects data stream information in the network and generates a capture task allocation scheme that maximally utilizes the performance of each node according to the network traffic information, so as to achieve the maximization of the utilization of traffic capture device resources in the network and effectively respond to emergencies. In addition, it is designed based on the DPDK technology, which can make full use of the performance of multi-core processors, thereby significantly improving the throughput of the system. It uses user space driver technology to achieve high-speed capture of data packets, bypasses the kernel network protocol stack, effectively reduces the delay of data packet processing, improves the efficiency and speed of data packet capture. At the same time, it adopts a memory-disk multi-level storage strategy to effectively improve the storage speed of data.

[0027] This method can perform data collection on a load balancing network, improve the throughput of the system, and improve the storage speed of data. Brief Description of the Drawings

[0028] Figure 1 It is a schematic diagram of the architecture of the node in the embodiment;

[0029] Figure 2 It is a schematic diagram of the data dump process in the embodiment;

[0030] Figure 3 It is a schematic diagram of the network architecture in the embodiment. Detailed Embodiment

[0031] The following further elaborates on the content of the present invention in conjunction with the drawings and embodiments, but it is not a limitation to the present invention.

[0032] Embodiment:

[0033] A distributed data capture method based on DPDK, which is applicable to a load balancing network. Among them, as Figure 1 shown, a single node consists of flow rule management, data stream information collection, task allocation, data stream authentication, data packet processing, and network data capture. The upper layer of this method is a blockchain network for communication and a distributed SDN controller based on the blockchain network, and the lower layer is a software router with traffic capture developed based on DPDK. The method includes the following steps:

[0034] 1) Data stream information collection: Network status collection is located on the router node, which periodically collects the status information of the current node and notifies other nodes through the blockchain network. The data stream information collection process collects the data packets passing through the current router during the current time period, classifies and counts the number of forwarded data packets and the total number of forwarded bytes during the current time period according to the data stream characteristics, and organizes them together with the timestamp, the length of the time period, and the system load time series data into a structured message, and broadcasts and publishes it through the blockchain network. First, the data stream forwarding module forwards the data stream according to the matching rules, and while forwarding the data stream, records the information of the data stream. The forwarding module obtains the data block storing the network data according to the matching rules and updates the relevant information of the data block. The data collection module regularly obtains the data stream status information dumped by the forwarding module, packs the status information into a specific data structure, and broadcasts it through the blockchain. Specifically, the data stream information collection process is as follows: First, obtain the current time, obtain the storage object storing the information of the current time period from the data stream information, and calculate the duration of the current time period. If the duration of the current time period exceeds the predefined time period length, update the data information of the current time period to the data stream information, generate a data stream status event, broadcast it to all router nodes through the blockchain network, and at the same time generate a storage object that clears the information of the current time period and reset the start time of the time period. Then, accumulate the number of received data packets and the received bytes during the current time period according to the data packet information;

[0035] 2) Task allocation: The task allocation algorithm subscribes to the status information of all nodes in the current network through the blockchain, and calculates the task arrangement of each current node according to the network topology information and the current status information of the network, based on the following task allocation algorithm. The task allocation algorithm obtains the network topology and data stream information through the blockchain network. First, the task allocation algorithm traverses the network topology, obtains the initial value of the load capacity of each router node, and stores the initial value of the load capacity in the array R load Secondly, the task allocation algorithm obtains the current network status information through the blockchain network, calculates the processing overhead of each data capture task, and stores the task processing overhead in F cost Then, according to the information collected in the data stream information collection process in step 1, traverse all the data stream information in the network, mark the nodes on the data stream path, indicating that the capture task of the data stream can be executed on the marked nodes, and store the traversal result in the two-dimensional array V. Finally, take the node load capacity R load , the overhead required for the data stream F cost , and the optional matrix V of task allocation as parameters and pass them into the capture task allocation algorithm to obtain the task allocation result F route, the task allocation monitors the data flow information in the network. When the data flow scale changes and exceeds the upper limit of the data that a single node can process, the task allocation algorithm part re-allocates the tasks and obtains a new task allocation scheme. The performance of each router node and the load required to process the data flow are quantified. Assuming that the maximum number of data packets that a single router node route can dump in a unit time is the load capacity of the node, as shown in formula (1), the set of router node load capacity is:

[0036]

[0037] Assume that the number of packets that a data flow needs to process per unit time is the load required by the flow, as shown in formula (2). The set of loads required by the data flow is:

[0038]

[0039] Assume that via i,j Indicates data flow j Passing through router route i , then the symbol V is used to represent the optional matrix of data capture task allocation, as shown in formula (3):

[0040]

[0041] When there is only one data flow in the entire network or the paths of all data flows in the network do not overlap, the load capacity of the nodes through which the data flow passes should be selected. The largest node is used as the execution node of the task. When the transmission path of the data flow in the entire network is the same, the data capture task is captured by any node through which the data flow passes. At this time, a greedy algorithm is used, that is, according to the overhead of the capture task and the load capacity of the node, the node with the largest load capacity is selected in turn to execute the task with the largest overhead, and the load capacity of the node is recalculated until all tasks are assigned. The tasks are evenly assigned to the nodes on the path. The task allocation algorithm gives priority to data flows with relatively few optional task execution nodes to prevent allocation failure problems caused by improper task allocation order. The task allocation process calls tasks with relatively few optional task execution nodes low-heat tasks and tasks with more optional task execution nodes high-heat tasks. The ratio of the maximum available load of the data flow to the data flow overhead indicates the hotness of the data flow, and the hotness is used as the priority for node selection. When the maximum available load of all tasks is the same, the task allocation process degenerates into priority sorting according to overhead:

[0042] Assume that flow is used path Indicates the router path that the data flow passes through, and the hotness and coldness of the data flow flowpop The calculation method is as shown in formula (4):

[0043]

[0044] The set of data flow hot and cold degrees is as shown in formula (5):

[0045]

[0046] The algorithm calculates the ratio of the maximum available load of all tasks to the data flow overhead, adds all data capture tasks to the priority queue, and sorts them in ascending order according to the ratio of the maximum available load required by the task to the data flow overhead, so that tasks with low heat are preferentially allocated, ensuring that the algorithm will not have the problem of insufficient node margin caused by high-heat tasks piling up on individual nodes, resulting in task allocation failure. Secondly, the optional nodes of each task are sorted according to the load capacity, so that high-performance nodes preferentially select tasks. Then, according to the priority queue, the candidate tasks of all nodes are traversed to select a suitable task arrangement plan. Select the task at the head of the task queue, select the optional nodes of this task according to the priority, and sequentially judge whether the nodes meet the conditions. If the load required by the task is less than or equal to the remaining load capacity of the node, the task is allocated to the current node, and the remaining load of the node is recalculated, ending the current loop. Otherwise, continue to traverse the next node. When the task queue is empty, end the loop and return the execution result of the algorithm;

[0047] 3) Data dump: As Figure 2 shown, the data dump uses the DPDK framework and the memory file system. The data dump utilizes the multi-core architecture of DPDK and the characteristics of CPU and thread binding, and assigns a specific logical core to each thread, so that the processing process of each data flow is carried out on a specific core. At the same time, the data dump uses a memory-hard disk multi-level cache for the collected data files. Workflow: First, when a data packet arrives at the current node, the data flow forwarding module obtains the permission information of the data flow according to the tuple information of the data flow. If the data flow is marked as a captured data flow, the data packet and the data structure storing the data flow are pushed into the queue. The data dump obtains the data packet and the data flow information from the queue, and obtains the file handle of the data dump file of the data flow according to the data flow information, and writes the data packet into the data dump file through the file handle. After writing the file, calculate the size of the data dump file after dumping the current data packet, and judge whether the file size exceeds the upper limit. In this example, it is 128M. If the file exceeds the upper limit, a new data dump file is created to replace the current file, the data dump object of the data flow is reconfigured, and the file is moved to the disk for persistent storage.

[0048] In this example, specifically, as Figure 3As shown, when there is partial overlap in the paths of data streams in a network, there will be a situation where some data streams have relatively more optional task execution nodes and some have relatively fewer. At this time, if the data capture overhead is used as the priority for selection, there may be a situation where some data streams cannot be captured due to lack of resources. The path of data stream flow1 is {route1, route2, route3, route5}, the path of data stream flow1 is {route6, route3, route2, route4}, and the path of data stream flow1 is {route3, route2}. The paths of the three data streams overlap at {route2, route3}. If the priorities of data streams flow1 and flow2 are relatively high, and the capture tasks of data streams flow1 and flow2 are respectively assigned to router nodes route2 and route3, resulting in the remaining loads of nodes route2 and route3 being insufficient to execute the capture task of flow3, leading to task allocation failure. At this time, the capture task of data stream flow3 should be preferentially allocated. The method in this example should preferentially allocate the data stream with relatively fewer optional task execution nodes to prevent problems such as allocation failure caused by improper task allocation order. The method in this example refers to the task with relatively fewer optional task execution nodes as a low-heat task and the task with relatively more optional task execution nodes as a high-heat task. The ratio of the maximum available load of a data stream to the data stream overhead represents the heat level of the data stream, and the heat level is used as the priority for node selection. When the maximum available loads of all tasks are the same, the method in this example degenerates into a method of sorting priorities according to the overhead.

Claims

1. A distributed data capture method based on DPDK, which is applicable to a load-balanced network, wherein, A single node consists of flow rule management, data flow information collection, task allocation, data flow authentication, packet processing, and network data capture. Above the method is a blockchain network for communication and a distributed SDN controller based on the blockchain network, and below is a software router with traffic capture developed based on DPDK. The method is characterized in that the method includes the following steps: 1) Data flow information collection: Network status collection is located on the router node, periodically collects the status information of the current node, and notifies other nodes through the blockchain network. The data flow information collection process collects the data packets passing through the current router during the current time period, classifies and counts the number of forwarded data packets and the total number of forwarded bytes during the current time period according to the data flow characteristics, and organizes them into a structured message together with the timestamp, the length of the time period, and the system load time series data, and broadcasts and publishes it through the blockchain network. First, the data flow forwarding module forwards the data flow according to the matching rule, and at the same time records the information of the data flow when forwarding the data flow. The forwarding module obtains the data block storing network data according to the matching rule and updates the relevant information of the data block. The data collection module periodically obtains the data flow status information dumped by the forwarding module, packs the status information into a specific data structure, and broadcasts it through the blockchain. Specifically, the data flow information collection process is as follows: First, obtain the current time, obtain the storage object storing the information of the current time period from the data flow information, and calculate the duration of the current time period. If the duration of the current time period exceeds the predefined time period length, update the data information of the current time period to the data flow information, generate a data flow status event, broadcast it to all router nodes through the blockchain network, and at the same time generate a storage object for clearing the information of the current time period and reset the start time of the time period. Then, accumulate the number of received data packets and the received bytes during the current time period according to the packet information; 2) Task Allocation: The task allocation algorithm subscribes to the status information of all nodes in the current network through the blockchain, and calculates the task arrangement of each current node according to the network topology information and the current status information of the network, based on the following task allocation algorithm. The task allocation algorithm obtains the network topology and data flow information through the blockchain network. First, the task allocation algorithm traverses the network topology, obtains the initial value of the load capacity of each router node, and stores the initial value of the load capacity in the array R load Secondly, the task allocation algorithm obtains the current network status information through the blockchain network, calculates the processing overhead of each data capture task, and stores the task processing overhead in F cost Then, according to the information collected in the data flow information collection process in step 1, traverse all the data flow information in the network, mark the nodes on the data flow path, indicating that the capture task of the data flow can be executed on the marked nodes, and store the traversal result in the two-dimensional array V. Finally, the node load capacity R load , the required overhead F of the data flow cost , and the optional matrix V of task allocation are passed as parameters into the capture task allocation algorithm to obtain the task allocation result F route . The task allocation monitors the data flow information in the network. When the change in the data flow scale exceeds the upper limit of the data that a single node can process, the task allocation algorithm partially re-performs the task allocation to obtain a new task allocation scheme, and quantifies the performance of each router node and the load required to process the data flow. Assume that the maximum number of packets that a single router node route can dump within a unit time is the load capacity of the node, as shown in formula (1). The set of router node load capacities is: Assume that the number of data packets that the data flow flow needs to process per unit time is the load required by the flow, as shown in formula (2). The set of loads required by the data flow is: Assume using via i,j to represent the data flow j passing through the router i , then the symbol V is used to represent the optional matrix for data capture task allocation, as shown in formula (3): When there is only one data stream in the entire network or the paths of all data streams in the network do not overlap, select the node with the maximum load capacity among the nodes passed by the data stream as the execution node of the task. When the transmission paths of the data streams in the entire network are the same, the data capture task is captured by any node passed by the data stream. At this time, the greedy algorithm is adopted, that is, according to the overhead of the capture task and the load capacity of the node, select the node with the maximum load capacity to execute the task with the largest overhead in turn, and recalculate the load capacity of the node until all tasks are allocated. The tasks are evenly allocated to the nodes on the path. The task allocation algorithm preferentially allocates the data stream with a relatively small number of optional task execution nodes. The task with a relatively small number of optional task execution nodes is called a low-heat task, and the task with a relatively large number of optional task execution nodes is called a high-heat task during the task allocation process. The ratio of the maximum available load of the data stream to the data stream overhead represents the heat level of the data stream, and the heat level is used as the priority for node selection. When the maximum available loads of all tasks are the same, the task allocation process degenerates into sorting by overhead according to the priority: When there is only one data stream in the entire network or the paths of all data streams in the network do not overlap, select the node with the maximum load capacity among the nodes passed by the data stream as the execution node of the task. When the transmission paths of the data streams in the entire network are the same, the data capture task is captured by any node passed by the data stream. At this time, the greedy algorithm is adopted, that is, according to the overhead of the capture task and the load capacity of the node, select the node with the maximum load capacity to execute the task with the largest overhead in turn, and recalculate the load capacity of the node until all tasks are allocated. The tasks are evenly allocated to the nodes on the path. The task allocation algorithm preferentially allocates the data stream with a relatively small number of optional task execution nodes. The task with a relatively small number of optional task execution nodes is called a low-heat task, and the task with a relatively large number of optional task execution nodes is called a high-heat task during the task allocation process. The ratio of the maximum available load of the data stream to the data stream overhead represents the heat level of the data stream, and the heat level is used as the priority for node selection. When the maximum available loads of all tasks are the same, the task allocation process degenerates into sorting by overhead according to the priority: Assume that flow is used path to represent the router path through which the data flow flow passes. Then, the calculation method of the hotness and coldness degree of the data flow flow pop is shown in formula (4) as follows: The set of the cold and hot degrees of the data flow is as shown in formula (5): The algorithm calculates the ratio of the maximum available load to the data flow overhead for all tasks, adds all data capture tasks to the priority queue, and sorts them in ascending order according to the ratio of the maximum available load to the data flow overhead required by the tasks, so that tasks with low heat are preferentially allocated, ensuring that the algorithm will not have the problem of insufficient node margin caused by high-heat tasks piling up on individual nodes, resulting in task allocation failure. Secondly, the optional nodes of each task are sorted according to the load capacity, so that high-performance nodes preferentially select tasks. Then, the candidate tasks of all nodes are traversed according to the priority queue to select a suitable task arrangement plan. Select the task at the head of the task queue, select the optional nodes of this task according to the priority, and judge whether the node meets the conditions in turn. If the load required by the task is less than or equal to the remaining load capacity of the node, the task is allocated to the current node, and the remaining load of the node is recalculated, and the current loop ends. Otherwise, continue to traverse the next node. When the task queue is empty, the loop ends and the execution result of the algorithm is returned; 3) Data dump: The data dump uses the DPDK framework and the memory file system. The data dump utilizes the multi-core architecture of DPDK and the characteristics of CPU and thread binding, and assigns a specific logical core to each thread, so that the processing process of each data flow is carried out on a specific core. At the same time, the data dump uses a memory-hard disk multi-level cache for the collected data files. Workflow: First, when a data packet arrives at the current node, the data flow forwarding module obtains the permission information of the data flow according to the tuple information of the data flow. If the data flow is marked as a captured data flow, the data packet and the data structure storing the data flow are pushed into the queue. The data dump obtains the data packet and the data flow information from the queue, and obtains the file handle of the data dump file of the data flow according to the data flow information, and writes the data packet into the data dump file through the file handle. After writing the file, calculate the size of the data dump file after dumping the current data packet, and judge whether the file size exceeds the upper limit, which is 64M or 128M. If the file exceeds the upper limit, a new data dump file is created to replace the current file, the data dump object of the data flow is reconfigured, and the file is moved to the disk for persistent storage.