Method and device for collecting and analyzing full-link network traffic in cloud native scene

Through bypass packet capture and eBPF technology to collect and analyze network traffic in a cloud-native environment, the insufficient monitoring of traditional methods and the complexity of eBPF are solved, efficient monitoring of full-link traffic and process information correlation are achieved, and network quality analysis capabilities are improved.

CN120281638APending Publication Date: 2025-07-08HANGZHOU CHENGYUN DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510485087.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art cannot fully and accurately collect network transmission data in cloud-native scenarios, especially in containerized and virtualized environments. Traditional methods cannot fully monitor network traffic, and eBPF technology will affect business logic complexity and performance when performance sensitive or system load is heavy.

Method used

Bypass packet capture method is used to capture traffic, combine eBPF technology to inject hook functions into Linux kernel and applications, obtain basic network connection information, processing delay and application data, and save analysis results through eBPF map objects to realize the collection and analysis of full-link network traffic.

Benefits of technology

It realizes efficient collection and analysis of all traffic in the cloud-native environment, solves the monitoring needs of different traffic types, provides process information association relationships, helps locate problems and observe network behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281638A_ABST
    Figure CN120281638A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of network flow collection and analysis, in particular to a full-link network flow collection and analysis method and device in a cloud native scene, and the main technical scheme comprises the step of carrying out the flow capture at all key points where the network flow is forwarded and processed, so that the fault occurrence position can be positioned more accurately. Meanwhile, during traffic capture, a bypass capture and eBPF mode is adopted, and data acquisition is realized at the minimum overhead. Besides, hook injection is carried out in a system kernel and a key processing function of a system library by utilizing an eBPF technology, in a hook function, analyzed information of the system can be conveniently extracted, and calculation of related performance time delay is carried out. In addition, in the hook function, context information of flow processing can be conveniently obtained, and great help is provided for positioning problems or observing behaviors of processes on the network level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of network traffic collection and analysis, and particularly to a method and device for full-link network traffic collection and analysis in a cloud-native scenario. Background Art

[0002] With the wide application of cloud-native microservice architectures, the communication and data flow between services have become increasingly complex and frequent, and the performance of traditional network monitoring tools in microservice architectures is limited. Existing monitoring methods are difficult to comprehensively and accurately collect network transmission data. Especially in containerized and virtualized environments, the demand for high-performance and low-overhead network data collection is gradually increasing.

[0003] In the solutions of the prior art, Patent CN 114666249A, named "Traffic collection method, device and computer-readable storage medium on a cloud platform", describes a full-image traffic collection method that screens traffic through configuration information to reduce the network traffic to be processed.

[0004] Patent CN 115664930A, named "Non-intrusive network fault diagnosis and prediction method in a cloud-native environment", describes that in a cloud-native environment, a technology based on eBPF is adopted, combined with kernel instrumentation and event triggering mechanisms for application monitoring to collect network data, and analyze the monitoring data to achieve the diagnosis and prediction of network faults.

[0005] Patent CN 114978897A, named "Network control method and system based on eBPF and application recognition technology", describes reading data packets from a virtual image traffic network interface, performing application recognition analysis on the data packets, and transmitting the recognition results to an eBPF-based control module to achieve control of specific network sessions.

[0006] Analyzing the existing technical solutions, we find that for network traffic collection in cloud-native scenarios, one way is to use traditional NPM products and collect and analyze data by mirroring traffic through separate physical devices without affecting the original system. Another relatively new way is to use the eBPF technology in the Linux kernel to insert instrumentation in the key processing functions of the protocol stack in the kernel state to collect relevant network data information.

[0007] The first way can only collect traffic between hosts. In a cloud-native environment, there is a large amount of communication data between different containers within the same host, and this data cannot be collected by the mirroring method.

[0008] In the second method, network data can be collected when the network protocol stack processes network data. However, there are dozens of key processing functions in the kernel protocol stack, which means that at least a dozen points need to be instrumented to track different states of processing, and the processing logic will become very complex. In addition, the instrumented eBPF program is in series in the business processes of the kernel / application. In the case of performance sensitivity or heavy system load, if too much logic is implemented in eBPF, it will have an immeasurable impact on the existing services.

[0009] In addition, in the cloud native scenario, network traffic has more forms, such as cross-host traffic, container-to-container traffic, host-to-container traffic, etc. All traffic needs to be monitored and analyzed, and different traffic needs to be collected and analyzed by different means to achieve more efficient and comprehensive analysis of the network quality in the cloud native environment. Summary of the Invention

[0010] In view of this, the purpose of the present invention is to propose a full-link network traffic collection and analysis method and device in the cloud native scenario to solve the problem that network traffic cannot be completely monitored and efficiently processed in the existing solutions.

[0011] Based on the above purpose, the present invention provides a full-link network traffic collection and analysis method in the cloud native scenario, including the following steps:

[0012] S1. Grab traffic on the network device in the host environment by means of bypass packet capture, and perform aggregation analysis on the traffic based on the TCP / UDP five-tuple to obtain basic network connection information;

[0013] S2. Inject a hook function into the relevant processing functions of the Linux kernel network protocol stack through eBPF technology, track the processing process of application data in the Linux operating system kernel, obtain network connection processing delay data, and calculate the performance overhead of each processing stage;

[0014] S3. Inject a hook function into the relevant processing functions of the standard library or SDK called by the application through eBPF technology, track the processing process of data in the application process and obtain network application data;

[0015] S4. Correlate the basic network connection information obtained in step S1, the network connection processing delay obtained in step S2, and the network application data obtained in step S3, analyze the processing process of the full link of the application request and the delay and performance overhead of each processing process, and save the correlation result.

[0016] Preferably, step S1 specifically includes:

[0017] Read network packets received and sent by the physical network card through the libpcap library;

[0018] Parse each read network packet to obtain the five-tuple information of the network connection, including source IP, source port, destination IP, destination port, and network protocol;

[0019] Taking the connection five-tuple as an object, count the basic network traffic information, including the number of received packets, the number of received bytes, the number of sent packets, the number of sent bytes, and the counting buckets of the packet size.

[0020] Create a system shared memory map object, using the five-tuple as the key and the counted traffic information as the value for storage.

[0021] Preferably, step S2 specifically includes:

[0022] Inject a hook function into the key tcp processing function in the linux network protocol stack through eBPF;

[0023] In the hook function, obtain the TCP information of the network data through the context parameter, and extract the connection five-tuple field from it;

[0024] Taking the connection five-tuple as an object, analyze the establishment delay, retransmission times, state transition delay, data reception delay, data transmission delay, and connection close delay of the tcp connection;

[0025] Create an eBPF map object, using the five-tuple as the key and the analyzed statistical data as the value for storage.

[0026] Preferably, the key tcp processing functions in the linux network protocol stack are kernel functions, including tcp_v4_connect, tcp_v6_connect, tcp_close, tcp_recvmsg, tcp_sendmsg, tcp_retransmit_skb, and the injected hook functions include kprobe and kretprobe type hook functions;

[0027] In the tcp_xx_connect hook function, record the timestamp when the connection is completed;

[0028] In the tcp_close hook function, record the timestamp when the connection ends;

[0029] In the tcp_recvmsg hook function, record the data reception delay;

[0030] In the tcp_sendmsg hook function, record the data transmission delay.

[0031] Preferably, step S3 specifically includes:

[0032] Inject hook functions into functions related to data reception and transmission in the system standard library through eBPF;

[0033] Inject hook functions into functions related to data reception and transmission in the third-party library through eBPF;

[0034] In the hook function, obtain the network socket handle through the context parameter, obtain the five-tuple information of the TCP connection through the handle, and obtain the data sent and received by the application through the context parameter, and extract business-related information from it;

[0035] Analyze the TCP connection latency, data reception latency, and data transmission latency with the connection five-tuple as the object;

[0036] Create an eBPF map object, and save the analyzed statistical data with the five-tuple as the key and the value;

[0037] Preferably, the functions related to data reception and transmission in the system standard library include connect, close, send, sendmsg, recv, recvmsg functions, and the functions related to data reception and transmission in the third-party library include SSL_read and SSL_write of openssl. The injected hook functions include uprobe and uretprobe type hook functions;

[0038] In the connect hook function, record the timestamp when the connection is completed;

[0039] In the close hook function, record the timestamp when the connection ends;

[0040] In the recv / recvmsg hook function, record the data reception latency and analyze the business information from the data memory;

[0041] In the send / sendmsg hook function, record the data transmission latency and analyze the business information from the data memory.

[0042] The present invention also provides a full-link network traffic collection and analysis device in a cloud-native scenario, including:

[0043] A traffic capture module, which is used to capture traffic on network devices in the host environment by means of bypass packet capture, perform aggregation analysis on the traffic based on TCP / UDP five-tuples, and obtain basic network connection information;

[0044] A calculation module, which is used to inject hook functions into relevant processing functions of the Linux kernel network protocol stack through eBPF technology, track the processing process of application data in the Linux operating system kernel, obtain network connection processing delay data, calculate the performance overhead of each processing stage, and inject hook functions into relevant processing functions of the standard library or SDK called by the application through eBPF technology, track the processing process of data in the application process, and obtain network application data;

[0045] A processing module, which is used to associate the basic network connection information obtained by the traffic capture module in the previous step, the network connection processing delay and network application data obtained by the calculation module, analyze the processing process of the full link of the application request, the delay and performance overhead of each processing process, and save the association result to the storage module.

[0046] The beneficial effects of the present invention: The traffic capture and analysis method proposed by the present invention not only captures all traffic in the cloud native environment, but also analyzes the data in a novel and efficient manner. In addition, it also solves the association relationship between applications and network data, so that when taking evidence or issuing warnings through network traffic, the corresponding process information can also be obtained, which provides great help for locating problems or observing the behavior of processes at the network level. Description of the Drawings

[0047] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only those of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0048] Figure 1 It is a schematic flowchart of the network traffic collection and analysis method according to the embodiment of the present invention;

[0049] Figure 2 It is a typical cloud native network environment diagram;

[0050] Figure 3 It is a schematic flowchart of a conventional traffic analysis method. Detailed Embodiments

[0051] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the following further elaborates on the present invention in conjunction with specific embodiments.

[0052] It should be noted that unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meanings understood by those with ordinary skills in the field to which the present invention belongs. The "first", "second" and similar words used in the present invention do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "comprising" or "including" mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Words such as "connected" or "linked" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right", etc. are only used to represent relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0053] As Figure 2 shown in a typical cloud-native network environment. As can be seen from the figure, in a cloud-native environment, traffic exists between containers, between containers and hosts, and also between hosts. The former traffic is exchanged through a virtual switch / virtual bridge inside the host, while the latter traffic is sent through the physical network card of the host. In the figure, all traffic can be obtained by using raw socket or directly calling the libpcap library. Traffic inside the container can be captured at point C1, traffic between containers and between containers and hosts can be captured at point C2, and traffic between the host and the external network can be captured at point C3.

[0054] After obtaining the traffic, analysis is required. The conventional analysis method is to start from the Ethernet frame header and perform protocol parsing layer by layer until all data is parsed. The main process is as Figure 3 shown.

[0055] It can be seen that the parsing of network data packets is a complex task and a CPU-intensive operation. For applications running in a cloud-native environment, multiple services run in the same environment. If the analysis program consumes too many resources, it is very likely to affect the normal operation of the business system.

[0056] Therefore, the embodiments of this specification adopt a segmented processing method, trying to reuse the functions and information that have been implemented in the system. In this way, not only can the information we need be obtained, but the introduced overhead is also relatively small. This method provides a network traffic collection and analysis method as Figure 1 shown, which specifically includes the following steps:

[0057] 1. Capture network traffic on network devices (C1, C2, C3) in the host environment through the method of bypass packet capture. After the traffic is captured, establish and count basic information.

[0058] 1.1. Allocate a system shared memory, create a map (named map-1) in this memory, with the key being the connection quintuple (source IP + source port + destination IP + destination port + network protocol), and the value being the traffic statistics information (number of received packets, number of received bytes, number of sent packets, number of sent bytes, count buckets of packet sizes, receive and send durations).

[0059] 1.2. Scan the network devices of the system, including physical network cards, virtual network bridges, virtual switches, etc., and identify the devices C1, C2, C3 that need to be collected.

[0060] 1.3. Capture network traffic on devices C1, C2, C3 by means of bypass packet capture.

[0061] 1.4. After the traffic is successfully captured, extract the basic connection information from the network packet: including source IP, source port, destination IP, destination port, network protocol (forming the key of the connection quintuple), look up this key in map-1, and if it does not exist, add an object.

[0062] 1.5. Calculate the statistical values of the connection: number of received packets, number of received bytes, number of sent packets, number of sent bytes, count buckets of packet sizes, receive and send durations.

[0063] 1.6. Update the statistical values to map-1.

[0064] 2. Inject hook functions into relevant functions of the linux kernel network protocol stack through the eBPF method, and analyze network data within the hook functions and save the analysis results to memory.

[0065] 2.1. Create an eBPF map (named map-2), with the key being the connection quintuple (source IP + source port + destination IP + destination port + network protocol), and the value being the connection statistical information (establishment delay, number of retransmissions, state transition delay, data reception delay, data transmission delay, connection close delay)

[0066] 2.2. Inject hook functions of the kprobe and kretprobe types into the kernel functions tcp_v4_connect, tcp_v6_connect, tcp_close, tcp_recvmsg, tcp_sendmsg, tcp_retransmit_skb.

[0067] 2.3. In the context of the hook function, obtain the struct sock object. From the struct sock object, obtain the tcp connection quintuple.

[0068] 2.4. Record the timestamp when the connection is completed in the tcp_xx_connect hook function.

[0069] 2.5. Record the timestamp when the connection ends in the tcp_close hook function.

[0070] 2.6. Record the latency of data reception in the tcp_recvmsg hook function.

[0071] 2.7. Record the latency of data transmission in the tcp_sendmsg hook function.

[0072] 2.8. Search for the tcp connection five-tuple key in map-2. If it doesn't exist, add an object and update the statistical data.

[0073] 3. Inject hook functions into relevant functions of the system standard library and third-party SDKs in the eBPF manner, and analyze network data within the hook functions and save the analysis results to memory.

[0074] 3.1. Create an eBPF map (named map-3), with the key being the connection five-tuple (source IP + source port + destination IP + destination port + network protocol), and the value being connection-related information, process information, and business information (establishment latency, data reception latency, data transmission latency, process name, process ID, transaction number, etc.).

[0075] 3.2. Inject uprobe and uretprobe type hook functions into connect, close, send, sendmsg, recv, and recvmsg in the system standard library libc.so.

[0076] 3.3. Inject uprobe and uretprobe type hook functions into relevant third-party libraries, such as SSL_read and SSL_write in openssl.

[0077] 3.4. In the context of the hook function, obtain the socket handle, and from this handle, obtain the tcp connection five-tuple. In addition, process information such as the process name and process ID can also be obtained.

[0078] 3.5. Record the timestamp when the connection is completed in the connect hook function.

[0079] 3.6. Record the timestamp when the connection ends in the close hook function.

[0080] 3.7. Record the latency of data reception in the recv / recvmsg hook function, and at the same time analyze business information such as the transaction number from the data memory.

[0081] 3.8. In the send / sendmsg hook function, record the latency of data transmission, and at the same time analyze service information such as transaction numbers from the data memory.

[0082] 3.9. In map-3, search for the TCP connection five-tuple key. If it does not exist, add an object and update the statistical data.

[0083] 4. Aggregate and correlate the data generated in each stage, so as to establish full-link state tracking for network requests.

[0084] 4.1. Traverse map-1 to obtain the five-tuple key and connection traffic information of each connection.

[0085] 4.2. Search for the key in map-2 to obtain information such as connection latency.

[0086] 4.3. Search for the key in map-3 to obtain connection application layer latency, process information, and service information.

[0087] 4.4. Merge the key and the obtained information and save it to the storage module.

[0088] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present invention is limited to these examples; under the concept of the present invention, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the present invention as described above, which are not provided in detail for the sake of brevity. Any omission, modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for collecting and analyzing all-link network traffic in a cloud-native scenario, characterized in that, It includes the following steps: S1. Capture traffic on network devices in the host environment through bypass packet capture, aggregate and analyze the traffic based on TCP / UDP quintuples to obtain basic network connection information; S2. Inject hook functions into relevant processing functions of the Linux kernel network protocol stack through eBPF technology, track the processing process of application data in the Linux operating system kernel, obtain network connection processing delay data, and calculate the performance overhead of each processing stage; S3. Inject hook functions into relevant processing functions of the standard library or SDK called by the application through eBPF technology, track the processing process of data in the application process and obtain network application data; S4. Associate the basic network connection information obtained in step S1, the network connection processing delay obtained in step S2, and the network application data obtained in step S3, analyze the processing process of the full link of the application request and the delay and performance overhead of each processing process, and save the association result.

2. The full-link network traffic collection and analysis method in the cloud-native scenario according to claim 1, wherein Step S1 specifically includes: Read network data packets received and sent by the physical network card through the libpcap library; Parse each read network data packet to obtain the quintuple information of the network connection, including source IP, source port, destination IP, destination port, and network protocol; Taking the connection quintuple as the object, count the basic network traffic information, including the number of received packets, the number of received bytes, the number of sent packets, the number of sent bytes, and the counting buckets of the packet size; Create a system shared memory map object, and save it with the quintuple as the key and the statistically traffic information as the value.

3. The full-link network traffic collection and analysis method in the cloud-native scenario according to claim 1, characterized in that Step S2 specifically includes: Inject hook functions into key tcp processing functions in the linux network protocol stack through eBPF; In the hook function, obtain the TCP information of the network data through the context parameter, and extract the connection quintuple field from it; Taking the connection quintuple as the object, analyze the connection establishment delay, retransmission times, state transition delay, data reception delay, data transmission delay, and connection close delay of the tcp connection; Create an eBPF map object, and save it with the quintuple as the key and the analyzed statistical data as the value.

4. The method for collecting and analyzing full-link network traffic in a cloud-native scenario according to claim 3, wherein The key tcp processing functions in the linux network protocol stack are kernel functions, including tcp_v4_connect, tcp_v6_connect, tcp_close, tcp_recvmsg, tcp_sendmsg, tcp_retransmit_skb, and the injected hook functions include kprobe and kretprobe type hook functions; In the tcp_xx_connect hook function, record the timestamp when the connection is completed; In the tcp_close hook function, record the timestamp when the connection ends; In the tcp_recvmsg hook function, record the data reception delay; In the tcp_sendmsg hook function, record the data transmission delay.

5. The full-link network traffic collection and analysis method in the cloud-native scenario according to claim 1, characterized in that Step S3 specifically includes: Inject hook functions into functions related to data reception and transmission in the system standard library through eBPF; Inject hook functions into functions related to data reception and transmission in a third-party library through eBPF; In the hook function, obtain the network socket handle through the context parameter, obtain the five-tuple information of the TCP connection through the handle, and obtain the data sent and received by the application through the context parameter, and extract the business-related information from it; Taking the connection five-tuple as the object, analyze the TCP connection delay, data reception delay, and data transmission delay; Create an eBPF map object, save the analyzed statistical data with the five-tuple as the key and the value; 6. The method for collecting and analyzing all-link network traffic in the cloud-native scenario according to claim 5, wherein The functions related to data reception and transmission in the system standard library include connect, close, send, sendmsg, recv, recvmsg functions, and the functions related to data reception and transmission in the third-party library include SSL_read and SSL_write of openssl. The injected hook functions include uprobe and uretprobe type hook functions; In the connect hook function, record the timestamp when the connection is completed; In the close hook function, record the timestamp when the connection ends; In the recv / recvmsg hook function, record the data reception delay and analyze the business information from the data memory; In the send / sendmsg hook function, record the data transmission delay and analyze the business information from the data memory.

7. A full-link network traffic collection and analysis device in a cloud-native scenario, characterized in that, It includes: A traffic capture module for capturing traffic on network devices in the host environment through the bypass packet capture method, aggregating and analyzing the traffic based on the TCP / UDP five-tuple to obtain basic network connection information; A calculation module for injecting hook functions into relevant processing functions of the Linux kernel network protocol stack through eBPF technology to track the processing process of application data in the Linux operating system kernel, obtain network connection processing delay data, calculate the performance overhead of each processing stage, and inject hook functions into relevant processing functions of the standard library or SDK called by the application through eBPF technology to track the processing process of data in the application process and obtain network application data; A processing module for associating the basic network connection information obtained by the traffic capture module in the previous step, the network connection processing delay and network application data obtained by the calculation module, analyzing the processing process of the full link of the application request, the delay and performance overhead of each processing process, and saving the association result to the storage module.