A data processing method, device, apparatus, and storage medium

By acquiring IP packet sets, parsing the 5-tuple and data traffic, dividing the data flow, and determining the target path and traffic, the problem of high system resource consumption and large disk overhead caused by packet collection rule matching is solved, thereby reducing resource consumption and improving performance.

CN115883217BActive Publication Date: 2026-01-30BEIJING NETTAI TECH DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211533820.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-01
Publication Date
2026-01-30
Estimated Expiration
2042-12-01

AI Technical Summary

Technical Problem

In existing technologies, data packet collection rule matching results in high system resource consumption, large disk overhead, and large storage space usage.

Method used

By acquiring IP packet sets, parsing the 5-tuple and data traffic, dividing the data flow, determining the target path and traffic, reducing the amount of data packets stored locally, and using DPDK for packet processing.

Benefits of technology

It reduces system resource and disk overhead, improves performance efficiency, and reduces hardware and network resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115883217B_ABST
    Figure CN115883217B_ABST
Patent Text Reader

Abstract

This invention discloses a data processing method, apparatus, device, and storage medium. The method includes: acquiring a set of IP data packets; parsing each IP data packet in the set to obtain data flow information, wherein the data flow information includes: a 5-tuple of the IP data packets and the data flow of the IP data packets; dividing the set of IP data packets according to the 5-tuple of the IP data packets to obtain at least one data flow and a 5-tuple for each data flow; and determining the target path of each data flow and the data flow of the target path of each data flow based on the 5-tuple of each data flow and the data flow of the IP data packets in each data flow. Through the technical solution of this invention, disk space overhead can be reduced, hardware and network resource consumption can be lowered, and performance efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of network security technology, and in particular to a data processing method, apparatus, device and storage medium. Background Technology

[0002] In existing technologies, data processing results are presented in data packets. Each data packet undergoes collection rule matching, which consumes significant system resources and places high demands on CPU and memory. Large datasets can lead to excessive system resource usage.

[0003] In existing technologies, the collected data packets are directly saved locally, which incurs significant disk overhead on the device, results in numerous stored files, and consumes a large amount of storage space. Summary of the Invention

[0004] This invention provides a data processing method, apparatus, device, and storage medium, which solves the problem of high system resource consumption and large disk overhead caused by directly saving the collected data packets locally after obtaining them through matching collection rules.

[0005] According to one aspect of the present invention, a data processing method is provided, comprising:

[0006] Get the IP packet set;

[0007] Each IP packet in the IP packet set is parsed to obtain data flow information, wherein the data flow information includes: the 5-tuple of the IP packet and the data flow of the IP packet;

[0008] The set of IP packets is divided according to the 5-tuple of the IP packets to obtain at least one data stream and a 5-tuple for each data stream;

[0009] The destination path and the data traffic of the destination path of each data stream are determined based on the 5-tuple of each data stream and the data traffic of the IP packets in each data stream.

[0010] According to another aspect of the present invention, a data processing apparatus is provided, the data processing apparatus comprising:

[0011] The collection acquisition module is used to acquire a collection of IP packets;

[0012] The set parsing module is used to parse each IP packet in the set of IP packets to obtain data flow information, wherein the data flow information includes: the 5-tuple of the IP packet and the data flow of the IP packet;

[0013] The set partitioning module is used to partition the set of IP packets according to the quintuples of the IP packets to obtain at least one data stream and a quintuple for each data stream;

[0014] The traffic determination module is used to determine the target path of each data stream and the data traffic of the target path of each data stream based on the five-tuple of each data stream and the data traffic of the IP packets in each data stream.

[0015] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0016] At least one processor; and

[0017] A memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the data processing method according to any embodiment of the present invention.

[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the data processing method described in any embodiment of the present invention.

[0020] This invention addresses the problem of high system resource consumption and large disk overhead caused by directly storing collected data packets locally after obtaining them through rule matching. This process reduces disk space consumption, lowers hardware and network resource consumption, and improves performance efficiency. The invention involves acquiring an IP packet set; parsing each IP packet in the set to obtain data flow information, including the IP packet's 5-tuple and data traffic; dividing the IP packet set based on the 5-tuple to obtain at least one data flow and the 5-tuple for each data flow; and determining the target path and data traffic of each data flow based on the 5-tuple and the data traffic of the IP packets within each data flow.

[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0022] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a flowchart of a data processing method according to Embodiment 1 of the present invention;

[0024] Figure 2 This is a schematic diagram of the structure of a data processing device according to Embodiment 2 of the present invention;

[0025] Figure 3 This is a schematic diagram of the structure of an electronic device according to Embodiment 3 of the present invention. Detailed Implementation

[0026] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0027] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0028] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0029] Example 1

[0030] Figure 1This is a flowchart of a data processing method according to Embodiment 1 of the present invention. This embodiment is applicable to the processing of data packets. The method can be executed by the data processing device in this embodiment of the invention, which can be implemented in software and / or hardware, such as... Figure 1 As shown, the method specifically includes the following steps:

[0031] S110, Obtain the set of IP packets.

[0032] The IP packet set is a collection of packets acquired periodically. This periodic acquisition can be performed at set intervals. It should be noted that the IP packets are stored in pcap file format.

[0033] Specifically, the method for obtaining the IP packet set can be as follows: use DPDK to periodically capture packets and generate the IP packet set based on the captured packets. DPDK (Data Plane Development Kit) is primarily based on Linux systems and is a collection of function libraries and drivers for fast packet processing, which can greatly improve data processing performance and throughput, and increase the efficiency of data plane applications.

[0034] By periodically acquiring sets of IP packets, the overhead of hardware resources such as CPU and memory can be reduced.

[0035] S120: Parse each IP packet in the IP packet set to obtain data flow information, which includes: the 5-tuple of the IP packet and the data flow of the IP packet.

[0036] The five-tuple of an IP packet includes the source IP address, source port, destination IP address, destination port, and transport layer protocol. The data traffic of an IP packet is the network traffic of IP packets.

[0037] Specifically, the method to parse each IP packet in the IP packet set to obtain data flow information can be as follows: after periodically acquiring the IP packet set, parse each IP packet in the IP packet set to obtain the source IP address, source port, destination IP address, destination port, transport layer protocol, and network traffic of each IP packet.

[0038] S130, the set of IP packets is divided according to the 5-tuple of the IP packets to obtain at least one data stream and a 5-tuple for each data stream.

[0039] The data stream consists of multiple IP packets with the same five-tuple.

[0040] Specifically, the method to divide the IP packet set according to the 5-tuple of the IP packet to obtain at least one data stream and the 5-tuple of each data stream can be as follows: obtain the 5-tuple of all IP packets in the IP packet set, and group IP packets with the same 5-tuple into data streams, thereby obtaining at least one data stream and the 5-tuple of each data stream.

[0041] Optionally, the IP packet set is partitioned based on the 5-tuple of the IP packets to obtain at least one data stream and a 5-tuple for each data stream, including:

[0042] The target IP packets in the IP packet set are deleted to obtain a filtered IP packet set, wherein the five-tuple of the target IP packet is the same as at least one five-tuple in the target list, and the target list includes at least one five-tuple of an abnormal packet.

[0043] The filtered set of IP packets is divided according to the quintuples of the IP packets to obtain at least one data stream and a quintuple for each data stream.

[0044] In this context, the 5-tuple in the target list represents the 5-tuple of the abnormal data packet, and the 5-tuple of the target IP data packet is identical to at least one 5-tuple in the target list. It should be noted that the 5-tuple of the target IP data packet can be identical to at least one of the 5-tuples of the abnormal data packet in the target list; for example, the source IP address of the target IP data packet can be the same as the source IP address of the abnormal data packet in the target list. Furthermore, the 5-tuple of the target IP data packet can also be completely identical to the 5-tuple of the abnormal data packet in the target list.

[0045] Specifically, the method for deleting target IP packets from the IP packet set to obtain a filtered IP packet set can be as follows: filter target IP packets that are identical or completely identical to at least one of the source IP address, source port, destination IP address, destination port, and transport layer protocol in the five-tuple of at least one abnormal packet in the target list, and delete the target IP packets from the IP packet set to obtain the filtered IP packet set.

[0046] Specifically, the method of dividing the filtered IP packet set according to the 5-tuple of the IP packet to obtain at least one data stream and the 5-tuple of each data stream can be as follows: obtain the 5-tuple of all IP packets in the filtered IP packet set, divide according to the 5-tuple of all IP packets in the filtered IP packet set, divide IP packets with the same 5-tuple into the same data stream, and obtain at least one data stream and the 5-tuple of each data stream.

[0047] IP packets are divided into 5-tuples to obtain at least one data stream and a 5-tuple for each data stream. The resulting data streams can be transmitted remotely over the network, reducing bandwidth requirements.

[0048] S140, determine the destination path of each data flow and the data traffic of the destination path of each data flow based on the five-tuple of each data flow and the data traffic of the IP packets in each data flow.

[0049] The destination path for each data flow is determined based on the source and destination IP addresses in the 5-tuple of each data flow. The data traffic of the destination path for each data flow is the sum of the data traffic of the IP packets in each data flow.

[0050] Specifically, the method for determining the target path of each data stream and the data traffic of the target path of each data stream based on the 5-tuple of each data stream and the data traffic of the IP packets in each data stream can be as follows: determine the target path of each data stream based on the 5-tuple of each data stream, obtain the data traffic of each IP packet in each data stream, and determine the data traffic of the target path of each data stream by summing the data traffic of each IP packet in each data stream.

[0051] Obtaining the corresponding target path through data streams eliminates the need for packet processing, reducing the processing load and consequently lowering system hardware resource consumption.

[0052] Optionally, the destination path of each data flow and the data traffic of the destination path of each data flow are determined based on the five-tuple of each data flow and the data traffic of the IP packets in each data flow, including:

[0053] The target path for each data stream is determined based on the quintuple of each data stream;

[0054] The sum of the data traffic of IP packets in each data stream is used to determine the data traffic of the target path of each data stream.

[0055] Specifically, the method for determining the target path of each data stream based on the 5-tuple of each data stream can be as follows: determine the target path of each data stream based on the source IP address and destination IP address of the 5-tuple of each data stream.

[0056] Optionally, the target path for each data stream is determined based on the 5-tuple of each data stream, including:

[0057] The target routing table is determined based on the five-tuple of the data stream;

[0058] The target path for each data stream is obtained by querying the target routing table based on the five-tuple of the data stream.

[0059] There can be multiple routing tables, and the target routing table is determined from multiple routing tables based on the five-tuple of the data flow.

[0060] Specifically, the method for determining the target routing table based on the five-tuple of the data flow can be as follows: obtain the source IP address and destination IP address from all routing tables and the five-tuple of the data flow, and select the routing table that has the same source IP address and destination IP address as the data flow as the target routing table.

[0061] Optionally, the five-tuple of the data stream includes: source IP address and destination IP address;

[0062] The target routing table is queried based on the five-tuple of the data stream to obtain the target path for each data stream, including:

[0063] The target routing table is queried based on the source IP address and destination IP address of each data stream to obtain the target path of each data stream, wherein the starting point of the target path is the source IP address of the data stream, and the ending point of the target path is the destination IP address of the data stream.

[0064] Specifically, the method for obtaining the target path of each data flow by querying the target routing table based on the source IP address and destination IP address of each data flow can be as follows: query the target routing table based on the source IP address and destination IP address of each data flow, and determine the path in the target routing table that starts at the source IP address of the data flow and ends at the destination IP address of the data flow as the target path.

[0065] Optionally, after determining the destination path of each data flow and the data traffic of the destination path of each data flow based on the five-tuple of each data flow and the data traffic of the IP packets in each data flow, the method further includes:

[0066] Store the target path of each data stream and the data traffic of the target path of each data stream in the database;

[0067] Obtain target information, wherein the target information includes: source IP address, source port, destination IP address, destination port, and at least one of transport layer protocols;

[0068] The database is queried based on the target information to obtain the data stream, the target path of the data stream, and the data traffic of the target path corresponding to the target information.

[0069] For example, if the source IP address of the target information is A, the data stream with source IP address A can be obtained by querying the database, along with the target path of the data stream with source IP address A and the data traffic of the target path of the data stream with source IP address A.

[0070] It should be noted that the database can be updated based on the set of IP packets obtained each time.

[0071] By storing the target path of each data stream and the data flow of the target path in the database, disk space overhead can be reduced.

[0072] The technical solution of this embodiment obtains an IP packet set; parses each IP packet in the IP packet set to obtain data flow information, wherein the data flow information includes: the 5-tuple of the IP packet and the data flow of the IP packet; divides the IP packet set according to the 5-tuple of the IP packet to obtain at least one data flow and the 5-tuple of each data flow; and determines the target path of each data flow and the data flow of the target path of each data flow based on the 5-tuple of each data flow and the data flow of the IP packets in each data flow. This solves the problem of high system resource consumption and large disk overhead caused by directly saving the collected data packets locally after obtaining them through matching collection rules. It can reduce disk space consumption, reduce hardware and network resource consumption, and improve performance efficiency.

[0073] Example 2

[0074] Figure 2 This is a schematic diagram of a data processing device according to Embodiment 2 of the present invention. This embodiment is applicable to the processing of data packets. The device can be implemented in software and / or hardware, and can be integrated into any device that provides data processing functions, such as... Figure 2 As shown, the data processing device specifically includes: a set acquisition module 210, a set parsing module 220, a set partitioning module 230, and a flow determination module 240.

[0075] Among them, the set acquisition module 210 is used to acquire the set of IP data packets;

[0076] The set parsing module 220 is used to parse each IP packet in the set of IP packets to obtain data flow information, wherein the data flow information includes: the five-tuple of the IP packet and the data flow of the IP packet;

[0077] The set partitioning module 230 is used to partition the set of IP packets according to the five-tuples of the IP packets to obtain at least one data stream and a five-tuple for each data stream;

[0078] The traffic determination module 240 is used to determine the target path of each data stream and the data traffic of the target path of each data stream based on the five-tuple of each data stream and the data traffic of the IP packets in each data stream.

[0079] Optionally, the flow determination module is specifically used for:

[0080] The target path for each data stream is determined based on the quintuple of each data stream;

[0081] The sum of the data traffic of IP packets in each data stream is used to determine the data traffic of the target path of each data stream.

[0082] Optionally, the flow determination module is specifically used for:

[0083] The target routing table is determined based on the five-tuple of the data stream;

[0084] The target path for each data stream is obtained by querying the target routing table based on the five-tuple of the data stream.

[0085] Optionally, the flow determination module is specifically used for:

[0086] The target routing table is queried based on the five-tuple of the data stream to obtain the target path for each data stream, including:

[0087] The target routing table is queried based on the source IP address and destination IP address of each data stream to obtain the target path of each data stream, wherein the starting point of the target path is the source IP address of the data stream, and the ending point of the target path is the destination IP address of the data stream.

[0088] Optionally, the set partitioning module is specifically used for:

[0089] The target IP packets in the IP packet set are deleted to obtain a filtered IP packet set, wherein the five-tuple of the target IP packet is the same as at least one five-tuple in the target list, and the target list includes at least one five-tuple of an abnormal packet.

[0090] The filtered set of IP packets is divided according to the quintuples of the IP packets to obtain at least one data stream and a quintuple for each data stream.

[0091] Optional, also includes:

[0092] A traffic storage module is used to store the target path of each data stream and the data traffic of the target path of each data stream into a database.

[0093] An information acquisition module is used to acquire target information, wherein the target information includes at least one of the following: source IP address, source port, destination IP address, destination port, and transport layer protocol;

[0094] The traffic acquisition module is used to query the database based on the target information to obtain the data stream corresponding to the target information, the target path of the data stream, and the data traffic of the target path of the data stream.

[0095] The above-described products can perform the methods provided in any embodiment of the present invention, and have the corresponding functional modules and beneficial effects for performing the methods.

[0096] The technical solution of this embodiment obtains an IP packet set; parses each IP packet in the IP packet set to obtain data flow information, wherein the data flow information includes: the 5-tuple of the IP packet and the data flow of the IP packet; divides the IP packet set according to the 5-tuple of the IP packet to obtain at least one data flow and the 5-tuple of each data flow; and determines the target path of each data flow and the data flow of the target path of each data flow based on the 5-tuple of each data flow and the data flow of the IP packets in each data flow. This solves the problem of high system resource consumption and large disk overhead caused by directly saving the collected data packets locally after obtaining them through matching collection rules. It can reduce disk space consumption, reduce hardware and network resource consumption, and improve performance efficiency.

[0097] Example 3

[0098] Figure 3 This is a schematic diagram of an electronic device according to Embodiment 3 of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0099] like Figure 3As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0100] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0101] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as data processing methods.

[0102] In some embodiments, the data processing method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the data processing method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the data processing method by any other suitable means (e.g., by means of firmware).

[0103] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0104] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0105] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0106] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0107] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0108] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0109] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0110] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A data processing method, characterized by, The method comprises the following steps: acquiring an IP packet set; parsing each IP packet in the IP packet set to obtain data flow information, wherein the data flow information comprises a five-tuple of the IP packet and data traffic of the IP packet; dividing the IP packet set according to the five-tuple of the IP packet to obtain at least one data flow and the five-tuple of each data flow; determining a target path of each data flow and data traffic of the target path of each data flow according to the five-tuple of each data flow and the data traffic of the IP packet in each data flow; wherein the IP packet set is acquired by setting a periodic time; determining a target path of each data flow and data traffic of the target path of each data flow according to the five-tuple of each data flow and the data traffic of the IP packet in each data flow, comprising: determining the target path of each data flow according to the five-tuple of each data flow; determining the data traffic of the target path of each data flow as the sum of the data traffic of the IP packet in each data flow; dividing the IP packet set according to the five-tuple of the IP packet to obtain at least one data flow and the five-tuple of each data flow, comprising: deleting target IP packets in the IP packet set to obtain a filtered IP packet set, wherein the five-tuple of the target IP packet is the same as at least one five-tuple in a target list, and the target list comprises the five-tuple of at least one abnormal packet; dividing the filtered IP packet set according to the five-tuple of the IP packet to obtain at least one data flow and the five-tuple of each data flow; the data flow is composed of a plurality of IP packets with the same five-tuple; after determining the target path of each data flow and the data traffic of the target path of each data flow according to the five-tuple of each data flow and the data traffic of the IP packet in each data flow, further comprising: storing the target path of each data flow and the data traffic of the target path of each data flow to a database.

2. The method of claim 1, wherein, determining the target path of each data flow according to the five-tuple of the data flow, comprising: determining a target routing table according to the five-tuple of the data flow; querying the target routing table according to the five-tuple of the data flow to obtain the target path of each data flow.

3. The method of claim 2, wherein, the five-tuple of the data flow comprises a source IP address and a destination IP address; querying the target routing table according to the five-tuple of the data flow to obtain the target path of each data flow, comprising: querying the target routing table according to the source IP address and the destination IP address of each data flow to obtain the target path of each data flow, wherein the starting point of the target path is the source IP address of the data flow, and the end point of the target path is the destination IP address of the data flow.

4. The method of claim 1, wherein, after determining the target path of each data flow and the data traffic of the target path of each data flow according to the five-tuple of each data flow and the data traffic of the IP packet in each data flow, further comprising: Obtaining target information, wherein the target information comprises at least one of a source IP address, a source port, a destination IP address, a destination port and a transport layer protocol; According to the target information, querying the database to obtain a data flow corresponding to the target information, a target path of the data flow and data traffic of the target path of the data flow.

5. A data processing apparatus, characterized by, Comprise: A set obtaining module is configured to obtain a set of IP data packets; A set analysis module is configured to analyze each IP data packet in the set of IP data packets to obtain data flow information, wherein the data flow information comprises a five-tuple of the IP data packet and data traffic of the IP data packet; A set division module is configured to divide the set of IP data packets according to the five-tuple of the IP data packet to obtain at least one data flow and a five-tuple of each data flow; A traffic determination module is configured to determine a target path of each data flow and data traffic of the target path of each data flow according to the five-tuple of each data flow and the data traffic of the IP data packet in each data flow; Wherein, the set of IP data packets is obtained by setting a periodic time; The traffic determination module is specifically configured to: Determine the target path of each data flow according to the five-tuple of each data flow; Determine the sum of the data traffic of the IP data packet in each data flow as the data traffic of the target path of each data flow; The set division module is specifically configured to: Delete target IP data packets in the set of IP data packets to obtain a filtered set of IP data packets, wherein the five-tuple of the target IP data packet is the same as at least one five-tuple in a target list, and the target list comprises a five-tuple of at least one abnormal data packet; Divide the filtered set of IP data packets according to the five-tuple of the IP data packet to obtain at least one data flow and a five-tuple of each data flow; the data flow is composed of a plurality of IP data packets with the same five-tuple; The device further comprises: A traffic storage module is configured to store the target path of each data flow and the data traffic of the target path of each data flow into a database.

6. An electronic device, comprising: The electronic device comprises: At least one processor; and The memory is in communication connection with the at least one processor; wherein The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the data processing method in any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for enabling the processor to execute the data processing method in any one of claims 1-4 when executed.

Citation Information

Patent Citations

  • Data packet transmission method and device, computer system and readable storage medium

    CN111200561A

  • Network data packet storage device and method based on user strategy

    CN114047881A