Data Stream Processing Method and System for Incomplete Matching
By using incompletely comparative data stream processing methods and Bloom filters in network switches, the problems of large storage capacity and heavy processing burden in the prior art are solved, and lower storage requirements and processing burden are achieved.
Patent Information
- Application Number
- CN202110777457.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-09
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-07-09
AI Technical Summary
In the prior art, the processing circuit of a network switch requires a large storage space to store data stream records, and an additional processing burden is caused when processing each packet.
Using incomplete comparison data stream processing method, incomplete comparison and look-up tables are achieved through packet filtering mechanism and Bloom filter, reducing the storage capacity requirements of data stream tables and the burden of analyzing data streams.
It effectively reduces the storage capacity requirements of data flow tables and the processing burden of analyzing data flows, reducing the overall cost and additional burden of the system.
Smart Images

Figure CN115604203B_ABST
Abstract
Description
Technical Field
[0001] The present invention proposes a data stream processing technology, in particular, a data stream processing method and system that adopt an incomplete comparison mechanism to reduce the processing burden. Background Art
[0002] Identifying network traffic types, thereby improving the quality of service (QoS), or enhancing network security has been a rather important differentiating requirement for network devices in recent years. For example, if a network switch can identify two different traffic types, namely video conference and file transfer, it is applicable to improve the quality of service (QoS), that is, it can give priority to the traffic of video conferences, which can enhance the user experience. As another example, if a network switch can identify malicious application traffic behaviors, such as the traffic behavior of a Trojan program, it can also block the occurrence of security vulnerabilities in a timely manner.
[0003] However, identifying network traffic types has always been a key issue. The most common way is for network administrators to set the priority order of various network protocol port numbers. For example, it can be set that TCP (Transmission Control Protocol) or UDP (User Datagram Protocol) port numbers are of high priority or low priority. However, in addition to causing inconvenience and a threshold for users, more and more applications use dynamic TCP or UDP port numbers, and more and more applications are hidden under these known TCP or UDP port numbers (such as port 80), and more and more application programs will be transmitted after being encrypted. All these will make traffic identification difficult to achieve.
[0004] To solve the above problems, in the prior art, a traffic feature-based identification method has emerged. It will identify the type of traffic according to the headers and statistical characteristics of the first few (N) packets of each data stream (flow). The statistical characteristics include: the length of each packet in one-way or two-way, the packet interval, the average packet length, the packet length variation, the average packet interval, the packet interval variation, etc. In this way, through the characteristics of the first few packets of each data stream, the prior art uses machine learning or deep learning technology to classify traffic types.
[0005] To achieve the purpose of examining the first few packets of each flow, reference can be made to Figure 1Schematic diagram of the prior art for processing data streams within a network switch. The network switch receives an input data stream 10. After being parsed by the processor within the network switch, a forwarding table 12 is formed. In this example, the forwarded data is the port number (port = Y) of the destination media access control (MAC). A flow table is provided within the network switch to record each data stream flowing through this network switch. As shown in the figure, generally, a data stream header can be expressed by a 5-tuple 14. The 5-tuple (in this example, showing DIP, SIP, SP, DP, Prot (Protocol)) includes: Destination IP (DIP), Source IP (SIP), Destination Layer 4 port (DP), Source Layer 4 port (SP), and communication protocol (Protocol). This example schematically shows that there are two data stream headers with different states recorded in the 5-tuple 14.
[0006] When the network switch receives a packet, it queries the flow table, such as the 5-tuple 14. If it is found that the flow entry corresponding to this data stream does not exist, it means this is a new data stream. Then, the packet is copied to the data stream analysis module 18 (data stream direction 101). The data stream analysis module 18 is a software module that can perform analysis and classification on the packet to identify the application program category to which the data stream belongs. After the data stream analysis module 18 receives the first few packets of this data stream, it starts to execute the traffic identification algorithm. Finally, the identified data stream is inserted into the flow table of the network switch (data stream direction 103), and the classification result is marked on the flow entry. For the subsequent packets of this data stream entering the switch, they can be queried in this flow table (i.e., the 5-tuple 14 in this example) and do not need to be processed by the data stream analysis module 18. Finally, the output data stream 16 is formed according to the destination recorded in the packets within the data stream.
[0007] However, according to the above prior art, specific processing needs to be performed on each data stream. All data streams need to be recorded in the data stream table, and new data streams also need to be copied to the data stream analysis module 18 and then inserted into the data stream table. The disadvantage is that the processing circuit (such as an application-specific integrated circuit, ASIC) in the network switch needs to have sufficient memory to store a large number of data stream records. Generally, the required capacity is about 100K level, which is quite large compared to the number of application data streams that users really care about. Moreover, since the packets of new data streams will be copied to the data stream analysis module 18, additional processing requirements are generated. And although the data stream analysis module 18 only needs the first few packets of the data stream, due to the time difference in processing, it may not be able to send the classified data stream back to the data stream table in time, resulting in packets exceeding the original required packet volume being copied to the data stream analysis module 18, generating an additional processing burden. Summary of the Invention
[0008] In view of the problems such as the processing circuit of the network switch in the prior art needing a large-capacity storage space to store data stream records and causing an additional processing burden when processing each packet, the disclosure proposes an incomplete comparison data stream processing method and system. Through a packet screening mechanism, in an incomplete comparison table lookup (incomplete comparison table) manner, the storage capacity requirement of the data stream table is reduced, and the burden of analyzing data streams is also reduced.
[0009] According to an embodiment, the data stream processing system that executes the incomplete comparison data stream processing method is provided in a network device, which includes a memory, a data stream table and a data stream filter are provided, and a data stream analysis module implemented by software or circuit. The data stream analysis module is used to perform analysis and classification on the packets of the input data stream and also identify the application program category to which the input data stream belongs.
[0010] The incomplete comparison data stream processing method executed in this system includes receiving an input data stream, parsing the input data stream, and then according to the result of parsing the input data stream, querying the data stream table to determine whether the input data stream conforms to any data stream entry in the data stream table. When the input data stream does not conform to any data stream entry in the data stream table, querying the data stream filter to determine whether the input data stream conforms to any filtering condition in the data stream filter.
[0011] Among them, according to the results of querying the data flow table and the data flow filter, one of the following steps is executed: when the input data flow conforms to any data flow entry in the data flow table, a corresponding processing policy is applied; when the input data flow does not conform to any data flow entry in the data flow table, the data flow filter is queried again to determine whether the input data flow conforms to any of the filtering conditions; when the input data flow conforms to any of the filtering conditions in the data flow filter, it means that the input data flow is an existing data flow that has been transmitted for more than multiple packets, and an action is executed according to the process set by the network device; and when the input data flow does not conform to any of the filtering conditions in the data flow filter, it also means that the input data flow does not conform to the conditions of the data flow table entry and the data flow filter, and the system directs the input data flow to the data flow analysis module to process the input data flow.
[0012] Preferably, the data flow processing system implements a data processing circuit in a network switch for processing data flows coming and going through the network switch.
[0013] Further, the data flow table is applicable to all forms of data flows. When the input data flow conforms to any data flow entry in the data flow table, one of the following processing policies is executed: setting the input data flow to a high priority order. In another embodiment, it is also allowed to set the priority to a low priority order and forward the input data flow to a destination transmission port; discarding the input data flow; and copying the input data flow to the data flow analysis module and forwarding the input data flow to the destination transmission port.
[0014] Further, the data flow table records the 5-tuple data in the header of the data flow, which may include a destination network address, a source network address, a destination layer 4 port, a source layer 4 port, and a communication protocol, and records the processing policies corresponding to each data flow entry.
[0015] Preferably, the data flow filter is implemented by a Bloom filter to perform an incomplete comparison look-up table and is used to query connection-oriented data flows. The Bloom filter performs k hash calculations on the input data flow to obtain hash values and determines whether they correspond to k 1-bit entries in the Bloom filter.
[0016] Further, when the input data flow is the first packet of a connection-oriented data flow and it is determined after querying the data flow table that it does not conform to any data flow entry in the data flow table, the first packet is directed to the data flow analysis module for analysis, classification, and identification of the application program category to which it belongs, and then, according to the process set by the network device, the first packet is forwarded according to the destination information recorded in the header of the first packet. When it is determined that the input data flow is distorted in the data flow filter, the data flow analysis module pre-inserts the input data flow into the data flow table and sets a corresponding processing policy according to the application program category to which the input data flow belongs.
[0017] Furthermore, when it is determined that the input data stream meets the filtering conditions of the data stream filter, the input data stream that meets the filter conditions is filled into the data stream filter, and the input data stream that was pre-inserted into the data stream table to prevent distortion of the input data stream in the data stream filter is removed.
[0018] To enable a further understanding of the features and technical content of the present invention, please refer to the following detailed description of the present invention and the accompanying drawings. However, the provided drawings are only for reference and illustration and are not used to limit the present invention. Description of the Drawings
[0019] Figure 1 Schematic diagram of the prior art for processing data streams in a network switch;
[0020] Figure 2 Schematic diagram of an example implementation of a Bloom filter;
[0021] Figure 3 Schematic diagram of an example of a Bloom filter using parallel k hash tables;
[0022] Figure 4 Schematic diagram of an architectural embodiment of a data stream processing system using a data stream filter;
[0023] Figure 5 Flowchart of an embodiment of a data stream processing method using an incomplete comparison with a data stream filter;
[0024] Figure 6 Flowchart of an embodiment for processing the first packet of a data stream;
[0025] Figure 7 Flowchart of an embodiment of a method for processing a distorted data stream in a data stream filter;
[0026] Figure 8 Flowchart of an embodiment of the operation of a data stream analysis module in a data stream processing method Figure 1 ; and
[0027] Figure 9 Flowchart of an embodiment of the operation of a data stream analysis module in a data stream processing method Figure 2 .
[0028] Symbol Explanation
[0029] 10: Input data stream
[0030] 12: Transfer table
[0031] 14: 5-tuple
[0032] 16: Output data stream
[0033] 18: Data flow analysis module
[0034] 101, 103: Data flow direction
[0035] 20: Bit array
[0036] 30: 5-tuple
[0037] 40: Input data stream
[0038] 42: Transfer table
[0039] 44: Data flow table
[0040] 45: Data flow filter
[0041] 46: Output data stream
[0042] 48: Data flow analysis module
[0043] 401 - 413: Data flow direction
[0044] 400: Memory
[0045] Data flow processing flow for incomplete comparison of S501 - S517
[0046] Flow for processing the first packet of data stream of S601 - S607
[0047] Processing flow for distorted data stream of S701 - S705
[0048] Data flow processing flow of S801 - S809
[0049] Data flow processing flow of S901 - S907 Specific implementation manners
[0050] The following are specific embodiments to illustrate the implementation manners of the present invention. Those skilled in the art can understand the advantages and effects of the present invention from the content disclosed in this specification. The present invention can be implemented or applied through other different specific embodiments, and various details in this specification can also be modified and changed based on different viewpoints and applications without departing from the concept of the present invention. Additionally, the drawings of the present invention are only for simple schematic illustration and are not drawn according to actual dimensions, which is hereby stated in advance. The following implementation manners will further detail the related technical content of the present invention, but the disclosed content is not used to limit the protection scope of the present invention.
[0051] It should be understood that although terms such as "first", "second", "third", etc. may be used herein to describe various elements or signals, these elements or signals should not be limited by these terms. These terms are mainly used to distinguish one element from another, or one signal from another. Additionally, the term "or" used herein should, depending on the actual situation, possibly include any one or a combination of more of the associated listed items.
[0052] The disclosure is about a data stream processing method and system for incomplete comparison, where the proposed data stream processing method is based on an incomplete comparison table - lookup method. The embodiments show that the method uses a data stream filter such as a bloom filter (halo filter), and the advantage is that it can avoid problems such as the need for a large - capacity memory and processing burden caused by processing each data stream. According to one of the embodiments, the bloom filter used is a probabilistic data structure that can quickly verify whether each data stream exists in the data stream table and only uses relatively little storage space.
[0053] In the data stream processing method for incomplete comparison, the concept of one of the embodiments is to simultaneously use an incomplete comparison table - lookup and a table - lookup that requires complete comparison in the known art (such as a data stream table). What is queried is, for example, Figure 1 the data stream header showing the 5 - tuple 14, whereby the size of the table - lookup for complete comparison can be reduced, further reducing the overall system cost.
[0054] Furthermore, the incomplete comparison table - lookup means that when storing the record of a data stream into this incomplete comparison table - lookup, only the characteristic value of the data stream is stored. The characteristic value can be the compressed information of the data stream, a digest, or a hash value. When the characteristic value of the data stream is recorded in this incomplete comparison table - lookup, since the complete data stream is not stored in the table, the capacity of such a table can be much smaller than the record content of the complete comparison table - lookup. However, using an incomplete comparison table - lookup still needs to solve the situation of aliasing. For example, a data stream that does not originally exist in the table may be misjudged as existing during the lookup.
[0055] Therefore, the disclosure proposes a data stream processing method using an incomplete comparison with a data stream filter, and realizes the look-up table of the incomplete comparison through the data stream filter. The data stream filter can be the above-mentioned Bloom filter, and the Bloom filter has an extremely compact k - hash table structure. The principle is that when an element is added to a set (i.e., a look-up table is implemented), this element is mapped to k points in a bit array through k hash functions, and these (k) points are set to 1. When filtering (checking) the data stream, as long as it is checked whether the positions mapping these points are all 1, it can be determined whether the data stream is included in the set. If any of these mapped points is 0, the filtered data stream is not in this set; if the mapped points are all 1, the filtered data stream is covered by this set.
[0056] According to the data stream processing method using the incomplete comparison with this Bloom filter, when a data stream is inserted into the Bloom filter, the positions addressed by the k - hash corresponding to this data stream (the width of the data stream entry is 1 bit) will be set to 1; when using a certain data stream to look up the Bloom filter to determine whether this data stream has been inserted into this Bloom filter, if the k 1 - bit entries corresponding to this data stream after k - hash are all 1, it means that it conforms to the k - hash of the Bloom filter, that is, the data stream that conforms to the Bloom filter is filtered out, and it is judged that this data stream was previously inserted into this Bloom filter; otherwise, it passes through the Bloom filter and is not filtered out.
[0057] For example, taking Figure 2 the schematic diagram of the Bloom filter implementation example shown as an example, which shows a bit array 20. Taking the filter formed by the data set {x, y, z} as an example, in the example of k = 3, each data element in the data set {x, y, z} is subjected to 3 hash operations as the characteristic value of each data element. In this example, the positions addressed by each data are set to 1. In the figure, it shows that x has three connections, respectively representing 3 1s (3 bits are set to 1) addressed to the bit array 20, y has three connections, respectively representing 3 1s addressed to the bit array 20, and z has three connections, respectively representing 3 1s addressed to the bit array 20, thus forming a Bloom filter.
[0058] Taking the input data w as an example in the example, calculating the characteristic value of the data w, and after mapping this bit array 20, it is found that the characteristic value of the data w has a position corresponding to 0 (not all 1s), indicating that it is not in the data set {x, y, z}. This is an example of using the Bloom filter to perform the data stream processing method of incomplete comparison.
[0059] Figure 3 Next, a Bloom filter using parallel k hash tables is enumerated. In this example, the 5-tuple 30 of the input data stream is hashed k times in the Bloom filter to obtain k (k = 4) hash tables. By comparing hash 0, hash 1, hash 2, and hash 3 respectively, it can be used to filter k eigenvalue in the data stream. The width of the data stream entry (flow entry) addressed by k hashes is 1 bit and is set to 1 in the figure. In application, when calculating the eigenvalue of a 5-tuple of an input data stream, that is, performing k hash calculations, and comparing with different hash values against this Bloom filter using parallel k hash tables, it can be determined whether this input data stream exists in the lookup table with incomplete comparison.
[0060] Figure 4 FIG. shows an architectural embodiment diagram of a data stream processing system using a data stream filter proposed in the disclosure. The shown data stream processing system implements a data processing circuit in a network device, for example, a data processing circuit in a network switch. This circuit executes an incomplete comparison data stream processing method to process the data stream transmitted back and forth in the network switch.
[0061] In the architectural embodiment of the data stream processing system shown in the figure, when the system receives the input data stream 40, after parsing the packet header therein, it is passed to a forwarding table 42. The forwarding table 42 is used to record the Media Access Control address (MAC address) of the second layer (L2) or the network address (IP address) of the third layer (L3) in the network communication protocol of the data stream. In this example, the forwarding table 42 records the destination Media Access Control address (DMAC) and the destination port number (Port = Y) in this input data stream 40.
[0062] After the input data stream 40 is parsed, the data of the data stream will be submitted to a data stream table (flow table) 44 and a data stream filter 45 implemented in the memory 400 in the system (data processing circuit) simultaneously. According to the embodiment, the data stream table 44 can record 5-tuple data obtained from the headers of multiple data streams, such as the destination network address (DIP), source network address (SIP), destination layer 4 port (DP), source layer 4 port (SP), and communication protocol (Prot (Protocol)) of each data stream, and record the processing policies corresponding to each data stream entry, such as setting the priority of the data stream, discarding the data stream, copying the data stream (copy) to the data stream analysis module, etc. The data stream filter 45 is a Bloom filter as described above to implement an incomplete comparison lookup table.
[0063] In addition to having a data flow table 44 and a data flow filter 45 in the memory 400, the data flow processing system can also implement a data flow analysis module 48 in software or circuitry. When the input data flow 40 does not match any data flow entry in the data flow table 44, it is determined as a new data flow and copied to the data flow analysis module 48. In the data flow analysis module 48, the packet can be analyzed and classified through the program therein, the application program category to which the data flow belongs can also be identified, and the packet can be forwarded according to the destination information recorded in each packet header in accordance with the process set by the switch or the network device of the application, such as forwarding to the destination transmission port number Y (port = Y).
[0064] The data flow processing system can implement the processing circuitry in the network switch, in which the operation can simultaneously refer to Figure 5 The flowchart of the embodiment of the data flow processing method that adopts the incomplete comparison using the Bloom filter as described.
[0065] The data flow processing system receives the input data flow 40, which can be, for example, a packet formed from a source transmission port number X (port = X) ( Figure 4 data flow direction 401, step S501). After the input data flow 40 is parsed, the header information therein is obtained (step S503). As recorded in the transfer table 42 in the example, a destination transmission port number Y (port = Y) is obtained. Then, the programs of querying the data flow table 44 and the data flow filter 45 are performed, and according to the results of querying the data flow table 44 and the data flow filter 45, one of the following steps is executed. Among them, the data flow table 44 is applicable to all forms of data flows, while the data flow filter 45 is used to query connection-oriented flows.
[0066] Next, the data flow table 44 is queried according to the result of parsing the input data flow (data flow direction 403, step S505). During the query process, it is determined whether the characteristics of the input data flow 40 match any data flow entry (flow entry) in the data flow table (step S507). If the input data flow 40 matches any data flow entry in the data flow table 44 (yes), the corresponding processing policy recorded therein is applied (step S509). For example, when the input data flow 40 matches any data flow entry in the data flow table 44, one of the following processing policies can be executed: the input data flow can be set to a high priority order (in another embodiment, it is also allowed to set the priority to a low priority order), and at the same time, relevant actions are performed in accordance with the original process in the network device (such as a network switch) applied by the data flow processing system, such as forwarding this packet. As shown in this flowchart, the data flow can be forwarded to the destination transmission port as recorded in step S515. For the above example, the data flow is transmitted to the destination transmission port number Y (port = Y). Figure 4Data flow direction 413). Alternatively, the input data stream 40 may be discarded according to the processing strategy, or this input data stream may be copied to the data stream analysis module 48 (step S517). In addition to analyzing, classifying, and identifying packets, this packet is also forwarded simultaneously (step S515).
[0067] However, if the input data stream 40 does not match any data stream entry in the data stream table 44 (No), the data stream filter 45 is queried based on the result of parsing the input data stream (data flow direction 405, step S511) to determine whether this input data stream 40 meets any of the filtering conditions. Taking the Bloom filter as an example, the elements in the input data stream are hashed k times to obtain a hash value, which is used as the characteristic value of the input data stream, and based on this, it is determined whether it corresponds to k 1-bit entries in the Bloom filter (step S513). If the query result is that one of the filtering conditions is met (Yes), it means that this input data stream 40 is an existing data stream that has transmitted more than multiple packets, but no additional action is required. Just directly execute an action according to the original process set in the network device (such as a network switch) applied by the data stream processing system. For example, forward the input data stream to the destination transmission port number Y (port = Y) according to the setting of the network device (step S515), forming the output data stream 46 in the figure. However, if the query result shows that the input data stream 40 does not meet one of the filtering conditions in the data stream filter 45 (No), it means that the input data stream 40 does not meet any information in the data stream table and the data stream filter. Then, this input data stream is directed to the data stream analysis module to process this data stream (data flow direction 407, step S517). At the same time, this input data stream 40 can also be forwarded to the destination transmission port according to the original process in the network device (such as a network switch) applied by the data stream processing system, forming the output data stream 46.
[0068] Figure 6 The flowchart of an embodiment showing the processing of the first packet of a data stream, which runs in the data stream filter. The data stream processed by this embodiment process is specifically for a connection-oriented flow that establishes a communication session before transmitting data, such as a data stream of the TCP protocol; conversely, the proposed data stream filter does not process non-connection-oriented data streams, such as data streams of the UDP protocol.
[0069] In this embodiment, it starts from when the first packet of the connection-oriented data stream is received by the network device to which the method is applied (step S601). Taking the TCP protocol as an example, under this protocol, the header of the first packet records that the SYN flag is 1 and the ACK flag is 0. Therefore, the first packet in the input data stream can be judged according to the content recorded in these headers. If it is the first packet, it indicates a new data stream and does not match any data stream entry in the data stream table, and it is directly directed to the data stream analysis module (data stream direction 407, step S603).
[0070] At this time, the data stream analysis module will analyze, classify, and identify the application program category to which the input data stream belongs (step S605). Then, the data stream analysis module records the analysis and forwards this packet to the destination transmission port number Y according to the process set by the network device according to the destination information recorded in the packet header ( Figure 4 data stream direction 413, step S607). It should be noted here that the work of the data stream analysis module is to obtain the packets that do not match the data stream entries in the data stream table according to the query and comparison results of the data stream table, and also receive the data streams that do not meet any filtering conditions in the data stream filter, and analyze, classify, and identify the application program categories to which these data streams belong.
[0071] Continue Figure 6 Process Figure 7 Show the process of the embodiment of the method for processing the distorted data stream in the data stream filter.
[0072] When the first packet of the connection-oriented data stream is directed to the data stream analysis module, the data stream analysis module can judge whether this new data stream will cause aliasing if it is put into the data stream filter. This judgment is also to judge whether there will be a conflict with any filtering condition (such as any data stream entry) already existing in the data stream filter. Taking the data stream under the TCP protocol as an example, if the first packet is received and it is found that the values of the k data stream entries associated with this data stream are not equal to 0, it means there is aliasing.
[0073] In this way, when it is judged that the input data stream is a new data stream and there will indeed be aliasing in the data stream filter (step S701), the data stream analysis module pre-inserts it into the data stream table ( Figure 4 data stream direction 409, step S703), and sets the corresponding processing strategy (step S705). For example, in the data stream table, the processing strategy of this new data stream is set to be copied to the data stream analysis module. The purpose of this is to enable the other (the 2nd to the Nth) packets of this data stream except the first packet to be copied to the data stream analysis module for analysis through the processing strategy of the data stream table, rather than being directly transferred to the output port of the network device due to wrongly matching the data stream filter.
[0074] Figure 8 Next, a flowchart of an embodiment of the operation of the data flow analysis module in the data flow processing method is shown. In this example, after the first N packets of the received input data flow are analyzed by the data flow analysis module, the data flow analysis module will determine the application program category of this input data flow (step S801), and give a corresponding processing strategy (step S803). For example, it can be set to a high priority order or discard this packet. Then, the process determines whether this input data flow already exists in the data flow table (step S805). If the data flow does not exist in the data flow table (there is no matching data flow entry) (No), the data flow that does not conform to the data flow entry defined in the data flow filter can be inserted into the data flow table (step S807). If the distortion factor generated by the conflict of the foregoing data flow filter (such as step S701) of this data flow has been pre-put into the data flow table, the data flow analysis module can directly change the processing strategy of the data flow entry (flow entry) recorded in the data flow table from the original copy to the data flow analysis module to a desired final strategy (step S809).
[0075] Another embodiment of the operation of the data flow analysis module in the data flow processing method can be referred to Figure 9 in the flowchart shown
[0076] When receiving an input data flow, the data flow is copied to the data flow analysis module. The data flow analysis module will analyze the first N packets, determine the application program category of this input data flow (step S901), and obtain the corresponding processing strategy therefrom (step S903). After comparing with the filtering conditions of the data flow filter, the data flow that meets the filter conditions can be directly filled into the data flow filter (step S905) to become a data flow entry therein.
[0077] At this time, referring to Figure 7 step S701, in order to avoid the distortion of the input data flow caused by the conflict in the data flow filter, the input data flow will be pre-placed in the data flow table, and its processing strategy is to copy the data flow to the data flow table. After step S905, if it meets the filtering conditions in the data flow filter after analyzing N packets, the data flow analysis module removes this data flow from the data flow table and then fills this data flow into the data flow filter( Figure 4 data flow direction 411, step S907).
[0078] In summary, according to the method and system for processing data streams with incomplete comparison described in the above embodiments, an incomplete comparison data stream filter (which can be implemented by a Bloom filter) and a data stream table for performing complete comparison are adopted simultaneously. Among them, data streams that meet the filtering conditions in the data stream filter are placed into the data stream filter. When a data stream that does not meet the filtering conditions is received, the data stream is copied to the data stream analysis module. In this way, the requirement for the existing data stream table is reduced, so that the data stream table only records data streams that require special processing (such as data streams with high priority or malicious data streams) and connection-less flows, rather than storing all data streams, which can reduce the memory space required by the existing method of only using complete comparison for table lookup. Therefore, the overall system cost and additional burden can be effectively reduced.
[0079] The content disclosed above is only the preferred feasible embodiment of the present invention, and does not limit the claims of the present invention. Therefore, all equivalent technical changes made by using the content of the specification and drawings of the present invention are included in the claims of the present invention.
Claims
1. An incomplete comparison data stream processing method, applied to a network device, includes: Receiving an input data stream and parsing the input data stream; Querying a data stream table according to the result of parsing the input data stream to determine whether the input data stream conforms to any data stream entry in the data stream table; Querying a data stream filter according to the result of parsing the input data stream to determine whether the input data stream conforms to any filtering condition in the data stream filter; Wherein, according to the results of querying the data stream table and the data stream filter, one of the following steps is executed: When the input data stream conforms to any data stream entry in the data stream table, apply a corresponding processing policy; When the input data stream does not conform to any data stream entry in the data stream table, query the data stream filter again to determine whether the input data stream conforms to any filtering condition therein; When the input data stream conforms to any filtering condition in the data stream filter, indicating that the input data stream is an existing data stream that has been transmitted for more than multiple packets, execute an action according to the process set by the network device; and When the input data stream does not conform to any filtering condition in the data stream filter, direct the input data stream to a data stream analysis module to process the input data stream, wherein the data stream filter uses a Bloom filter to implement an incomplete comparison look-up table and is used to query the data stream for connection orientation.
2. The incomplete comparison data stream processing method according to claim 1, wherein the data stream table is applicable to all forms of data streams. When the input data stream conforms to any data stream entry in the data stream table, one of the following processing policies is executed: Set the input data stream to a high priority order and forward the input data stream to a destination transmission port; Discard the input data stream; and Copy the input data stream to the data stream analysis module and then forward the input data stream to the destination transmission port.
3. The incomplete comparison data stream processing method according to claim 2, wherein the data stream table records the 5-tuple data in the header of the data stream, including a destination network address, a source network address, a destination layer 4 port, a source layer 4 port, and a communication protocol, and records the processing policies corresponding to each data stream entry.
4. The incomplete comparison data stream processing method according to claim 1, wherein the Bloom filter performs k hash calculations on the input data stream to obtain hash values and determines whether they correspond to k 1-bit entries in the Bloom filter.
5. The incomplete comparison data stream processing method according to claim 4, wherein the Bloom filter performs k hash calculations on the 5-tuple of the input data stream to obtain k hash values.
6. The incomplete comparison data stream processing method according to any one of claims 1 to 5, wherein the data stream analysis module is used to perform analysis and classification on the packets of the input data stream and also identify the application program category to which the input data stream belongs.
7. The data stream processing method with incomplete comparison as described in claim 6, wherein, When the received input data stream is the first packet of a connection-oriented data stream, after querying the data stream table and determining that there is no data stream entry in the data stream table that matches it, the first packet is directed to the data stream analysis module for analysis, classification, and identification of the application program category to which it belongs, and then the first packet is forwarded according to the destination information recorded in the header of the first packet in accordance with the process set by the network device.
8. The data stream processing method with incomplete comparison as claimed in claim 6, wherein, When it is determined that the input data stream is distorted in the data stream filter, the data stream analysis module pre-inserts the input data stream into the data stream table and sets a corresponding processing policy according to the application program category to which the input data stream belongs.
9. A data stream processing system, provided in a network device, includes: A memory, in which a data stream table and a data stream filter are provided; And A data stream analysis module that performs analysis and classification on the packets of an input data stream and also identifies the application program category to which the input data stream belongs; Wherein the data stream processing system executes an incomplete matching data stream processing method, including: Receiving the input data stream and parsing the input data stream; According to the result of parsing the input data stream, querying the data stream table to determine whether the input data stream matches any data stream entry in the data stream table; According to the result of parsing the input data stream, querying the data stream filter to determine whether the input data stream matches any filtering condition in the data stream filter; Wherein, according to the results of querying the data stream table and the data stream filter, one of the following steps is executed: When the input data stream matches any data stream entry in the data stream table, a corresponding processing policy is applied; When the input data stream does not match any data stream entry in the data stream table, query the data stream filter again to determine whether the input data stream matches any of the filtering conditions; When the input data stream matches any filtering condition in the data stream filter, indicating that the input data stream is a data stream that already exists and has been transmitted for more than multiple packets, an action is executed according to the process set by the network device; and When the input data stream does not match any filtering condition in the data stream filter, direct the input data stream to the data stream analysis module to process the input data stream, wherein the data stream filter uses a Bloom filter to implement an incomplete matching look-up table and is used to query connection-oriented data streams.
Citation Information
Patent Citations
NUMA aware network interface
US20140122634A1