Network traffic analysis method, device and storage medium

By pre-matching identifying the location of characteristic fragments in network traffic and merging adjacent fragments, the problems of waste of resources and poor real-time performance caused by full scanning are solved, and efficient network traffic analysis is achieved.

CN120263542BActive Publication Date: 2025-09-02HUNAN RONGTENG NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510727223.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-09-02
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

In the existing network traffic analysis, full scanning payload leads to waste of resources and poor real-time performance, especially scanning of invalid payload areas occupies a large amount of resources, affecting real-time business performance.

Method used

By pre-matching, identify the potential hit position of the featured fragment in the rule library, extract the initial fragment from the payload, and merge the adjacent initial fragments into detection fragments according to the merge conditions. Only the merged detection fragment and the unmerged initial fragment are matched byte byte to avoid scanning of invalid areas.

Benefits of technology

While ensuring detection accuracy, it reduces the amount of redundant data processing, improves processing efficiency and real-time performance, optimizes resource utilization, and improves the overall performance of network traffic analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263542B_ABST
    Figure CN120263542B_ABST
Patent Text Reader

Abstract

The present invention discloses a network traffic analysis method, device, and storage medium, which relate to the field of data transmission and solve the problem of resource waste caused by full payload scanning. In this solution, only the potential hit positions of feature fragments in the rule library are identified through pre-matching, and the payload areas without threats are quickly filtered out; initial fragments are extracted based on the potential hit positions and the maximum rule length, and adjacent initial fragments are dynamically merged into detection fragments through merging conditions, reducing the amount of redundant data to be processed; the algorithm is only performed on the merged detection fragments and the unmerged initial fragments. The byte-by-byte matching is thus avoided, thereby avoiding the scanning of invalid payload areas. While ensuring detection accuracy, the processing efficiency is improved, and the problems of high resource usage and poor real-time performance caused by forced full scanning in traditional solutions are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data transmission, and in particular to a network traffic analysis method, device and storage medium. Background Art

[0002] Deep Packet Inspection (DPI) technology identifies and manages network traffic by analyzing packet headers and payloads. It utilizes the Aho-Corasick (AC) multi-pattern matching algorithm to identify traffic characteristics, detect security threats, and implement service policy control. The existing detection process involves extracting the entire payload from the message. The finite state machine built by the AC algorithm then scans each byte, recording the rule that triggers the final state (hit node) and its location. However, in existing solutions, the AC algorithm is forced to scan the entire payload regardless of whether it contains valid threat signatures, resulting in wasted resources. For example, if only the first 10 bytes of a 1KB payload contain threat signatures, the subsequent 1014 bytes of ineffective scanning will consume resources, impacting real-time service performance. Summary of the Invention

[0003] The purpose of the present invention is to provide a network traffic analysis method, device and storage medium to avoid scanning of invalid payload areas, improve processing efficiency while ensuring detection accuracy, and solve the problems of high resource usage and poor real-time performance caused by forced full scanning in traditional solutions.

[0004] To solve the above technical problems, the present invention provides a network traffic analysis method, comprising: extracting a payload from a network message, scanning the payload, and identifying a potential hit position of a feature fragment including a rule in a rule library; the feature fragment is a plurality of consecutive bytes in the rule; extracting at least one initial fragment from the payload based on the potential hit position and the maximum rule length in the rule library; judging whether there is an initial fragment that meets a merging condition in at least one initial fragment, and if so, merging the initial fragments that meet the merging condition to obtain a detection fragment; matching the merged detection fragment and / or the initial fragment that does not meet the merging condition with the rules in the rule library byte by byte to determine a final hit rule and a final hit position.

[0005] To solve the above technical problems, the present invention also provides a network traffic analysis device, comprising: a memory for storing a computer program; and a processor for implementing the steps of the network traffic analysis method described above when executing the computer program.

[0006] To solve the above technical problems, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the network traffic analysis method as described above are implemented.

[0007] The present invention provides a network traffic analysis method, device, and storage medium, relating to the field of data transmission, and solving the problem of resource waste caused by full payload scanning. In this solution, pre-matching is used to identify only the potential hit positions of feature fragments in the rule library, quickly filtering out non-threatening payload areas; initial fragments are extracted based on the potential hit positions and the maximum rule length, and adjacent initial fragments are dynamically merged into detection fragments through merging conditions, reducing the amount of redundant data to be processed; an algorithmic byte-by-byte matching is performed only on the merged detection fragments and the unmerged initial fragments, thereby avoiding the scanning of invalid payload areas. This improves processing efficiency while ensuring detection accuracy, solving the problems of high resource usage and poor real-time performance caused by forced full scanning in traditional solutions. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] In order to more clearly illustrate the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0009] Figure 1 A schematic diagram of a conventional network traffic analysis method;

[0010] Figure 2 A schematic diagram of an improved network traffic analysis method;

[0011] Figure 3 A schematic diagram of a network traffic analysis method provided by the present invention;

[0012] Figure 4 A flow chart of a network traffic analysis method provided by the present invention;

[0013] Figure 5 This is a schematic diagram of identification information of an initial segment provided by the present invention;

[0014] Figure 6 A schematic diagram of a specific flow chart of a network traffic analysis method provided by the present invention;

[0015] Figure 7 A schematic diagram of identification information of a merged segment provided by the present invention;

[0016] Figure 8 A schematic diagram of a network traffic analysis device provided by the present invention;

[0017] Figure 9 A schematic diagram of a computer-readable storage medium provided by the present invention. DETAILED DESCRIPTION

[0018] The core of the present invention is to provide a network traffic analysis method, device and storage medium to avoid scanning invalid payload areas, improve processing efficiency while ensuring detection accuracy, and solve the problems of high resource usage and poor real-time performance caused by forced full scanning in traditional solutions.

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0020] like Figure 1 As shown in the figure, a network traffic analysis method extracts the complete payload from the network message and then directly feeds the entire payload into the DPI AC matching engine, which performs a byte-by-byte match on the payload to obtain a final hit result. Regardless of whether the payload to be matched ultimately matches, the engine must retrieve the entire payload. If only a partial hit is found, this significantly wastes the DPI engine's matching performance, resulting in low overall DPI performance given limited FPGA resources.

[0021] like Figure 2 As shown, after extracting the complete payload, a pre-match is performed on the payload to obtain identifiers for some pre-matching segments. Based on these identifiers, corresponding pre-matching payload segments are extracted from the complete payload. These pre-matching payload segments are then fed into the DPI AC matching engine for byte-by-byte matching to obtain hit results. In some cases, this can reduce the computational load. However, in other cases, the computational load increases. For example, a 128-byte complete payload {{16{32'haabbccdd}},512'h0} matches the 5-byte keyword rule 0xaabbccddee. The 4-byte pre-matching filter for 0xaabbccdd matches the hit position of 4×n, where n is an increasing integer from 0 to 15. This scheme ultimately extracts 16 64-byte hit segments, resulting in a total of 1024 bytes of segments fed into the DPI engine, far exceeding the 128 bytes of the original complete payload. In this case, the computational load increases significantly.

[0022] To solve the above problems, Figure 3 and Figure 4 As shown, the present invention provides a network traffic analysis method, comprising:

[0023] S11: extracting a payload from a network message, scanning the payload, and identifying potential hit locations of feature fragments of rules in a rule base; the feature fragments are multiple consecutive bytes in the rules.

[0024] In this step, the payload is first extracted from the network message, excluding the network message header and protocol control information, focusing on the actual application data. The payload is the main content transmitted in the network message, containing the transmitted business data or potential characteristics of network security threats. After the payload is extracted, it is scanned to identify areas that may contain threat characteristics for further processing.

[0025] This process uses a rule-based feature fragment identification method. The rule base contains multiple network security rules, each of which defines a set of feature fragments. These feature fragments consist of multiple consecutive bytes and represent specific threat patterns or security events. Specifically, a feature fragment can be the first two, three, four, or eight consecutive bytes in a rule. For example, a rule's feature fragment might have the first two bytes representing a certain protocol type, with the following bytes representing specific attributes of the packet. By extracting these feature fragments, potential threat areas can be identified more quickly, improving matching efficiency.

[0026] By scanning the payload, we search for characteristic fragments in these rules and mark potential hit locations of these characteristic fragments. Potential hit locations are areas that may contain characteristic fragments in the rules. Although not every potential hit location will trigger a rule hit, a preliminary scan can quickly identify these potential areas.

[0027] The advantage of this scanning method is that, through a preliminary scan of the payload and the identification of regular feature fragments, it can effectively screen out areas that may contain threats, avoiding the waste of resources required to fully scan the entire payload. This not only improves data processing efficiency but also reduces the burden of scanning invalid data, providing an important basis for subsequent more accurate matching operations.

[0028] S12: Extract at least one initial segment from the payload according to the potential hit position and the maximum rule length in the rule base.

[0029] Based on the potential hit positions identified above, combined with the maximum rule length of each rule in the rule base, the initial fragment is further extracted from the payload. Specifically, the potential hit position refers to the area in the payload that may contain the characteristic fragment in the rule base, and the maximum rule length is the fragment length of the longest expected matching rule among all the rules in the rule base. Through these two parameters, it is possible to determine which areas to extract from the payload as initial fragments for further analysis. Specifically, by identifying the starting position of the characteristic fragment of the potential hit position in the payload, combined with the maximum rule length, an initial fragment can be extracted. The length of this initial fragment at least covers the possible rule characteristic fragments, thereby ensuring that the extracted fragment can contain complete potential matching information. The purpose of extracting the initial fragment is to perform effective preliminary verification of the potential threat area in the payload in a shorter time, without having to wait for the byte-by-byte scan of the entire payload to be completed.

[0030] The extraction of initial segments is based not only on the distribution of potential hit locations but also on the maximum rule length limit in the rule base, avoiding the extraction of irrelevant areas and ensuring that the initial segments are long enough to cover all possible signature segments. These initial segments become the basic units for subsequent merging and matching operations, ensuring a more efficient and targeted matching process. This approach allows for the rapid identification of preliminary candidate regions that match signature segments in the rule base without scanning the entire payload, laying the foundation for further byte-by-byte matching, reducing the amount of redundant data processing, and thus improving overall processing efficiency.

[0031] S13: Determine whether there is an initial segment that meets the merging condition in at least one initial segment. If so, merge the initial segments that meet the merging condition to obtain a detection segment.

[0032] Next, the extracted initial fragments need to be judged based on whether there are initial fragments that can be merged. The merging condition means that among multiple initial fragments, if they are adjacent and meet certain rule requirements (for example, continuity between bytes or specific matching logic), these initial fragments can be merged into a larger detection fragment. By merging, the amount of redundant data in subsequent byte-by-byte matching operations can be reduced, thereby improving processing efficiency. The merged detection fragment usually contains the complete information of multiple initial fragments and can cover more potential threat feature areas. The merging operation not only optimizes resource utilization and avoids repeated processing of invalid data, but also significantly improves the real-time performance and overall efficiency of network traffic analysis while ensuring detection accuracy.

[0033] S14: Match the merged detection segments and / or the initial segments that do not meet the merging conditions with the rules in the rule library byte by byte to determine the final hit rule and the final hit position.

[0034] Finally, a byte-by-byte match is performed on the merged detection fragments and the initial fragments that did not meet the merge criteria to determine the final hit rule and hit location. Since the merged detection fragment contains information from multiple initial fragments, it can be matched over a wider range, increasing matching accuracy and efficiency. The unmerged initial fragments are processed separately to ensure that each fragment is accurately matched. During the byte-by-byte matching process, these fragments are compared with the rules defined in the rule library to check for patterns that match the rule feature fragments, and the specific location where the rule was triggered is recorded.

[0035] The byte-by-byte matching approach can employ the AC (Aho-Corasick) algorithm, a highly efficient multi-pattern matching algorithm. The AC algorithm handles the matching of multiple patterns by constructing a finite state machine (FSM). First, it organizes the feature fragments from all rules into a Trie tree and then constructs a failure pointer to accelerate the matching process. During the actual matching process, the AC algorithm scans the input data byte by byte, moving along the Trie tree path. When a match is no longer possible, the failure pointer is used to backtrack to a valid state. This algorithm can match multiple rules simultaneously, and its time complexity is linear in the input data length and the number of patterns. Therefore, it is highly efficient in large-scale data matching, enabling fast and accurate byte-by-byte matching.

[0036] Through this precise byte-by-byte matching, the final hit rule and hit location can be determined, thereby ensuring the accuracy and reliability of traffic analysis, while avoiding redundant scanning of invalid data and improving processing efficiency.

[0037] like Figure 5 As shown, in an exemplary embodiment, the identification information of the initial fragment includes a head identifier, a tail identifier, a hit identifier and a hit position offset; the head identifier indicates whether the current initial fragment is the first initial fragment in the network message, the tail identifier indicates whether the current initial fragment is the last initial fragment in the network message, the hit identifier indicates whether the current initial fragment is the fragment corresponding to the potential hit position, and the hit position offset indicates the offset address of the starting position of the current initial fragment in the payload.

[0038] In this embodiment, the identification information of the initial fragment includes a head identifier, a tail identifier, a hit identifier, and a hit position offset. These identification information are used to effectively mark and track the characteristics of each initial fragment and its position in the payload. Specifically, the head identifier is used to indicate whether the current initial fragment is the first initial fragment in the network message, that is, it identifies the initial fragment corresponding to the first potential hit position. If the head identifier is "1", it means that the fragment is the beginning of a potential hit sequence. The tail identifier indicates whether the current initial fragment is the last initial fragment in the network message, that is, it identifies the initial fragment corresponding to the last potential hit position. If the tail identifier is "1", the fragment marks the end of the potential hit sequence. The hit identifier is used to indicate whether the current initial fragment actually corresponds to the potential hit position, that is, to determine whether the fragment contains a feature fragment that matches the rule in the rule base. If the hit identifier is "1", it means that the initial fragment matches the feature fragment of a certain rule and is a valid hit fragment. The hit position offset represents the offset address of the starting position of the current initial fragment in the payload. This offset value can accurately indicate the position of the initial fragment in the entire payload, thereby providing a position reference for subsequent matching and analysis.

[0039] In this embodiment, by combining these identification information, the initial fragments can be accurately managed and processed, ensuring the identification of potential hit positions and the efficient execution of subsequent matching processes, while reducing redundant calculations during large-scale data processing and improving overall processing efficiency.

[0040] In an exemplary embodiment, determining whether there is an initial fragment that meets the merging condition in at least one initial fragment includes: obtaining the values ​​of the head identifier and the tail identifier of the current initial fragment; if the head identifier and the tail identifier of the current initial fragment are both first preset values, determining that the current initial fragment is the first initial fragment and the last initial fragment, and does not meet the merging condition.

[0041] Specifically, the head identifier and the tail identifier are respectively used to indicate whether the current initial fragment is the first or last initial fragment in the network message. When the current initial fragment is detected, the values ​​of the head identifier and the tail identifier of the fragment are first obtained. These two identifiers respectively identify the position of the fragment in the potential hit sequence. If the head identifier and the tail identifier of the current initial fragment are both the first preset value (that is, the value is 1), it means that the current initial fragment is both the first and last initial fragment in the network message. Based on the identification information, it is determined that the initial fragment does not meet the merging conditions. In other words, when the head identifier and the tail identifier are 1, the current initial fragment has neither the previous fragment to be merged nor the subsequent fragment to be merged, so it is regarded as a separate matching result and does not need to be merged with other fragments. At this time, the matching result of the initial fragment is directly output, that is, it is not compressed and merged, but the current matching state is maintained.

[0042] In this embodiment, this judgment mechanism avoids redundant merging operations on separate matching results with the same first and last beats, thereby improving processing efficiency and ensuring that the final matching result can be correctly output in specific circumstances (such as when there is only a single matching result) without the need for further calculations.

[0043] In an exemplary embodiment, after obtaining the values ​​of the head identifier and the tail identifier of the current initial fragment, it also includes: if the head identifier of the current initial fragment is a first preset value, and the tail identifier is a second preset value, it is determined that the current initial fragment is the first initial fragment and not the last initial fragment, and the current initial fragment meets the merging condition.

[0044] Specifically, when the head identifier of the current initial fragment is the first preset value (i.e., 1) and the tail identifier is the second preset value (i.e., 0), the current initial fragment is determined to be the first initial fragment in the network message, but not the last initial fragment. In this case, the current initial fragment meets the merging condition, that is, the fragment can be merged with the subsequent initial fragment. The first preset value of 1 indicates that the fragment is the beginning of a potential hit sequence, and the tail identifier of the second preset value of 0 indicates that the fragment is not the end of the sequence. Therefore, the current fragment is a header fragment, that is, it is the starting part of a sequence, but there may be other initial fragments connected to it. In this case, the header fragment needs to be compressed by default, that is, the initial fragment and the subsequent initial fragments are merged into a larger detection fragment, thereby reducing redundant data and improving processing efficiency.

[0045] This mechanism effectively identifies and processes consecutive matching segments, and through merging operations, reduces the number of byte-by-byte matches, improving overall processing efficiency. During the merging process, subsequent initial segments are concatenated with the current initial segment to form a new merged segment, ensuring a complete scan of potential threat areas while avoiding unnecessary repetitive processing.

[0046] In an exemplary embodiment, after obtaining the values ​​of the head identifier and the tail identifier of the current initial fragment, it also includes: if the head identifier of the current initial fragment is a second preset value, determining whether the length between the starting position of the current initial fragment and the starting position of the previous initial fragment is less than the maximum rule length; if less than, determining that the current initial fragment and the previous initial fragment meet the merging condition; if greater than or equal to, determining that the current initial fragment and the previous initial fragment do not meet the merging condition.

[0047] Specifically, when the header identifier of the current initial fragment is the second preset value (i.e., the value is 0), it will be further determined whether the length between the starting position of the current initial fragment and the starting position of the previous initial fragment is less than the maximum rule length. This judgment is intended to determine whether there is overlap between the current fragment and the previous fragment, thereby deciding whether to perform a merge operation. If the length between the starting position of the current initial fragment and the starting position of the previous initial fragment is less than the maximum rule length, it means that the two initial fragments have overlapping parts in the payload, and the range they jointly cover will not exceed the maximum rule length specified in the rule library. Therefore, the two initial fragments can be merged into a larger detection fragment, thereby improving matching efficiency. At this time, it is determined that the current initial fragment and the previous initial fragment meet the merge condition, that is, there is an overlapping part, and the merge will reduce redundant matching.

[0048] Conversely, if the distance between the start position of the current initial segment and the start position of the previous initial segment is greater than or equal to the maximum rule length, it means that the distance between the two initial segments is large and there is no overlap, so there is no need to merge them. In this case, it is determined that the current initial segment and the previous initial segment do not meet the merge condition, that is, there is no overlap, and the two initial segments must be processed separately.

[0049] It can be seen that through this judgment mechanism, it is possible to effectively distinguish which initial fragments can be merged and which cannot be merged, ensuring that the merging operation is only performed under reasonable conditions, thereby improving processing efficiency and avoiding unnecessary calculations.

[0050] In an exemplary embodiment, the identification information of the detection segment includes at least two parameters: a merged first position and a merged length. The merged first position is used to represent the starting position of the merged detection segment in the payload, and the merged length is used to represent the total length of the merged detection segment.

[0051] Specifically, the merged first position represents the starting position of the merged detection segment in the payload. It indicates the specific location of the merged segment within the entire network packet payload and determines where the merged segment should begin matching scans. The merged length represents the total length of the merged detection segment, indicating the coverage of the merged segment, that is, the number of bytes in the merged segment.

[0052] These two parameters allow for accurate location of the merged detection segments and byte-by-byte matching, enabling efficient detection of potential threats. The identification of the merged first position and merged length is crucial for merging multiple initial segments, helping to avoid redundant computations and improve overall processing efficiency.

[0053] In an exemplary embodiment, the initial segments that meet the merging conditions are merged to obtain a detection segment, including: if the current initial segment is the first initial segment and not the last initial segment, the merged first position of the detection segment is initialized to the starting position of the current initial segment, and the merged length of the detection segment is initialized to the length of the current initial segment; the length of the current initial segment is the maximum rule length.

[0054] In this embodiment, when initial segments that meet the merging conditions need to be merged, the merged detection segments are determined according to the merging rules. Specifically, if the current initial segment is the first initial segment and not the last initial segment, the merged first position of the detection segments is first initialized to the starting position of the current initial segment. This means that the merged detection segments will be scanned starting from the starting position of the current initial segment. Next, the merged length of the detection segments is initialized to the length of the current initial segment, which ensures that the merged segment can completely cover the byte range of the current initial segment at the beginning.

[0055] In addition, the length of the detection segment is adjusted based on the maximum rule length, setting the length of the current initial segment to the maximum rule length. This design ensures that the scan range in the merged segment does not exceed the maximum match length defined in the rule base, thus avoiding scanning unnecessary byte areas and ensuring an efficient matching process.

[0056] Through this initialization process, the scanning range of the merged segments can be effectively managed, and each merged detection segment can be matched byte by byte within a reasonable length range, thereby improving the efficiency and accuracy of network traffic analysis.

[0057] In an exemplary embodiment, the initial segments that meet the merging conditions are merged to obtain a detection segment, including: if the current initial segment is the i-th initial segment, and the length between the starting position of the i-th initial segment and the starting position of the i-1-th initial segment is less than the maximum rule length, the merged length of the detection segment is updated to the cumulative length of the previous i initial segments; the total number of initial segments > i > 1, where i is an integer.

[0058] In this embodiment, when the initial segments that meet the merging conditions need to be merged, the length of the merged detection segment is dynamically updated according to the position of the current initial segment and its relative position to the previous initial segment. Specifically, if the current initial segment is the i-th initial segment, and the length between the starting position of the i-th initial segment and the starting position of the i-1-th initial segment is less than the maximum rule length, it means that the current initial segment and the previous initial segment overlap or are closely connected in the payload, so they can be merged into a larger detection segment. In this case, the merged length of the detection segment is updated to the cumulative length of the previous i initial segments, that is, the length of the current initial segment is added to the length of all the previous merged segments to form a continuous detection segment. This update process can ensure that the byte range covered by the merged detection segment includes all merged initial segments, thereby improving the matching efficiency.

[0059] Furthermore, this process applies when i is greater than 1 and less than or equal to the total number of initial segments, where i represents the index of the currently processed initial segment. In other words, merging occurs only when the distance between the start position of the current initial segment and the start position of the previous initial segment is less than the maximum rule length. The merge length gradually increases as the number of segments increases, ensuring that merging only occurs under reasonable conditions, avoiding unnecessary redundant operations and improving the efficiency of traffic analysis.

[0060] In addition, if Figure 6 As shown, a register may also be provided for storing the merged first position and merged length of the merged segment. When an initial segment that does not meet the merge condition is detected, the merged first position and merged length of the merged segment stored in the register are output.

[0061] Specifically, when the current initial segment is the first initial segment and the last initial segment, there is no need to merge or store them in a register, and the identification information of the current initial segment can be directly output.

[0062] When the current initial fragment is the first initial fragment and not the last initial fragment, it is assumed that the current initial fragment meets the merge condition. At this time, the starting address and length of the first initial fragment are initialized to the merged first address and merged length of the merged fragment and stored in the register.

[0063] When the current initial fragment is the middle initial fragment (that is, it is neither the first initial fragment nor the last initial fragment), calculate the interval between the starting position of the current initial fragment and the starting position of the previous initial fragment, and compare it with the maximum rule length. If it is less than the maximum rule length, it means that the current initial fragment meets the merging condition. At this time, the length of the fragments that meet the merging condition is accumulated, and the merged length stored in the register is updated; if it is greater than the maximum rule length, it means that the current initial fragment does not meet the merging condition. At this time, the parameter information (merged first position and merged length) of the merged fragment merged from the initial fragments that originally met the merging condition stored in the register is output. At the same time, the merged first position of the merged fragment is initialized to the starting position of the current initial fragment, and the merged length of the merged fragment is initialized to the length of the current initial fragment, and the judgment of the next initial fragment continues.

[0064] When the current initial segment is the last initial segment, calculate the interval between the starting position of the current initial segment and the starting position of the previous initial segment, and compare it with the maximum rule length. If it is less than the maximum rule length, it means that the current initial segment meets the merging condition. At this time, the length of the segments that meet the merging condition is accumulated, and the merged length stored in the register is updated; if it is greater than the maximum rule length, it means that the current initial segment does not meet the merging condition. At this time, the parameter information (merged first position and merged length) of the merged segment merged from the initial segments that originally met the merging condition stored in the register is output, and the non-merged result of the last initial segment is output in the next beat.

[0065] Figure 6 The specific process is as follows: initialize the merge head position M = 0, the merge length P = 0, and K is the maximum rule length. If the current initial segment has the head identifier = S(i), the tail identifier = E(i), and the hit position = O(i), determine whether S(i) is 1.

[0066] If S(i) is 1, it is determined to be the first initial fragment, and then further determine whether E(i) is 1. If E(i) is 1, it is determined that no merging is required, and output M(i)=O(i), P(i)=K; if E(i) is not 1, initialize the merged fragment identifier M(i)=O(i), P(i)=K, set i++; and enter the processing of the next initial fragment.

[0067] If S(i) is not 1, it is determined that this is not the first initial segment. At this time, the difference between the starting position of the current initial segment and the starting position of the previous initial segment is calculated (B(i) = O(i) - O(i-1)), and further judgment is made as to whether E(i) is 1.

[0068] If E(i) is not 1, it is determined that it is not the last initial segment, and whether B(i) is less than K. If it is, the merge condition is met, the cumulative merge length P(i) = P(i-1) + B(i), and i++ is set to start processing the next initial segment; if it is not less than K, it is determined that the merge condition is not met, the merge result identifier is output, and the merge segment identifier M(i) = O(i), P(i) = K is initialized, i++ is set to start processing the next initial segment.

[0069] If E(i) is 1, it is determined to be the last initial fragment, and whether B(i) is less than K is determined. If it is, it is determined that the merging condition is met, and the cumulative merging length P(i) = P(i-1) + B(i) is output, and the identifier of the merged fragment is output; if it is not less than K, it is determined that the merging condition is not met, and the identifier of the merged fragment stored in the register is output. In the next cycle, the non-merged result of the tail fragment (the last initial fragment) is output.

[0070] The information stored in the merged register may include not only the merged first position and the merged length, but also the head identifier, tail identifier, and hit identifier, such as Figure 7 shown.

[0071] Taking the above extreme case as an example, the merged pre-matching results are as follows. After the initial segment of the pre-matching results, the 16 matching result inputs are finally merged into one output, as shown in the following table:

[0072] Table 1 Schematic diagram of merged matching results

[0073]

[0074] Based on the offset and length of the merged first position result of the merged fragment, the corresponding payload fragment is taken out. At this time, the fragment sent to the DPI matching engine has been greatly compressed and will not exceed the original payload length. Taking the above extreme case as an example, when K=64, the length of the merged fragment is only 124 bytes, which is less than the original message length of 128 bytes, greatly improving the overall performance of the DPI engine.

[0075] In order to solve the above technical problems, Figure 8 As shown, the present invention also provides a network traffic analysis device, including: a memory 81 for storing a computer program; a processor 82 for implementing the steps of the above-mentioned network traffic analysis method when executing the computer program.

[0076] For further introduction to the network traffic analysis device, please refer to the above embodiments, and this application will not go into details here.

[0077] In order to solve the above technical problems, Figure 9As shown, the present invention further provides a computer-readable storage medium 91 on which a computer program 92 is stored. When the computer program 92 is executed by a processor, the steps of the network traffic analysis method described above are implemented.

[0078] For other introductions to the computer-readable storage medium 91 , please refer to the above embodiments, and this application will not go into details here.

[0079] It should also be noted that, in this specification, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.

[0080] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A network traffic analysis method, characterized in that: include: Extracting a payload from a network message, scanning the payload to identify potential hit locations of a characteristic segment of a rule in a rule library; the characteristic segment is a plurality of consecutive bytes in the rule; Extracting at least one initial fragment from the payload according to the potential hit position and the maximum rule length in the rule base; the identification information of the initial fragment includes a head identifier, a tail identifier, a hit identifier, and a hit position offset; the head identifier indicates whether the current initial fragment is the first initial fragment in the network message, the tail identifier indicates whether the current initial fragment is the last initial fragment in the network message, the hit identifier indicates whether the current initial fragment is the fragment corresponding to the potential hit position, and the hit position offset indicates the offset address of the starting position of the current initial fragment in the payload; determining whether there is an initial segment in at least one initial segment that meets a merging condition; and if so, merging the initial segments that meet the merging condition to obtain a detection segment; wherein the identification information of the detection segment includes at least two parameters: a merged first position and a merged length; the merged first position is used to indicate the starting position of the merged detection segment in the payload; and the merged length is used to indicate the total length of the merged detection segment; The merged detection segments and / or the initial segments that do not meet the merging conditions are matched byte by byte with the rules in the rule base to determine the final hit rule and the final hit position.

2. The network traffic analysis method according to claim 1, wherein: Determining whether there is an initial segment in at least one initial segment that meets the merging condition includes: Get the values ​​of the head and tail identifiers of the current initial segment; If the head identifier and the tail identifier of the current initial segment are both the first preset value, it is determined that the current initial segment is the first initial segment and the last initial segment, and the merging condition is not met.

3. The network traffic analysis method according to claim 2, wherein: After obtaining the values ​​of the head and tail identifiers of the current initial segment, it also includes: If the head identifier of the current initial segment is a first preset value and the tail identifier is a second preset value, it is determined that the current initial segment is the first initial segment and not the last initial segment, and the current initial segment meets the merging condition.

4. The network traffic analysis method according to claim 2, wherein: After obtaining the values ​​of the head and tail identifiers of the current initial segment, it also includes: If the header identifier of the current initial segment is a second preset value, determining whether the length between the starting position of the current initial segment and the starting position of the previous initial segment is less than the maximum rule length; If less than, it is determined that the current initial segment and the previous initial segment meet the merging condition; If it is greater than or equal to, it is determined that the current initial segment and the previous initial segment do not meet the merging condition.

5. The network traffic analysis method according to any one of claims 1 to 4, characterized in that: Merging the initial segments that meet the merging conditions to obtain detection segments includes: If the current initial segment is the first initial segment and not the last initial segment, initializing the merged first position of the detection segments to the starting position of the current initial segment, and initializing the merged length of the detection segments to the length of the current initial segment; The length of the current initial segment is the maximum rule length.

6. The network traffic analysis method according to claim 5, wherein: Merging the initial segments that meet the merging conditions to obtain detection segments includes: If the current initial segment is the i-th initial segment, and the length between the starting position of the i-th initial segment and the starting position of the i-1-th initial segment is less than the maximum rule length, the combined length of the detection segment is updated to the cumulative length of the previous i initial segments; the total number of initial segments > i > 1, where i is an integer.

7. A network traffic analysis device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the network traffic analysis method according to any one of claims 1 to 6 when executing a computer program.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the network traffic analysis method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Method, device and storage device for accurately detecting network traffic

    CN107426049A

  • Message rule matching method and device and electronic equipment

    CN118944958A