Regular expression matching method and device for detecting double-dictionary compressed data

By building a regular matching engine and preprocessing module, using the state equivalence of the finite state automaton, extracting the metadata of the double dictionary compressed data and skipping redundant data detection, the problem of slow detection of double dictionary compressed data in the prior art is solved, and efficient regular expression matching is achieved.

CN116484069BActive Publication Date: 2025-05-09ANHUI UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202310459944.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-25
Publication Date
2025-05-09
Estimated Expiration
2043-04-25

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently detect and match compressed data based on double dictionary, especially under the limited bandwidth of mobile networks, resulting in a significant reduction in detection speed.

Method used

By building a regular matching engine and preprocessing module, the metadata of the double dictionary compressed data is extracted, and the state equivalence of the finite state automaton skips detection redundant data to achieve accelerated detection.

Benefits of technology

It significantly improves the speed of detecting regular expression matching based on double dictionary compressed data, reduces memory overhead, and improves the performance and efficiency of the detection system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116484069B_ABST
    Figure CN116484069B_ABST
Patent Text Reader

Abstract

The present invention discloses a regular expression matching method and device for detecting compressed data based on a double dictionary. The method can skip detecting most of the compressed data with minimal overhead, and can effectively improve the detection speed. The method mainly includes two stages: preprocessing and matching. The preprocessing stage decompresses the compressed data and generates metadata information. The matching stage reads the metadata information and skips detecting most of the data represented by compressed coding in combination with the state equivalence of a finite state automaton. The technical solution of the present invention improves the basic theory of the compressed data detection method, significantly improves the compressed data detection speed, provides technical support for the detection system based on regular expression matching, and broadens the application scope of compressed data detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep packet inspection, and specifically relates to data compression, pattern matching and other related technologies. Specifically, it focuses on accelerating the detection of compressed data based on a dual dictionary, such as data compressed by the Brotli compression algorithm. Background Art

[0002] The rise and development of mobile Internet has generated a massive amount of network traffic. In order to improve data transmission efficiency on the limited bandwidth of mobile networks, enhance user experience and reduce the billing traffic generated by users, more and more network services use compression technology to compress the transmitted data, which brings new challenges to related tools and systems based on Deep Packet Inspection (DPI) technology. Taking HTTP network traffic as an example, the HTTP1.1 protocol uses Gzip as the default compression encoding, and the compression rate of transmitted data is about 20% (the compression rate is the ratio of the volume of data after compression to that before compression). When detecting such compressed traffic, the DPI system usually needs to decompress and detect all the decompressed data. The expansion of the data volume makes the system performance only 1 / 5 of that when detecting uncompressed traffic.

[0003] Existing methods for accelerating the detection of network compressed data have achieved good results in terms of detection speed. However, they mainly focus on the compressed data generated by the compression algorithm using a single adaptive dictionary. For example, reference [1] Pattern Matching in LZW Compressed Files, IEEE Transactions on Computers, 2005, 54(8): 929-938; reference [2] Accelerating Multi-pattern Matching on Compressed HTTP Traffic, IEEE / ACM Transactions on Networking 2012, 20(3): 970-983; reference [3] Accelerating regular expression matching over compressed http, IEEE Conference on Computer Communications, 2015, 540-548; reference [4] Efficient Regular Expression Matching over Compressed Traffic, Computer Networks, 2020, 168: 106996; reference [5] US Patent: US8458354, Multi-pattern matching in compressed communication traffic; Reference [6] Chinese Patent: ZL201710354909.0, a multi-string matching method for compressed traffic; Reference [7] Chinese Patent: ZL201810420111.6, a Pairs method for accelerating regular expression matching of compressed traffic; Reference [8] Chinese Patent: ZL201810419466.3, a Twins method for accelerating regular expression matching of compressed traffic; Reference [9] Chinese Patent: ZL201910950410.5, a compressed traffic pattern matching engine and pattern matching method based on FPGA platform. These methods focus on the data generated by compression algorithms such as LZ77 and LZW, and cannot be applied to compressed data based on dual dictionaries.

[0004] Dual-dictionary based compressed data is compressed using a compression algorithm that uses an adaptive dictionary and a static dictionary. For example, Brotli encoding proposed by Google has obvious advantages over existing encodings in terms of compression rate, compression or decompression speed, etc. Brotli has been used by Google as the default compression encoding for Chrome browser and Google-related services due to its lower compression rate and faster decompression speed. Currently, mainstream web browsers and web servers also support this compression encoding. This type of compressed data accounts for an increasing proportion of network traffic, but there is currently no efficient method for accelerating the detection of dual-dictionary based compressed data. Summary of the invention

[0005] The present invention provides a regular expression matching method and device for detecting double-dictionary compressed data, which can effectively improve the speed of detecting double-dictionary compressed data.

[0006] To achieve the above object, the present invention provides a method for detecting regular expression matching based on double-dictionary compressed data, comprising the following steps:

[0007] Step 1: Build a regular matching engine; call the regular matching engine to scan the static dictionary, store the returned finite state automaton state into the state area, and reset the active state state of the regular matching engine to the initial state;

[0008] Step 2: pre-process the double-dictionary compressed data to obtain decompressed data; at the same time, parse metadata from the double-dictionary compressed data to be detected and store it, wherein the metadata includes the length of the uncompressed data, which is recorded as len1, the compression code, which is recorded as<dist,len2> ;

[0009] Step 3: Read a metadata structure;

[0010] Step 4: read len1 bytes of data from the decompressed data, use the active state state as input, call the regular matching engine to scan, and update the state; this process saves the state obtained by scanning each data in the state area, and checks whether each state is the receiving state of the automaton, and outputs the receiving state and the corresponding character position as the matched pattern information as the detection result;

[0011] Step 5: Compression encoding according to metadata<dist,len2> Locate the data area corresponding to the compression code in the dynamic dictionary or static dictionary used by the compression algorithm, and locate the area corresponding to the code in the status area and the reference area of ​​the code according to the information of the compression code, record the position of the previous character of the code area in the status area as curPos, and record the position of the previous character of the reference area as refPos;

[0012] Step 6. Check whether the state saved at the refPos position is equivalent to the active state:

[0013] If they are equivalent, copy the state in the area referenced by the code to the current area, check whether the received state exists in the received copy state, if so, output the mode information, and then jump to step 3 to read the next metadata structure; otherwise, jump to step 7;

[0014] Step 7: Call the regular expression matching engine to scan the character at the curPos position in the code, update the state and write it to the state area synchronously, and then move refPos and curPos backward by one character each; if curPos is not the end of the code, jump to step 6; otherwise, jump to step 3.

[0015] Furthermore, in step 1, a regular matching engine is constructed according to the regular expression rule set.

[0016] Further, in step 2, preprocessing the dual-dictionary compressed data includes the following steps:

[0017] Parse the double-dictionary compressed data to obtain all compressed data blocks. Each data block corresponds to a string of compressed data. For each data block, parse the information of three fields, namely insert-copy-length, literal, and distance. Based on the information of the three fields, restore the uncompressed data and write it into text. Additionally, record insert-copy-length and distance as metadata for storage.

[0018] Further, step 5 includes the following steps:

[0019] Read the decompressed data and metadata to determine the data type:

[0020] If it is a common character, directly call the automaton to scan it character by character;

[0021] If it is encoded data, first determine whether the reference string of the current redundant data is taken from the dynamic dictionary or the static dictionary according to dist:

[0022] If it is a static dictionary, calculate the previous position refPos of the reference string subscript in the static dictionary according to dist;

[0023] If it is a dynamic dictionary, calculate the previous position refPos of the corresponding position of the reference string in the dynamic dictionary;

[0024] When processing encoded data, first check whether the current activation state and the state at the corresponding position of refPos are equivalent:

[0025] If they are equivalent, copy the state directly from the reference area, record the matching result, and update the current activation state;

[0026] If they are not equivalent, call the automaton scan, update the current activation state, and continue to compare with the corresponding state of the reference area until the encoding data processing is completed or it is equivalent and skipped directly.

[0027] A regular expression matching device for accelerated detection based on double-dictionary compressed data comprises a preprocessing module, a Rainbow detection module, an auxiliary information storage module and a regular matching engine; the regular matching engine is constructed by regular expression rules, the preprocessing module is used to parse the double-dictionary compressed data to be detected, and the Rainbow detection module is used to implement detection and output detection results.

[0028] Furthermore, the dual-dictionary compressed data to be detected is generated by a compression algorithm using an adaptive dynamic dictionary and a static dictionary; the preprocessing module parses the input dual-dictionary compressed data in combination with the static dictionary, outputs decompressed data, and sends the output decompressed data to the Rainbow detection module; the preprocessing module also outputs metadata to the auxiliary information storage module.

[0029] Furthermore, the regular expression matching engine is implemented using a finite state automaton-based implementation.

[0030] Furthermore, the auxiliary information storage module (103) is used to store auxiliary information, including a state area and a metadata area, wherein the auxiliary information includes a finite state automaton state returned by matching a static dictionary and decompressing data using a regular matching engine, and metadata capable of identifying the original composition of the compressed data, each metadata including the length of the uncompressed data and the compression code.

[0031] Furthermore, the Rainbow detection module identifies the decompressed data after being parsed by the preprocessing module based on the metadata stored in the metadata area, and distinguishes between the uncompressed data in the original state and the data represented by the compressed code; then the regular matching engine is called to match the uncompressed data character by character, and combined with the finite state automaton state saved by the auxiliary information, the Rainbow detection algorithm is used to skip the detection of the data represented by the compressed code.

[0032] Compared with the prior art, the present invention has at least the following beneficial technical effects:

[0033] (1) The theory is more complete

[0034] Existing work skips detecting redundant data based on the condition of equal state of finite state automaton, and its advantages can only be fully exerted when the regular matching engine adopts minimized DFA. The present invention proposes a regular expression matching method for accelerated detection based on double-dictionary compressed data, which skips detecting redundant data based on the state equivalence in the finite state automaton theory as the judgment condition. State equality is only a subset of state equivalence. The basic theory relied on by the present invention is more universal and the theory is more complete.

[0035] (2) Low overhead and fast detection speed

[0036] The present invention proposes a regular expression matching method for accelerating detection of double-dictionary compressed data, which makes full use of the principle of memory space locality and cache to reduce random memory access when using a finite state automaton; in the preprocessing stage, only metadata is output while outputting decompressed data, metadata information is read in the matching stage, and most of the data represented by compressed coding is skipped in combination with the state equivalence of the finite state automaton. The existing method for accelerating detection of double-dictionary compressed coding VCDIFF expands each byte to 4 bytes to identify different compressed data. Compared with this method, the metadata designed by the present invention reduces memory overhead and makes more full use of cache, thereby obtaining a performance improvement of multiple detection speed. In addition, the additional memory overhead introduced by the entire detection process only includes the metadata in the auxiliary information and the stored automaton state, and the required memory space is very small, and is not affected by the regular expression rule set, the regular matching engine and the data to be detected.

[0037] (3) High scalability and good compatibility

[0038] The method proposed in the present invention uses a regular matching engine constructed by a standard finite state automaton. The designed acceleration detection algorithm has no direct coupling relationship with the regular matching engine, which facilitates the method to be embedded in an existing system or application and has high scalability and compatibility. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 This is a schematic diagram of the data encoding process based on the dual-dictionary compression algorithm using Brotli as an example. The compression algorithm constructs an adaptive dynamic dictionary from the original text, and combines the static dictionary to compress the original text into a mixture of the original text that cannot be compressed and the compressed encoded data, and then performs Huffman coding on it.

[0040] Figure 2 This is a schematic diagram of the positional relationship between metadata, various types of compression codes, and uncompressed original text after data decompression, using Brotli as an example;

[0041] Figure 3It is a state transition diagram of a deterministic finite state automaton constructed using the regular expression "ab+cd|bc+d";

[0042] Figure 4 It is a schematic diagram of the Rainbow detection algorithm accelerating the detection of data represented by dynamic coding and static coding;

[0043] Figure 5 It is a functional framework diagram of a regular expression matching device for accelerating detection based on double-dictionary compressed data proposed by the present invention;

[0044] Figure 6 is a schematic diagram of the process of processing example compressed data by the algorithm of the present invention;

[0045] Figure 7 It is a performance comparison between the embodiments of the present invention and the existing method in terms of detection speed. DETAILED DESCRIPTION

[0046] In order to make the purpose and technical solution of the present invention clearer and easier to understand, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. The specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0047] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first" and "second" are used only for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, "multiple" means two or more. In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0048] The present invention proposes a regular expression matching method for accelerated detection of compressed data based on dual dictionaries, which outputs decompressed data and metadata information that can distinguish between uncompressed data in the original state and data represented by compressed codes through preprocessing; then, the context-free characteristics and state equivalence of finite state automata are used to skip the detection of most of the data represented by compressed codes, thereby achieving the purpose of accelerated detection. The technical solution of the present invention improves the basic theory of compressed data detection methods, significantly improves the speed of compressed data detection, provides technical support for detection systems based on regular expression matching, and broadens the application scope of compressed data detection.

[0049] To further illustrate the specific content of the present invention, the implementation part takes Brotli as an example for introduction, and other data compressed and encoded by compression algorithms using dual dictionaries can be implemented with reference to this embodiment. Next, the terms and related technologies involved in the present invention are first introduced.

[0050] (1) Brotli

[0051] Brotli is a new compression codec proposed by Google in 2011, and its data format has been standardized by RFC7932. Due to Brotli's performance advantages in compression rate, decompression speed and resources used, it has been supported by mainstream web browsers. Google uses it as the default compression codec when Chrome transmits HTTP data, and deploys it in a variety of network services. Brotli combines the traditional adaptive dynamic dictionary with a shared static dictionary to compress data. The adaptive dynamic dictionary is implemented based on the LZ77 variant algorithm, and the shared static dictionary contains words, paragraphs, and code blocks that appear frequently in HTTP data packets. The data encoded using dynamic and static dictionaries is then Huffman encoded to output the final Brotli compressed data.

[0052] Brotli compresses data by referring to both the dynamic dictionary and the static dictionary, and compresses a continuous sequence of characters with three instructions, namely insert-copy-length, literals, and distance. insert-copy-length gives two lengths, namely the literals length and the encoding length. In the field of data compression, literals represent characters or character sequences that cannot be compressed, and encodings represent repeated character sequences represented by encodings. The literals instruction is a sequence of characters that cannot be compressed. The distance instruction indicates finding the distance between the encoding and its corresponding reference string.

[0053] Figure 1A schematic diagram of Brotli's process of processing the character sequence "xbcccdabcccdabcd" is given. Brotli first searches the dynamic dictionary and the static dictionary for the existence of a substring based on the LZ77 algorithm. Then it encodes the found repeated substrings, and keeps the incompressible character sequences in their original state. Finally, Huffman coding is used to re-encode the incompressible character sequences and codes.

[0054] Figure 2 The compressed character sequence "xbcccdabcccdabcd" is taken as an example to illustrate the positional relationship between the compressed character sequence and the Brotli instruction. In this example, the original character sequence is compressed and encoded into two groups of instructions. The first instruction "(7,5)xbcccda(6)" will output the character sequence "xbcccdabcccd" after decompression, in which the 7-character sequence "xbcccda" at the beginning cannot be compressed, and the 5-character sequence "bcccd" following it is a substring of the previous 7-character sequence. (7,5) is an insert-copy-length instruction, indicating that a string of 7-byte characters that cannot be compressed will be output immediately after it, that is, "xbcccda" represented by the literals instruction; and then there is a code, indicating that the length of the character sequence to be output is 5, and the distance between the character sequence and the reference string (the first "bcccd") is 6 bytes. Since the original character sequence referenced by the code is located in the adaptive dynamic dictionary constructed by the compression process, the present invention calls it a dynamic code.

[0055] The second instruction "(0, 4)(16)" will output "abcd" after decompression. The insert-copy-length instruction is "(0, 4)", which means that there is no character sequence that cannot be compressed, that is, the literals instruction does not output any characters, and the subsequent encoding length is 4, and the distance from the reference string is 16 bytes. In this example, we assume that the maximum offset length is 10. This distance exceeds the maximum offset, indicating that the reference string is located in the static dictionary, so the second instruction only copies the output character sequence "abcd" from the static dictionary. The present invention refers to the encoding of the reference static dictionary as static encoding.

[0056] (2) Regular expression matching

[0057] A regular expression is an algebraic notation of a regular language. It is a string that can be used to describe and match a set of strings that conform to a certain syntactic rule, that is, a regular language. Regular expressions are widely used in tools that use pattern matching methods, such as text editors, to describe the pattern string to be found. Since regular expressions and finite state automata (FSA) are both described in regular languages, when computers use regular expressions to match strings, they usually convert them into corresponding FSAs first, and then use the FSAs for matching.

[0058] FSA is defined as a 5-tuple A = (Q, Σ, δ, q0, F), where: Q is a non-empty finite set of states; Σ is a non-empty finite set of characters, usually called the input alphabet; δ is the transition function Q×Σ * →Q; q0∈Q is the initial state; is the set of received states. According to the number of states returned by the transition function δ of FSA, FSA is divided into deterministic finite automata (DFA) and non-deterministic finite automata (NFA). The transition function of DFA only returns a single state, while NFA returns a set of states. In addition, DFA has a fast detection speed, but a large memory overhead; NFA has a small memory overhead, but a slow detection speed. The best practice is usually to convert NFA to DFA and use DFA as the preferred underlying engine.

[0059] Take DFA as an example. When matching an input string, DFA starts from the start state, reads the string to be matched character by character, and obtains the next state according to the given transition function until all characters are checked. During the matching process, if a state obtained belongs to the receiving state of F, it means that the DFA matches a pattern. Figure 3 The DFA constructed using the regular expression "ab+cd|bc+d" shown in the figure has an initial state of 0, a receiving state set of 6, an input alphabet of {a, b, c, d}, and a state set of {0, 1, 2, 3, 4, 5, 6}. When the input character sequence is "abccd", the DFA will transition along the state "012356" and obtain the receiving state 6, indicating that a pattern that meets the regular expression "ab+cd|bc+d" is found in the input character sequence "abccd".

[0060] (3) Technical solution

[0061] Based on the above basic concepts and technologies, the basic theory on which the method proposed in the present invention relies and the technical solution of the present invention are introduced below. The reason why data can be compressed is that there is a lot of redundancy in the data itself. Therefore, the compressed data contains information that can restore the original data. At present, the relevant methods mainly use this information to skip the detection of redundant data in the compressed data as much as possible, while reducing the additional overhead generated by skipping the detection, so as to achieve the effect of accelerating the detection.

[0062] Finite state automata are context-free grammars with a context-free property, that is, the next transition state is only related to the active state and the next input character, and is independent of any other information. This property can show that starting from two equal states, if the same input character sequence is matched, the resulting automaton state will also be the same. Based on this, previous methods use equal automaton states to design accelerated detection methods, which can avoid scanning more redundant data.

[0063] According to the state equivalence in the finite automaton theory, the present invention finds that if any identical character sequence is matched starting from two equivalent states, the resulting automaton states will also be equivalent, and state equivalence does not require the states to be equal. Figure 3 Taking the DFA shown in the figure as an example, there are three state sets {0, 1}, {2, 4}, {3, 5}, and the states in the sets are not equal, but they are equivalent. Based on this, the present invention proposes a more complete basic theory, and designs a detection method using equivalent automaton states, which can skip scanning more redundant data.

[0064] like Figure 4 As shown, in the data row a n The following character sequence w0w1…w n and z0z1…z n The character sequences in the adaptive dynamic dictionary and the static dictionary are referenced respectively. In the compressed form, these data are replaced by compression codes, which are called dynamic codes and static codes; the states in the status line are the states of the automaton obtained after scanning the characters corresponding to their positions. When the characters represented by the compression code are to be scanned, for example, a n The following character sequence w0w1…w n , you can first determine the scan a n and a m The status u returned when n and u mAre the two states equivalent? If the two states are not equivalent, continue to scan the characters represented by the dynamic code in sequence to obtain pattern matching results until the dynamic code is equivalent to the state stored in the corresponding position of the dynamic dictionary it refers to, or all characters represented by the code are scanned. Without loss of generality, it is assumed that the state q0 obtained after scanning the first character w0 is equivalent to the state p0 stored in the corresponding position of the dictionary. According to the above theory, even if the subsequent character sequence w1…w is scanned character by character, n , the obtained state q1…q n will be respectively n Therefore, we can directly convert w1…w n The state p1…p n Copy to the original q1…q n The location of the redundant characters after the equivalent state in the compression encoding can be obtained without scanning each character one by one.

[0065] In addition, the existing methods are mainly designed for compression algorithms using a single adaptive dictionary, and do not consider compression algorithms based on dual dictionaries such as Brotli. To this end, the present invention pre-scans the characters in the static dictionary used by the compression algorithm in the preprocessing stage, and stores the automaton state corresponding to the static dictionary data part after scanning; then, when matching the redundant data referencing the static dictionary, the automaton state saved when scanning the static dictionary is used to achieve accelerated detection. Figure 4 For example, when scanning the character sequence z0z1…z represented by the static code n When , first determine the position and the state obtained by pre-scanning in the static dictionary, that is, k1t0…t n and q n s0…s n Whether they are equivalent, once an equivalent state is found, there is no need to scan the character sequence represented by the subsequent static encoding character by character.

[0066] A regular expression matching method for accelerating detection based on double-dictionary compressed data comprises the following steps:

[0067] Step 1: construct a regular matching engine 104 according to the regular expression rule set 107, allocate and initialize the memory space required for the auxiliary information; call the regular matching engine 104 to scan the static dictionary 106, and store the returned finite state automaton state into the state area, and finally reset the active state state of the regular matching engine 104 to the initial state;

[0068] Step 2, the preprocessing module 101 reads the input dual-dictionary compressed data 105, allocates memory space to store the decompressed data output by the preprocessing module 101, and transmits the first address of the memory space to the Rainbow detection module 102; at the same time, the metadata that can identify the original composition of the compressed data is parsed from the dual-dictionary compressed data 105 to be detected and stored in the metadata area 1032;

[0069] Step 3: The length of the uncompressed data contained in the metadata is recorded as len1, and the compressed code is recorded as<dist,len2> ; The Rainbow detection module 102 reads a metadata structure from the metadata area 1032 each time until all metadata output by the preprocessing module 101 are processed;

[0070] Step 4, the Rainbow detection module 102 reads len1 bytes of data from the decompressed data output by the preprocessing module 101, takes the active state state as input, calls the regular matching engine 104 to scan the data, and updates the state; the process saves the state obtained by scanning each data into the state area 1031 of the auxiliary information storage module 103, and checks whether each state is the receiving state of the automaton, and takes the receiving state and the corresponding character position as the matched pattern information, and outputs it as the detection result 108;

[0071] Step 5: Rainbow detection module 102 detects the compressed<dist,len2> Locate the data area corresponding to the compression code in the dynamic dictionary or static dictionary used by the compression algorithm, and locate the area corresponding to the code in the status area 1031 and the reference area of ​​the code according to the information of the compression code, record the position of the previous character of the code area in the status area 1031 as curPos, and record the position of the previous character of the reference area as refPos;

[0072] Step 6. Check whether the state saved at the refPos position is equivalent to the active state:

[0073] If they are equivalent, copy the state in the area referenced by the code to the current area in the state area 1031, check whether there is a receiving state in the receiving copy state, if so, output the mode information, and after completion, jump to step 3 to read the next metadata structure; otherwise, jump to step 7;

[0074] Step 7, call the regular matching engine 104 to scan the character at the curPos position in the code, update the state and synchronously write it into the state area 1031, and then move refPos and curPos backward by one character respectively; if curPos is not the end of the code, jump to step 6 and continue the next state equivalence judgment; otherwise jump to step 3 and read the next metadata structure.

[0075] Based on the above basic idea, the present invention proposes a regular expression matching method for accelerating detection based on double-dictionary compressed data, and its functional framework is as follows: Figure 5 As shown, it is based on a detection system, which mainly includes a preprocessing module 101, a Rainbow detection module 102, an auxiliary information storage module 103 and a regular matching engine 104.

[0076] The method constructs a regular matching engine 104 according to regular expression rules 107, and parses input dual-dictionary compressed data 105 through a preprocessing module 101, and then implements accelerated detection through a Rainbow detection algorithm, and outputs a detection result 108.

[0077] The preprocessing module 101 parses the input dual-dictionary compressed data 105 in conjunction with the static dictionary 106, and directly hands the output decompressed data to the Rainbow detection algorithm; the preprocessing module 101 also outputs metadata to assist the Rainbow detection algorithm in accelerating detection. Algorithm 1 uses Brotli as an example to explain in detail the processing process of the preprocessing module 101 except for scanning the static dictionary. The preprocessing algorithm includes the following steps: First, parse the Brotli compressed data to obtain all compressed data blocks Block (as shown in line 3 of Algorithm 1). Each data block corresponds to a string of compressed data. For each data block, the information of three fields can be parsed, namely insert-copy-length, literal, and distance. Based on the information of the three fields, the uncompressed data can be restored and written into text; insert-copy-length and distance are additionally recorded as metadata and written into the metadata area metadata (as shown in lines 5-11 of Algorithm 1).

[0078]

[0079]

[0080] The Rainbow detection algorithm identifies the decompressed data after preprocessing and parsing based on the metadata, and distinguishes the uncompressed data in the original state and the data represented by the compressed code. Then, the regular matching engine 104 is called to match the uncompressed data character by character, and combined with the automaton state saved by the auxiliary information, the Rainbow algorithm is used to skip the detection of the data represented by the compressed code. Algorithm 2 uses Brotli as an example to explain the processing of the Rainbow detection algorithm in detail, including the following steps:

[0081] Read the decompressed data and metadata information, and perform different operations according to different types:

[0082] If it is a common character, directly call the automaton to scan character by character (line 5, 9-10 of the algorithm);

[0083] If it is encoded data, first determine whether the reference string of the current redundant data is taken from a dynamic dictionary or a static dictionary based on the length of dist: if it is a static dictionary, calculate the previous subscript refPos of the reference string in the static dictionary based on dist; if it is a dynamic dictionary, directly calculate the previous position refPos of the corresponding position of the reference string in the dynamic dictionary (algorithm 12-13 lines).

[0084] When processing the encoded data, first check whether the current activation state (the state corresponding to curPos) and the state of the corresponding position of refPos are equivalent (line 16 of the algorithm). If they are equivalent, copy the state directly from the reference area, record the matching results (lines 17-19 of the algorithm), and update the current activation state. If they are not equivalent, directly call the automaton scan (line 23 of the algorithm). After updating the current activation state, continue to compare with the corresponding state of the reference area refPos until the encoded data processing is completed or an equivalent state is found to directly skip scanning subsequent data.

[0085]

[0086] (4) Method example

[0087] In order to more intuitively illustrate the processing of the Rainbow detection algorithm proposed in the present invention, the present invention uses Figure 3 The DFA detection shown Figure 2 The sample data shown in the figure, the detection process is as follows Figure 6 shown. Figure 6 The index in it is the index position of each character in the decompressed character sequence; the input character is the character to be detected; the metadata information and type give the detailed information of the compression encoding corresponding to the character sequence, including uncompressed common characters, static dictionary, dynamic encoding and static encoding; each subsequent line lists the automaton state information saved by the corresponding process.

[0088] Initialization (process 0): At this point, the preprocessing process has been completed, the decompressed character sequence and metadata information have been obtained, and the operation of pre-scanning the static dictionary (index [0,4]) has also been completed.

[0089] Process 1: The Rainbow detection algorithm starts scanning the input data from the initial state 0, that is, starting from the index 5. First, the metadata information (7,5,6) is read to obtain an uncompressed character sequence and a compressed encoded information, that is, the index [5,11] interval is a common character with a length of 7 bytes; the index [12,16] is a character sequence with a length of 5 bytes dynamically encoded, and the character sequence is 6 bytes away from the character in the reference dictionary, that is, the area where the index [6,10] is located. Therefore, the Rainbow detection algorithm directly scans the common characters in the index [5,11] interval and obtains the state "0455561", and then stores the obtained state in the state area in the auxiliary information. In this process, a receiving state 6 (the second 6 in the corresponding line of process 1) is found, indicating that a pattern is matched, and the algorithm records the matched pattern information. After the process is completed, the active state of the automaton is 1.

[0090] Process 2-1: The Rainbow detection algorithm starts processing the character sequence represented by the dynamic encoding at index [12, 16], first checking whether the active state 1 and the state 0 at index 5 are equivalent (the state indicated by the arrow).

[0091] Process 2-2: The Rainbow detection algorithm finds that state 0 and state 1 are equivalent states. There is no need to scan the character sequence represented by the dynamic code character by character. The state at index [6,11] is directly copied to index [12,16] (the state circled by the dotted box). The copying process finds the receiving state 6, records the matched pattern information, and sets the active state of the automaton to the last one in the copied state, that is, state 6.

[0092] Process 3-1: The Rainbow detection algorithm reads the second metadata information (0,4,16). The data represented by this metadata contains only a static code and no ordinary characters. The index [17,20] is a 4-byte character sequence represented by the static code. The character sequence it refers to is in the static dictionary, and the offset distance between the two is 16 bytes. Therefore, the algorithm checks whether the active state 6 and the state 0 at index 0 are equivalent.

[0093] Process 3-2: Since state 0 and state 6 are not equivalent, the algorithm needs to call DFA to scan the character at index 17, obtain the active state 1, and continue to compare whether the active state 1 is equivalent to the state at index 1.

[0094] Process 3-3, at this time the two states are equal and must be equivalent, so the algorithm directly copies the state at index [2,4] to index [18,20]. During the copying process, the receiving state 6 is found, the matched pattern information is recorded, and the active state is set to 6.

[0095] At this point, the Rainbow detection algorithm completes the detection process of all input characters. The total number of decompressed characters is 16. 8 characters are skipped and 3 patterns are finally matched.

[0096] (5) Performance evaluation

[0097] In order to illustrate the actual effect, the present invention selects real network compression data and regular expression rule sets to evaluate the performance of the Rainbow detection algorithm. Among them, the data set is the web page data obtained by the crawler program after searching keywords using Google. These web pages all use Brotli as their compression code. The comparison method decompresses the data and then re-compresses it using Gzip. The characteristics of the data set are shown in Table 2. The regular expression rule sets are the three rule sets of Snort24, Snort31 and Snort34 used in references [3] and [4].

[0098] Table 2 Characteristics of the collected dataset

[0099] Number of pages Decompressed size Brotli compressed size Gzip compressed size 873 218.59MB 52.15MB 62.22MB

[0100] The embodiments of the present invention are evaluated on a Xeon 4214R and 128GB RAM (DDR4 3200MHz) platform and compared with a baseline method (Baseline) that matches the data after decompression, and the best existing method (Twins) that can only accelerate the detection of single-dictionary compressed data. The evaluation process uses the time consumed to detect the same volume of uncompressed data as the evaluation indicator. The detection time of the three methods is as follows: Figure 7 As shown in the figure, it can be seen that the detection time of the present invention is significantly shorter than that of the baseline method and the Twins method.

[0101] The above content is a further detailed description of the present invention in combination with a specific preferred embodiment. It cannot be determined that the specific embodiments of the present invention are limited to this. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as belonging to the present invention and the scope of patent protection determined by the submitted claims.

Claims

1. A regular expression matching method for accelerating detection based on double-dictionary compressed data, characterized in that: The following steps are involved: Step 1, construct a regular matching engine (104); call the regular matching engine (104) to scan the static dictionary (106), store the returned finite state automaton state into the state area, and reset the active state state of the regular matching engine (104) to the initial state; Step 2: pre-process the double-dictionary compressed data to obtain decompressed data; at the same time, parse metadata from the double-dictionary compressed data to be detected and store it, wherein the metadata includes the length of the uncompressed data, which is recorded as len1, the compression code, which is recorded as<dist,len2> ; Step 3: Read a metadata structure; Step 4: read len1 bytes of data from the decompressed data, use the active state state as input, call the regular matching engine (104) to scan, and update the state; this process saves the state obtained by scanning each data in the state area (1031), and checks whether each state is the receiving state of the automaton, and outputs the receiving state and the corresponding character position as the matched pattern information as the detection result (108); Step 5: Compression encoding according to metadata<dist,len2> Locate the data area corresponding to the compression code in the dynamic dictionary or static dictionary used by the compression algorithm, and locate the area corresponding to the compression code in the status area (1031) and the reference area of ​​the compression code according to the information of the compression code, record the position of the previous character of the area corresponding to the compression code in the status area (1031) as curPos, and record the position of the previous character of the reference area as refPos; Step 6. Check whether the state saved at the refPos position is equivalent to the active state: If they are equivalent, copy the state in the area referenced by the compression code to the current area, check whether the received state exists in the received copy state, if so, output the mode information, and then jump to step 3 to read the next metadata structure; otherwise, jump to step 7; Step 7: Call the regular expression matching engine (104) to scan the character at the curPos position in the code, update the state and write it synchronously to the state area (1031), and then move refPos and curPos backward by one character respectively; if curPos is not the end of the code, jump to step 6; otherwise, jump to step 3.

2. According to claim 1, a regular expression matching method for accelerating detection based on double-dictionary compressed data is characterized in that: In step 1, a regular matching engine (104) is constructed according to a regular expression rule set (107).

3. The method for accelerating detection of regular expression matching based on double-dictionary compressed data according to claim 1, characterized in that: In step 2, preprocessing the dual-dictionary compressed data includes the following steps: Parse the double-dictionary compressed data to obtain all compressed data blocks. Each data block corresponds to a string of compressed data. For each data block, parse the information of three fields, namely insert-copy-length, literal, and distance. Based on the information of the three fields, restore the uncompressed data and write it into text. Additionally, record insert-copy-length and distance as metadata for storage.

4. The method for accelerating detection of regular expression matching based on double-dictionary compressed data according to claim 1, characterized in that: The step 5 comprises the following steps: Read the decompressed data and metadata to determine the data type: If it is a common character, directly call the automaton to scan it character by character; If it is coded data, first determine whether the reference string of the current redundant data is taken from the dynamic dictionary or the static dictionary according to dist: If it is a static dictionary, calculate the previous position refPos of the reference string subscript in the static dictionary according to dist; If it is a dynamic dictionary, calculate the previous position refPos of the corresponding position of the reference string in the dynamic dictionary; When processing encoded data, first check whether the current activation state and the state at the corresponding position of refPos are equivalent: If they are equivalent, copy the state directly from the reference area, record the matching result, and update the current activation state; If they are not equivalent, call the automaton scan, update the current activation state, and continue to compare with the corresponding state of the reference area until the encoding data processing is completed or it is equivalent and skipped directly.

5. A regular expression matching device for accelerating detection based on double-dictionary compressed data, used to implement the method described in any one of claims 1 to 4, characterized in that: The invention comprises a preprocessing module (101), a Rainbow detection module (102), an auxiliary information storage module (103) and a regular matching engine (104); the regular matching engine (104) is constructed by a regular expression rule (107); the preprocessing module (101) is used to parse the double dictionary compressed data (105) to be detected; and the Rainbow detection module (102) is used to implement detection and output the detection result.

6. The device for accelerating detection of regular expression matching based on double-dictionary compressed data according to claim 5, characterized in that: The dual-dictionary compressed data (105) to be detected is generated by a compression algorithm using an adaptive dynamic dictionary and a static dictionary (106); the pre-processing module (101) parses the input dual-dictionary compressed data (105) in combination with the static dictionary (106), outputs decompressed data, and sends the output decompressed data to a Rainbow detection module (102); the pre-processing module (101) also outputs metadata to an auxiliary information storage module.

7. The device for accelerating detection of regular expression matching based on double-dictionary compressed data according to claim 5, characterized in that: The regular expression matching engine (104) is implemented using a finite state automaton-based implementation.

8. The device for accelerating detection of regular expression matching based on double-dictionary compressed data according to claim 5, characterized in that: The auxiliary information storage module (103) is used to store auxiliary information, including a state area (1031) and a metadata area (1032). The auxiliary information includes a finite state automaton state returned by matching a static dictionary (106) and decompressed data using a regular matching engine (104), and metadata capable of identifying the original composition of compressed data, each metadata including the length of uncompressed data and compression coding.

9. A regular expression matching device for accelerating detection based on double-dictionary compressed data according to claim 8, characterized in that: The Rainbow detection module (102) identifies the decompressed data parsed by the preprocessing module (101) based on the metadata stored in the metadata area (1032), and distinguishes between the uncompressed data in the original state and the data represented by the compressed code; then calls the regular matching engine (104) to match the uncompressed data character by character, and uses the Rainbow detection algorithm to skip the detection of the data represented by the compressed code in combination with the finite state automaton state stored in the auxiliary information.

Citation Information

Patent Citations

  • Multi-string matching method for compressed traffic

    CN107277109B

  • Pairsmethod for accelerating regular expression matching of compressed traffic

    CN108563795A

  • Twins method for accelerating regular expression matching of compression flow

    CN108573069A

  • Compressed flow pattern matching engine and pattern matching method based on FPGA platform

    CN110865970A

  • Multi-pattern matching in compressed communication traffic

    US8458354B2