A ReDoS vulnerability detection method and system based on dynamic and static combination
Through a ReDoS vulnerability detection method based on the combination of dynamic and static methods, and utilizing fine-grained representation structure and anti-interference attack string generation, the accuracy and recall rate problems of ReDoS vulnerability detection in the existing technology are solved, and efficient vulnerability detection and repair assistance are achieved.
Patent Information
- Application Number
- CN202310562533.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-18
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2043-05-18
AI Technical Summary
Existing ReDoS vulnerability detection technologies suffer from high missed detection and false positive rates, mainly due to incomplete ReDoS vulnerability model construction and attack string generation that does not consider the interference between regular expression substructures, resulting in inaccurate detection results.
A ReDoS vulnerability detection method based on the combination of dynamic and static methods is adopted. By analyzing the semantic features of pathological LS, a fine-grained representation structure is constructed, including prefixes, infixes, suffixes and interference items, to generate interference-resistant attack strings. Static detection and dynamic verification are then performed to achieve detection with high accuracy and high recall rate.
It achieves ReDoS vulnerability detection with high accuracy and high recall rate, which can assist software developers in discovering potential vulnerabilities early, improve the security and reliability of software systems, and provide a basis for subsequent vulnerability repair.
Smart Images

Figure CN116614270B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security, and in particular to a ReDoS vulnerability detection method and system based on a combination of dynamic and static methods. Background Art
[0002] Regular expressions, as a core component and foundational technology in software and system development, are seeing their application areas and scale expand, but this also presents security challenges. Regular expression denial of service (ReDoS) is an algorithmic complexity attack. For vulnerable regular expressions, attackers can carefully construct strings to trigger the superlinear time complexity of the regular expression matching algorithm, causing a denial of service attack and impacting system availability. This vulnerability is characterized by low attack cost, simple exploitation, and high defense.
[0003] In recent years, various ReDoS vulnerability detection technologies have been proposed at home and abroad. Among them, hybrid detection technology that combines the advantages of static detection technology and dynamic detection technology has a good performance. However, the current hybrid detection technology still has the following two significant defects: (1) The construction of ReDoS vulnerability patterns is not perfect, resulting in a high false negative rate in existing hybrid detection technologies. The existing vulnerability modeling method is summarized based on the structural characteristics of the vulnerability loop sub-regular expression (i.e., the sub-regular expression with quantifiers, hereinafter referred to as pathological LS), and lacks a global understanding of LS features. For example, for LS^(r0|r1)+, the existing vulnerability pattern only analyzes the structural characteristics of r0 and r1, while ignoring the various possible permutations and combinations between them, such as r0r0, r1r1, r0r1, r1r0, etc. This ReDoS vulnerability pattern modeling method based on local understanding will lead to inaccurate detection results; (2) The attack string generation fails to consider the interference between the regular expression substructures, resulting in the inability to generate effective attack strings, which in turn leads to high false positive and false negative rates. Interference refers to the fact that the actual matching behavior of the attack string does not meet the expectations of the attack string generator. The existing attack string generator ignores the interference of other branches of the regular expression on the vulnerability branch. The prefix of the generated attack string may be accidentally received by other branches, resulting in the inability to trigger the pathological sub-regular expression of the regular expression, and thus the attack effect disappears. Summary of the Invention
[0004] The purpose of this invention is to propose a ReDoS vulnerability detection method and system based on the combination of dynamic and static methods. By analyzing the semantic characteristics of pathological LS, the ReDoS vulnerability is modeled and the structural characteristics of the attack string are accurately characterized. Through static detection and dynamic verification of ReDoS vulnerabilities, high accuracy and high recall rate of ReDoS vulnerability detection are achieved.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] A ReDoS vulnerability detection method based on a combination of dynamic and static methods includes the following steps:
[0007] Model the ReDoS vulnerability based on the semantic characteristics of pathological LS and predefine several ReDoS vulnerability patterns;
[0008] A fine-grained regular expression representation structure for ReDoS vulnerabilities is created. This structure consists of a three-part structure consisting of three sub-regular expressions: prefix, infix, and suffix, and two interference terms. The infix sub-regular expression is a pathological sub-regular expression, and the interference term is an interference sub-regular expression that potentially affects the generation of attack strings. Based on this fine-grained representation structure, all possible (or potential) interference types are enumerated and attack string generation constraints that eliminate the corresponding interference are constructed.
[0009] Convert the regular expression to be tested into a unified abstract syntax tree, determine all candidate pathological LSs in the regular expression based on the fine-grained representation structure, and determine which ReDoS vulnerability pattern the candidate pathological LSs conform to, thereby extracting candidate vulnerability information;
[0010] For each candidate vulnerability information, corresponding attack string generation constraints are added according to its interference type, and an attack string that meets the constraints is generated; the regular expression to be tested is dynamically verified. If the number of matching steps of the attack string exceeds a threshold, a ReDoS vulnerability is determined to exist.
[0011] Preferably, the ReDoS vulnerability is modeled based on the semantic feature that there is a common matching string between all sub-regular expressions in the pathological LS expansion process.
[0012] Preferably, three ReDoS vulnerability modes are predefined:
[0013] EOLS, exponential vulnerability, a pathological LS exists in a regular expression that contains two different sub-regular expressions, and the two sub-regular expressions have a common matching string;
[0014] POLS, polynomial-level vulnerability, a regular expression starts with a sub-regular expression r1 without a line-start anchor, when r1 = αβ, where α and β are sub-regular expressions, and β is a pathological LS, and meets one of the following two conditions: (1) α is empty; (2)
[0015] PTLS, a polynomial-level vulnerability, for the regular expression αβγ, where α and γ are pathological LS, meets one of the following two conditions: (1) β is empty and α and γ have a common matching string; (2) α, γ, αβ or βγ have a common matching string.
[0016] Preferably, the fine-grained representation structure is represented as in represents a prefix subregular expression, represents a pathological subregular expression, represents the suffix sub-regular expression, θ1 and θ2 are two interference terms.
[0017] Preferably, the step of determining which ReDoS vulnerability pattern the candidate pathological LS conforms to includes:
[0018] First, extract all candidate pathological LS of the regular expression and regard each LS as a pathological sub-regular expression. And extract the corresponding prefix and suffix sub-regular expressions of each LS and two interference terms θ1, θ2;
[0019] Secondly, obtain the set of sub-regular expressions of each LS, i.e., k, that contains the common matching string. If the set is not empty, then k is considered to conform to EOLS;
[0020] Again, judge Is it nullable? If it is nullable, it satisfies the POLS condition (1), then Get the relevant sub-regular expression set in , if the set is not empty, then each LS, that is, k, is considered to comply with POLS, otherwise, determine whether the condition (2) of POLS is met, if so, then k is considered to comply with POLS;
[0021] Finally, for every two different and non-nested LS, namely k1 and k2, first determine whether the two are adjacent. If they are adjacent, the PTLS condition (1) is satisfied, and then the relevant sub-regular expression set is obtained from k1 and k2; if the set is not empty, k1 and k2 are considered to comply with PTLS, otherwise the connection of k1, k2 and the sub-regular expression ρ between them is taken as A sub-regular expression set containing a common matching string is obtained therefrom. If the set is not empty, it is considered that k1 and k2 comply with PTLS.
[0022] Preferably, the candidate vulnerability information includes a fine-grained representation structure, a vulnerability pattern type, and a non-empty sub-regular expression set.
[0023] A ReDoS vulnerability detection system based on a combination of dynamic and static methods is implemented based on the above method, including:
[0024] A ReDoS vulnerability static detection module includes a regular expression parser and a ReDoS static analyzer. The regular expression parser converts the regular expression to be tested into a unified abstract syntax tree. The ReDoS static analyzer determines all candidate pathological LSs in the regular expression based on the fine-grained representation structure, and determines which ReDoS vulnerability pattern the candidate pathological LSs conform to, thereby extracting candidate vulnerability information.
[0025] The ReDoS vulnerability dynamic verification module includes an anti-interference attack string generator and a dynamic verifier. The anti-interference attack string generator adds corresponding attack string generation constraints to each candidate vulnerability information based on its interference type and generates an attack string that meets the constraints. The dynamic verifier dynamically verifies the regular expression to be tested. If the number of matching steps of the attack string exceeds a threshold, it is determined that a ReDoS vulnerability exists.
[0026] The advantages of the present invention over the prior art are:
[0027] 1. Semantic-based vulnerability modeling: This paper models ReDoS vulnerabilities based on the semantic characteristics of pathological LS and predefines multiple vulnerability patterns (such as EOLS, POLS, and PTLS). This allows for better description of different types of ReDoS vulnerabilities and the design of corresponding detection methods for each vulnerability pattern.
[0028] 2. Fine-grained representation structure: This invention introduces a fine-grained regular expression representation structure, including a three-part structure of prefix, infix, and suffix and interference terms. This representation structure can more comprehensively express the components of the regular expression and provide more precise detection constraints.
[0029] 3. Interference-resistant attack string generation: The present invention determines corresponding interference-eliminating attack string generation constraints for each candidate vulnerability information based on the vulnerability pattern and interference type, so as to reduce missed reports and improve detection accuracy.
[0030] 4. Unified Abstract Syntax Tree: The present invention converts the regular expression to be tested into a unified abstract syntax tree, from which candidate vulnerability information is extracted. This unified representation can simplify the analysis process and facilitate the extraction and processing of vulnerability information.
[0031] In summary, this paper predefines multiple ReDoS vulnerability patterns based on pathological LS semantic features, establishes a fine-grained representation structure for interference-resistant attack string generation, and integrates static detection and dynamic verification of ReDoS vulnerabilities to achieve high-accuracy and high-recall ReDoS vulnerability detection. It can assist software developers in discovering potential ReDoS vulnerabilities in software systems as early as possible, thereby improving the security and reliability of software systems. In addition, accurate ReDoS vulnerability detection results can provide a solid foundation for the subsequent development of automated ReDoS vulnerability remediation systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 This is a schematic diagram of a ReDoS vulnerability detection system framework based on a combination of dynamic and static methods proposed by the present invention. DETAILED DESCRIPTION
[0033] In order to make the various technical features and advantages or technical effects of the above technical solutions of the present invention more obvious and easy to understand, they are described in detail below with reference to the accompanying drawings.
[0034] This paper proposes a ReDoS vulnerability detection method based on a combination of dynamic and static methods. The method mainly includes four parts: ReDoS vulnerability pattern modeling, interference-resistant attack string generation modeling, ReDoS vulnerability static detection, and ReDoS vulnerability dynamic verification. Among them, ReDoS vulnerability pattern modeling based on semantic features is the basis of ReDoS vulnerability detection, and interference-resistant attack string generation modeling is the core of ReDoS vulnerability verification.
[0035] 1. ReDoS vulnerability pattern modeling based on semantic features
[0036] This paper identifies the semantic feature of common matching strings between all sub-regular expressions during the pathological LS expansion process as the root cause of the ReDoS vulnerability and uses this to model the ReDoS vulnerability. Table 1 shows the three ReDoS vulnerability patterns modeled by this paper: EPLS (exponential vulnerability with one pathological LS), POLS (polynomial vulnerability with one pathological LS), and PTLS (polynomial vulnerability with two pathological LSes).
[0037] Table 1 ReDoS vulnerability patterns
[0038]
[0039]
[0040] 2. Modeling of Anti-interference Attack String Generation
[0041] The existing vulnerability regular expression structure based on the three-segment structure (using Represents a prefix sub-regular expression, Represents a pathological sub-regular expression, Represents a suffix regular expression; attack string w = xy n z, where n represents the number of times y is repeated, ) has not yet considered the impact of interference on attack string generation. In order to systematically analyze the interference between sub-regular expressions, this paper proposes a fine-grained representation structure of regular expressions. in This is the original three-segment structure, where θ1 and θ2 are interfering subregular expressions that may affect attack string generation. Based on this representation structure, the present invention summarizes five types of interference situations, as shown in Table 2, and the attack string generation constraints required to eliminate interference.
[0042] Table 2 Five types of interference
[0043]
[0044] Based on the above method, the present invention also proposes a ReDoS vulnerability detection system based on dynamic and static combination. The system includes a ReDoS vulnerability static detection module and a ReDoS vulnerability dynamic verification module, which are used to implement the static detection and dynamic verification of ReDoS vulnerabilities in the method of the present invention. The entire ReDoS vulnerability detection framework is as follows: Figure 1 Since the method and system proposed in the present invention have the same content in the static detection and dynamic verification of ReDoS vulnerabilities, they are described together as follows.
[0045] 3. Static detection of ReDoS vulnerabilities
[0046] ReDoS vulnerability static detection searches for predefined vulnerability patterns in the abstract syntax tree of a regular expression and reports the LS sets that match the vulnerability patterns. This is implemented by the ReDoS vulnerability static detection module, which mainly consists of two parts: a regular expression parser and a ReDoS static analyzer.
[0047] The regular expression parser first builds a unified abstract syntax tree for the regular expression features supported by most programming languages (including extended features such as lookaround assertions, backreferences, and named capture groups), and converts regular expressions with different syntaxes into the above abstract syntax tree.
[0048] ReDoS static analyzer can circle all candidate pathological sub-regular expressions in a regular expression. In the fine-grained representation structure proposed in this invention, the infix sub-regular expression is a pathological subregular expression. Each LS that satisfies the EOLS and POLS vulnerability patterns is considered an infix subregular expression. Furthermore, for the PTLS vulnerability pattern, the infix subregular expression consists of two LSs. Specifically, the processing is as follows:
[0049] a) The ReDoS static analyzer first extracts all LS K in the regular expression, i.e., candidate pathological LSs, and then considers each LS k∈K in the regular expression as a pathological sub-regular expression. And extract the prefix and suffix regular expressions accordingly and interference terms θ1 and θ2.
[0050] b) After that, the analyzer will detect Whether the vulnerability pattern EOLS described in Table 1 is satisfied, and the sub-regular expression sets IR containing the common matching string are obtained EOLS If the set is not empty, it is considered that k meets the vulnerability pattern EOLS and the candidate vulnerability information (fine-grained representation structure quintuple μ, vulnerability pattern type k, IR EOLS ) is added to the final candidate vulnerability set S.
[0051] c) Next, the analyzer will search for POLS vulnerability patterns. If If it is nullable, then the triggering condition (1) of the POLS vulnerability pattern in Table 1 is met, then Get the relevant sub-regular expression set IR POLS If the set is not empty, it is considered that k meets the vulnerability pattern POLS, and the candidate vulnerability information (fine-grained representation structure quintuple μ, vulnerability pattern type τ, IR POLS ) is added to S, otherwise the analyzer will judge Is it true? If so, the handling method is the same as above.
[0052] d) Finally, the analyzer searches for PTLS vulnerability patterns. For every two different and non-nested LSk1 and k2 in K, the analyzer first determines whether the two are adjacent. If they are adjacent, the triggering condition (1) of the PTLS vulnerability pattern in Table 1 is satisfied. Then, the relevant sub-regular expression set IR is obtained from k1 and k2. PTLS If the set is not empty, k1 and k2 are considered to conform to the vulnerability pattern PTLS, and the candidate vulnerability information (fine-grained representation structure quintuple μ, vulnerability pattern type τ, IR PTLS) is added to S, otherwise the analyzer takes the connection of k1, k2 and the sub-regular expression ρ between them as the infix sub-regular expression, and obtains the sub-regular expression set IR that contains the common matching string from it PTLS If the set is not empty, k1 and k2 are considered to conform to the vulnerability pattern PTLS, and the candidate vulnerability information (fine-grained representation structure quintuple μ, vulnerability pattern type τ, IR PTLS ) is added to S.
[0053] 4. Dynamic Verification of ReDoS Vulnerabilities
[0054] ReDoS vulnerability dynamic verification first generates an anti-interference attack string for each candidate vulnerability output by the ReDoS vulnerability detection component. It then filters out the actual vulnerability information through a dynamic verifier. This is implemented by the ReDoS vulnerability dynamic verification module. The ReDoS vulnerability dynamic verification module mainly consists of two parts: an anti-interference attack string generator and a dynamic verifier.
[0055] For each element in S output by the ReDoS static analyzer, the generator first extracts each part of the fine-grained representation structure from the five-tuple μ and creates an attack string set W and a constraint set CS. Initialize W to be empty and CS to be Then, the generator adds corresponding constraints according to the interference types shown in Table 2 to prevent the attack string generation from being affected by interference, for example, when (θ1≠ε and When , the generator will constrain Add to CS. Then, for the set IR (i.e. IR EOLS IR POLS and IR PTLS ), the generator creates a copy CS' of CS and adds y∈L(ir) to CS', and constrains |y according to the severity of the vulnerability pattern type τ. n | for N E or N P , where N E is the predefined number of repetitions of the exponential ReDoS vulnerability, N P It refers to the predefined number of repetitions of the polynomial-level ReDoS vulnerability. Finally, the generator will generate an attack string w=xy that satisfies the constraint. n z, add it to W.
[0056] The dynamic verifier is used to detect whether the vulnerability reported by the ReDoS static analyzer is real. For each attack string w in W, if the number of matching steps required for the regular expression to match w exceeds 10 5 , then the vulnerability is considered to be real.
[0057] Experimental test:
[0058] In order to evaluate the detection performance of the technical solution of the present invention (hereinafter referred to as RENGAR), this experiment conducted a comparative experiment with 9 currently most advanced ReDoS vulnerability detection tools on a data set of 304,701 regular expressions. The experimental results are shown in Table 3. RENGAR is significantly better than the comparison tools in all evaluation indicators. The number of regular expressions vulnerable to ReDoS attacks detected by RENGAR is 4 to 244 times that of other tools. Specifically, RENGAR detected 161,592 regular expressions vulnerable to ReDoS attacks, which is 132,420 more than the second-best ReDoSHunter. It can be seen that RENGAR has the best ReDoS vulnerability detection performance compared with these tools.
[0059] Table 3 Experimental evaluation results
[0060] Comparison Tools True positive False positive False negative Accuracy Recall RXXR2 1096 58 160496 94.97% 0.67% NFAA 2860 616 158732 82.27% 1.76% Rexploiter 5473 2786 156119 66.26% 3.38% Safe-regex 8297 17255 153295 32.47% 5.13% Regexploit 659 113 160933 85.36% 0.40% ResCue 756 0 160836 100% 0.46% Regulator 15017 0 146575 100% 9.29% Revealer 3352 0 158240 100% 2.07% ReDoSHunter 29172 0 132420 100% 18.05% Rengar 161592 0 0 100% 100%
[0061] Although the present invention has been disclosed as above by way of embodiments, they are not intended to limit the present invention. Any appropriate modification or equivalent substitution of the technical solution of the present invention by a person skilled in the art should be included in the protection scope of the present invention. The protection scope of the present invention shall be based on that defined in the claims.
Claims
1. A ReDoS vulnerability detection method based on a combination of dynamic and static methods, characterized in that: The following steps are involved: ReDoS vulnerabilities are modeled based on the semantic features of pathological LS, and several ReDoS vulnerability patterns are predefined. The pathological LS is a sub-regular expression with quantifiers. A fine-grained regular expression representation structure for ReDoS vulnerabilities was created. This structure consists of a three-part structure consisting of three sub-regular expressions: prefix, infix, and suffix, and two interference terms. The infix sub-regular expression is a pathological sub-regular expression, and the interference term is an interference sub-regular expression that potentially affects the generation of attack strings. Based on this fine-grained representation structure, all potential interference types are enumerated and attack string generation constraints that eliminate the corresponding interference are constructed. Convert the regular expression to be tested into a unified abstract syntax tree, determine all candidate pathological LSs in the regular expression based on the fine-grained representation structure, and determine which ReDoS vulnerability pattern the candidate pathological LSs conform to, thereby extracting candidate vulnerability information; For each candidate vulnerability information, corresponding attack string generation constraints are added according to its interference type, and an attack string that meets the constraints is generated; the regular expression to be tested is dynamically verified. If the number of matching steps of the attack string exceeds a threshold, a ReDoS vulnerability is determined to exist.
2. The method according to claim 1, wherein The ReDoS vulnerability is modeled based on the semantic feature that there are common matching strings between all sub-regular expressions in the pathological LS expansion process.
3. The method according to claim 1 or 2, wherein: Three predefined ReDoS vulnerability modes: EOLS, exponential vulnerability, a pathological LS exists in a regular expression that contains two different sub-regular expressions, and the two sub-regular expressions have a common matching string; POLS, polynomial-level vulnerability, a regular expression starts with a sub-regular expression r1 without a line-start anchor, when r1 = αβ, where α and β are sub-regular expressions, and β is a pathological LS, and meets one of the following two conditions: (1) α is empty; (2) L() represents the language accepted by the regular expression, and ∑ represents the alphabet; PTLS, a polynomial-level vulnerability, for the regular expression αβγ, where α and γ are pathological LS, meets one of the following two conditions: (1) β is empty and α and γ have a common matching string; (2) α, γ, αβ or βγ have a common matching string.
4. The method according to claim 3, wherein The fine-grained representation structure is expressed as in represents a prefix subregular expression, represents a pathological subregular expression, represents the suffix sub-regular expression, θ1 and θ2 are two interference terms.
5. The method according to claim 4, wherein The attack string is represented as w=xy n z, where n represents the number of times y is repeated, and the attack string L() represents the language accepted by the regular expression.
6. The method according to claim 5, wherein The steps to determine which ReDoS vulnerability pattern a candidate pathological LS matches include: First, extract all candidate pathological LSs of the regular expression, regard each LS as a pathological sub-regular expression φ2, and extract the corresponding prefix and suffix sub-regular expressions of each LS and two interference terms θ1, θ2; Secondly, obtain the set of sub-regular expressions of each LS, i.e., k, that contains the common matching string. If the set is not empty, then k is considered to conform to EOLS; Again, judge Is it nullable? If it is nullable, it satisfies the POLS condition (1), then Get the relevant sub-regular expression set in , if the set is not empty, then each LS, that is, k, is considered to comply with POLS, otherwise, determine whether the condition (2) of POLS is met, if so, then k is considered to comply with POLS; Finally, for every two different and non-nested LS, namely k1 and k2, first determine whether the two are adjacent. If they are adjacent, the PTLS condition (1) is satisfied, and then the relevant sub-regular expression set is obtained from k1 and k2; if the set is not empty, k1 and k2 are considered to comply with PTLS, otherwise the connection of k1, k2 and the sub-regular expression ρ between them is taken as A sub-regular expression set containing a common matching string is obtained therefrom. If the set is not empty, it is considered that k1 and k2 comply with PTLS.
7. The method according to claim 6, wherein The candidate vulnerability information includes a fine-grained representation structure, a vulnerability pattern type, and a non-empty sub-regular expression set.
8. The method according to claim 5, wherein The interference types include the following five: I.When When θ1 interferes with II. When When θ1 interferes with III. When When θ1 interferes with IV. When hour, Interference V. When hour, Interference Where L() represents the language accepted by the regular expression and ∑ represents the alphabet.
9. The method according to claim 8, wherein Add corresponding attack string generation constraints based on the interference type, including: For interference type I, add conditions For interference type II, add conditions For interference type III, add conditions For interference type IV, add conditions For interference type V, add conditions 10. A ReDoS vulnerability detection system based on a combination of dynamic and static methods, implemented based on the method described in any one of claims 1 to 9, characterized in that: include: A ReDoS vulnerability static detection module includes a regular expression parser and a ReDoS static analyzer. The regular expression parser converts the regular expression to be tested into a unified abstract syntax tree. The ReDoS static analyzer determines all candidate pathological LSs in the regular expression based on the fine-grained representation structure, and determines which ReDoS vulnerability pattern the candidate pathological LSs conform to, thereby extracting candidate vulnerability information. The ReDoS vulnerability dynamic verification module includes an anti-interference attack string generator and a dynamic verifier. The anti-interference attack string generator adds corresponding attack string generation constraints to each candidate vulnerability information based on its interference type and generates an attack string that meets the constraints. The dynamic verifier dynamically verifies the regular expression to be tested. If the number of matching steps of the attack string exceeds a threshold, it is determined that a ReDoS vulnerability exists.