A Secure Data Penetration Testing Method

By establishing a race condition vulnerability feature library and dynamic monitoring technology, the race condition vulnerability in multi-threaded applications is identified and verified, and the problem of difficult-to-trigger race condition is solved, efficient vulnerability detection and utilization is achieved, and system security is improved.

CN119377969BActive Publication Date: 2025-06-10JIANGXI LIANCHUANG ELECTROACOOUSTIC CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411522157.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-29
Publication Date
2025-06-10
Estimated Expiration
2044-10-29

AI Technical Summary

Technical Problem

In multithreaded applications, race conditions are a common and difficult problem, and attackers have difficulty triggering race conditions reliably, and the impact of different programming languages ​​and concurrency models on race conditions is complex.

Method used

By pre-establishing a race condition vulnerability feature library, static analysis is carried out to identify potential vulnerabilities, dynamic instrumentation monitoring threads’ access to shared resources, capture suspected race condition events, use symbolic execution technology to explore trigger conditions, and construct attack payloads to evaluate the hazards of vulnerabilities.

Benefits of technology

It realizes efficient detection, verification and utilization of race condition vulnerabilities for multi-threaded applications, improves system security evaluation capabilities, and provides common utilization primitives for different programming languages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119377969B_ABST
    Figure CN119377969B_ABST
Patent Text Reader

Abstract

The present application provides a secure data penetration testing method, including: pre-establishing a multi-threaded application race condition vulnerability feature library, performing static analysis on the source code of the target application to identify potential race condition vulnerability points and obtaining a preliminary set of vulnerability candidates; for the monitored suspected race condition events, abstracting the event context information into constraint conditions, using a constraint solver to explore the input combinations of the suspected static conditions, and determining the minimum input set that stably triggers the race condition; for the memory models and synchronization primitives of different programming languages, obtaining race condition anti-patterns and extracting general exploitation primitives covering various race condition vulnerabilities; during the penetration testing process, applying static analysis, dynamic monitoring, and symbolic execution to mine race condition vulnerabilities in the target application system, verifying the exploitability of the vulnerabilities, and assisting in completing the evaluation of the system security.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular, to a secure data penetration testing method. Background Art

[0002] In multithreaded applications, race conditions are a common and tricky problem. When multiple threads concurrently access shared resources, without proper synchronization mechanisms, it may lead to data inconsistency and abnormal program behavior. When using race conditions for attacks in penetration testing, there are challenges in reliably triggering race conditions. Due to the uncertainty of thread scheduling, race conditions have a certain degree of randomness and unpredictability. Attackers need to develop precise techniques to operate on shared resources within a specific time window to increase the probability of race conditions occurring. At the same time, the impact of different programming languages and concurrent models on race conditions is also worthy of in-depth study. For example, in languages using explicit locks, if the lock granularity is too large or the lock is used improperly, new race conditions may be introduced instead. In message-passing-based concurrent models, if the sending and receiving order of messages is not guaranteed, race conditions may also occur. Therefore, when conducting penetration testing, it is necessary to comprehensively consider the technology stack used by the target application and conduct in-depth analysis of its concurrent mechanisms to discover and exploit potential race condition vulnerabilities. Summary of the Invention

[0003] The present invention provides a secure data penetration testing method, mainly including:

[0004] Pre-establish a feature library of race condition vulnerabilities in multithreaded applications, perform static analysis on the source code of the target application to identify potential race condition vulnerability points, and obtain a preliminary set of vulnerability candidates;

[0005] According to the set of vulnerability candidates, establish and construct a multithreaded interaction scenario that can trigger race conditions, and through dynamic instrumentation, monitor the access of multiple threads to shared resources in real time during the running of the application, and capture suspected race condition events;

[0006] For the monitored suspected race condition events, abstract the event context information into constraint conditions, and use a constraint solver to explore the input combinations of the suspected race conditions to determine the minimum input set that can stably trigger race conditions;

[0007] Monitor whether race conditions are triggered. If race conditions are triggered, analyze the impact of race conditions on the data integrity and behavior correctness of the application, construct attack payloads that use race conditions to execute arbitrary code or elevate privileges, and evaluate the harmfulness of race condition vulnerabilities;

[0008] For the memory models and synchronization primitives of different programming languages, obtain race condition anti-patterns, and extract general exploitation primitives that cover various race condition vulnerabilities;

[0009] During the penetration testing process, apply static analysis, dynamic monitoring, and symbolic execution to the target application system to discover race condition vulnerabilities and verify the exploitability of the vulnerabilities, assisting in the assessment of the system's security;

[0010] Obtain typical race condition vulnerability cases of multi-threaded applications, summarize the causes, triggering conditions, exploitation methods, and repair solutions of the vulnerabilities, improve the race condition vulnerability knowledge base, and integrate the summarized race condition vulnerability knowledge and exploitation primitives into automated vulnerability discovery and penetration testing tools.

[0011] The technical solution provided by the embodiments of the present invention may include the following beneficial effects:

[0012] The present invention discloses a secure data penetration testing method. The method first establishes a race condition vulnerability feature library and identifies potential vulnerability points through static analysis. Then, it designs multi-threaded interaction scenarios, monitors shared resource access using dynamic instrumentation technology, and captures suspected race condition events. For the captured events, it uses symbolic execution technology to abstract the context information into constraint conditions and uses a constraint solver to explore the input combinations that trigger race conditions. Further analyze the impact of race conditions on the application program and construct attack payloads to evaluate the harmfulness of the vulnerabilities. The present invention also studies the concurrent patterns of different programming languages and extracts general exploitation primitives. During penetration testing, the present invention comprehensively applies static analysis, dynamic monitoring, and symbolic execution technologies to fully discover race condition vulnerabilities. By analyzing typical cases, the present invention continuously improves the vulnerability knowledge base and integrates relevant knowledge and exploitation primitives into automated tools. The present invention realizes the efficient detection, verification, and exploitation of race condition vulnerabilities in multi-threaded applications, providing strong support for improving system security. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 It is a flowchart of a secure data penetration testing method of the present invention.

[0014] Figure 2 It is a schematic diagram of a secure data penetration testing method of the present invention.

[0015] Figure 3 It is another schematic diagram of a secure data penetration testing method of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0016] To further understand the content of the present invention, the present invention will be described in detail in combination with the accompanying drawings and embodiments. The following further elaborates on the present application in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. Additionally, it should be noted that for the sake of description, only the parts related to the invention are shown in the drawings.

[0017] As Figures 1-3 , a security data penetration testing method in this embodiment may specifically include:

[0018] Step S101, pre - establish a multi - thread application race condition vulnerability feature library, perform static analysis on the source code of the target application, identify potential race condition vulnerability points, and obtain a preliminary vulnerability candidate set.

[0019] Extract key features according to the pre - established multi - thread application race condition vulnerability feature library. The key features include vulnerability patterns, code structures, and variable dependency relationships; receive the source code of the target application, perform static analysis on the source code using a matching method based on regular expressions and syntax analysis, identify potential race condition vulnerability points, and obtain a preliminary vulnerability candidate set; conduct in - depth semantic analysis on the preliminary vulnerability candidate set, perform code structure analysis through an abstract syntax tree, simulate the program execution path, extract the context information of each candidate vulnerability point, and generate vulnerability feature vectors; perform data pre - processing and feature selection on the vulnerability feature vectors, use the support vector machine algorithm to map the feature vectors to a high - dimensional feature space, construct an optimal classification hyperplane, and obtain a preliminary race condition vulnerability list.

[0020] Exemplarily, a multi-threaded application race condition vulnerability feature library is established in advance. Known race condition vulnerability samples are collected, key features are extracted, a feature representation method is designed, and a matching algorithm based on regular expressions and syntax analysis is used to perform static analysis on the source code of the target application to identify potential race condition vulnerability points and obtain a preliminary set of vulnerability candidates. According to the feature patterns defined in the pre-established vulnerability feature library, in-depth semantic analysis is performed on the preliminarily screened set of vulnerability candidates. The abstract syntax tree is used for code structure analysis, and symbolic execution technology is used to simulate the program execution path. The context information, variable dependency relationships, and control flow graph of each candidate vulnerability point are extracted. At the same time, multi-threaded synchronization primitives are identified and analyzed to generate detailed vulnerability feature vectors. Data preprocessing and feature selection are performed on the generated vulnerability feature vectors. Principal component analysis is used for feature dimensionality reduction, and then the support vector machine algorithm is used to map the feature vectors to a high-dimensional feature space, construct an optimal classification hyperplane, accurately classify and score each potential vulnerability point, and obtain a preliminary list of race condition vulnerabilities. Based on the risk score, the preliminary vulnerability list is screened to determine the high-risk vulnerabilities that need in-depth analysis. For the screened high-risk race condition vulnerabilities, code path analysis is performed. Taint analysis technology is used to track the propagation of shared variables, construct a data flow graph and a control flow graph, trace the process of variable transfer between multiple threads, locate key shared resource access points, and generate a detailed vulnerability location report. When establishing the multi-threaded application race condition vulnerability feature library in advance, 100 known race condition vulnerability samples are collected, and key features such as shared variable access patterns, thread synchronization operations, and resource release order are extracted. The feature representation method is designed using the vector space model, and each feature is represented by a binary value. A matching method based on regular expressions and syntax analysis is used to perform static analysis on the source code of the target application, and the abstract syntax tree is traversed using depth-first search to identify potential race condition vulnerability points. In-depth semantic analysis is performed on the preliminarily screened set of vulnerability candidates. The abstract syntax tree is used for code structure analysis to construct a tree structure containing node types, variable scopes, and control dependency relationships. Symbolic execution technology is used to simulate the program execution path and generate path constraint conditions. The context information of each candidate vulnerability point is extracted, including the code fragments before and after 5 lines. The variable dependency relationships are analyzed to construct a data dependency graph. A control flow graph is generated to identify key branches and loop structures. At the same time, multi-threaded synchronization primitives such as mutexes, semaphores, and condition variables are identified and analyzed. These information are integrated to generate a 200-dimensional vulnerability feature vector. Data preprocessing is performed on the generated vulnerability feature vector, including normalization and missing value processing. Principal component analysis is used for feature dimensionality reduction, retaining 95% of the variance information and reducing the feature dimension to 50 dimensions. Then, the support vector machine algorithm is used, with a radial basis kernel function, to map the feature vectors to a high-dimensional feature space and construct an optimal classification hyperplane through iterative optimization.Precisely classify and score each potential vulnerability point, with the score ranging from 0 to 100, to obtain a preliminary list of race condition vulnerabilities. Screen the preliminary vulnerability list based on the risk score, set the threshold at 80 points, and determine the high-risk vulnerabilities that require in-depth analysis. Conduct code path analysis for the screened high-risk race condition vulnerabilities, use taint analysis techniques to track the propagation of shared variables, construct forward slices and backward slices, and mark the taint sources and sinks. Construct a data flow graph to represent the variable assignment and usage relationships. Generate a control flow graph to identify conditional branches and loop structures. Trace the transfer process of variables between multiple threads and identify the inter-thread communication points. Locate the key shared resource access points, including read-write operations and synchronization operations. Finally, generate a detailed vulnerability location report, including the vulnerability type, severity, number of affected code lines, and repair suggestions.

[0021] Step S102: Based on the vulnerability candidate set, establish and construct a multi-threaded interaction scenario that can trigger a race condition, and use dynamic instrumentation to monitor the access of multiple threads to shared resources in real time during the application runtime, and capture suspected static condition events.

[0022] Generate test cases according to the multi-threaded interaction scenario, where the test cases include thread creation, resource allocation, and concurrent operation information; perform code injection on the target application, and insert monitoring probes at shared resource access points, thread synchronization operation calls, and key control flow nodes; obtain the execution status and resource access information of multiple threads through the monitoring probes, where the resource access information includes thread identifiers, operation types, access addresses, and timestamps; perform data filtering, deduplication, and classification processing on the resource access information to obtain a standardized event sequence; analyze the thread interaction patterns in the event sequence using temporal logic rules, and the analysis includes: modeling the thread behavior using a finite state machine; determining whether there are concurrent accesses that violate the synchronization constraints; if there are concurrent accesses that violate the synchronization constraints, then determine the type, occurrence location, and related thread information of the race condition.

[0023] Exemplarily, according to the shared resource access patterns in the vulnerability candidate set, a multi-threaded interaction scenario is designed and constructed. Test cases are generated through thread creation, resource allocation, and concurrent operations. In the test cases, read and write operations of multiple threads on shared variables, access to critical sections, and resource lock contention are simulated, including complex interaction situations such as concurrent read and write, read-write interleaving, and lock competition. The dynamic instrumentation technique is used to inject code into the target application. According to the pre-set instrumentation point selection criteria, monitoring probes are inserted at shared resource access points, thread synchronization operation calls, and key control flow nodes to achieve tracking of thread creation, resource allocation, variable read and write, and the use of synchronization primitives. The LLVM framework is used for code injection to ensure the flexibility and scalability of the instrumentation process. During the runtime of the application, the execution states and resource access situations of multiple threads are captured in real time through the inserted monitoring probes. Thread identifiers, operation types, access addresses, and timestamps are recorded to construct the resource access sequence and timing relationship graph among threads, and a directed graph data structure is used to represent the interaction relationship and resource access order among threads. The captured event data is preprocessed, including data filtering, deduplication, and categorization, to construct a standardized event sequence. For the processed resource access sequence, the interaction patterns among threads are analyzed using temporal logic rules, and the behavior of threads is modeled using a finite state machine to detect concurrent accesses that violate synchronization constraints. Combining with a pre-set race condition pattern library, it is judged whether there are suspected race condition events. The specific judgment criteria include time overlap of resource access, write operation conflicts of the same resource by different threads, etc. Finally, a detailed event report is generated, including the type of race condition, the occurrence location, and relevant thread information. According to the shared resource access patterns in the vulnerability candidate set, a multi-threaded interaction scenario is designed, creating 10 concurrent threads, including 5 read threads and 5 write threads, which operate on 3 shared variables respectively. 100 test cases are generated, and each test case contains 50 random read and write operations and 10 resource lock contentions. The LLVM framework is used for dynamic instrumentation, and 200 monitoring probes are inserted into the target application to cover all shared variable access points and synchronization operation calls. The probe trigger threshold is set to 1 millisecond, and thread IDs, operation types (read / write), memory addresses, and timestamps are recorded. The red-black tree data structure is used to store the collected event data, and 10,000 events can be processed per second. A directed graph is constructed to represent the interaction among threads, where nodes represent threads, edges represent resource access, and the weight is the access time interval. The depth-first search algorithm is used to traverse the graph to identify loops as potential deadlock situations. The sliding window algorithm is used in the event preprocessing stage, with the window size set to 100 ms, to merge duplicate events. The behavior of threads is modeled using a finite state machine, defining 5 states including idle, waiting, acquiring lock, accessing resource, and releasing lock. The state transitions are controlled by 20 rules, including lock acquisition timeout, resource access conflict, etc.The race condition judgment criterion is set such that the time interval between write operations of two threads to the same memory address is less than 5 ms, or the read-write operations overlap by more than 2 ms. The finally generated event report includes the race condition type, such as data race, atomicity violation, occurrence location (source code line number), thread IDs and variable addresses involved, as well as the proposed repair solutions.

[0024] In step S103, for the suspected race condition events monitored, the event context information is abstracted into constraint conditions, and a constraint solver is used to explore the input combinations of the suspected static conditions to determine the minimum input set that can stably trigger the race condition.

[0025] Obtain the program path and variable information related to the suspected race condition events monitored, perform variable symbolic processing according to the program path and variable information to obtain the initial path conditions including variable value ranges and operation sequences; perform breadth-first traversal on the initial path conditions to judge all executable paths, and if an execution path is judged, convert the branch conditions, variable assignments, and thread operations on the execution path into logical expressions; use a constraint solver to solve the logical expressions to obtain specific input value combinations that satisfy the constraint conditions; perform simulation execution according to the specific input value combinations, and determine the high-potential input combinations with a triggering probability greater than a preset threshold by calculating the probability of each input combination triggering the race condition; perform verification and refinement on the high-potential input combinations, and screen out the minimum input subset that can stably trigger the race condition from the high-potential input combinations by repeatedly executing the target program and adjusting the thread scheduling strategy.

[0026] Exemplarily, for the monitored suspected race condition event, extract the program paths and variable information related to the event, and perform symbolic processing on the variables. Specifically, replace integer variables with symbolic expressions, convert pointer variables to symbolic addresses in the memory model, construct the symbolic execution path of the event, and obtain the initial path condition containing the variable value ranges and operation orders. According to the initial path condition, use symbolic execution technology to perform breadth-first traversal of the program, explore all possible execution paths, convert the branch conditions, variable assignments, and thread operations on each path into logical expressions to form a complete set of path constraint conditions. At the same time, use path pruning strategies such as loop unrolling limit and state merging to cope with the state space explosion problem caused by multi-threading. Optimize the generated path constraints, merge similar constraints, eliminate redundant conditions, and simplify the constraint expressions. Use the Z3 constraint solver to solve the optimized set of path constraint conditions, set the maximum solving time to 60 seconds, and the maximum number of solutions to 1000, generate specific input value combinations that satisfy the constraint conditions, and construct a candidate input set for the race condition. Conduct a preliminary evaluation of the generated candidate input set, use fast simulation execution technology to calculate the probability of each input combination triggering the race condition, and filter out the high-potential input combinations with a triggering probability greater than 0.5. Verify and streamline the filtered high-potential input combinations. By repeatedly executing the target program and adjusting the thread scheduling strategy, such as using priority inversion and round-robin scheduling, filter out the smallest input subset that can stably trigger the race condition, where the "smallest input subset" is defined as the subset with the smallest input scale on the premise of ensuring that the triggering probability is not less than 0.8, and determine the triggering conditions and necessary inputs for the race condition. For the monitored suspected race condition event, extract the relevant program paths and variable information, replace 10 integer variables with symbolic expressions such as x1, x2,..., x10, convert 5 pointer variables to symbolic addresses in the memory model such as m1, m2,..., m5, construct the symbolic execution path, and obtain the initial path condition. Adopt a breadth-first search strategy, set the maximum search depth to 100, explore the execution path, convert 20 branch conditions, 30 variable assignments, and 15 thread operations into logical expressions to form a set of path constraint conditions. Use a loop unrolling limit of 3 times and a state merging threshold of 0.9 to cope with the state space explosion. Optimize the constraints, merge 8 similar constraints, eliminate 12 redundant conditions, and simplify the constraint expressions. Use the Z3 solver, with a maximum solving time of 60 seconds and a maximum number of solutions of 1000, to generate 800 input value combinations that satisfy the constraints. Rapidly simulate the execution of each input combination 100 times, calculate the triggering probability, and filter out 350 high-potential inputs with a triggering probability greater than 0.5. Repeatedly execute the target program, test each input 500 times, use priority inversion, reverse once every 10 ms, and round-robin scheduling with a time slice of 5 ms, and filter out 50 inputs that can stably trigger the race condition.Define the minimum input subset as the one with a triggering probability not lower than 0.8 and the smallest input scale. Finally, an input subset containing 3 integers and 2 pointers is determined, with a triggering probability of 0.85.

[0027] Record and classify the captured suspected race condition events respectively, extract the key context information, abstract the variables, conditions, and synchronization operations involved in the events into symbolic expressions, and generate constraint conditions and a constraint system; use a constraint solver to solve the constraint system to obtain the input combinations that satisfy the constraint conditions, perform equivalence class partitioning on the input combinations, and evaluate the stability and reproducibility of different input combinations triggering race conditions by testing different input combinations, and select the target input set as the basis for subsequent vulnerability exploitation and verification.

[0028] According to the captured suspected race condition events, extract the program execution path, variable status, and thread interaction information at the time of the event occurrence, replace the relevant variables with symbolic expressions, convert the conditional statements into logical constraints, identify the synchronization operations and convert them into timing constraints, and construct a symbolic execution path; analyze the constructed symbolic execution path to generate path constraint conditions, combine the thread - to - thread interaction constraints and resource access constraints to form a constraint system; solve the constraint system, set the maximum solution time and the maximum number of solutions, and obtain the input value combinations that satisfy the constraint conditions; perform clustering analysis on the obtained input combinations using the K - means algorithm, set the number of clusters as the square root of the number of input combinations, partition the equivalence classes, and select the center point of each equivalence class as the representative input; perform test execution on the selected representative inputs, record the race condition triggering situation of each execution, calculate the triggering probability and reproducibility index, and screen out the input combinations that stably trigger according to the set preset threshold as the target input set.

[0029] Exemplarily, according to the captured suspected race condition events, extract the program execution path, variable states, and thread interaction information when the events occur. Replace the relevant variables with symbolic expressions, convert the conditional statements into logical constraints using predicate logic expressions, identify synchronization operations and convert them into timing constraints, and construct a complete symbolic execution path. Use the KLEE symbolic execution tool to analyze the constructed symbolic execution path, generate path constraint conditions, combine the interaction constraints and resource access constraints between threads to form a complete constraint system, optimize the complexity of the constraint system through constraint simplification and merging, use the constraint graph analysis method to evaluate the complexity of the constraint system, and calculate the number of constraint nodes and edge density. Use the Z3 solver to solve the optimized constraint system, set the maximum solving time to 300 seconds and the maximum number of solutions to 10,000, generate specific input value combinations that satisfy the constraint conditions, use the K-means algorithm to perform clustering analysis on the generated input combinations, set the number of clusters to the square root of the number of input combinations, and divide the equivalent classes. For the divided equivalent classes, select the center point of each class as the representative input for 100 test executions, record the race condition triggering situation of each execution, use the Monte Carlo method to calculate the triggering probability and reproducibility index. The triggering probability is defined as the number of successful triggers divided by the total number of executions, and the reproducibility index is defined as the maximum number of consecutive successful triggers divided by the total number of executions. Filter out the input combinations that stably trigger according to the preset threshold set by historical data statistics analysis (triggering probability greater than 0.8 and reproducibility index greater than 0.6) as the target input set. For the 50 captured suspected race condition events, extract the program execution path of each event, which on average contains 20 basic blocks, record 100 variable states and 30 thread interactions. Replace 50 integer variables with symbolic expressions x1 to x50, and 20 pointer variables with symbolic addresses p1 to p20. Use predicate logic to convert 80 conditional statements into logical constraints, such as converting "if(x>0)" to "x>0". Identify 40 synchronization operations and convert them into timing constraints, such as "thread1_lock<thread2_lock". Use the KLEE tool to analyze the symbolic execution path, generate 500 path constraints, combine 60 thread interaction constraints and 80 resource access constraints to form a system consisting of 640 constraints. By merging 30 similar constraints and eliminating 50 redundant constraints, the optimized constraint system contains 560 constraints. Using constraint graph analysis, 560 nodes are calculated, and the edge density is 0.15. Use the Z3 solver, set the maximum solving time to 300 seconds and the maximum number of solutions to 10,000, and generate 8,000 input combinations that satisfy the constraints. Use the K-means algorithm, set the number of clusters to 90 (approximately equal to the square root of 8,000), cluster the 8,000 inputs, and divide 90 equivalent classes. Select the 90 class center points and perform 100 test executions for each, recording the triggering situation.Use the Monte Carlo method to calculate the triggering probability and reproducibility. For example, if a certain input is triggered 80 times with a probability of 0.8 and a maximum consecutive trigger of 30 times, the reproducibility is 0.3. Set the thresholds of 0.8 and 0.6 according to historical data, and finally screen out 25 stable triggering input combinations as the target input set.

[0030] In step S104, monitor whether a race condition is triggered. If a race condition is triggered, analyze the impact of the race condition on the data integrity and behavior correctness of the application program, construct an attack payload that uses the race condition to execute arbitrary code or elevate privileges, and evaluate the harmfulness of the race condition vulnerability.

[0031] Obtain the key execution point information of the application program, insert a monitoring probe at the key execution point, and the monitoring probe is used to capture the thread execution state, shared variable access sequence, and synchronization operation information in real time; according to the information captured by the monitoring probe, compare the resource access order and timing relationship between threads to determine whether a race condition is triggered; if the race condition is triggered, perform a snapshot comparison of the data states before and after program execution to obtain the tampered data items, and trace the propagation path of the tampered data; according to the obtained propagation path, analyze the attack vectors and select the most matching attack method from them; for the selected attack method, construct the corresponding input sequence and thread scheduling scheme, and introduce the tampered data into the control flow decision point or memory operation of the program, and the control flow decision point or memory operation is used to implement control flow hijacking or buffer overflow.

[0032] Exemplarily, at the critical execution points of the application, dynamic binary instrumentation technology is used to insert monitoring probes to capture the thread execution state, shared variable access sequences, and synchronization operation information in real time. The vector clock algorithm is used to compare the resource access order and timing relationship between threads to determine whether a race condition is triggered. If a race condition is triggered, the data states before and after the program execution are snapshot-compared to identify the tampered data items. The propagation path of the tampered data is traced through taint analysis methods, and the impact on the critical business logic is evaluated. The affected memory regions and control flow branches are marked. Based on the taint analysis results, possible attack vectors are analyzed, including memory corruption, type confusion, and use-after-free, etc., and the most suitable attack method is selected. According to the selected attack vector, a specific input sequence and thread scheduling scheme are constructed to try to introduce the tampered data into the program's control flow decision points or memory operations to achieve control flow hijacking or buffer overflow, specifically including techniques such as ROP chain construction, shellcode injection, and stack overflow. Using the constructed attack payload, the operation that triggers the race condition is repeatedly executed 100 times in the target environment, and the success rate and impact range are recorded. The CVSS scoring system is used to evaluate the severity level of the vulnerability based on the stability and destructiveness of the attack effect. At the same time, attempts are made to obtain a higher-level system privilege through privilege escalation operations. For multi-threaded applications, the Pin dynamic binary instrumentation tool is used to insert monitoring probes at 50 critical execution points. Each probe captures 10 types of thread state parameters, 20 access sequences of shared variables, and 5 types of synchronization operation information. The vector clock algorithm is used to maintain a vector of length 8 for each thread to record the logical timestamps of accessing shared resources. By comparing the vector clocks of different threads, 3 resource accesses that violate the happened-before relationship are found, and it is determined that a race condition is triggered. Memory snapshots of the program state 100 ms before and after the trigger point are taken, and 15 tampered data items are identified using a binary difference comparison tool. Applying the dynamic taint analysis method, the tampered data is marked as a taint source, and 40 program basic blocks are traced. It is found that the taint data affects 3 critical control flow branches and 2 memory allocation operations. The analysis results show that there are two possible attack vectors: type confusion and use-after-free. A special input sequence of 100 bytes is constructed, including a 32-byte ROP chain and a 68-byte shellcode. By adjusting the thread priorities, the probability of the taint data being read in the critical section is increased to 80%. Repeatedly executed 100 times in the target environment, the success rate reaches 65%, of which 30% causes the program to crash and 35% achieves arbitrary code execution. Using the CVSS3.1 scoring system, the base score is 9.8 (low attack complexity, wide impact range), corresponding to the "critical" level.The successfully executed shellcode attempts to obtain root privileges through privilege escalation operations, exploiting the kernel vulnerability CVE-2021-3156 on the Linux system, with a privilege escalation success rate of 20%.

[0033] Step S105: For the memory models and synchronization primitives of different programming languages, obtain race condition anti-patterns and extract general exploitation primitives that cover various race condition vulnerabilities.

[0034] Obtain the memory model characteristics and synchronization primitive implementation information of mainstream programming languages, and construct a language feature comparison matrix based on this information. The language feature comparison matrix includes dimensions such as memory consistency models, atomic operation support, lock mechanisms, and memory barriers; analyze the application of language features in actual code through the language feature comparison matrix to obtain the usage frequency and distribution data of different synchronization mechanisms; extract concurrent programming-related code snippets from the open-source project code library, and use static code analysis tools to identify critical sections and resource competition points in the code snippets, and judge the usage frequency and distribution characteristics of different types of synchronization operations; construct a test case set containing race conditions according to the concurrent programming pattern, and trigger potential race conditions through dynamic execution and concurrent scheduling control. The concurrent scheduling control includes thread priority adjustment and time slice control; classify and summarize race condition cases, extract common characteristics and trigger patterns, and construct a race condition anti-pattern library, which is organized according to the dimensions of vulnerability type, trigger condition, and impact scope.

[0035] Exemplarily, for mainstream programming languages such as C / C++, Java, Python, and Go, analyze their memory model characteristics and synchronization primitive implementations, construct a language feature comparison matrix, including dimensions such as memory consistency model, atomic operation support, lock mechanism, and memory barrier, to identify the advantages and potential risk points of each language in concurrent programming. Based on the constructed language feature comparison matrix, analyze the application of specific language features in actual code, and count the usage frequency and distribution of different synchronization mechanisms. Extract concurrent programming-related code snippets from the open-source project code repository, use static code analysis tools such as Clang Static Analyzer and Coverity to identify critical sections and resource contention points, and count the usage frequency and distribution characteristics of different types of synchronization operations. Based on the identified concurrent programming patterns, construct a test case set containing race conditions, and through dynamic execution and concurrent scheduling control, such as using thread priority adjustment and time slice control, trigger potential race condition problems, and record the triggering conditions and impact scope. Classify and summarize the collected race condition cases, extract common features and triggering patterns, construct a race condition anti-pattern library, organize it according to dimensions such as vulnerability type, triggering condition, and impact scope, establish a code template generation mechanism based on the abstract syntax tree, construct general race condition exploitation primitives, and cover different types of concurrent vulnerabilities. For the four languages C / C++, Java, Python, and Go, construct a 10x10 language feature comparison matrix, including memory consistency models such as sequential consistency and relaxed consistency, atomic operation support such as CAS, memory barriers, and lock mechanisms such as mutex locks and read-write locks in 10 dimensions, and assign a score of 1-5 to each feature. Analyze the top 1000 projects on GitHub by the number of stars, extract 1 million lines of concurrent-related code, and statistically find that the usage frequency of synchronized in Java projects is the highest, accounting for 35%, while the usage of channels in Go projects reaches 40%. Use Clang Static Analyzer to analyze C / C++ projects and identify 2000 potential data contention points; use Coverity to analyze Java projects and find 1500 possible deadlock scenarios. Based on the identified concurrent patterns, generate 10,000 test cases containing race conditions, covering types such as read-write conflicts, atomicity violations, and order violations. By adjusting the thread priority from level 1 to 10 and the time slice from 1ms to 100ms, execute the test cases 1 million times, trigger 8000 race conditions, and the average trigger rate is 8%. Conduct cluster analysis on the triggered race conditions, extract 50 typical anti-patterns, and construct a 5-layer deep decision tree to organize the anti-pattern library. Based on the abstract syntax tree, design 100 code templates, covering 20 common concurrent vulnerability types in 5 languages, generate 1000 general exploitation primitives, and the applicability rate in different programming languages reaches 85%.

[0036] Step S106, during the penetration testing process, apply static analysis, dynamic monitoring, and symbolic execution to the target application system to mine race condition vulnerabilities and verify the exploitability of the vulnerabilities, assisting in completing the evaluation of the system security.

[0037] Perform static code analysis on the target application system according to the preset rule set to obtain a list of potential race condition risk points including unsynchronized shared variable access and improper lock usage order; implant dynamic monitoring probes to obtain key event information on thread creation, resource allocation, and lock operations, and construct a concurrent behavior model of the system during runtime; analyze the constructed concurrent behavior model to determine whether there are potential deadlocks, resource competitions, and synchronization anomalies; if there are potential anomalies, perform symbolic execution on the risk points and analysis results, simulate different thread scheduling scenarios, generate input sequences and scheduling strategies that trigger race conditions; select matching exploitation primitives according to the input sequences and scheduling strategies, construct targeted attack payloads, and verify the exploitability of the vulnerabilities in a controlled environment configured with sandbox and virtualization methods.

[0038] Exemplarily, for the target application system, use the Coverity Static Application Security Testing (SAST) tool and the Fortify Static Code Analysis tool for static code analysis to identify multi-threaded operations, shared resource access, and the usage of synchronization mechanisms. Mark potential race condition risk points according to a preset rule set, including unsynchronized shared variable access, improper lock usage order, etc., and generate a preliminary list of vulnerability candidates. During the operation of the target system, use AspectJ for bytecode instrumentation, implant dynamic monitoring probes, capture key events such as thread creation, resource allocation, and lock operations, record the interaction sequence between threads and the resource access pattern, and construct a concurrent behavior model of the system during runtime. Use Petri nets to analyze the constructed concurrent behavior model to identify potential deadlocks, resource competitions, and synchronization anomalies. For the marked risk points and the results of Petri net analysis, use the KLEE symbolic execution tool to explore possible program execution paths, simulate different thread scheduling scenarios, generate input sequences and scheduling strategies that trigger race conditions, and combine fuzz testing techniques to explore boundary cases to expand the list of vulnerability candidates. Match the generated vulnerability candidates with a pre-built general exploitation primitive library based on common vulnerability types and attack vectors, select suitable exploitation primitives to construct targeted attack payloads, verify the exploitability of the vulnerabilities in a controlled environment configured with sandbox and virtualization technologies, and use the CVSS scoring system to quantitatively evaluate the discovered vulnerabilities to assess the severity and potential impact of the vulnerabilities. For the target application system, use Coverity and Fortify for static code analysis, scan 1 million lines of code, identify 500 multi-threaded operations, 200 shared resource access points, and 300 synchronization mechanism usage scenarios. Mark 150 potential race condition risk points according to 15 preset rule sets, including 80 unsynchronized shared variable accesses and 70 improper lock usage orders, and generate a preliminary list of vulnerability candidates containing 150 items. Use AspectJ to implant dynamic monitoring probes at the entrances of 1000 key methods, capture 50,000 thread creations, 100,000 resource allocations, 200,000 lock operations and other key events during the system operation, and record and form a 10GB thread interaction sequence log. Use a Petri net modeling tool to construct a concurrent behavior model containing 500 places and 1000 transitions, and identify 20 potential deadlock scenarios, 50 resource competitions, and 30 synchronization anomalies through reachability analysis. For these 100 high-risk points, use the KLEE symbolic execution tool to generate 10,000 test cases, covering 1000 different program execution paths, and simulate 500 thread scheduling scenarios. Combining fuzz testing techniques, additionally generate 5000 boundary case test cases to expand the list of vulnerability candidates to 300 items. From a pre-built library containing 1000 exploitation primitives, select suitable exploitation primitives for 300 vulnerability candidates and construct 250 targeted attack payloads.In a sandbox environment configured with 10 virtual machine instances, each of the 250 attack payloads was executed 100 times, successfully reproducing 180 vulnerabilities. The 180 vulnerabilities were quantitatively evaluated using the CVSS 3.1 scoring system, obtaining an average base score of 7.5. Among them, 30 vulnerabilities had a score higher than 9.0, belonging to the severe level.

[0039] Before penetration testing, collect and analyze the architecture, components, and communication protocol information of the target system. For code segments that are prone to introducing race conditions due to multithreaded concurrency, shared resource access, and synchronization mechanisms, monitor the running state of the application and resource access in real time. Further analyze and verify the captured suspected race condition events. Construct targeted attack payloads based on the symbolic execution results, verify the exploitability of the race condition vulnerabilities, evaluate the impact of the race condition vulnerabilities on the system, and propose corresponding risk levels and repair suggestions.

[0040] Conduct a comprehensive scan of the target system to obtain the system architecture, component information, and communication protocol data of the target system. According to the obtained system information, select a static code analysis tool that supports multiple programming languages and concurrent mode analysis capabilities, and scan the source code or decompiled code of the target system to mark potential race condition risk points. Monitor the running applications of the target system in real time, and inject delays and error conditions at the risk points marked by static analysis. For the captured suspected race condition events, use a symbolic execution tool for analysis to generate input sequences and thread scheduling schemes that can trigger race conditions. Combine with a pre-built general exploitation primitive library to construct targeted attack payloads and verify the exploitability of the race condition vulnerabilities in a controlled environment. If it is verified that the race condition vulnerabilities can be exploited, evaluate the scope of the vulnerability impact and give the risk level.

[0041] Exemplarily, use Nmap and Wireshark to conduct a comprehensive scan of the target system, collect system architecture, component information, and communication protocol data, construct a system topology diagram and a component dependency graph, and identify potential concurrent processing modules and shared resource access points. Based on the collected system information, select a static code analysis tool that supports multiple programming languages and concurrent mode analysis capabilities, such as Coverity or Fortify, to scan the source code or decompiled code of the target system, focusing on code snippets related to multithreaded concurrency, shared resource access, and synchronization mechanisms, and mark potential race condition risk points. Process the static analysis results to generate a risk report and a list of suspicious code locations, which serve as the input for dynamic analysis. Utilize dynamic instrumentation techniques and fault injection tools, such as PIN and FAIL*, to monitor the running application in real-time, inject delays and error conditions at the risk points marked by static analysis, capture key events such as thread creation, resource allocation, and lock operations, record the interaction sequence between threads and the resource access pattern, and construct a concurrent behavior model of the system during runtime. For the captured suspected race condition events, use the KLEE symbolic execution tool for in-depth analysis, generate input sequences and thread scheduling schemes that may trigger race conditions, combine with a general exploitation primitive library pre-built based on common vulnerability types and attack vectors, construct targeted attack payloads, verify the exploitability of the race condition vulnerability in a controlled environment, use the CVSS3.1 scoring system to evaluate the vulnerability impact scope and give the risk level, and utilize an automated repair suggestion generation tool to provide targeted repair solutions. Use Nmap to conduct a comprehensive port scan of the target system, discover 1000 open TCP ports and 50 UDP ports, capture 10GB of network traffic through Wireshark, and identify 15 different communication protocols. Based on the scan results, construct a system topology diagram with 50 nodes and a component graph with 100 dependencies, locate 20 potential concurrent processing modules and 30 shared resource access points. Select the Coverity static analysis tool, configure 50 concurrency-related inspection rules, scan 1 million lines of source code and 200,000 lines of decompiled code, and mark 500 potential race condition risk points. Generate an 80-page PDF report containing risk details and a list of 2000 lines of suspicious code locations. Use the PIN dynamic instrumentation tool to insert monitoring probes at 1000 key execution points, and inject random delays of 10 - 100ms and 5 error conditions at 200 high-risk points marked by static analysis through the Failpoints fault injection tool. Run the application 100 times, each lasting 10 minutes, capture 500,000 thread creation events, 1 million resource allocation events, and 2 million lock operation events, and record 1GB of thread interaction sequence logs.Use the KLEE symbolic execution tool to conduct in-depth analysis on 100 of the most suspicious events. For each event, explore 1000 possible execution paths to generate 10,000 input sequences that may trigger race conditions. From a pre-built library containing 1000 exploitation primitives, select 5 most matching primitives for each suspicious event to construct 500 attack payloads. Verify the attack payloads in a sandbox environment configured with 10 virtual machine instances, and successfully reproduce 80 race condition vulnerabilities. Evaluate the 80 vulnerabilities using the CVSS 3.1 scoring system, obtaining an average base score of 7.5. Among them, 15 vulnerabilities have a score higher than 9.0, belonging to the severe level. Use an automated repair suggestion generation tool based on deep learning to provide 3 - 5 specific code-level repair suggestions for each vulnerability.

[0042] Step S107: Obtain typical race condition vulnerability cases of multi-threaded applications, summarize the causes, trigger conditions, exploitation methods, and repair solutions of the vulnerabilities, improve the race condition vulnerability knowledge base, and integrate the summarized race condition vulnerability knowledge and exploitation primitives into automated vulnerability mining and penetration testing tools.

[0043] Obtain race condition vulnerability cases of multi-threaded applications in the database. The vulnerability cases include fields such as vulnerability type, scope of influence, severity, trigger conditions, and repair suggestions; conduct static code analysis and dynamic execution tracking based on the vulnerability cases; obtain the complete trigger process of the vulnerability through the static code analysis and dynamic execution tracking. The trigger process includes key thread interaction points and resource competition scenarios; establish race condition vulnerability exploitation primitives for the trigger process. The exploitation primitives include a thread scheduling control module, a memory access sequence construction module, and an abnormal state trigger module; integrate the exploitation primitives into an automated vulnerability mining tool. The automated vulnerability mining tool performs code pattern matching; if the code pattern matching result matches a preset threshold, conduct a heuristic search to obtain the identification and verification results of new race condition vulnerabilities.

[0044] Exemplarily, typical race condition vulnerability cases of multi-threaded applications are collected from GitHub and CVE databases. Each case is analyzed in depth to extract the code features, triggering conditions, and scope of influence of the vulnerability, and a structured vulnerability description model is constructed that includes fields such as vulnerability type, scope of influence, severity level, triggering conditions, and repair suggestions. For the collected vulnerability cases, Coverity is used for static code analysis, combined with the Pin tool for dynamic execution tracking to restore the complete triggering process of the vulnerability, identify key thread interaction points and resource competition scenarios, and summarize the general causes and typical patterns of race conditions. Based on the analysis results of the vulnerability cases, a set of race condition vulnerability exploitation primitives is established and implemented, including modules such as thread scheduling control, memory access sequence construction, and abnormal state triggering. It is tested under different operating systems and hardware environments to verify the effectiveness and stability of the exploitation primitives. The implemented exploitation primitives are systematically sorted and classified to establish an exploitation primitive library, and the Neo4j graph database is used to store the association relationships between vulnerability knowledge and exploitation primitives. The analyzed vulnerability knowledge and exploitation primitives are integrated into an automated vulnerability mining tool. The KMP algorithm is used for code pattern matching, combined with the A* algorithm for heuristic search to improve the tool's ability to identify and verify new race condition vulnerabilities. Finally, the developed functional modules are integrated into the Metasploit penetration testing framework. From 1000 open-source projects with more than 10,000 stars on GitHub and the records in the CVE database in the past five years, 500 typical race condition vulnerability cases of multi-threaded applications are screened out. Each case is analyzed in depth to extract 20 key code features, 10 triggering condition patterns, and 5 types of scope of influence, and a structured vulnerability description model with 15 fields is constructed. Coverity is used to perform static analysis on 5 million lines of source code to identify 2000 potential race condition points. Combined with the Pin tool for 100 hours of dynamic execution tracking, 100,000 thread interaction events are captured, and 300 complete vulnerability triggering processes are restored. 50 key thread interaction patterns and 30 resource competition scenarios are identified, and 10 general causes of race conditions are summarized. A set of race condition vulnerability exploitation primitives with 5 core modules and 20 sub-functions is designed and implemented. 1000 tests are carried out on three operating systems of Windows, Linux, and macOS and two hardware architectures of x86 and ARM to verify that the effectiveness of the exploitation primitives reaches 95% and the stability reaches 90%. The Neo4j graph database is used to construct a vulnerability knowledge graph containing 1000 nodes and 5000 relationships to store the association relationships between vulnerability cases, triggering conditions, and exploitation primitives. The KMP algorithm is integrated into the automated vulnerability mining tool, 50 feature patterns are set, and the target code is scanned, with the matching efficiency increased by 30%. Combined with the A* algorithm for heuristic search, 8 heuristic functions are set, the number of explored paths is reduced by 50%, and the accuracy is increased by 20%.Finally, the three core modules and ten auxiliary functions to be developed will be integrated into the Metasploit framework, and twenty new exploit modules will be extended.

[0045] Although the present invention has been described in detail above with general descriptions and specific embodiments, based on the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of protection required by the present invention.

Claims

1. A security data penetration testing method, characterized in that: The method comprises: Pre-establish a multi-threaded application race condition vulnerability feature library, perform static analysis on the target application source code, identify potential race condition vulnerability points, and obtain a preliminary set of vulnerability candidates; Based on the vulnerability candidate set, a multi-threaded interaction scenario that can trigger race conditions is constructed. Through dynamic instrumentation, the access of multiple threads to shared resources is monitored in real time during the application runtime to capture suspected race condition events. For suspected race condition events monitored, the event context information is abstracted into constraints, and the constraint solver is used to explore the input combination of suspected race conditions to determine the minimum input set that stably triggers the race conditions. Monitor whether a race condition is triggered. If a race condition is triggered, analyze the impact of the race condition on the data integrity and behavior correctness of the application, construct an attack payload that uses the race condition to execute arbitrary code or elevate permissions, and evaluate the harmfulness of the race condition vulnerability. Targeting the memory models and synchronization primitives of different programming languages, we can obtain race condition anti-patterns and extract common exploit primitives that cover various race condition vulnerabilities. During the penetration test, static analysis, dynamic monitoring, and symbolic execution are used to discover race condition vulnerabilities in the target application system and verify the exploitability of the vulnerabilities to assist in completing the assessment of system security. Obtain typical race condition vulnerability cases of multi-threaded applications, summarize the causes, triggering conditions, exploitation methods and repair solutions of the vulnerabilities, improve the race condition vulnerability knowledge base, and integrate the summarized race condition vulnerability knowledge and exploitation primitives into automated vulnerability mining and penetration testing tools.

2. The method according to claim 1, characterized in that: The multi-threaded application race condition vulnerability feature library is pre-established, the source code of the target application is statically analyzed, potential race condition vulnerability points are identified, and a preliminary vulnerability candidate set is obtained, including: Extracting key features based on a pre-established multi-threaded application race condition vulnerability feature library, wherein the key features include vulnerability patterns, code structures, and variable dependencies; Receiving source code of a target application, performing static analysis on the source code using a matching method based on regular expressions and syntax analysis, identifying potential race condition vulnerabilities, and obtaining a preliminary set of vulnerability candidates; Performing deep semantic analysis on the preliminary vulnerability candidate set, performing code structure analysis through an abstract syntax tree, simulating the program execution path, extracting context information of each candidate vulnerability point, and generating a vulnerability feature vector; Data preprocessing and feature selection are performed on the vulnerability feature vector, and the feature vector is mapped to a high-dimensional feature space using a support vector machine algorithm to construct an optimal classification hyperplane to obtain a preliminary race condition vulnerability list.

3. The method according to claim 1, characterized in that The method constructs a multi-threaded interaction scenario that can trigger a race condition based on the vulnerability candidate set, monitors the access of multiple threads to shared resources in real time during the application runtime through dynamic instrumentation, and captures suspected race condition events, including: Generate test cases according to multi-threaded interaction scenarios, wherein the test cases include thread creation, resource allocation, and concurrent operation information; Performing code injection on the target application, wherein the code injection inserts monitoring probes at shared resource access points, thread synchronization operation calls, and key control flow nodes; Acquire the execution status and resource access information of multiple threads through monitoring probes, wherein the resource access information includes thread identification, operation type, access address and timestamp; Performing data filtering, deduplication and classification processing on the resource access information to obtain a standardized event sequence; Analyzing the thread interaction pattern in the event sequence using temporal logic rules, the analysis comprising: modeling the thread behavior using a finite state machine; Determine whether there is concurrent access that violates synchronization constraints; If there is concurrent access that violates synchronization constraints, the type of race condition, the location where it occurs, and related thread information are determined.

4. The method according to claim 1, characterized in that: For the monitored suspected race condition events, the event context information is abstracted into constraint conditions, and the constraint solver is used to explore the input combination of the suspected race condition to determine the minimum input set that stably triggers the race condition, including: Obtaining the program path and variable information related to the monitored suspected race condition event, performing variable symbolization processing according to the program path and variable information, and obtaining an initial path condition including a variable value range and an operation sequence; Performing breadth-first traversal on the initial path condition to determine all executable paths, and if an execution path is determined, converting the branch conditions, variable assignments, and thread operations on the execution path into logical expressions; Solving the logical expression using a constraint solver to obtain a specific input value combination that satisfies the constraint condition; Perform simulation execution according to the specific input value combination, and determine a high-potential input combination whose triggering probability is greater than a preset threshold by calculating the probability of each input combination triggering the race condition; Verifying and streamlining the high-potential input combinations, and screening out a minimum input subset that can stably trigger a race condition from the high-potential input combinations by repeatedly executing the target program and adjusting the thread scheduling strategy; It also includes: recording and classifying the captured suspected race condition events, extracting key context information, abstracting the variables, conditions, and synchronization operations involved in the events into symbolic expressions, and generating constraints and constraint systems; Use the constraint solver to solve the constraint system and obtain the input combinations that meet the constraints. Divide the input combinations into equivalence classes. By testing different input combinations, evaluate the stability and reproducibility of different input combinations triggering race conditions, and select the target input set as the basis for subsequent vulnerability exploitation and verification.

5. The method according to claim 4, characterized in that The captured suspected race condition events are recorded and classified respectively, key context information is extracted, variables, conditions, and synchronization operations involved in the events are abstracted into symbolic expressions, and constraints and constraint systems are generated; Use the constraint solver to solve the constraint system and obtain the input combinations that meet the constraint conditions. Divide the input combinations into equivalence classes. By testing different input combinations, evaluate the stability and reproducibility of different input combinations triggering race conditions. Select the target input set as the basis for subsequent vulnerability exploitation and verification, including: Based on the captured suspected race condition events, the program execution path, variable status and thread interaction information when the event occurs are extracted, the relevant variables are replaced with symbolic expressions, the conditional statements are converted into logical constraints, the synchronization operations are identified and converted into timing constraints, and the symbolic execution path is constructed; Analyze the constructed symbolic execution path, generate path constraints, and combine the thread interaction constraints and resource access constraints to form a constraint system; Solving the constraint system, setting a maximum solution time and a maximum number of solutions, and obtaining a combination of input values ​​that meets the constraint conditions; The obtained input combinations are clustered using the K-means algorithm. The number of clusters is set to the square root of the number of input combinations. Equivalence classes are divided and the center point of each equivalence class is selected as the representative input. The test is executed for the selected representative inputs, the race condition triggering situation of each execution is recorded, the triggering probability and reproducibility index are calculated, and the input combination with stable triggering is screened out according to the preset threshold as the target input set.

6. The method according to claim 1, characterized in that The monitoring determines whether a race condition is triggered. If a race condition is triggered, the impact of the race condition on the data integrity and behavior correctness of the application is analyzed, an attack payload that uses the race condition to execute arbitrary code or elevate permissions is constructed, and the harmfulness of the race condition vulnerability is evaluated, including: Acquire key execution point information of the application program, insert monitoring probes at the key execution points, and the monitoring probes are used to capture thread execution status, shared variable access sequence, and synchronization operation information in real time; According to the information captured by the monitoring probe, the resource access sequence and timing relationship between threads are compared to determine whether a race condition is triggered; If the race condition is triggered, a snapshot comparison is performed on the data state before and after the program is executed to obtain the tampered data item and track the propagation path of the tampered data; According to the obtained propagation path, analyze the attack vector and select the most matching attack method; For the selected attack method, a corresponding input sequence and thread scheduling scheme are constructed to introduce the tampered data into the control flow decision points or memory operations of the program, and the control flow decision points or memory operations are used to achieve control flow hijacking or buffer overflow.

7. The method according to claim 1, characterized in that The memory models and synchronization primitives of different programming languages ​​are used to obtain race condition anti-patterns and extract general exploit primitives covering various race condition vulnerabilities, including: Obtain memory model characteristics and synchronization primitive implementation information of mainstream programming languages, and construct a language characteristic comparison matrix based on the information, wherein the language characteristic comparison matrix includes memory consistency model, atomic operation support, lock mechanism, and memory barrier dimensions; Analyze the application of language features in actual code through the language feature comparison matrix to obtain the usage frequency and distribution data of different synchronization mechanisms; Extract concurrent programming related code snippets from the open source project code base, use the competitive code analysis tool to identify the critical sections and resource competition points in the code snippets, and determine the usage frequency and distribution characteristics of different types of synchronization operations; Constructing a test case set containing race conditions according to a concurrent programming model, triggering potential race conditions through dynamic execution and concurrent scheduling control, wherein the concurrent scheduling control includes thread priority adjustment and time slice control; The race condition cases are classified and summarized, common features and triggering patterns are extracted, and a race condition anti-pattern library is constructed. The race condition anti-pattern library is organized according to the dimensions of vulnerability type, triggering condition, and impact scope.

8. The method according to claim 1, characterized in that During the penetration test, static analysis, dynamic monitoring, and symbolic execution are used to discover race condition vulnerabilities in the target application system and verify the exploitability of the vulnerabilities to assist in completing the assessment of system security, including: Perform race code analysis on the target application system according to the preset rule set to obtain a list of potential race condition risk points including unsynchronized shared variable access and improper lock usage order; Implant dynamic monitoring probes to obtain key event information on thread creation, resource allocation, and lock operations, and build a concurrent behavior model during system runtime; Analyze the constructed concurrent behavior model to determine whether there are potential deadlocks, resource competition, and synchronization anomalies; If there are potential abnormal situations, symbolic execution is performed on risk points and analysis results to simulate different thread scheduling scenarios and generate input sequences and scheduling strategies that trigger race conditions. Select matching exploit primitives based on input sequences and scheduling strategies, construct targeted attack payloads, and verify the exploitability of vulnerabilities in a controlled environment configured with sandbox and virtualization methods; It also includes: before the penetration test, collect and analyze the architecture, components, and communication protocol information of the target system, monitor the application running status and resource access in real time for code snippets that are prone to race conditions due to multi-threaded concurrency, shared resource access, and synchronization mechanisms, further analyze and verify the captured suspected race condition events, construct targeted attack payloads based on the symbolic execution results, verify the exploitability of race condition vulnerabilities, evaluate the impact of race condition vulnerabilities on the system, and propose corresponding risk levels and repair suggestions.

9. The method according to claim 8, characterized in that Before the penetration test, the target system's architecture, components, and communication protocol information are collected and analyzed. For code snippets that are prone to race conditions in multi-threaded concurrency, shared resource access, and synchronization mechanisms, the application's running status and resource access are monitored in real time. The captured suspected race condition events are further analyzed and verified. According to the symbolic execution results, a targeted attack payload is constructed to verify the exploitability of the race condition vulnerability, evaluate the impact of the race condition vulnerability on the system, and propose corresponding risk levels and repair suggestions, including: Perform a comprehensive scan on the target system to obtain the system architecture, component information and communication protocol data of the target system; According to the obtained system information, a race condition code analysis tool that supports multiple programming languages ​​and concurrent pattern analysis capabilities is selected to scan the source code or decompiled code of the target system to mark potential race condition risk points; Performing real-time monitoring of running applications of the target system, injecting delays and error conditions at risk points flagged by static analysis; For captured suspected race condition events, use symbolic execution tools to analyze and generate input sequences and thread scheduling schemes that can trigger race conditions; Combined with the pre-built common exploit primitive library, targeted attack payloads are constructed to verify the exploitability of race condition vulnerabilities in a controlled environment. If it is verified that the race condition vulnerability can be exploited, the impact scope of the vulnerability is evaluated and a risk level is given.

10. The method according to claim 1, characterized in that The method of obtaining typical race condition vulnerability cases of multi-threaded applications, summarizing the causes, triggering conditions, exploitation methods and repair solutions of the vulnerabilities, improving the race condition vulnerability knowledge base, and integrating the summarized race condition vulnerability knowledge and exploitation primitives into automated vulnerability mining and penetration testing tools includes: Obtain race condition vulnerability cases of multi-threaded applications in the database, wherein the vulnerability cases include vulnerability type, impact scope, severity, triggering condition, and repair suggestion fields; Perform race code analysis and dynamic execution tracking based on the vulnerability cases described; The complete triggering process of the vulnerability is obtained through the competitive code analysis and dynamic execution tracking, and the triggering process includes key thread interaction points and resource competition scenarios; Establishing a race condition vulnerability exploitation primitive for the triggering process, wherein the exploitation primitive includes a thread scheduling control module, a memory access sequence construction module, and an abnormal state triggering module; Integrating the exploit primitive into an automated vulnerability mining tool, the automated vulnerability mining tool performing code pattern matching; If the code pattern matching result matches a preset threshold, a heuristic search is performed to obtain identification and verification results of the new race condition vulnerability.

Citation Information

Patent Citations

  • Code vulnerability detection method and device, medium and equipment

    CN110363004A

  • Container penetration method based on escape attack model

    CN117094002A