Method, device, equipment and medium for multi-thread undefined behavior detection

By building a multi-threaded execution path model and performing causal graph analysis, the problem of undefined behavior in multi-threaded programs is solved, the detection accuracy and coverage are improved, and the system stability and security are enhanced.

CN120803727APending Publication Date: 2025-10-17镁佳(北京)科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510945251.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Undefined behavior in multi-threaded programs causes problems such as uncertain execution, data contention, and deadlock, which affect system stability and security.

Method used

By constructing control flow graphs and data flow graphs, creating a multi-threaded execution path model, allocating it to computing nodes for iterative and hierarchical processing, building a causal relationship graph, generating operation dependency trajectories and mapping them to the memory model for execution sequence simulation, and performing behavioral analysis.

Benefits of technology

Improves the accuracy and coverage of undefined behavior detection in multi-threaded programs, and enhances system stability and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803727A_ABST
    Figure CN120803727A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and discloses a method, a device, equipment and a medium for multi-thread undefined behavior detection, and the method comprises the following steps: constructing a control flow diagram and a data flow diagram based on a multi-thread program code, and creating a multi-thread execution path model based on the control flow diagram and the data flow diagram; distributing a thread fragment data set in the multi-thread execution path model to a plurality of computing nodes, and extracting an operation sequence in the distributed thread fragment data set in each computing node; performing iteration and layering processing on the operation sequence, and constructing a causal relationship graph; based on the causal relationship graph, generating an operation dependency track, mapping the operation dependency track to different memory models, performing execution sequence simulation to obtain a plurality of execution paths, and performing behavior analysis based on the execution paths to obtain a behavior analysis result. According to the method, the undefined behaviors in the multi-thread program can be systematically detected, and the stability and the safety of the system can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a method, apparatus, device and medium for detecting multi-threaded undefined behavior. Background Art

[0002] As modern computing systems continue to grow in complexity, multithreaded programming has become increasingly popular in high-performance computing, distributed systems, and real-time applications. Multithreading allows programs to utilize multiple processor cores simultaneously, improving computing efficiency.

[0003] However, multithreaded programs often face complex issues during execution, such as inter-thread synchronization, shared resource access, and memory consistency, making them prone to undefined behavior. This undefined behavior can lead to nondeterministic execution, data races, deadlocks, and other anomalies, posing serious risks to system stability and security. Summary of the Invention

[0004] In view of this, the present invention provides a method, apparatus, device and medium for detecting undefined behavior in multi-threaded programs to solve the problem that undefined behavior in multi-threaded programs may lead to uncertain execution of the program, data contention, deadlock and other anomalies, posing serious hidden dangers to the stability and security of the system.

[0005] In a first aspect, the present invention provides a method for detecting multi-threaded undefined behavior, the method comprising: constructing a control flow graph and a data flow graph based on multi-threaded program code, and creating a multi-threaded execution path model based on the control flow graph and the data flow graph; distributing the thread shard data set in the multi-threaded execution path model to multiple computing nodes, and extracting the operation sequence in the distributed thread shard data set in each computing node; iterating and hierarchically processing the operation sequence to construct a causal relationship graph; generating an operation dependency trajectory based on the causal relationship graph, mapping the operation dependency trajectory to different memory models, performing execution order simulation, obtaining multiple execution paths, performing behavior analysis based on the execution paths, and obtaining behavior analysis results.

[0006] In an optional implementation, the constructing a control flow graph and a data flow graph based on the multi-threaded program code, creating a multi-threaded execution path model based on the control flow graph and the data flow graph comprises: performing static analysis on the multi-threaded program code, extracting instruction sequences and basic blocks in the threads, constructing a control flow graph based on the basic blocks, identifying data operation nodes in the instruction sequences, and constructing a data flow graph based on the data operation nodes; marking shared variable access instructions and synchronization operations in the control flow graph and the data flow graph respectively to obtain access labels, determining inter-thread dependency relationships based on the control flow graph, the data flow graph, and the access labels; dividing the threads into a plurality of thread slices based on a path allocation strategy, the control flow graph, the data flow graph, and the dependency relationships, and determining the multi-threaded execution path model based on the thread slices in each thread and the inter-thread dependency relationships.

[0007] In an optional implementation, the allocating thread slice data sets in the multi-threaded execution path model to a plurality of computing nodes, and extracting operation sequences in the allocated thread slice data sets in each of the computing nodes comprises: allocating the thread slice data sets in the multi-threaded execution path model to a plurality of the computing nodes based on a mapping function; and extracting the operation sequences in each of the computing nodes based on a time sequence.

[0008] In an optional implementation, the performing step-by-step iteration and hierarchical processing on the operation sequences, and constructing a causal relationship graph comprises: fusing operation sequences corresponding to the time sequence and label information corresponding to the time sequence based on the time sequence to obtain global time sequence operation sequences; mapping the global time sequence operation sequences into a partial order relationship, and representing initial causal constraints based on the partial order relationship; and correcting the initial causal constraints to obtain the causal relationship graph.

[0009] In an optional implementation, the generating an operation dependency track based on the causal relationship graph comprises: extracting operations of each operation sequence on shared variables based on nodes and edges in the causal relationship graph to obtain memory state mappings of each operation node; performing constraint modeling based on the memory state mappings, determining path priorities based on constraint modeling results; and performing recursive expansion on paths with a path priority higher than a preset path priority threshold to generate the operation dependency track.

[0010] In an optional implementation, the mapping of the operation dependency track to different memory models, performing execution order simulation, and obtaining a plurality of execution paths comprises: abstracting each of the memory models into a triple, wherein the triple comprises a set of event types, a set of ordered relationships, and a set of consistency constraints; performing event type mapping and relationship mapping of the operation dependency track to the memory models; constructing a constraint system, and performing constraint solving based on the constraint system; and performing topological sorting on a partial order relationship graph after the constraint solving, to obtain a plurality of the execution paths.

[0011] In an optional implementation, the behavior analysis based on the execution paths and obtaining a behavior analysis result comprises: analyzing whether there is a data race, atomicity violation, or memory order exception in the execution paths, and representing the behavior analysis result based on the analysis result.

[0012] In a second aspect, the present application provides a device for multi-thread undefined behavior detection, the device comprising: a first module configured to construct a control flow graph and a data flow graph based on multi-thread program code, and create a multi-thread execution path model based on the control flow graph and the data flow graph; a second module configured to assign thread slice data sets in the multi-thread execution path model to a plurality of computing nodes, and extract operation sequences in the assigned thread slice data sets in each of the computing nodes; a third module configured to perform iteration and hierarchical processing on the operation sequences, and construct a causal relationship graph; and a fourth module configured to generate an operation dependency track based on the causal relationship graph, map the operation dependency track to different memory models, perform execution order simulation, obtain a plurality of execution paths, and perform behavior analysis based on the execution paths, and obtain a behavior analysis result.

[0013] In a third aspect, the present application provides a computer device, comprising: a memory and a processor, which are communicatively connected to each other, and the memory stores computer instructions; the processor executes the computer instructions, thereby performing the method for multi-thread undefined behavior detection of the first aspect or any of the corresponding embodiments thereof.

[0014] In a fourth aspect, the present application provides a computer readable storage medium, which stores computer instructions for causing a computer to perform the method for multi-thread undefined behavior detection of the first aspect or any of the corresponding embodiments thereof.

[0015] In a fifth aspect, the present application provides a computer program product, which comprises computer instructions for causing a computer to perform the method for multi-thread undefined behavior detection of the first aspect or any of the corresponding embodiments thereof.

[0016] The method for multi-thread undefined behavior detection provided by the embodiment can systematically detect the undefined behavior in the multi-thread program through the process architecture of multi-stage modeling, distributed parallel processing, causal analysis and memory model simulation, capture the causal relationship between the shared variable access and the synchronization operation between threads by constructing the multi-thread execution path model and explicitly expressing the control flow and data flow information between threads, clearly represent the dependency relationship between different threads based on the causal relationship diagram, and provide a solid foundation for subsequent anomaly detection. In addition, the application also reasonably fragments and allocates the path, and combines the distributed computing framework, so that the analysis of large-scale programs becomes more efficient. In this way, the precision and coverage of the multi-thread program undefined behavior detection are effectively improved, and the stability and security of the system can be improved. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the specific embodiments or related art, the following will briefly introduce the drawings needed to be used in the specific embodiments or related art description. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.

[0018] Figure 1 The flowchart of the method for multi-thread undefined behavior detection of the embodiment of the present application is shown;

[0019] Figure 2 Another flowchart of the method for multi-thread undefined behavior detection of the embodiment of the present application is shown;

[0020] Figure 3 The flowchart of the Map stage and the Reduce stage of the embodiment of the present application is shown;

[0021] Figure 4 The structural diagram of the device for multi-thread undefined behavior detection provided by the embodiment of the present application is shown;

[0022] Figure 5 The hardware structure diagram of the computer device of the embodiment of the present application is shown. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical scheme and advantages of the embodiments of the present application more clear, the technical scheme in the embodiments of the present application will be clearly and completely described below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.

[0024] The complexity of multi-thread undefined behavior is reflected in the following aspects. First, thread-shared variable access conflicts: in a multi-thread program, concurrent access to shared variables by multiple threads can cause race conditions, especially without proper use of synchronization primitives, which can easily trigger undefined behaviors such as data races.

[0025] Second, the diversity of memory models: hardware platforms such as x86, ARM, etc. usually provide different memory consistency models. Weak memory models, such as ARM architecture, allow a certain degree of instruction reordering to optimize performance, but also increase the unpredictability of multi-thread program execution, which can trigger undefined behaviors at the program semantic level.

[0026] Third, the implicitness of synchronization operations and causality: synchronization mechanisms in multi-thread programs, such as locking, unlocking, and condition variables, usually implicitly establish causality between threads, but these causality relationships are difficult to be explicitly captured in program execution path analysis, making it difficult for related detection methods to locate potential undefined behaviors.

[0027] In related technologies, methods for detecting multi-thread undefined behaviors include:

[0028] First, static analysis method: static analysis scans source code to find potential problems such as data races or deadlocks. These methods usually rely on control flow and data flow analysis techniques, but due to the inability to obtain runtime context information, they are prone to a large number of false positives and false negatives.

[0029] Second, dynamic analysis method: dynamic analysis runs the program to capture actual interaction information between threads, such as memory access traces. Although it can provide higher detection accuracy, its coverage is limited by the quality of test cases, and the overhead is large, which is not suitable for large-scale programs.

[0030] Third, weak memory model simulation: some memory model simulation tools in related technologies attempt to simulate the weak memory behavior of hardware platforms by adjusting the execution order. However, these tools are usually limited to simple program logic and are difficult to deal with complex thread interactions and synchronization operations in actual programs.

[0031] Fourth, causality reasoning: some methods based on causality reasoning attempt to analyze the precedence relationship between threads, but due to the inability to effectively combine distributed computing and path expansion mechanisms, they are difficult to deal with the path explosion problem of large multi-thread programs, limiting the analysis scale and accuracy.

[0032] According to the embodiment of the present application, a method for multi-thread undefined behavior detection is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.

[0033] In the present embodiment, a method for multi-thread undefined behavior detection is provided, which can be used in terminals such as mobile phones, tablets, laptops, etc. Figure 1 The flowchart of the method for multi-thread undefined behavior detection according to the embodiment of the present application is shown as follows. Figure 1 As shown in the figure, the flow includes the following steps:

[0034] Step S101, based on the multi-thread program code, a control flow graph and a data flow graph are constructed, and based on the control flow graph and the data flow graph, a multi-thread execution path model is created.

[0035] In this step, the execution path information under the multi-thread running environment can be extracted from the source code of the multi-thread program, and by constructing the control flow graph and the data flow graph, the construction and preprocessing of the multi-thread execution path model are realized, thereby laying a data foundation for the extraction of subsequent operation sequences and the analysis of causal relationships.

[0036] Specifically, step S101 includes:

[0037] Step S1011, thread-level control flow and data flow analysis;

[0038] Step S1012, shared variable access information and synchronization operation identification;

[0039] Step S1013, inter-thread dependency analysis and preliminary path decomposition;

[0040] Step S1014, path allocation strategy and path slice organization;

[0041] Step S1015, formation of the preliminary multi-thread execution path model.

[0042] Step S102, the thread slice dataset in the multi-thread execution path model is allocated to a plurality of computing nodes, and in each computing node, the operation sequence in the allocated thread slice dataset is extracted.

[0043] In this step, the multi-threaded execution path model is fragmented by thread or code segment, the fragmented results are distributed to cluster nodes for parallel processing, each node independently analyzes the fragmented data, extracts the read-write operation sequence of the shared variable, and records the operation type and context. Based on the causal relationship graph, the causal relationship between the operations of the multi-threaded program is represented. The causal relationship graph includes: cross-thread shared variable access, synchronization sequence, and effective order relationship affected by the memory model.

[0044] Specifically, step S102 includes:

[0045] Step S1021, distributed path task decomposition and scheduling is performed;

[0046] Step S1022, operation sequence extraction and classification annotation is performed;

[0047] Step S1023, timestamp synchronization and distributed lock coordination are performed;

[0048] Step S1024, abnormal fragmentation and high-coupling path marking are performed;

[0049] Step S1025, operation sequence dataset is generated.

[0050] Step S103, operation sequences are iterated and hierarchically processed to construct a causal relationship graph.

[0051] In this step, the dependency relationship between operation sequences is analyzed layer by layer, such as data dependency or control dependency, which can be constructed into a causal relationship graph through a transitive closure algorithm.

[0052] Specifically, step S103 includes:

[0053] Step S1031, global timestamp normalization and sequence preliminary fusion are performed;

[0054] Step S1032, causal basis construction based on partial order relationship is performed;

[0055] Step S1033, recursive refinement and matrix representation of causal partial order constraints are performed;

[0056] Step S1034, out-of-order execution correction and causal loop resolution are performed;

[0057] Step S1035, the final causal relationship graph is constructed.

[0058] Step S104, based on the causal relationship graph, operation dependency trajectories are generated, the operation dependency trajectories are mapped to different memory models, execution order simulation is performed, a plurality of execution paths are obtained, behavior analysis is performed based on the execution paths, and behavior analysis results are obtained.

[0059] In this step, based on the causal relationship diagram, the shared variable access path and the memory value change history are extended to generate operation dependency tracks. Based on the scenario analysis strategy and the multiple filtering mechanism, the key path can be determined in a progressive manner, so as to construct the key operation dependency track which can truly reflect the multi-thread interaction and the memory value change. In this way, the problems of path scale explosion and irrelevant path over-expansion can be avoided.

[0060] The operation dependency tracks of multiple threads are mapped into different memory model constraint systems, wherein the memory model includes x86 strong consistency, ARM weak consistency, POWER model, etc. Through progressive constraint integration and topological processing, the possible execution sequences under various hardware platforms can be simulated to obtain multiple execution paths. Based on the multiple execution paths, an execution path set is determined. Based on the execution path set, the potential abnormal behaviors are analyzed and classified. Through multi-level statistical measurement, symbolic constraint refinement, time difference calculation and complex scoring mechanism based on logarithm and exponential function, the detection and classification of abnormal behaviors can be realized. The abnormal behaviors include ABA problem, memory consistency problem, operation sequence conflict, etc. The feature of the ABA problem is that after a thread checks whether the variable value meets the expectation, the variable value experiences a process from A to B and then back to A before the modification operation is performed, which leads to the misjudgment of the modification operation that the variable has not been modified, thereby causing a logic error.

[0061] The method for multi-thread undefined behavior detection provided by the embodiment can systematically detect the undefined behavior in the multi-thread program through the process architecture of multi-stage modeling, distributed parallel processing, causal analysis and memory model simulation. By constructing the multi-thread execution path model, the control flow and data flow information between threads are explicitized, and the causal relationship between the shared variable access and the synchronization operation between threads is captured. Based on the causal relationship diagram, the dependency relationship between different threads can be clearly represented, which provides a solid foundation for subsequent abnormal detection. In addition, the path is reasonably fragmented and allocated, and combined with the distributed computing framework, so that the analysis of large-scale programs becomes more efficient. In this way, the precision and coverage of the multi-thread program undefined behavior detection are effectively improved, and the stability and security of the system can be improved.

[0062] In some optional embodiments, based on the multi-threaded program code, a control flow graph and a data flow graph are constructed, and based on the control flow graph and the data flow graph, a multi-threaded execution path model is created, including: performing static analysis on the multi-threaded program code, extracting instruction sequences and basic blocks within threads, constructing a control flow graph based on the basic blocks, identifying data operation nodes in the instruction sequences, constructing a data flow graph based on the data operation nodes; marking shared variable access instructions and synchronization operations in the control flow graph and the data flow graph respectively to obtain access labels, determining inter-thread dependency relationships based on the control flow graph, the data flow graph and the access labels; based on a path allocation strategy, the control flow graph, the data flow graph and the dependency relationships, dividing the threads into a plurality of thread slices, and determining the multi-threaded execution path model based on the thread slices in each thread and the inter-thread dependency relationships.

[0063] In the embodiment, in step S1011, for each thread T i in the multi-threaded program, static analysis is performed, where i is a thread index, and i∈{1,2,…,M}, M is the number of threads, basic blocks and instruction sequences within the threads are extracted, and a control flow graph and a data flow graph are constructed. The thread set is as follows:

[0064]

[0065] For each thread T i , the instruction sequence I i is defined as follows:

[0066]

[0067] where n i is used to represent the number of instructions within the thread T i .

[0068] Based on the basic blocks, a control flow graph G C (T i )=(V C,i ,E C,i ) is constructed, where V C,i is a basic block set, and E C,i is a control flow transition edge set.

[0069] Data operation nodes in the instruction sequences are identified, and based on the data operation nodes, a data flow graph G D (T i )=(V D,i ,E D,i ) is constructed, where V D,i represents a set of operation nodes related by data dependency, and E D,i represents a data dependency relationship edge set.

[0070] In this way, a single thread level program structure representation can be obtained, and thus subsequent stages can locate shared variable access points and inter-thread interaction points.

[0071] Specifically, in step S1012, shared variable access instructions and synchronization operations are marked in the control flow graph and the data flow graph respectively, to obtain access marks. The shared variable access instructions and the synchronization operations can be locking, unlocking, barrier, conditional variable waiting and waking up, atomic operation compare-and-swap (CAS), etc. The shared variable set V is as follows:

[0072] V = {v1, v2, …, vK} K}

[0073] Wherein, K is used to represent the number of shared variables.

[0074] In the instruction sequence I i , if I i,j is a read-write or CAS operation on a shared variable v k ∈ V, a shared variable mark M(I i,j ) = v k and a corresponding access type mark (read, write, CAS) are added to the instruction. For synchronization instructions (such as locking and unlocking), a synchronization mark S(I i,j ) is added. Based on this processing, subsequent steps can establish dependency mapping between multiple threads based on this identification information.

[0075] Specifically, in step S1013, based on the control flow, data flow and shared variable access marks of each thread, the potential dependency relationship between multiple threads is initially characterized. The inter-thread dependency mapping relationship set R = {(T i , T j , v k , τ) | T i} is defined. An operation in the inter-thread dependency mapping relationship set has a dependency relationship on a shared variable v k with an operation in T j , and τ is relative timing information. By traversing all shared variable access points of the threads, if a T i write operation is found to occur before a T j read or write operation, or a synchronization primitive is detected, a time sequence causal chain is formed between the two threads, and the corresponding relationship is added to R. At the same time, these dependency relationships are integrated into the thread path model, to ensure that each thread path fragment contains correct preconditions and postconditions and key synchronization points, so that the path fragment can be correctly mapped to the execution order and timing dependency in the subsequent Map stage.

[0076] Specifically, in step S1014, after the inter-thread dependency analysis is completed, the control flow paths of each thread are appropriately sliced by the path assignment strategy to ensure that the analysis accuracy is not reduced due to path merging errors or missing dependencies in the subsequent processing stage.

[0077] Set the path assignment strategy function F path (G C (T i ), G D (T i ), R), according to the control flow and data flow structure and the dependency set R, the thread execution path is divided into several slices:

[0078]

[0079] where p i,x represents the xth path slice of thread T i , which contains consecutive basic blocks and related instructions within the slice, and maintains the continuity of key dependencies.

[0080] The slicing principle includes but is not limited to: each path slice contains at least one shared variable access operation or synchronization operation to maintain the effectiveness of the slice in the subsequent causal analysis stage.

[0081] For the detected high-coupling inter-thread dependency points, the slice length can be appropriately shortened to ensure accurate positioning of the interaction relationship in the subsequent Map processing and causal graph construction.

[0082] Specifically, in step S1015, based on the processing results of S1011 to S1014, a preliminary multi-thread execution path model M exec is formed:

[0083] M exec = {(T i , P i , R) | i = 1, 2, …, M}

[0084] The model records the path slice set P i of each thread and the inter-thread dependency set R, and ensures that in the subsequent steps, the operation sequence extraction, causal relationship construction and out-of-order execution scenario simulation can be directly based on this.

[0085] Thus, by constructing the control flow graph and the data flow graph, and cross-annotating the shared variable access, the data dependency and the control dependency between threads can be accurately captured, and the limitations of single graph analysis can be avoided. Meanwhile, based on the annotated synchronization operation and the shared variable access, the dependency relationship between threads is explicitly constructed, which can provide a theoretical basis for subsequent path generation. In addition, the threads are divided into independent slices according to the dependency relationship, which supports distributed parallel processing and breaks through the performance bottleneck of traditional single-thread analysis.

[0086] In some optional embodiments, the thread slice dataset in the multi-thread execution path model is distributed to a plurality of computing nodes, and in each computing node, an operation sequence in the distributed thread slice dataset is extracted, including: distributing the thread slice dataset in the multi-thread execution path model to the plurality of computing nodes based on a mapping function; and in each computing node, extracting the operation sequence based on a time sequence.

[0087] In the present embodiment, after the construction and preprocessing of the multi-thread execution path model are completed, the present step aims to distribute the sliced multi-thread path to a plurality of computing nodes for parallel processing through the Map phase of the distributed computing framework (Map-Reduce), so as to extract the operation sequence containing the shared variable access, the atomic operation, the synchronization operation and the corresponding timestamp information.

[0088] In step S1021, based on the multi-thread execution path model M exec = {(T i , P i , R) | i = 1, 2, …, M} created in step S101, the thread slice dataset of each thread is distributed to an independent computing node according to a predefined strategy.

[0089] The set of Map tasks W is defined as:

[0090] W = {W1, W2, …, W L}

[0091] Wherein, L is used to represent the number of Map tasks. The distribution principle is completed through a mapping function F map (M exec ):

[0092]

[0093] The slices of different threads are dispersed to different Map tasks to fully utilize the parallel computing resources. The distribution strategy needs to consider the dependency relationship R between the slices, and can achieve high parallelism under the condition of ensuring that the critical dependency condition is not destroyed.

[0094] Specifically, in step S1022, in each Map computing node W​l Internally, the allocated path fragments are analyzed instruction by instruction to extract the following instructions: shared variable read and write, CAS operation, lock and unlock, memory fence, condition variable operation, i.e. synchronization semantics related instructions.

[0095] Define the operation set O:

[0096] O = {o1, o2, …, o α}

[0097] Each operation o r ∈ O corresponds to a specific instruction and contains the following information: thread identification T i , fragment identification p i,x , operation type (read, write, CAS, lock, unlock, fence, etc.), associated shared variable v k ∈ V, timestamp information τ(o r ), operation context (local relationship of previous and subsequent instructions).

[0098] Further, when extracting the operation sequence, all operations are recorded in chronological order as:

[0099]

[0100] Where O i,x is the operation sequence in fragment p i,x , and u i,x is the number of operations extracted in the fragment.

[0101] Specifically, in step S1023, further to ensure that the generated operation sequence maintains global timing consistency when the multi-Map task is processed in a parallel environment, a timestamp synchronization and distributed lock coordination mechanism is adopted.

[0102] Define a global time synchronization function F time , which assigns global timestamps to different fragments in the following way:

[0103] τ global (o i,x,j ) = F time (o i,x,j )

[0104] It can be ensured that in the subsequent Reduce phase, the operations can be correctly merged and sorted in time according to τ global .

[0105] When it involves atomicity guarantee or cross-fragment competitive access to global synchronization resources, such as when restoring inter-thread causal dependency moments, the present application proposes a distributed lock mechanism L to ensure the consistency of updating global timing, marking queues or associated global counters.

[0106] Represents the operation of locking resource R

[0107] With the help of timestamp synchronization and locking mechanism, multiple map tasks will not produce uncontrollable timing deviations due to network delays and distributed communication uncertainties during the extraction of operation sequences.

[0108] Specifically, in step S1024, in order to meet the needs of subsequent complex causal analysis, abnormal features or high-coupling paths in the sharding operation sequence are specially marked at this stage.

[0109] If there is a high frequency CAS operation or obvious shared variable fast alternating write in a shard, add a high priority tag x to it. high , so that it can be processed first in subsequent Reduce integration.

[0110] If there is a clear synchronization operation mode (such as Lock / Unlock sequence) in the slice that is tightly coupled with other thread slices, give it a dependency mark x dep , to ensure that the subsequent causal diagram construction focuses on checking the timing of these fragments. high and χ dep The symbol is χ normal .

[0111] Define the fragment labeling function F tag :

[0112] F tag (p i,x ,O i,x )→{χ high ,χ dep ,χ normal ,…}

[0113] Specifically, in step S1025, after completing the above steps, each Map task W l Both produce local operation sequence datasets:

[0114]

[0115] These data sets are stored in a distributed file system for causal analysis and global integration in the subsequent Reduce phase.

[0116] In this way, based on the processing flow from step S1021 to step S1025, the multi-thread execution path model M is realized. execThe path is divided into fragments, distributed to different computing nodes for parallel processing, and the operation sequence of multi-threading is extracted. By means of timestamp synchronization, distributed lock mechanism and feature marking method, the operation sequence generated is ensured to have timing consistency and identifiable characteristics in the global range. At the same time, by means of timestamp synchronization mechanism and distributed parallel processing strategy, the operation sequence is ensured to maintain timing consistency in the global range, which can lay a reliable data foundation for subsequent causal relationship construction and abnormal behavior detection.

[0117] In some optional embodiments, the operation sequence is iterated and processed step by step to construct a causal relationship graph, including: based on the time sequence, fusing the operation sequence corresponding to the time sequence and the marking information corresponding to the time sequence to obtain a global timing operation sequence; mapping the global timing operation sequence into a partial order relation, and representing an initial causal constraint based on the partial order relation; correcting the initial causal constraint to obtain a causal relationship graph.

[0118] In this embodiment, on the basis of the distributed operation sequence dataset output in the Map phase of step S102, a method of step-by-step iteration and hierarchical processing is adopted to construct a causal relationship graph (CRG) of multi-thread shared variable access and synchronization dependency. Through continuous correction and refinement of global timestamp, cross-thread dependency chain and potential out-of-order execution mode, this step realizes a gradual construction process from local timing segment to global causal structure.

[0119] Specifically, in step S1031, in the Reduce node, first, the datasets from each Map task are uniformly processed. Through the global timestamp τ global determined in step S1023, each fragment operation sequence O i,x is preliminarily fused with the corresponding marking information χ tags in time sequence to form a set of global timing operation sequences:

[0120]

[0121] In this set, each operation o i,x,j has a global timestamp τ global (o i,x,j ), and has been preliminarily linearly spliced with the operation sequences of other threads in time as the main axis.

[0122] Specifically, in step S1032, the causal dependency of multi-thread execution satisfies a partial order relation. According to the constraint of the order of shared variable access and synchronization operation, the global timing data is mapped into a set of partial order constraints. For any two operations o a ,b ∈O global :

[0123] If o a is a write operation on a shared variable v k , and o b is a read or write operation on the same variable v k , and according to global timestamps have τ global (o a )<τ global (o b ), then a preliminary partial order constraint is generated:

[0124]

[0125] If o a and o b belong to different threads and there exists a synchronization primitive, such as Lock / Unlock, CAS or memory barrier, etc., which establishes an explicit precedence relationship, then the corresponding synchronization partial order is defined:

[0126]

[0127] By extracting the above partial order relation set P = {<,< s}, a plurality of original causal constraint relations are obtained. However, the constraints at this time may be in conflict due to potential reordering and excessive reliance on local timing information.

[0128] Specifically, in step S1033, in order to further eliminate local partial order conflicts and improve the global consistency of the partial order relationship, the partial order relationship is converted into a matrix representation for recursive refinement and global consistency processing.

[0129] The thread-operation causal relationship matrix M is defined with a dimension of N x N, where N = |O global | is the total number of global operations. The matrix element M a,b is defined as follows:

[0130]

[0131] By performing a transitive closure operation on the matrix M, hidden transitive dependencies are represented and made explicit:

[0132] M * = M ∨ (M · M) ∨ (M · M · M) ∨...

[0133] In this process, the transitive closure of M is updated iteratively to gradually eliminate mutual exclusion, loops or contradictory relationships. When a possible partial order conflict is detected, such as o a < o b and o b < oa Forming loops, further causal rule correction is performed.

[0134] Specifically, in step S1034, logical loops are generated in the preliminary causal relationship based on the above-mentioned out-of-order execution phenomenon introduced by hardware and compiler optimization. To eliminate these loops, the application further corrects step by step based on the semantic properties of the synchronization operation and the memory model knowledge. The correction process is divided into the following progressive steps:

[0135] Micro-correction: For a single loop, for example, o a <o b With o b <o a , local partial order conflict resolution is performed using known synchronization primitive semantics (fence instruction forces order). For the detected loop, the weight of the part of the edge is reduced, or the wrong dependence is cancelled according to the synchronization rule, as follows:

[0136] M b,a :=0

[0137] Based on the above formula, if it is judged that the reverse dependence needs to be removed based on the fence memory model constraint.

[0138] Macro condensation: based on the strong connected component decomposition of the directed graph, the loop or strong connected structure is condensed into a single node, and the graph is acyclic based on the following method:

[0139]

[0140] Wherein, is the causal directed graph representation between operations (determined by the matrix M * ), and SCC_Condensation is a function that condenses the strongly connected component into a single node.

[0141] Based on this step, it is ensured that an acyclic directed graph structure is finally obtained, so that the scheme strictly satisfies the partial order property.

[0142] The multi-level correction and condensation process can effectively overcome the problems that cannot be solved by directly using timestamp ordering or single-pass transitive closure, and realize ordering under weak memory model and complex synchronization semantics.

[0143] Specifically, in step S1035, after completing the out-of-order correction and loop elimination, the matrix transitive closure result and the acyclic directed graph converted into the final state of the causal relationship graph. The nodes of the graph represent global operation events, and the directed edges represent the determined causal dependence relationship, wherein the node set is as follows:

[0144] V CRG ={o′1,o′2,…,o′N′}

[0145] wherein N'≤N, and N'N if there is SCC condensation, and o" represents the Nth node in the graph.

[0146] The edge set is as follows:

[0147]

[0148] In this way, based on the causal graph, the out-of-order access and complex synchronization structure that are originally difficult to determine are solved through iterative refinement and condensation; at the same time, the edges and nodes retained in the graph can reflect the data and control dependencies between threads, and provide accurate and operable data-driven representations for path expansion and weak memory model simulation in subsequent steps; in addition, by combining causal reasoning and distributed computing, not only can the complex dependencies in the multi-threaded program be efficiently captured, but also higher accuracy can be provided for subsequent simulation of different memory models; at the same time, by using distributed analysis of multi-threaded execution paths, complex concurrent scenarios can be adapted, and the performance bottleneck of the solutions in related technologies when processing large-scale programs can be avoided.

[0149] In some optional embodiments, based on the causal graph, an operation dependency track is generated, including: based on the nodes and edges in the causal graph, extracting the operations of each operation sequence on the shared variables to obtain a memory state mapping of each operation node; based on the memory state mapping, constraint modeling is performed, and based on the constraint modeling result, a path priority is determined; and for a path whose path priority is higher than a preset path priority threshold, recursive expansion is performed to generate the operation dependency track.

[0150] In this embodiment, the path dependency information can be expanded based on the causal graph, including:

[0151] In step S1041, the memory value state space and the dependency mapping are initially expanded.

[0152] In this step, based on the node set V CRG ={o'1, o'2, …, o' N ′} and the directed edge set E CRG in the CRG, the access history of the shared variable set V is extracted. For each operation node o' a , a set of memory state markers m(o' a ) is assigned to describe the memory mapping generated after the operation reads (R), writes (W), and CAS updates (C) of the shared variable values.

[0153] The memory state function is defined as:

[0154]

[0155] in, Represents a mapping space from shared variables to their value sets.

[0156] For example, if is a shared variable, denoted by Val(v k ,o′ a ) is the operation o′ a After execution k By calculating the value of each dependency edge (o′ a →o′ b ) Calculate node by node to obtain the transfer relationship and final state of each operation on the shared variable value.

[0157] Step S1042: Symbolic execution and dynamic information fusion.

[0158] In this step, in order to avoid the situation caused by conditional branches or complex data dependencies in path expansion,

[0159] The state is uncertain, and the expansion process is optimized based on symbolic execution. For each operation o a If it contains conditional judgment, CAS comparison conditions or read operations based on past write results, then introduce symbolic constraint variables surface

[0160] Show v k In operation a State constraints before and after:

[0161]

[0162] In this phase, symbolic constraints are replaced with observed actual values ​​based on actual dynamic traces or test logs, thereby reducing the size of the state space. When no dynamic information is available, symbolic execution provides flexible data retention space for subsequent scenario selection.

[0163] Step S1043: Scenario-based selection function and key path priority expansion.

[0164] In this step, since CRG may contain a large number of scalable paths, in order to prevent the number of paths from expanding excessively during expansion, a scenario selection function can be introduced. Priority evaluation is performed on the extension path based on features such as memory value change pattern, operation type, synchronization point density, and ABA suspicion.

[0165] Define the scene selection function:

[0166]

[0167] Among them, ρ a To score the path expansion priority, we can use the multidimensional feature vector xa The multi-dimensional feature vector x is calculated a The calculation method is as follows:

[0168] x a = (δ CAS (o′ a ), δ ABA (o′ a ), Δ val (o′ a ), γ sync (o′ a ))

[0169] δ CAS (o′ a ) is used to indicate whether the operation is a CAS operation, and the larger the value is, the more critical the operation is; δ ABA (o′ a ) is used to indicate whether there is an ABA suspicious scenario, for example, the same variable value sequence repeatedly appears; Δ val (o′ a ) is used to represent the amplitude of the shared variable value change caused by the operation; and γ sync (o′ a ) is used to represent the synchronization point density and complexity.

[0170] The comprehensive score function is defined as:

[0171]

[0172] Wherein, w is used to represent the weight vector. In the expansion path, only the nodes and adjacent paths whose ρ a exceed the threshold θ are selected for further expansion, so as to ensure that high-value scenarios are focused on and blind growth of irrelevant paths is reduced.

[0173] Step S1044, multi-level operational dependency tracking (ODT) construction and scenario aggregation.

[0174] In this step, after filtering based on the scenario selection function, the selected key paths and upstream and downstream nodes thereof are recursively expanded to obtain a preliminary operational dependency tracking set, as follows:

[0175]

[0176] Each ODT represents a dependency sequence from the initial memory state to the key operation occurrence point, and records all shared variable value transitions, symbolic constraints and synchronization semantics between o′ start and o′ end . In order to reduce the redundancy between paths, the ODTs are clustered according to the similarity of the shared variable value transitions and the synchronization points, and the representative ODTs are selected as the final operational dependency tracking set. Performing scenario aggregation and redundancy elimination.

[0177] Defining an ODT similarity measure function σ(ODT p ,ODT q ) to identify high overlap patterns between different traces based on the similarity measure function σ(ODT p ,ODT q ). The similarity measure function σ(ODT p ,ODT q ) is as follows:

[0178]

[0179] where V p and V q are the CRG node sets corresponding to two ODTs, respectively.

[0180] If σ(·,·) is higher than a certain threshold, the ODTs are merged, and the value change relationship of the merged nodes is unified through symbolic constraints to reduce the repeated analysis overhead.

[0181] After scenario selection and aggregation, the generated operation-dependent trace set is as follows:

[0182]

[0183] Each trace in the set is a high-value path, and key operations such as CAS, potential ABA, complex synchronization, memory value evolution history, and node sequences satisfying global causality constraints are retained, and redundancy is reduced through aggregation.

[0184] Step S1045, re-verification based on alignment of front and back states.

[0185] In this step, to ensure that the final ODT set not only reflects the real causal relationship in structure but also is strictly consistent in memory state level, in this phase, the memory state continuity verification is performed on each ODT k .

[0186] For each pair of adjacent operations (o′ a ,o′ b ) in the sequence, the following conditions are checked:

[0187]

[0188] If state inconsistency or violation of reachability is found, the constraints generated by symbolic execution are used to solve again to decide the correction path or eliminate invalid traces. Based on this step, it can be ensured that the output traces are not only logically feasible (no loop, satisfying causal order), but also self-consistent in the memory value change level.

[0189] In this way, based on the nodes and edges of the causal graph, that is, operations and dependencies, shared variable operations are extracted to ensure that the memory state mapping fully conforms to the actual data flow logic of the program, and can avoid state misjudgments caused by ignoring dependencies in traditional static analysis; at the same time, based on the constraint modeling results, such as operation type, variable change amplitude, synchronization density, etc., the path priority is dynamically calculated, and only high-risk paths, such as CAS operations, frequent shared variable accesses, etc., are recursively expanded. Compared with full path enumeration, the amount of invalid paths generated can be reduced, significantly improving analysis efficiency.

[0190] In some optional implementations, the operation dependency trace is mapped to different memory models, and execution order simulation is performed to obtain multiple execution paths, including: abstracting each memory model into a triple, where the triple includes an event type set, an ordered relationship set, and a consistency constraint set; performing event type mapping and relationship mapping of the operation dependency trace to the memory model; constructing a constraint system, and performing constraint solving based on the constraint system; and topologically sorting the partially ordered relationship graph after constraint solving to obtain multiple execution paths.

[0191] In this embodiment, multiple levels of sub-steps can be deeply combined with the formulated description to simulate out-of-order execution under different memory models. The weak memory model is a model used in computer architecture to describe the order of processor memory operations. Compared with the strong memory model, it has looser constraints on the order of memory operations, allowing the processor to reorder certain memory operations to optimize performance. Step S105, simulating the out-of-order execution scenario under the weak memory model, includes:

[0192] Step S1051: Memory model definition and multi-element ordered set initialization.

[0193] In this step, define the target memory model parameter set Memory Model Expressed as a constraint triple:

[0194]

[0195] Where: E represents the event type set, such as read R, write W, CAS, Fence, Barrier, etc. O={< po ,< rf ,< mo{…} is a set of multi-class order relations, including program order, read-write visibility relation, memory order and related extensions, etc. C represents a set of global consistency constraints, such as linear consistency of sequential consistency (SC), weaker read-after-write order requirement under total store order (TSO), and rules of read-write disorder but preserving synchronization order under advanced RISC machines (ARM). For different platforms (such as x86, ARM), There are different sets of C, and different degrees of relaxation of order relation constraints.

[0196] Step S1052, ODT is embedded in the memory model framework.

[0197] In this step, the constraints fused into the memory model are marked with event types and pre-mapped with order relations for each operation sequence in the ODT k

[0198] Define the operation event set ε ODT =∪ k V(ODT k ), where V(ODT k ) is the operation node contained in the track, and the operation node comes from the subset of the reduced CRG node set.

[0199] Assign an event type label λ(e)∈E to each event e∈ε ODT , for example:

[0200] λ(e)=R(v k ),W(v k ),CAS(v k ),Fence,…

[0201] On this basis, the explicit causal relationship in the track is mapped to the initial program order <e po > candidate set

[0202]

[0203] This preliminary mapping provides a structured input for the injection of memory model consistency constraints.

[0204] Step S1053, constraint injection of read-write visibility and reordering conditions.

[0205] ​​In this step, in the weak memory model, where the read operation reads the value (reads-from), and the order in which the write operation appears in the global memory sorting, the confirmation process has certain problems in the related art, so the direct read-write visibility relationship (< rf ) and memory order relationship (< mo ) and solves it according to the memory model constraints.

[0206] Read-write visibility constraints:

[0207] For any read operation r∈ε ODT Read variable v k , determine the unique write operation w∈ε corresponding to the read ODT , so that λ(w)=W(v k ) or λ(w)=CAS(v k ), and meet the following requirements:

[0208] w< rf r

[0209] And there is no such thing, under the memory model requirements, that a later visible write overwrites the read. rf Determining the relationship requires solving a set of selection constraints to ensure that read operations retrieve values ​​from the correct write source.

[0210] Memory order (mo) relationship solution:

[0211] To meet the requirements of the memory model, all write operations on the same variable are sorted linearly. mo For any variable v k , if there is a write set {w1,w2,…,w n}, then there must exist a linear order:

[0212] w1< mo w2< mo …< mo w n

[0213] And meet < rf and mo The consistency requirements between them are Single-Writer, Multi-Reader conditions under storage consistency.

[0214] Step S1054: cross-level progressive solution and constraint optimization.

[0215] In this step, to determine rf and mo and other necessary ordered relations, such as < scTotal order relation, feasibility under the premise of satisfying memory model constraint C, this application directly adopts multi-stage progressive solving and optimization strategy:

[0216] Initial candidate graph generation: the candidate set and < mo initial guess sequence combination, form a directed candidate graph

[0217] To perform transitive closure calculation and loop detection, if there is a loop or structure conflicting with the memory model requirements, eliminate the contradictory edges according to the specific constraint rules of the memory model:

[0218]

[0219] Among them, characterize if the edge is incompatible with the memory model constraints in C. Perform secondary modification on each relation, and update read-write matching and write ordering step by step until no loop and constraint conflict is generated.

[0220] This progressive solving process is similar to solving constraint satisfaction problems or constraint optimization problems. Under high complexity, with the help of Boolean satisfiability solver (SAT) / satisfiability modulo theories solver (SMT), the Boolean condition set of < is solved to find a memory consistency solution that satisfies rf , mo .

[0221] Step S1055, model topological feasible execution sequence enumeration.

[0222] In this step, after successfully determining a set of relations that satisfy the constraints of , the final acyclic directed graph is converted into a set of feasible execution sequences. Through topological sorting, a set of linear extensions is extracted from :

[0223]

[0224] Among them, each is a total order sequence that satisfies the target memory model constraint, and all operation events are arranged in a legal order. When facing a weak memory model, there are multiple legal linear extensions.

[0225] Step S1056, multi-platform simulation output and differential analysis.

[0226] ​In this step, for a more comprehensive evaluation of multi-threaded programs, potential execution patterns under different hardware architectures, a variety of memory models The processes of S1051-S1055 are executed respectively to obtain the corresponding feasible execution sequence set:

[0227]

[0228] By comparing the differences between them, the out-of-order read-write, hidden ABA dangerous scenarios and synchronization failure problems that may occur under different hardware platforms are identified.

[0229] This step fuses the operation dependency track with a variety of memory models. Starting from the definition of the memory model and the event relationship, through read-write visibility and memory ordering relationship injection, progressive constraint solving, multi-version feasible sequence topological ordering, and multi-platform difference analysis, a series of execution path sets constrained by memory models are formed. Through this simulation process, the multi-threaded interaction mode originally contained in a single CRG can be expanded under a variety of hardware memory models, providing more realistic and diversified execution scenarios for further anomaly detection and optimization.

[0230] In this way, the present scheme further simulates the out-of-order execution phenomenon under different hardware platforms through weak memory model simulation technology, solving the problem of insufficient adaptation of weak memory models in related technologies. At the same time, by combining progressive constraint solving and multi-platform analysis, it can accurately handle memory consistency problems under x86, ARM and other platforms, ensuring the consistency of program execution order and memory access under different architectures. In addition, it not only can deal with synchronization problems in traditional platforms, but also can effectively analyze and simulate potential abnormalities under weak consistency memory models, thereby improving the detection capability of undefined behaviors and avoiding omissions caused by hardware differences.

[0231] In some optional embodiments, behavior analysis is performed based on the execution path to obtain a behavior analysis result, including: analyzing whether there is a data race, atomicity violation or memory order exception in the execution path, and representing the behavior analysis result based on the analysis result.

[0232] ​In this embodiment, after completing the execution order simulation under the weak memory model, this step analyzes and classifies potential abnormal behaviors, such as ABA problem, memory consistency problem, operation order conflict, etc., based on the set of feasible execution paths of multiple platforms and multiple scenarios. The results of the previous steps of multi-thread execution path, causal dependence, operation trajectory expansion and weak memory model simulation are integrated, and a new type of progressive abnormality recognition strategy is proposed. The abnormality recognition strategy realizes the detection and classification of abnormal behaviors through multi-level statistical measurement, symbolic constraint refinement, time sequence difference calculation and complex scoring mechanism based on logarithmic and exponential functions. In step S106, it is analyzed whether there is data competition, atomicity violation or memory order exception in the execution path. The behavior analysis result is represented based on the analysis result, including:

[0233] Step S1061, cross-scenario execution sequence aggregation and benchmark distribution construction.

[0234] In this step, the set of feasible execution sequences obtained in step S1056 is denoted as ε MM :

[0235]

[0236] Each set has several topologically sorted sequences The sequences from different models and scenarios are merged and labeled to form a unified execution distribution benchmark:

[0237]

[0238] By statistically analyzing the operation event distribution in , the benchmark event frequency mapping π(e) and the benchmark causal order mode distribution Π(o a ,o b ) are constructed.

[0239] Step S1062, definition of abnormality measurement function.

[0240] In this step, in order to quantitatively measure the significance of abnormal behavior in an operation sequence, the present application directly defines a multi-dimensional abnormality scoring function

[0241] Let the event sequence in a given execution sequence be (e1, e2, …, e m ), consider the following factors:

[0242] Deviation from benchmark distribution: if the event combination (e a , e b ) has low probability in the benchmark Π(o a , o b ), it has abnormal tendency.

[0243] The order that cannot be repeated in multiple simulation scenarios, such as special out-of-order caused by memory model relaxation, will lead to higher abnormal scores.

[0244] Define the abnormal score:

[0245]

[0246] where w(e a ,e b ) is a weight function, and the event type and the previous scene selection function are weighted according to the extension result. Since -logΠ amplifies the weight of rare sequences, it highlights abnormal scenarios.

[0247] Further introduce an exponential decay factor to modulate the sequence length and the density of abnormal points:

[0248]

[0249] where α>0, represents the density of specific abnormal features, such as ABA pattern repetition, memory barrier absence, etc. Through the combination of logarithm and exponential, the abnormal score function increases sharply in rare event patterns, making weak probability abnormal behaviors that cannot be easily captured by standard methods stand out.

[0250] Step S1063, re-verification of abnormal patterns based on symbolic constraint refinement.

[0251] In this step, high abnormal score sequences extract the sub-interval of interest, which contains the paragraph of potential ABA or memory consistency problems, and reload the operation dependency track and its symbolic execution constraints corresponding to the sub-interval of interest to the SMT solver for constraint re-verification:

[0252]

[0253] where is the symbolic constraint of a single operation, such as access value condition, CAS comparison condition, synchronization requirement, etc. If is not satisfied (unsat), it means that the abnormal sequence cannot be established at the symbolic level, and it may be a false positive produced only based on frequency statistics in timing analysis; if is satisfiable (sat) and consistent with the high abnormal score, it confirms the realizability of its abnormal scenario.

[0254] Step S1064, iterative optimization and block fusion of timing difference degree.

[0255] In this step, after confirming that part of the high-score abnormal sequences are feasible, the difference degree analysis is performed on the key operation paragraphs in these sequences. The difference degree function is defined:

[0256]

[0257] Difference degree For measuring the deviation of the sub-sequence in the causal relationship, i.e., the probability of double-event combination, from the single-event benchmark frequency. When the sub-sequence is extremely rare in causality and logically feasible, it has a significant abnormal feature. For high-difference-degree segments, the block fusion method is used to combine and repeatedly calculate them with adjacent sub-sequences to eliminate noise or small abnormalities, so as to refine more stable abnormal blocks.

[0258] Step S1065, multi-level abnormal fingerprint extraction and category separation.

[0259] In this step, based on the abnormal block set after block fusion For each abnormal block, the feature vector is extracted f(b j ) based on the operation type distribution (read-write ratio or CAS number), memory model difference degree (the frequency of the block appearing under different memory models), and symbol constraint complexity. To realize category separation (ABA, memory consistency or sequential conflict), the classification auxiliary metric based on Kullback-Leibler divergence is directly defined in this application:

[0260]

[0261] Among them, f(class) is used to represent the typical feature distribution of a certain abnormal category (such as ABA). By comparing KL(f(b j )||f(class)), the category with smaller KL divergence is selected as the preliminary classification result.

[0262] Step S1066, abnormal category maximum likelihood division under non-linear optimization solution.

[0263] In this step, after obtaining the preliminary classification result based on KL divergence, the maximum likelihood classification strategy is introduced for each abnormal block b j The abnormal category set For each category The likelihood value is calculated respectively:

[0264] L(c|b j )∝exp.-KL(f(b j )||f(c))+β·S(b j ,c) /

[0265] Among them, S(bj c) is the category-adapted score based on the scene and anomaly score weighting (based on the anomaly score S1062 ), and β is a tuning parameter.

[0266] The maximum likelihood solution in the category space is sought by a nonlinear optimization solver, either based on gradient descent or quasi-Newton methods:

[0267]

[0268] Step S1067, anomaly behavior classification result output and trigger condition calibration.

[0269] In this step, after completing the maximum likelihood classification, the determination result of each anomaly block is merged into the global view, and the trigger condition of the associated thread, shared variable and key operation (CAS point, read-after-write point, read-without-synchronization point) is calibrated. The final output result of this step includes:

[0270] The precise definition and instance list of each type of anomaly (ABA, memory consistency problem, operation sequence conflict, etc.); The trigger condition (such as instruction reordering under a specific memory model, specific thread interaction mode, specific value change sequence) and operation conflict point are marked for each anomaly point.

[0271] In this way, through the anomaly score function and the symbolic constraint mechanism in this step, various undefined behaviors can be accurately located and classified, providing strong support for developers and helping them effectively optimize programs and improve the stability and security of the system.

[0272] Figure 2 Another flowchart of the method for multi-thread undefined behavior detection of the embodiment of the application is shown. As shown in Figure 2 Another flow of the method for multi-thread undefined behavior detection of the embodiment of the application includes:

[0273] Step S201, constructing a multi-thread execution path model.

[0274] For the specific implementation of this step, please refer to the content of step S101.

[0275] Step S202, extracting operation sequences based on the Map stage.

[0276] For the specific implementation of this step, please refer to the content of step S102.

[0277] Step S203, constructing a causal relationship graph based on the Reduce stage.

[0278] For the specific implementation of this step, please refer to the content of step S103.

[0279] Step S204, extending the path dependence based on the causal relationship graph.

[0280] For the specific implementation of this step, please refer to the content of step S104.

[0281] Step S205, simulating out-of-order execution under weak memory model.

[0282] For the specific implementation of this step, please refer to the content of step S105.

[0283] Step S206, abnormal behavior analysis and classification.

[0284] For the specific implementation of this step, please refer to the content of step S106.

[0285] Figure 3 The flowchart of the Map phase and the Reduce phase of the embodiment of the application is shown. As shown in the figure, the data processing flow of the Map phase and the Reduce phase includes: Figure 3

[0286] Step S301, constructing a multi-thread execution path model.

[0287] Step S302, based on the multi-thread execution path model, performing path task decomposition and sharding organization.

[0288] Step S303, distributed task allocation.

[0289] Step S304, operation sequence extraction and labeling.

[0290] Step S305, timestamp synchronization and lock coordination.

[0291] The above steps S301 to S305 are the operation sequence extraction of the Map phase.

[0292] Step S306, based on steps S301 to S305, obtaining an operation sequence dataset.

[0293] Step S307, based on the operation sequence dataset, performing global timestamp normalization.

[0294] Step S308, based on the normalization result, performing partial order relationship construction.

[0295] Step S309, based on the partial order relationship, performing conflict correction and out-of-order correction.

[0296] Step S310, based on the correction and correction results, obtaining a causal relationship graph.

[0297] ​An apparatus for multi-thread undefined behavior detection is also provided in the embodiments, which is configured to implement the above-described embodiments and preferred embodiments, and will not be described herein again. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, implementation in hardware or a combination of software and hardware is also possible and contemplated.

[0298] An apparatus for multi-thread undefined behavior detection is also provided in the embodiments, which is configured to implement the above-described embodiments and preferred embodiments, and will not be described herein again. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, implementation in hardware or a combination of software and hardware is also possible and contemplated. Figure 4 A structural schematic diagram of the apparatus for multi-thread undefined behavior detection provided by the embodiments of the present application is shown in FIG. 4, which includes: Figure 4 As shown in FIG. 4, the apparatus includes:

[0299] A first module 401 is configured to construct a control flow graph and a data flow graph based on multi-thread program code, and create a multi-thread execution path model based on the control flow graph and the data flow graph.

[0300] A second module 402 is configured to assign thread slice data sets in the multi-thread execution path model to a plurality of computing nodes, and extract operation sequences in the assigned thread slice data sets in each computing node.

[0301] A third module 403 is configured to perform iterative and hierarchical processing on the operation sequences, and construct a causal relationship graph.

[0302] A fourth module 404 is configured to generate operation dependency tracks based on the causal relationship graph, map the operation dependency tracks to different memory models, perform execution order simulation to obtain a plurality of execution paths, and perform behavior analysis based on the execution paths to obtain behavior analysis results.

[0303] In some optional embodiments, the first module 401 includes:

[0304] A first unit of the first module is configured to perform static analysis on the multi-thread program code, extract instruction sequences and basic blocks within threads, construct a control flow graph based on the basic blocks, identify data operation nodes in the instruction sequences, construct a data flow graph based on the data operation nodes, mark shared variable access instructions and synchronization operations in the control flow graph and the data flow graph respectively to obtain access labels, determine inter-thread dependency relationships based on the control flow graph, the data flow graph and the access labels, divide the threads into a plurality of thread slices based on a path allocation strategy, the control flow graph, the data flow graph and the dependency relationships, and determine a multi-thread execution path model based on the thread slices in each thread and the inter-thread dependency relationships.

[0305] In some optional embodiments, the second module 402 includes:

[0306] The second module first unit is configured to assign, based on the mapping function, the thread slice data set in the multi-thread execution path model to the plurality of computing nodes; and extract, in each computing node, the operation sequence based on the time sequence.

[0307] In some optional embodiments, the third module 403 includes:

[0308] The third module first unit is configured to fuse, based on the time sequence, the operation sequence corresponding to the time sequence and the marking information corresponding to the time sequence to obtain a global time sequence operation sequence; map the global time sequence operation sequence to a partial order relation, and represent the initial causal constraint based on the partial order relation; and correct the initial causal constraint to obtain a causal relationship graph.

[0309] In some optional embodiments, the fourth module 404 includes:

[0310] The fourth module first unit is configured to extract, based on the nodes and edges in the causal relationship graph, the operations of each operation sequence on the shared variable to obtain a memory state mapping of each operation node; perform constraint modeling based on the memory state mapping, determine the path priority based on the constraint modeling result, and perform recursive expansion on the path whose path priority is higher than a preset path priority threshold to generate an operation dependency track.

[0311] In some optional embodiments, the fourth module 404 further includes:

[0312] The fourth module second unit is configured to abstract each memory model into a triple, wherein the triple includes a set of event types, a set of ordered relations, and a set of consistency constraints; perform event type mapping and relation mapping of the operation dependency track to the memory model; construct a constraint system, and perform constraint solving based on the constraint system; and perform topological sorting on the partial order relation graph after the constraint solving to obtain a plurality of execution paths.

[0313] In some optional embodiments, the fourth module 404 further includes:

[0314] The fourth module third unit is configured to analyze whether there is a data race, an atomicity violation, or a memory order exception in the execution path, and represent the behavior analysis result based on the analysis result.

[0315] Further function descriptions of each of the above modules and units are the same as those of the corresponding embodiments described above, and will not be described here.

[0316] The device for detecting multi-threaded undefined behavior in this embodiment is presented in the form of a functional unit, where the unit refers to an application-specific integrated circuit (ASIC) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above-mentioned functions.

[0317] The embodiment of the present invention also provides a computer device having the above Figure 4 The apparatus for detecting multi-threaded undefined behavior is shown.

[0318] See also Figure 5 , Figure 5 is a structural diagram of a computer device provided by an optional embodiment of the present invention, such as Figure 5 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components utilize different buses to communicate with each other and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of a graphical user interface on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Equally, multiple computer devices can be connected, and each device provides part of the necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 5 A processor 10 is taken as an example.

[0319] The processor 10 may be a central processing unit, a network processor, or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.

[0320] The aforementioned memory 20 stores instructions that can be executed by at least one processor 10, so that the aforementioned at least one processor 10 executes the method shown in the above embodiment.

[0321] The memory 20 can include a program storage area and a data storage area. The program storage area can store an operating system, application programs required for at least one function, and the like. The data storage area can store data created according to usage of the computer device, and the like. In addition, the memory 20 can include a high-speed random access memory, and can further include a non-transitory memory such as at least one of a magnetic disk storage device, a flash memory device, or other non-transitory solid state memory device. In some alternative embodiments, the memory 20 can optionally include a memory that is remotely located with respect to the processor 10, and these remote memories can be connected to the computer device through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0322] The memory 20 can include a volatile memory such as a random access memory, and can further include a non-volatile memory such as a flash memory, a hard disk, or a solid state disk. The memory 20 can also include a combination of the above-mentioned types of memories.

[0323] The computer device also includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30, and the output device 40 can be connected through a bus or other means, Figure 5 The connection through the bus is used as an example.

[0324] The input device 30 can receive input digital or character information, and generate key signal inputs related to user settings and function controls of the computer device. Examples of the input device 30 include, but are not limited to, a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, and the like. The output device 40 can include a display device, an auxiliary lighting device (such as a light emitting diode), a tactile feedback device (such as a vibration motor), and the like. Examples of the display device include, but are not limited to, a liquid crystal display, a light emitting diode, a display, and a plasma display. In some alternative embodiments, the display device can be a touch screen.

[0325] The embodiments of the present application further provide a computer readable storage medium, and the method according to the embodiments of the present application can be implemented in hardware, firmware, or recorded in a storage medium, or stored in a remote storage medium or a non-transitory machine readable storage medium and downloaded to a local storage medium through network, so that the method described herein can be processed by such software on a storage medium using a general purpose computer, a special purpose processor, or programmable or special hardware. The storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid state disk, etc. Further, the storage medium can also include a combination of the above-mentioned memories. It can be understood that the computer, the processor, the microprocessor controller, or the programmable hardware includes a storage component that can store or receive software or computer code, when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.

[0326] Part of the present application can be applied as a computer program product, for example, computer program instructions, when executed by a computer, through the operation of the computer, the method and / or technical solutions according to the present application can be called or provided. Those skilled in the art should understand that the form of computer program instructions in a computer readable medium includes but is not limited to source files, executable files, installation package files, etc. Correspondingly, the way of executing computer program instructions by computer includes but is not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Here, the computer readable medium can be any available computer readable storage medium or communication medium accessible to the computer.

[0327] Although the embodiments of the present application are described in conjunction with the accompanying drawings, various modifications and changes can be made by those skilled in the art without departing from the spirit and scope of the present application, and such modifications and changes fall within the scope defined by the appended claims.

Claims

1. A method for detecting multi-threaded undefined behavior, characterized in that: The method comprises: Constructing a control flow graph and a data flow graph based on the multi-threaded program code, and creating a multi-threaded execution path model based on the control flow graph and the data flow graph; Allocating the thread-sliced ​​data set in the multi-threaded execution path model to a plurality of computing nodes, and extracting the operation sequence in the allocated thread-sliced ​​data set in each computing node; Iterating and hierarchically processing the operation sequence to construct a cause-effect relationship graph; Based on the causal relationship graph, an operation dependency trajectory is generated, the operation dependency trajectory is mapped to different memory models, and execution sequence simulation is performed to obtain multiple execution paths. Behavior analysis is performed based on the execution paths to obtain behavior analysis results.

2. The method according to claim 1, characterized in that The method of constructing a control flow graph and a data flow graph based on the multi-threaded program code, and creating a multi-threaded execution path model based on the control flow graph and the data flow graph, includes: Performing static analysis on the multi-threaded program code, extracting instruction sequences and basic blocks within the threads, constructing a control flow graph based on the basic blocks, identifying data operation nodes in the instruction sequences, and constructing a data flow graph based on the data operation nodes; Marking shared variable access instructions and synchronization operations in the control flow graph and the data flow graph respectively to obtain access tags, and determining dependencies between threads based on the control flow graph, the data flow graph, and the access tags; Based on the path allocation strategy, the control flow graph, the data flow graph and the dependency relationship, the thread is divided into multiple thread slices, and based on the thread slices in each thread and the dependency relationship between the threads, the multi-thread execution path model is determined.

3. The method according to claim 1 or 2, characterized in that The step of distributing the thread slice data set in the multi-threaded execution path model to a plurality of computing nodes and extracting the operation sequence in the distributed thread slice data set in each computing node includes: Based on a mapping function, the thread-sharded data set in the multi-threaded execution path model is distributed to the plurality of computing nodes; In each of the computing nodes, the operation sequence is extracted based on a time sequence.

4. The method according to claim 3, characterized in that The stepwise iteration and hierarchical processing of the operation sequence to construct a causal relationship graph includes: Based on the time sequence, fusing the operation sequence corresponding to the time sequence and the tag information corresponding to the time sequence to obtain a global time sequence operation sequence; Mapping the global temporal operation sequence into a partial order relationship, and representing the initial causal constraint based on the partial order relationship; The initial causal constraints are corrected to obtain the causal relationship graph.

5. The method according to claim 1, wherein Generating an operation dependency trace based on the causal relationship graph includes: Extracting operations on shared variables by each operation sequence based on the nodes and edges in the causal relationship graph to obtain a memory state map of each operation node; Performing constraint modeling based on the memory state mapping, and determining path priority based on the constraint modeling result; Recursively expand the paths whose path priorities are higher than a preset path priority threshold to generate the operation dependency trajectory.

6. The method according to claim 1, characterized in that The operation dependency trace is mapped to different memory models, and execution sequence simulation is performed to obtain multiple execution paths, including: Abstracting each of the memory models into a triple, wherein the triple includes an event type set, an ordered relationship set, and a consistency constraint set; Performing event type mapping and relationship mapping from the operation dependency trace to the memory model; Constructing a constraint system, and performing constraint solving based on the constraint system; The partial order relationship graph after the constraint solution is topologically sorted to obtain a plurality of execution paths.

7. The method according to claim 1, characterized in that The performing behavior analysis based on the execution path to obtain a behavior analysis result includes: The execution path is analyzed to determine whether data contention, atomicity violation, or memory order anomaly exists, and the behavior analysis result is characterized based on the analysis result.

8. A device for detecting multi-threaded undefined behavior, characterized in that: The device comprises: The first module is used to construct a control flow graph and a data flow graph based on the multi-threaded program code, and create a multi-threaded execution path model based on the control flow graph and the data flow graph; The second module is used to distribute the thread slice data set in the multi-threaded execution path model to multiple computing nodes, and extract the operation sequence in the distributed thread slice data set in each computing node; The third module is used to iterate and hierarchically process the operation sequence to construct a causal relationship graph; The fourth module is used to generate an operation dependency trajectory based on the causal relationship diagram, map the operation dependency trajectory to different memory models, perform execution sequence simulation, obtain multiple execution paths, perform behavioral analysis based on the execution paths, and obtain behavioral analysis results.

9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method for detecting multi-threaded undefined behavior according to any one of claims 1 to 7 by executing the computer instructions.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method for detecting multi-threaded undefined behavior according to any one of claims 1 to 7.