Method and system for confronting symbolic execution attack based on dual opaque predicates

By introducing double opaque predicates in symbol execution attacks, combining feedforward neural network and discrete logarithmic problems, multi-level path judgment conditions are constructed, and the problem of poor protection effect of single opaque predicates is solved, which significantly improves the protection effect of symbol execution attacks.

CN120145387APending Publication Date: 2025-06-13HENAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510168022.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-15
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the prior art, when facing complex conditions or nonlinear logic, the calculation efficiency of symbol execution tools is seriously affected, and a single highly nonlinear opaque predicate cannot provide a comprehensive protection effect.

Method used

Using a method based on double opaque predicates, combined with the non-interpretation of feedforward neural networks and the computational complexity of discrete logarithmic problems, a multi-level path judgment condition is constructed, so that symbol execution tools face extremely high computational costs and complexity in the parsing process.

Benefits of technology

It effectively improves the protection effect of symbolic execution attacks, increases the resolution resistance of the path, limits the attacker's complete resolution of the path, and significantly enhances the anti-symbol execution ability of the program.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145387A_ABST
    Figure CN120145387A_ABST
Patent Text Reader

Abstract

The invention discloses a method and system for confronting symbolic execution attacks based on double opaque predicates, and the method comprises the steps: 1, constructing a feedforward neural network which comprises an input layer, a plurality of hidden layers, and an output layer; 2, constructing a discrete logarithm problem of formalized linear transformation; and 3, setting a dual opaque predicate scheme, and inserting the feedforward neural network and the discrete logarithm problem of formalized linear transformation into a target program according to the dual opaque predicate scheme. According to the method, multi-level path judgment conditions are constructed by combining the non-interpretability of the feedforward neural network and the calculation complexity of the discrete logarithm problem, so that a symbolic execution tool faces extremely high calculation cost and complexity in the analysis process, and the protection effect on symbolic execution attacks is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of software security technology, and particularly to a method and system for combating symbolic execution attacks based on dual opaque predicates. Background Art

[0002] With the increasing requirements for software security, symbolic execution has gradually become an important program analysis method, which is widely used in fields such as software testing, vulnerability discovery, and program verification. The symbolic execution technology analyzes different paths of a program by treating program variables as symbolic variables, so as to discover potential vulnerabilities. However, when the program scale increases, symbolic execution will face the "path explosion" problem, that is, the number of paths grows exponentially, making the computational cost of comprehensively exploring program paths unacceptable. In addition, when encountering complex conditions or non-linear logics, even with the introduction of advanced solvers, the computational efficiency of symbolic execution tools will be severely affected. Facing these challenges, program obfuscation technology, especially opaque predicates, has gradually become an effective means to combat symbolic execution.

[0003] As a common program obfuscation method, an opaque predicate aims to hinder the path resolution of symbolic execution tools by constructing complex conditions that cannot be parsed, thereby increasing the anti-analysis property of the program. Traditional opaque predicates are usually based on simple mathematical complexity or non-linear expressions. However, with the development of symbolic execution tools, these techniques have gradually become insufficiently powerful and difficult to provide sufficient defense effects. Although a single highly non-linear opaque predicate can prevent the resolution of some paths, it still cannot provide comprehensive protection in the context of the continuous progress of symbolic execution tools. Summary of the Invention

[0004] In order to at least partially solve the problem that the protection effect of a single highly non-linear opaque predicate against symbolic execution attacks is not good, the present invention provides a method and system for combating symbolic execution attacks based on dual opaque predicates. By combining the inexplicability of a feed-forward neural network and the computational complexity of the discrete logarithm problem, a multi-level path judgment condition is constructed, making the symbolic execution tool face extremely high computational costs and complexities during the parsing process, and effectively improving the protection effect against symbolic execution attacks.

[0005] To achieve the above object, the technical solution of the present invention is:

[0006] The first aspect of the present invention proposes a method for combating symbolic execution attacks based on dual opaque predicates, including:

[0007] Step 1: Construct a feed-forward neural network, where the feed-forward neural network includes an input layer, multiple hidden layers, and an output layer, and is used to make it difficult for a symbolic execution tool to parse its path conditions according to the output characteristics of the feed-forward neural network;

[0008] Step 2: Construct a discrete logarithm problem for formal linear transformation to consume the computational resources of the symbolic execution tool according to the computational complexity of the discrete logarithm;

[0009] Step 3: Set up a double opaque predicate scheme. Insert the feedforward neural network and the discrete logarithm problem of formal linear transformation into the target program according to the double opaque predicate scheme. Through the double opaque predicate scheme, the parsing difficulty of the path condition increases exponentially, effectively improving the anti-parsing property of the path and enhancing the protection effect against symbolic execution attacks.

[0010] Furthermore, the input layer, multiple hidden layers, and output layer in the feedforward neural network all include multiple neurons and ReLU activation functions.

[0011] Furthermore, the construction of the discrete logarithm problem of formal linear transformation specifically includes:

[0012] Set large prime numbers, generators, target values, results, and unknowns;

[0013] Set up a loop for solving the unknown according to the large prime number, generator, target value, result, and unknown to consume the computational resources of the symbolic execution tool.

[0014] Furthermore, the setting up of a loop for solving the unknown according to the large prime number, generator, target value, initialized result, and initialized unknown specifically includes:

[0015] Step 1: Determine whether the unknown is less than the large prime number. If the unknown is greater than the large prime number, end the loop; if the unknown is less than the large prime number, continue the loop;

[0016] Step 2: Calculate and update the result according to a preset formula for formalizing the transformation of the discrete logarithm condition, making it seemingly a linear structure on the surface, thereby inducing the symbolic execution tool to directly perform symbolic solution and fall into the dilemma of exhausting computational resources;

[0017] Step 3: Determine whether the result is equal to the target value. If the result is equal to the target value, jump out of the loop; if the result is not equal to the target value, increment the unknown by 1 and return to Step 1.

[0018] Furthermore, the preset formula is specifically expressed as follows:

[0019] result += (result + g - 1) % p

[0020] where result is the result, g is the generator, % is the modulo operation, and p is the large prime number.

[0021] Furthermore, the double opaque predicate scheme specifically includes:

[0022] Embed a feedforward neural network and a formally linearly transformed discrete logarithm problem into the target program to construct a path condition;

[0023] Connect the feedforward neural network and the formally linearly transformed discrete logarithm problem through a logical "AND" operator. When a tool for symbolic execution attack attacks the target program, it needs to parse the feedforward neural network. After obtaining the input that satisfies the output condition of the feedforward neural network, it is also required to solve the formally linearly transformed discrete logarithm problem to complete the path parsing, which is convenient for improving the protection effect against symbolic execution attacks;

[0024] Set multiple false branches and logical short circuits in the path condition to increase the computational burden of the symbolic execution tool.

[0025] Furthermore, the conditions for the output of the feedforward neural network include:

[0026] The output value of the feedforward neural network is greater than 1. By restricting the output value range, the parsing burden of the symbolic execution tool is increased, hindering its further analysis of the path.

[0027] The second aspect of the present invention proposes a system for countering symbolic execution attacks based on double opaque predicates, including:

[0028] A feedforward neural network module for constructing a feedforward neural network, which includes an input layer, multiple hidden layers, and an output layer, and is used to make it difficult for a symbolic execution tool to parse its path condition according to the output characteristics of the feedforward neural network;

[0029] A discrete logarithm module for constructing a formally linearly transformed discrete logarithm problem and consuming the computational resources of the symbolic execution tool according to the computational complexity of the discrete logarithm;

[0030] An embedding module for setting a double opaque predicate scheme, inserting the feedforward neural network and the formally linearly transformed discrete logarithm problem into the target program according to the double opaque predicate scheme, and making the parsing difficulty of the path condition increase exponentially through the double opaque predicate scheme, effectively improving the anti-parsing property of the path and enhancing the protection effect against symbolic execution attacks.

[0031] The third aspect of the present invention proposes an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the method for countering symbolic execution attacks based on double opaque predicates as described in the first aspect above.

[0032] A fourth aspect of the present invention provides a computer-readable storage medium, which includes a stored computer program. When the computer program runs, it controls the device where the storage medium is located to execute the method for countering symbolic execution attacks based on dual opaque predicates as described in the first aspect above.

[0033] Advantages of the present invention:

[0034] (1) The present invention proposes a method for countering symbolic execution attacks based on dual opaque predicates. By combining a feedforward neural network and the discrete logarithm problem, a path condition that is extremely difficult to analyze is constructed. The "black box" characteristic of the neural network makes it difficult for symbolic execution tools to deduce its output, forming a complex condition that is difficult to solve. At the same time, the computational complexity of the discrete logarithm problem further increases the difficulty of solving by the tool, effectively restricting the attacker's complete analysis of the path and effectively increasing the path analysis difficulty of the symbolic execution tool.

[0035] (2) The present invention constructs a formalized linear transformation of the discrete logarithm problem. By disguising the original multiplication, division, and modulo operations in the discrete logarithm problem as addition, subtraction, and loop operations, the symbolic execution tool will misjudge it as a linear structure when symbolizing and directly try to solve it instead of calling a non-linear solver. In this way, when the tool processes this transformed condition, it still faces the problem of being difficult to solve, but due to misjudgment, the computational overhead is increased. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a flowchart of the method for countering symbolic execution attacks based on dual opaque predicates provided by an embodiment of the present invention.

[0037] Figure 2 It is a schematic diagram of the code of the standard discrete logarithm condition before transformation provided by an embodiment of the present invention.

[0038] Figure 3 It is a schematic diagram of the code of the formalized linear transformation of the discrete logarithm problem provided by an embodiment of the present invention.

[0039] Figure 4 It is an architecture diagram of the system for countering symbolic execution attacks based on dual opaque predicates provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.

[0041] Example 1

[0042] As Figure 1 shown, a method for countering symbolic execution attacks based on double opaque predicates includes:

[0043] S101: Construct a feedforward neural network, which includes an input layer, multiple hidden layers, and an output layer.

[0044] Specifically, due to its powerful learning ability and complex non-linear structure, neural networks have been widely used in various tasks in recent years, including image recognition, natural language processing, etc. Research shows that neural networks can approximate almost any non-linear function, and the more hidden layers there are, the stronger its modeling ability. However, a significant defect of neural networks is their "black box" property, that is, their internal decision-making process is difficult to explain and understand. Due to the complex internal architecture of the neural network model, the mapping relationship between input and output is difficult to intuitively explain by traditional logical or mathematical means. Existing technologies have proved that the rule extraction of neural networks is NP-hard (a problem to which all NP problems can be reduced within polynomial time complexity).

[0045] This black box property provides new possibilities for the construction of opaque predicates. Introducing neural network-based opaque predicates into a program can construct extremely complex and difficult-to-parse judgment conditions. For example, when a symbolic execution tool encounters a complex neural network function, it is difficult to understand the decision logic of the neural network through symbolic solution methods, thus increasing the difficulty of path parsing and computational cost. Even if the symbolic execution tool uses a non-linear solver, it is difficult to obtain useful information from the highly non-linear structure of the neural network, resulting in the tool's difficulty in judging the actual control flow direction of the program.

[0046] Therefore, the feedforward neural network structure constructed by the present invention includes:

[0047] Input layer: Accept symbolic random variables, which represent different inputs or states of the program.

[0048] Hidden layer: Design multiple hidden layers, at least 5 layers, each layer contains 64 neurons, and use the ReLU activation function to enhance the non-linear expression ability of the neural network.

[0049] Output layer: Use the Sigmoid activation function to limit the output of the feedforward neural network within the range of (0, 1).

[0050] Use randomly generated symbolic inputs for training, so that the feedforward neural network can generate complex non-linear outputs under diverse input conditions, ensuring that the dependence of the feedforward neural network output on the input is very complex and difficult to be parsed by the symbolic execution tool.

[0051] During the training process, the cross-entropy loss function and the Adam optimizer are used to optimize the network parameters to ensure that the model can exhibit sufficient complexity under different inputs.

[0052] Preferably, L2 regularization is added during training to avoid overfitting of the model and ensure that it can maintain complex non-linear characteristics on new data.

[0053] The dropout technique can be used to increase the randomness of the model and further enhance its complexity.

[0054] S102: Construct a discrete logarithm problem for formal linear transformation.

[0055] S103: Set a double opaque predicate scheme, and insert the feedforward neural network and the discrete logarithm problem for formal linear transformation into the target program according to the double opaque predicate scheme.

[0056] By combining the feedforward neural network and the discrete logarithm problem, the present invention sets a double opaque predicate scheme. The "black box" characteristic of the neural network makes it difficult for the symbolic execution tool to deduce its output, forming complex conditions that are difficult to solve. At the same time, the computational complexity of the discrete logarithm problem further increases the solving difficulty of the tool, effectively restricting the attacker's complete parsing of the path and effectively improving the protection effect against symbolic execution attacks.

[0057] Embodiment 2

[0058] Based on the above embodiments, the embodiment of the present invention provides a construction process of a discrete logarithm problem for formal linear transformation, specifically including:

[0059] The discrete logarithm problem is a mathematical problem with extremely high computational complexity and is widely used in the field of modern cryptography. Its formal definition is: given a large prime number p, a generator g, and a target value y, solve the unknown x such that g x mod p = y. Where mod is the modulo operation.

[0060] Currently, there is no effective algorithm that can solve the general discrete logarithm problem in polynomial time. This computational complexity makes the discrete logarithm problem the theoretical basis of many cryptographic algorithms (such as Diffie-Hellman key exchange, DSA digital signature, etc.).

[0061] Embedding the discrete logarithm condition in the program path judgment can greatly increase the solving difficulty of the symbolic execution tool. For example, when the symbolic execution tool needs to solve the path condition in the form of discrete logarithm, even with the assistance of an advanced solver, it still requires extremely high computational overhead, thus hindering the effective parsing of the tool. Compared with traditional mathematical complexity expressions, the discrete logarithm problem is extremely difficult to solve and can significantly increase the unsolvability of opaque predicates. The present invention constructs an opaque predicate with higher computational complexity by combining the discrete logarithm with the neural network output, thereby enhancing the anti-reverse ability of the program.

[0062] When the symbolic execution tool faces the discrete logarithm condition, if it recognizes the non-linear multiplication and modulo operations therein, it will usually hand them over to the non-linear solver for processing. However, in order to confuse the symbolic execution tool and increase its computational cost, the present invention formalizes the discrete logarithm condition to construct a formally linearly transformed discrete logarithm problem, making it seemingly a linear structure on the surface, thereby inducing the symbolic execution tool to directly perform symbolic solving and fall into the dilemma of exhausting computational resources.

[0063] As shown in the formally linearly transformed discrete logarithm problem, it specifically includes:

[0064] Set a large prime number, a generator, a target value, a result, and an unknown quantity.

[0065] According to the large prime number, the generator, the target value, the result, and the unknown quantity, set a loop for solving the unknown quantity, and this loop includes:

[0066] Step 1: Judge whether the unknown quantity is less than the large prime number. If the unknown quantity is greater than the large prime number, end the loop; if the unknown quantity is less than the large prime number, continue the loop.

[0067] Step 2: Calculate and update the result according to a preset formula. The preset formula is specifically expressed as the following formula:

[0068] result += (result + g - 1) % p

[0069] where result is the result, g is the generator, % is the modulo operation, and p is the large prime number.

[0070] Step 3: Judge whether the result is equal to the target value. If the result is equal to the target value, jump out of the loop; if the result is not equal to the target value, increment the unknown quantity by 1 and return to Step 1.

[0071] Specifically, by disguising the original multiplication, division, and modulo operations as addition, subtraction, and loop operations, the symbolic execution tool may misjudge them as linear structures during symbolization and directly attempt to solve them instead of calling a non-linear solver. In this way, when the tool processes this transformed condition, it still faces the problem of being difficult to solve, but due to the misjudgment, the computational overhead is increased. The following shows the standard discrete logarithm condition before transformation (as shown in Figure 2 ), code snippet 1) and the transformed linear disguise condition (the formalized linear transformation of the discrete logarithm problem, as shown in Figure 3 ), code snippet 2). In code snippet 1, the program attempts to gradually calculate g x mod p to find x that satisfies g x mod p = y. If the symbolic execution tool recognizes multiplication and modulo operations, it usually hands this condition over to a non-linear solver for parsing. In code snippet 2, the multiplication and modulo operations are formalized as addition and loop operations, causing the symbolic execution tool to mistakenly think that this condition is linear during symbolic solution and thus directly attempt to solve it instead of using a non-linear solver. During the solution process, the tool still needs to handle addition and loops, but because it is essentially a complex discrete logarithm problem, the symbolic execution tool will get stuck in a situation of exhausting computational resources when parsing this condition.

[0072] To make the formalized linear transformation of the discrete logarithm problem achieve the expected effect, the present invention analyzes the probability of the equation holding, that is, randomly selects qualified x and y and details the probability of satisfying g x mod p = y. According to number theory principles, the result of g x mod p = y is randomly distributed in the interval [0, p - 1]. Therefore, the possibility of satisfying the condition g x mod p = y is like randomly selecting a number in the interval that is consistent with the target. Since the result generated by the discrete logarithm is evenly distributed between [0, p - 1], the probability of this equation holding is approximately 1 / (p - 1). Therefore, in practical applications, only need to select a sufficiently large p, for example, p = 2 63 , then the probability of g x mod p = y is almost close to zero. When the symbolic execution tool solves this condition, it calculates a large number of possible values but cannot find the input value that satisfies the condition, resulting in a sharp increase in computational resources and even computational collapse, thus effectively preventing the symbolic execution tool from further parsing the program path.

[0073] Example 3

[0074] Based on the above embodiments, the embodiment of the present invention provides a dual opaque predicate scheme, which specifically includes:

[0075] To further improve the anti-symbolic execution ability, the dual opaque predicate scheme combines the output of a feedforward neural network (hereinafter referred to as neural network) and the discrete logarithm condition into a nested path judgment structure. Specifically, the two are connected by a logical "AND" operator, so that after the symbolic execution tool satisfies the neural network condition, it must further solve the discrete logarithm condition to complete the path parsing, thus forming a highly complex path condition structure.

[0076] During the normal execution of the program, if the first condition is not satisfied, the program skips the second condition, avoiding unnecessary complex calculations, thereby reducing the performance overhead of actual operation. Based on the above logic, the core of the dual opaque predicate scheme of the present invention lies in utilizing the short-circuit effect of the logical "AND" operator. During the normal operation of the program, if the neural network condition is false (the output is less than or equal to 1), the program will directly skip the second condition judgment and proceed with subsequent operations, avoiding additional computational overhead. This short-circuit mechanism ensures that while the scheme improves anti-parsing ability, it reduces the performance burden of normal execution. The specific steps are as follows:

[0077] Take the output of the neural network as the first-layer judgment condition of the path, and connect it with the discrete logarithm condition using the logical "AND" operator. If the neural network condition is true (the output is greater than 1), enter the discrete logarithm judgment; otherwise, skip directly. This design ensures that the symbolic execution tool gradually increases the computational burden during the parsing process.

[0078] Specifically, use the output of the feedforward neural network as the path judgment condition in the program. The specific judgment method is to compare the output of the neural network with the set threshold 1:

[0079] If the output is greater than 1, the program solves the discrete logarithm problem of formal linear transformation.

[0080] If the output is less than or equal to 1, the program enters other paths (executes the subsequent program in the target program).

[0081] The output range of the Sigmoid function is (0, 1). However, due to the non-linearity and deep structure of the neural network, it is difficult for the symbolic execution tool to effectively parse the specific input that satisfies the condition. Therefore, the symbolic execution tool will repeatedly parse the feedforward neural network to meet the output condition, which increases the insolubility of the path condition.

[0082] This design ensures the security of the program, enhances the unpredictability of the path condition, and effectively hinders the symbolic execution tool from parsing the program path.

[0083] With this design, the output of the neural network can be compared with a fixed threshold of 1 while maintaining the complexity and non-analyzability of the network, making it more difficult for symbolic execution tools to parse the path judgment conditions, thereby enhancing the security of program analysis.

[0084] Preferably, the present invention adds false branches and redundant conditions to the path structure of the dual opaque predicate scheme to further interfere with the path parsing of symbolic execution tools, making them encounter more unnecessary branches and increasing the computational cost.

[0085] With this design of the present invention, the combined conditions of dual opaque predicates add multi-level computational obstacles when symbolic execution tools parse paths. The tool constantly faces high computational amounts and complex path structures when judging conditions layer by layer, thereby significantly enhancing the anti-symbolic execution ability of the program.

[0086] Embodiment 4

[0087] Based on the above embodiments, the embodiments of the present invention provide a process of how to insert a feedforward neural network and a formally linearized discrete logarithm problem into a target program, specifically including:

[0088] In order to implement the call of the neural network model in C++ code and use it as a path judgment condition, the trained neural network model needs to be converted into the TorchScript format. The model is saved as a.pt file through PyTorch so that it can be loaded through the TorchScript API in C++ code and forward propagated. The converted TorchScript file will be called in LLVM IR to generate the neural network judgment condition in the opaque predicate. The model contains a multi-layer fully connected structure, and the nonlinearity of the output is enhanced through ReLU and Sigmoid activation functions to ensure that the neural network judgment is difficult to be parsed by symbolic execution tools.

[0089] After preparing the TorchScript model, use clang in the LLVM toolchain to compile the target C++ code into an LLVM IR file (i.e., a.ll file). Through the -emit-llvm option, the C++ source code is converted into the LLVM IR format, and the generated IR file is stored in text format. This file provides the basis for subsequent Pass insertion and allows direct insertion and modification operations at the IR level.

[0090] By customizing the LLVM Pass, dual opaque predicate logic can be inserted into the IR file of the target program. The task of the Pass is to traverse the functions and basic blocks in the IR file, find suitable insertion positions, and generate the required IR code. The IR code includes the combination of neural network judgment and linearized discrete logarithm conditions. The above process specifically includes:

[0091] Locate the insertion point: The custom Pass traverses the IR structure to locate the basic blocks and instructions suitable for inserting opaque predicates, such as conditional branches or loop positions.

[0092] Construct the opaque predicate logic: At the insertion point, use LLVM's IRBuilder to generate IR code, call the TorchScript model for forward propagation to generate the neural network output as the first-layer condition. Subsequently, construct a linearized discrete logarithm judgment in the form of addition and loops, disguising it as a linear structure to mislead symbolic execution tools. Connect the two with the logical "AND" operator to form a nested condition of nn(x)>1&&findDiscreteLogLinear(x, y).

[0093] Insert the opaque predicate logic: Insert the generated opaque predicate IR code at the appropriate position to make the parsing of the target program path more complex, thereby enhancing the anti-symbolic execution ability.

[0094] After the custom Pass is written, use LLVM's opt tool to apply the Pass to the target IR file. By loading the Pass and applying it to the IR file, a new IR file containing the opaque predicate logic is generated. This process ensures that opaque predicates are inserted in the critical path of the program, thereby enhancing the obfuscation effect.

[0095] The generated new IR file contains dual opaque predicate logic and can be further compiled into the final executable file through llc and clang. Based on the opaque predicate logic, the anti-symbolic execution ability of the final program is enhanced.

[0096] After completing the insertion of the opaque predicate logic, use llc to compile the IR file into assembly code, and then link it with the Torch library through clang to generate a binary executable file containing the opaque predicate. This file calls the TorchScript model at runtime to implement complex path judgment of dual conditions, forming an effective parsing obstacle for symbolic execution tools.

[0097] The dual opaque predicate scheme proposed by the present invention cleverly combines the non - linear output of the feed - forward neural network and the complexity of the discrete logarithm problem. The feed - forward neural network part is responsible for generating unpredictable outputs, increasing the complexity of the path condition; while the linearization of the discrete logarithm problem misleads the symbolic execution tool into misjudging the path condition as a linear structure. In addition, by utilizing the short - circuit characteristic of the logical AND statement in programming, the present invention ensures that even if the discrete logarithm problem is difficult to solve, it will not affect the normal operation of the program, and at the same time effectively interferes with the attacks of the symbolic execution tool. This combination method not only enhances the insolubility of the path condition but also improves the anti - analysis ability of the program. Different from the traditional methods that rely only on a single mathematical problem or static logical conditions, the dual opaque predicate introduces two distinct obfuscation strategies, achieving a more powerful and diverse obfuscation effect.

[0098] Embodiment 5

[0099] Based on the above - mentioned embodiments, the embodiment of the present invention provides an experimental evaluation process of the present invention, which specifically includes:

[0100] To ensure the scientific nature of the experiment and the repeatability of the results, this experiment is carried out in a multi - platform environment. The compilation and instruction count statistics experiment are carried out in the Ubuntu 20.04 system, equipped with 4GB of memory and 60GB of storage, and LLVM12.0.1 is installed to generate different versions of binary files and perform IR insertion. The Intel PIN tool also runs in this environment to record the changes in the code instruction count. The use of the LLVM compiler ensures the stability of code generation before and after inserting opaque predicates.

[0101] The symbolic execution resistance test is also carried out on the Ubuntu system, using the Angr tool. Angr is a powerful symbolic execution engine used to verify the resistance of the program after inserting opaque predicates to symbolic execution parsing. This test environment ensures that the Angr tool can obtain accurate data when processing the obfuscated and non - obfuscated codes to test the anti - parsing effect of the opaque predicates.

[0102] The similarity analysis and reverse - engineering experiment are carried out in the Windows 11 system, using IDAPro 7.5 and BinDiff 5.0 to perform detailed disassembly and comparison on the binary codes before and after obfuscation. The experimental system under the Windows platform is also configured with 4GB of memory and 60GB of storage to support the analysis of complex codes and similarity measurement.

[0103] Six classic sorting algorithms were selected as the experimental benchmarks: bubble sort, selection sort, insertion sort, Shell sort, merge sort, and quick sort. For each algorithm, both the original version and the obfuscated version embedded with dual opaque predicates were generated. In the obfuscated version, a path determination combining a non-linear condition generated by a neural network and the solution of discrete logarithm was introduced to increase the complexity of the code logic and path uncertainty.

[0104] The experimental scheme evaluated the effects of opaque predicates from multiple dimensions, including program running time, binary similarity, instruction count statistics, and symbolic execution resistance testing.

[0105] The program running time test aimed to quantify the impact of opaque predicates on program execution efficiency. The running time of each sorting algorithm was measured using the time command, and the execution times before and after obfuscation were compared. The average value was taken from multiple experiments to analyze whether the computational overhead brought by introducing the neural network and discrete logarithm solution was within an acceptable range.

[0106] Binary similarity analysis evaluated the degree of change in the code structure caused by the obfuscation process. The IDA Pro was used to disassemble the code, and BinDiff was used to compare the similarity of the assembly code before and after obfuscation to determine whether the obfuscation technique successfully changed the code structure and increased the difficulty of reverse engineering.

[0107] Instruction count statistics recorded the instruction counts before and after obfuscation through the Intel PIN tool, reflecting the increase in code complexity. The change in the instruction count can intuitively show how the dual opaque predicates increase the complexity of the code path and make reverse engineering analysis more difficult.

[0108] The symbolic execution resistance test used the Angr tool to explore the paths of the code before and after obfuscation to verify the effect of opaque predicates in preventing symbolic execution from resolving paths. The official example code of Angr and the obfuscated version were used in the experiment to compare the success rates of the symbolic execution tool in resolving paths.

[0109] The results of the running time test are shown in Table 1. The introduction of dual opaque predicates led to an increase in the program running time, but it was still within a reasonable range. Although the neural network calculation and discrete logarithm solution brought additional overhead, the test results showed that the performance overhead was controllable.

[0110] Table 1 Comparison of running times of sorting algorithms (unit: ms)

[0111]

[0112]

[0113] Binary similarity analysis shows that opaque predicates effectively change the code structure, significantly reducing the matching rate between the obfuscated code and the original code. Table 2 shows the function matching rate, basic block matching rate, and overall similarity of the sorting algorithm before and after obfuscation.

[0114] Table 2 Comparison of Similarities of Sorting Algorithm before and after Obfuscation

[0115] Sorting algorithm Function matching rate (%) Basic block matching rate (%) Overall similarity Bubble sort 6.13 4.90 0.107 Selection sort 6.13 4.90 0.107 Insertion sort 6.02 4.83 0.106 Shell sort 6.02 4.93 0.109 Merge sort 6.95 6.06 0.129 Quick sort 6.14 5.19 0.114

[0116] The decrease in similarity indicates that the obfuscation technique effectively increases the complexity of the code structure, making it more difficult for reverse analysis tools to parse.

[0117] The results of the instruction count statistics are shown in Table 3. The number of instructions in the obfuscated code has increased significantly, with the highest increase approaching a thousand times. This result verifies the effectiveness of opaque predicates in enhancing code complexity.

[0118] Table 6-34 Comparison of Instruction Count Statistics of Sorting Algorithm

[0119] Sorting algorithm Number of unobfuscated instructions Number of obfuscated instructions Bubble sort 2,216,112 1,969,931,785 Selection sort 2,214,726 1,969,927,659 Insertion sort 2,214,027 1,969,940,467 Shell sort 2,213,944 1,969,935,602 Merge sort 2,234,414 2,051,398,635 Quick sort 2,214,437 2,012,166,669

[0120] The increase in the number of instructions indicates that the introduction of opaque predicates has greatly increased the complexity of the code path, interfering with the reverse analysis work.

[0121] The symbolic execution resistance test shows that Angr encounters significant difficulties in parsing the obfuscated code, especially in path explosion and complex condition parsing. By comparing the path exploration success rates of Angr for the code before and after obfuscation, the experimental results show that opaque predicates effectively prevent the path parsing of symbolic execution tools. Due to the insolubility of multi-layer complex judgments and discrete logarithm problems, the parsing success rate of the symbolic execution tool in the obfuscated version has decreased significantly.

[0122] The experimental results show that the double opaque predicate technique shows significant effects in increasing code complexity, reducing binary similarity, increasing the number of instructions, and combating symbolic execution tools. Although the introduction of the feedforward neural network and discrete logarithm conditions has increased the program execution time, the performance overhead is still within an acceptable range.

[0123] The present invention proposes a method for combating symbolic execution attacks based on double opaque predicates. By integrating the non-linear characteristics of the feedforward neural network and the complexity of the discrete logarithm problem, the anti-reverse engineering ability of the code is significantly enhanced. The experiment verifies the excellent performance of this technology in improving the complexity of program paths, reducing code similarity, increasing the number of instructions, and effectively resisting symbolic execution tools. Especially the application to six classic sorting algorithms shows that this technology can significantly improve the code protection strength while keeping the performance overhead under control, and has obvious advantages in security and anti-parsing compared with traditional methods.

[0124] Example 5

[0125] Based on the above embodiments, as Figure 4 shown, the embodiment of the present invention provides a system for countering symbolic execution attacks based on dual opaque predicates, specifically including:

[0126] A feedforward neural network module for constructing a feedforward neural network, where the feedforward neural network includes an input layer, multiple hidden layers, and an output layer.

[0127] A discrete logarithm module for constructing a discrete logarithm problem with a formal linear transformation.

[0128] An embedding module for setting a dual opaque predicate scheme and inserting the feedforward neural network and the discrete logarithm problem with a formal linear transformation into the target program according to the dual opaque predicate scheme.

[0129] It should be noted that the system for countering symbolic execution attacks based on dual opaque predicates provided by the embodiment of the present invention is to implement the above method for countering symbolic execution attacks based on dual opaque predicates. For its functions, reference can be specifically made to the above method embodiments, and details are not described herein again.

[0130] Example 6

[0131] Based on the above embodiments, the embodiment of the present invention further provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the method for countering symbolic execution attacks based on dual opaque predicates in the above embodiments.

[0132] The present invention also provides a computer-readable storage medium, where the storage medium includes a stored computer program. When the computer program runs, it controls the device where the storage medium is located to execute the method for countering symbolic execution attacks based on dual opaque predicates in the above embodiments.

[0133] In summary, the present invention proposes a method for countering symbolic execution attacks based on dual opaque predicates. By combining a feedforward neural network and the discrete logarithm problem, a path condition that is extremely difficult to parse is constructed. The "black box" characteristic of the neural network makes it difficult for the symbolic execution tool to deduce its output, forming a complex condition that is difficult to solve. At the same time, the computational complexity of the discrete logarithm problem further increases the difficulty of solving for the tool, effectively restricting the attacker's complete parsing of the path and effectively increasing the path analysis difficulty of the symbolic execution tool. The present invention constructs a formalized linear transformation of the discrete logarithm problem. By disguising the original multiplication, division, and modulo operations in the discrete logarithm problem as addition, subtraction, and loop operations, the symbolic execution tool will misjudge it as a linear structure during symbolization and directly attempt to solve it instead of calling a non-linear solver. In this way, when the tool processes this transformed condition, it still faces the problem of being difficult to solve, but the computational overhead is increased due to misjudgment.

[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for countering symbolic execution attacks based on double opaque predicates, characterized in that: include: Step 1: construct a feedforward neural network, wherein the feedforward neural network includes an input layer, multiple hidden layers and an output layer; Step 2: Construct the discrete logarithm problem of formal linear transformation; Step 3: Set up a double opaque predicate scheme, and insert the feedforward neural network and the discrete logarithm problem of formal linear transformation into the target program according to the double opaque predicate scheme.

2. The method for countering symbolic execution attacks based on double opaque predicates according to claim 1, characterized in that: The multiple hidden layers in the feedforward neural network each include multiple neurons and a ReLU activation function.

3. The method for countering symbolic execution attacks based on double opaque predicates according to claim 1, characterized in that: The discrete logarithm problem of constructing a formalized linear transformation specifically includes: Set up large prime numbers, generators, target values, results, and unknowns; Set up a loop to solve for the unknowns based on a large prime number, a generator, a target value, a result, and the unknowns.

4. The method for countering symbolic execution attacks based on double opaque predicates according to claim 3, characterized in that: The step of setting a loop for solving unknown numbers according to the large prime number, the generator, the target value, the initialized result and the initialized unknown number specifically includes: Step 1: Determine whether the unknown number is smaller than the large prime number. If the unknown number is larger than the large prime number, end the loop. If the unknown number is smaller than the large prime number, continue the loop. Step 2: Calculate and update the results according to the preset formula; Step 3: Determine whether the result is equal to the target value. If the result is equal to the target value, exit the loop. If the result is not equal to the target value, add 1 to the unknown number and return to step 1.

5. The method for countering symbolic execution attacks based on double opaque predicates according to claim 4, characterized in that: The preset formula is specifically expressed as follows: result+=(result+g-1)%p Among them, result is the result, g is the generator, % is the modulo operation, and p is a large prime number.

6. The method for countering symbolic execution attacks based on double opaque predicates according to claim 1, characterized in that: The dual opaque predicate scheme specifically includes: The feedforward neural network and the discrete logarithm problem of formalized linear transformation are embedded into the target program to construct the path conditions; The feedforward neural network and the discrete logarithm problem of formal linear transformation are connected through the logical "AND" operator, so that when the symbolic execution tool attacks the target program, it is necessary to parse the feedforward neural network. After obtaining the input that meets the output conditions of the feedforward neural network, it is also necessary to solve the discrete logarithm problem of formal linear transformation to complete the path parsing; Set up multiple false branches and logical short-circuits in path conditions.

7. The method for countering symbolic execution attacks based on double opaque predicates according to claim 6, characterized in that: The conditions for the output of the feedforward neural network include: The output value of a feedforward neural network is greater than 1.

8. A system for countering symbolic execution attacks based on double opaque predicates, characterized in that include: A feedforward neural network module, used to construct a feedforward neural network, wherein the feedforward neural network includes an input layer, multiple hidden layers and an output layer; Discrete logarithm module, used to construct discrete logarithm problems for formal linear transformation; The embedding module is used to set up a double opaque predicate scheme, and insert the discrete logarithm problem of feedforward neural network and formalized linear transformation into the target program according to the double opaque predicate scheme.

9. An electronic device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the method for countering symbolic execution attacks based on double opaque predicates as described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that: The storage medium includes a stored computer program, wherein when the computer program is executed, the device where the storage medium is located is controlled to execute the method for countering symbolic execution attacks based on double opaque predicates as described in any one of claims 1 to 7.