Code deobfuscation method, device, electronic device and storage medium

Through the Token-level and AST-level correction methods, the code is deobfused, which solves the problem of identifying incomplete and missing obfuscated statements in the prior art, and improves the antiobfuscated effect and reliability.

CN114186233BActive Publication Date: 2025-05-23QI AN XIN TECHNOLOGY GROUP INC +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111521024.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-13
Publication Date
2025-05-23
Estimated Expiration
2041-12-13

AI Technical Summary

Technical Problem

Existing code antiobfuscation technology has the problem of identifying incomplete obfuscation statements and missing a large number of obfuscation statements, resulting in poor antiobfuscation effect.

Method used

By obtaining the token list of the pending code, performing token-level anti-obfuscation processing, obtaining AST, and then simulated and executed and replaced the content of the target node of the AST to achieve token-level and statement fragment-level corrections.

Benefits of technology

It reduces the risk of omission of confusing statements and improves the anti-confusion effect, making the resulting AST content more readable and the anti-confusing reliability is higher.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114186233B_ABST
    Figure CN114186233B_ABST
Patent Text Reader

Abstract

The present application provides a code deobfuscation method, device, electronic device and storage medium, the method comprising: obtaining a Token list of a code to be processed; each Token in the Token list is a code word constituting the code to be processed; performing deobfuscation processing on each Token in the Token list to obtain a first deobfuscation code; parsing the first deobfuscation code to obtain an AST of the first deobfuscation code; simulating the execution of the content in the target node of the AST to obtain a first execution result, and replacing the content of the target node according to the first execution result; the target node is a node of a set type in the AST; and obtaining the final deobfuscation code according to the replaced AST. The above scheme reduces the risk of missing obfuscated statements and improves the deobfuscation effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a code deobfuscation method, device, electronic device and computer-readable storage medium. Background Art

[0002] Code obfuscation refers to converting program code into a form that remains functional but is difficult to read and understand. Most existing script codes are obfuscated to prevent source code leakage. However, the obfuscated scripts are quite different from normal scripts, which results in many antivirus software being unable to identify the scripts normally, and thus making it impossible to effectively detect and kill malicious scripts.

[0003] Code anti-obfuscation technology can help security personnel analyze the attack methods and intentions of malicious scripts more quickly, and effectively improve the antivirus software's ability to identify malicious scripts.

[0004] Currently, code deobfuscation techniques are mainly divided into regular expression-based deobfuscation methods and machine learning-based deobfuscation methods.

[0005] The regular expression-based deobfuscation method is a common obfuscation location and processing technology. The string-level obfuscation method has specific text features. By designing specific regular expressions, such obfuscated statements can be quickly identified and deobfuscated using pre-designed strategies. However, regular expressions do not consider the syntax and semantics of the statements, which often results in the obfuscated statements being syntactically incomplete. For complex coding obfuscation, it is difficult to design a suitable pattern for regular expressions to match, resulting in the omission of a large number of obfuscated statements, which in turn leads to poor deobfuscation effects.

[0006] The machine learning-based deobfuscation method uses the AST (Abstract Syntax Tree) features of the code to identify obfuscated subtrees, and restores the obfuscated statements corresponding to the obfuscated subtrees by executing them through the simulator. However, this solution can only process PipeLine nodes in the AST (i.e. nodes with complete obfuscated statements), which will also miss many obfuscated statements, resulting in poor deobfuscation effects. Summary of the invention

[0007] The purpose of the embodiments of the present application is to provide a code deobfuscation method, device, electronic device and computer-readable storage medium to improve the deobfuscation effect.

[0008] An embodiment of the present application provides a code deobfuscation method, comprising: obtaining a Token list of a code to be processed; each Token in the Token list is a code word constituting the code to be processed, and the code to be processed is an obfuscated code; performing deobfuscation processing on each Token in the Token list to obtain a first deobfuscation code; parsing the first deobfuscation code to obtain an AST of the first deobfuscation code; simulating execution of the content in a target node of the AST to obtain a first execution result, and replacing the content of the target node according to the first execution result; the target node is a node of a set type in the AST; and obtaining a final deobfuscation code according to the replaced AST.

[0009] In the above implementation process, by first correcting each code word of the code to be processed (i.e., correction at the Token level), and then simulating the execution of the content in the target node that may be confused based on AST, correction at the Token level and the sentence fragment level (the content in the node is the sentence fragment of the code) is achieved. Compared with the existing anti-obfuscation method based on regular expressions, the obfuscation correction at the code statement level is considered to reduce the risk of missing obfuscated statements and improve the anti-obfuscation effect. Compared with the anti-obfuscation method based on machine learning, on the one hand, the Token-level correction is considered, so that the content in the obtained AST is more readable, easier to be simulated and executed to perform correction at the sentence fragment level, and the anti-obfuscation reliability is higher; on the other hand, the inventor found that the nodes that are obfuscated in the code are usually specific types of nodes in the AST (such as "binaryexpression" nodes, "PipeLine" nodes, etc.), so by simulating the execution of these nodes to achieve the correction of the sentence fragment, compared with the existing solution that can only process the PipeLine nodes in the AST, it can also achieve more fine-grained corrections, and also reduce the risk of missing obfuscated statements, and improve the anti-obfuscation effect.

[0010] Furthermore, the content in the target node of the AST is simulated and executed to obtain a first execution result, and the content of the target node is replaced according to the first execution result, including: post-order traversal of each node of the AST; when the currently traversed node belongs to the target node of the set type, the content in the currently traversed node is simulated and executed to obtain the first execution result of the currently traversed node; according to the first execution result of the currently traversed node, the content in the currently traversed node is replaced until all nodes of the AST are traversed.

[0011] In the above implementation process, each node is identified and the target node is simulated and executed by post-order traversal, and the content in the target node is replaced according to the first execution result of the target node, thereby achieving strict in-place replacement, which can ensure the effective correction of the content of all target nodes and ensure the semantic consistency of the deobfuscation result with the code to be processed.

[0012] Furthermore, in the process of post-order traversal of each node of the AST, the method also includes: obtaining the content in the child node of the currently traversed node; and replacing the content corresponding to the child node in the currently traversed node with the content in the child node.

[0013] Through the above implementation process, it can be ensured that the replacement made in a target node can be effectively transmitted to the root node, thereby achieving reliable replacement on the entire chain from the target node to the root node, thereby ensuring that the restored deobfuscated code is accurate and reliable. In addition, since the post-order traversal is a traversal starting from the lower-level child nodes, when traversing to the parent node, since the content corresponding to the child node in the parent node will be replaced, the parent node is also the target node at this time, so the simulation execution can be realized more quickly during the simulation execution, improving the simulation execution efficiency, thereby improving the efficiency of the deobfuscation processing, and reducing the overhead generated by the deobfuscation processing.

[0014] Furthermore, according to the first execution result of the currently traversed node, replacing the content in the currently traversed node includes: determining whether there is a preset obfuscation execution command in the first execution result; the obfuscation execution command is a command required for code obfuscation; if so, obtaining the parameter part of the obfuscation execution command; performing deobfuscation processing on each code word in the parameter part to obtain a processed parameter part; performing simulation execution on the processed parameter part to obtain a second execution result; replacing the parameter part in the first execution result with the second execution result; and when the obfuscation execution command does not exist in the second execution result, replacing all the content in the currently traversed node with the first execution result.

[0015] In actual applications, there may be multiple layers of obfuscation. In the above implementation process, by judging whether there is an obfuscated execution command in the first execution result, and then extracting the parameter part of the obfuscated execution command when it exists, and then performing deobfuscation processing and simulated execution of the code words. In this way, the continuous deobfuscation of multiple layers of obfuscation can be effectively achieved, and a more reliable and trustworthy deobfuscated code can be obtained.

[0016] Furthermore, when there are variables in the content of the target node, before simulating the execution of the content in the target node, the method also includes: obtaining the variable value of the variable in the scope to which the target node belongs from a pre-established record table; and writing the variable value into the variable.

[0017] It should be understood that in the existing deobfuscation method based on machine learning, the processing is based on the obfuscated statements in the independent PipeLine node. However, due to the lack of obfuscated statement context in each node, it will lead to the inability to obtain the correct deobfuscation result by directly executing the obfuscated statement through the simulator. In the above implementation process, by pre-establishing a record table, variable tracking is realized, and the context of the obfuscated statement fragments during simulation execution is improved, so that the obfuscated statement fragments can be executed to obtain the correct restoration result. Compared with the existing deobfuscation method based on machine learning, the deobfuscation result is more accurate and has a better deobfuscation effect.

[0018] Furthermore, after obtaining the AST of the first deobfuscated code and before simulating the execution of the content in the target node of the AST, the method also includes: obtaining variable values ​​and the scope of each variable value from the assignment class node of the AST; and associating and recording each variable value and the scope of each variable value in the record table.

[0019] In the above implementation process, by obtaining the variable values ​​and the scope of each variable value from the assignment class node of AST, a record table is constructed, so that the variable values ​​and the scope of each variable value can be pre-associated, which facilitates the subsequent simulation execution of the content in the target node.

[0020] Furthermore, each Token in the Token list is deobfuscated to obtain a first deobfuscation code, including: traversing the Token list in reverse order; correcting the currently traversed Token according to a preset Token correction rule until the traversal of the Token list is completed; and obtaining the first deobfuscation code according to the corrected Token list.

[0021] In the above implementation process, the Token list is traversed in reverse order to correct the Token of the code to be processed from back to front. This ensures that when the number of characters in the corrected Token changes compared to the number of characters in the original Token, the position of the uncorrected Token in the Token list will not be affected. Therefore, each time a Token is corrected, there is no need to reposition the position of the uncorrected Token, which reduces processing overhead and improves processing efficiency.

[0022] Furthermore, obtaining the final deobfuscation code according to the replaced AST includes: converting the replaced AST into a second deobfuscation code; renaming the variable names and function names in the second deobfuscation code, and rearranging the second deobfuscation code to obtain the final deobfuscation code.

[0023] In the above implementation process, by renaming the variable names and function names in the second deobfuscation code and rearranging the second deobfuscation code, the final deobfuscation code obtained can be more convenient to identify and read, which is convenient for antivirus software to detect and kill and for engineers to review.

[0024] Furthermore, before renaming the variable names and function names in the second deobfuscation code, the method further includes: determining that a ratio between vowel characters and consonant characters in a character string set consisting of all non-repeated variables and functions in the second deobfuscation code is not within a preset ratio range, so as to rename the variable names and function names in the second deobfuscation code when the ratio is not within the preset ratio range.

[0025] In the above implementation process, by detecting the ratio between vowel characters and consonant characters in the string set composed of all non-repeated variables and functions in the second deobfuscation code, the randomness of the variable names and function names in the second deobfuscation code can be determined based on the ratio. If the ratio between vowel characters and consonant characters in the string set composed of all non-repeated variables and functions in the second deobfuscation code is not within the preset ratio range, it can be considered that the variable names and function names in the second deobfuscation code are currently randomly named, and thus renamed to ensure that the variable names and function names have higher readability.

[0026] An embodiment of the present application also provides a code deobfuscation device, including: an acquisition module, used to obtain a Token list of a code to be processed, wherein the code to be processed is an obfuscated code; a processing module, used to perform deobfuscation processing on each Token in the Token list to obtain a first deobfuscation code, parse the first deobfuscation code to obtain an AST of the first deobfuscation code, simulate execution of the content in a target node of the AST to obtain a first execution result, and replace the content of the target node according to the first execution result, and obtain a final deobfuscation code according to the replaced AST; wherein each Token in the Token list is a code word constituting the code to be processed; and the target node is a node of a set type in the AST.

[0027] An embodiment of the present application also provides an electronic device, including a processor, a memory and a communication bus; the communication bus is used to realize connection and communication between the processor and the memory; the processor is used to execute one or more programs stored in the memory to implement any of the above-mentioned code deobfuscation methods.

[0028] A computer-readable storage medium is also provided in an embodiment of the present application. The computer-readable storage medium stores one or more programs. The one or more programs can be executed by one or more processors to implement any of the above-mentioned code deobfuscation methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0030] Figure 1 A basic flow chart of a code deobfuscation method provided in an embodiment of the present application;

[0031] Figure 2 An example node diagram provided in an embodiment of the present application;

[0032] Figure 3 A schematic diagram of a process for processing multiple obfuscations provided in an embodiment of the present application;

[0033] Figure 4 An exemplary process diagram provided for an embodiment of the present application;

[0034] Figure 5 A schematic diagram of the structure of a code anti-obfuscation device provided in an embodiment of the present application;

[0035] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0036] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.

[0037] Embodiment 1:

[0038] In order to improve the anti-obfuscation effect, a code anti-obfuscation method is provided in the embodiment of the present application. Figure 1 As shown, Figure 1 The flowchart of the code deobfuscation method provided in the embodiment of the present application includes:

[0039] S101: Obtain a token list of codes to be processed.

[0040] In the embodiment of the present application, the code to be processed refers to the obfuscated code that currently needs to be deobfuscated, which can be PowerShell code or other types of code, and is not limited in the embodiment of the present application.

[0041] Obfuscation refers to converting program code into a form that remains functional but is difficult to read and understand. Code obfuscation can be achieved using a code obfuscator.

[0042] Deobfuscation is the reverse process of obfuscation, which refers to the process of reverse analysis of the obfuscated code to restore its original and more readable source code logic.

[0043] In the embodiment of the present application, by Figure 1 The provided solution de-obfuscates the obfuscated code to be processed to obtain a more readable de-obfuscated code, which is convenient for anti-virus software to identify and engineers to review.

[0044] In the embodiment of the present application, the official library of the code can be used to perform lexical analysis on the code to be processed to obtain a token list of the code to be processed. For example, for PowerShell code, the official library of PowerShell can be used to perform lexical analysis on the code to be processed to obtain the individual code words that make up the PowerShell code.

[0045] It should be noted that, in the embodiment of the present application, each Token in the Token list is a code word constituting the code to be processed.

[0046] S102: Deobfuscation processing is performed on each Token in the Token list to obtain a first deobfuscation code.

[0047] In the embodiment of the present application, Token modification rules can be pre-constructed, so that each Token can be deobfuscated using the Token modification rules.

[0048] In the embodiment of the present application, the normal token form corresponding to different tokens can be defined in the token correction rule, so that each token in the token list can be corrected to a normal form through the token correction rule. For example, the abbreviation or abbreviation of a word in the obfuscated code can be corrected to the full name of the word; for another example, a character string with mixed capital letters in the obfuscated code can be corrected to a code word with the first letter capitalized and the rest lowercase.

[0049] Exemplarily, the above-mentioned Token modification rules can be implemented by means such as regular expressions, preset correspondence tables, etc., but this is not a limitation.

[0050] In the embodiment of the present application, the Token list can be traversed in reverse order, and then the currently traversed Token can be corrected according to the preset Token correction rules until the Token list is completely traversed. In this way, by traversing the Token list in reverse order, the Token of the code to be processed is corrected from back to front, so that when the number of characters of the corrected Token changes compared to the number of characters of the original Token, the position of the uncorrected Token in the Token list will not be affected, so that each time a Token is corrected, there is no need to reposition the position of the uncorrected Token, which reduces processing overhead and improves processing efficiency.

[0051] Of course, in the embodiment of the present application, the Token list can also be traversed in forward order, or in other orders. In this case, it is only necessary to re-parse the code to be processed after each correction of a Token to obtain a new Token list, thereby re-positioning the next Token to be corrected to ensure that each Token can be corrected correctly.

[0052] S103: Parse the first deobfuscated code to obtain an AST of the first deobfuscated code.

[0053] It should be understood that ATS is a tree formed when the program code is deduced according to the grammatical rules of the program statement after grammatical analysis, and AST represents the derivation result of the program code.

[0054] In an embodiment of the present application, by parsing the first deobfuscated code after Token correction rather than parsing the original code to be processed, the content in the obtained AST can be made more readable, and the obtained AST can be closer to the structure of the code to be processed before obfuscation, so that when deobfuscation is performed based on AST, it can be easier to simulate and execute, and a more accurate deobfuscation result can be obtained.

[0055] S104: Simulate the execution of the content in the target node of the AST to obtain a first execution result, and replace the content of the target node according to the first execution result.

[0056] It should be noted that in the embodiment of the present application, the target node is a node of a set type in the AST. The set type here can be set by the engineer according to the type of node that may be confused. For example, a node of type "BinaryExpression", a node of type "PipeLine", etc. can be set as a node of a set type.

[0057] In the embodiment of the present application, each node of the AST can be traversed in sequence by post-order traversal.

[0058] For each node, when it is traversed, it can be determined whether the node (that is, the currently traversed node) belongs to a target node of a set type.

[0059] If the currently traversed node belongs to the target node of the set type, the content in the target node can be simulated and executed to obtain the first execution result of the target node. Furthermore, the content in the target node can be replaced according to the first execution result of the target node until all nodes of the AST are traversed.

[0060] In the embodiment of the present application, a simulator or various execution functions may be used to simulate and execute the content in the target node. For example, the Invoke function may be used for execution.

[0061] It should be noted that for each node in the AST, the nodes have a direct superior-subordinate relationship, and the parent node will have all the content of the child node, and may also have other content. In order to ensure that the replaced content can be effectively restored to code, after replacing the content of any child node, the corresponding content in all nodes in the entire link from the child node to the root node also needs to be replaced. That is, assuming that the content "A+B" is replaced with "AB" in a child node, then "A+B" needs to be replaced with "AB" in all nodes in the entire link from the child node to the root node to ensure the effectiveness of the replacement.

[0062] In order to ensure the effectiveness of the replacement, in a feasible implementation of the embodiment of the present application, after simulating the execution of the content in a target node, all nodes in the entire link from the child node to the root node can be replaced while replacing the content of the target node.

[0063] For example, after simulating the execution of the content in a target node, the content in the target node may be replaced, and the content identical to the target node in all nodes in the link may be replaced in sequence from the target node to the root node.

[0064] For example, Figure 2As shown, assuming that the currently traversed node is node 1, and node 1 is the target node, the content is A+B, and the simulation execution result is AB, then the content of node 1 can be replaced with AB, and the content "A+B" in all nodes (i.e. node 1, node 3, node 5) in the link from node 1 to the root node (i.e. node 5) can be replaced with "AB".

[0065] In another feasible implementation of the embodiment of the present application, after simulating the execution of the content in a target node, while replacing the content of the target node, only the parent node of the target node is replaced. Then, after traversing to the parent node, the parent node of the parent node is replaced. In this way, the effectiveness of the replacement can also be guaranteed, and since there are fewer replacement operations each time, the time of each simulation execution can be shortened, which can provide overall simulation execution efficiency.

[0066] It should be noted that in the above feasible implementation, when performing node traversal, for each node, it is necessary to determine whether the content of the currently traversed node has been replaced. If the content has been replaced, the content of the replaced part is replaced in the parent node of the node. In this way, when the parent node of the target node is traversed, the corresponding content can be further replaced upward, thereby achieving the effect of replacing all nodes in the entire link from the target node to the root node.

[0067] For example, Figure 2 As shown, assuming that the currently traversed node is node 1, and node 1 is the target node, the content is A+B, and the simulation execution result is AB. Assume that the content of node 1's parent node (node ​​3) is A+B+C. After simulating the execution of the content of node 1, the content of node 1 can be replaced with AB, and the "A+B" in the content of node 3 can be replaced with "AB", and the new content is "AB+C". When traversing to node 3, since it has been replaced, regardless of whether it is the target node, the "A+B" in node 5 is replaced with "AB".

[0068] In another feasible implementation of the embodiment of the present application, after simulating the execution of the content in a target node, only the content of the target node can be replaced. In the process of post-order traversal of each node of the AST, when the parent node of the target node is traversed, the content of the target node is replaced with the parent node.

[0069] That is, when performing post-order traversal, for the currently traversed node, whether it is the target node or not, the content of the child node of the node can be obtained, and the content corresponding to the child node in the currently traversed node can be replaced with the content of the child node. In this way, the effectiveness of the replacement can also be guaranteed, and the effect of replacing all nodes in the entire link from the target node to the root node can be achieved.

[0070] It should be noted that, for the lowest-level child node, since it has no lower-level child node, it may not be replaced.

[0071] In this feasible implementation, if the currently traversed node is the target node, the content in the currently traversed node corresponding to the child node can be replaced with the content in the child node, and then the content in the currently traversed node can be simulated and executed.

[0072] For example, Figure 2 As shown, assuming that the currently traversed node is node 1, and node 1 is the target node, the content is A+B, and the simulation execution result is AB. Assume that the content of node 1's parent node (node ​​3) is A+B+C. After simulating the execution of the content of node 1, the content of node 1 can be replaced with AB. When traversing to node 3, regardless of whether it is the target node, first replace "A+B" in the content of node 3 with "AB", and the new content is "AB+C". If node 3 is also the target node, simulate and execute "AB+C". If node 3 is not the target node, continue traversing. When traversing to node 5, assuming that both node 3 and node 4 are not the target nodes, the content of node 3 is "AB+C", and the content of node 4 is "D+E", then replace the content of node 5 with "AB+C+D+E".

[0073] It should be noted that in actual applications, many code statements in target nodes may contain variables. If the code statement is directly executed without assigning values ​​to the variables in the code statement, the result of the simulation execution may be inaccurate, thus affecting the overall deobfuscation effect.

[0074] To this end, in an embodiment of the present application, after obtaining the AST of the first deobfuscated code, before simulating the execution of the content in the target node of the AST, it is also possible to first obtain the variable values ​​and the scope of each variable value from the assignment class node of the AST; and then record each variable value and the scope of each variable value in a record table.

[0075] Then, when there are variables in the content of the target node, before simulating the execution of the content in the target node, the variable value of the variable in the scope to which the target node belongs is obtained from the pre-established record table, and then the variable value is written into the variable, and then the simulation execution is performed. In this way, variable tracking is achieved, and the context of the obfuscated statement fragment during the simulation execution is improved, so that the obfuscated statement fragment can be executed to obtain the correct restoration result.

[0076] It should also be noted that in actual applications, there may be a situation where the statement fragments in the target node are too scattered, resulting in the simulation execution failing to output results, that is, there may be a situation where there is no first execution result. In this case, the target node may not be replaced, and the next node in the AST may be traversed.

[0077] It should also be noted that in actual applications, there may be multiple layers of obfuscation. If the content in a target node is simulated and executed, the first execution result may not be able to completely remove the obfuscation. For this reason, in the embodiment of the present application, after the content in the target node of the AST is simulated and executed to obtain the first execution result, the following can be executed: Figure 3 The operation process shown:

[0078] S301: Determine whether there is a preset obfuscated execution command in the first execution result. If yes, go to step S302. If not, go to step S307.

[0079] It should be noted that the obfuscated execution command refers to the command required for code obfuscation. For example, in PowerShell code, the obfuscated execution command can be the Invoke-Expression command, PowerShell command, etc.

[0080] S302: Obtain the parameter portion of the obfuscated execution command.

[0081] S303: Deobfuscation processing is performed on each code word (ie, Token) in the parameter part to obtain a processed parameter part.

[0082] The anti-obfuscation processing method here is consistent with the Token anti-obfuscation processing described above, so I will not go into details here.

[0083] S304: Perform simulation execution on the processed parameter part to obtain a second execution result.

[0084] S305: Replace the parameter portion in the first execution result with the second execution result.

[0085] S306: Determine whether the second execution result still contains the preset obfuscated execution command. If so, go to step S302. If not, go to step S307.

[0086] S307: Replace all the content in the target node with the first execution result.

[0087] It should be understood that the above process continuously performs Token-level deobfuscation processing on the parameter part of the obfuscated execution command and re-simulates the execution until the obfuscated execution command no longer exists, thereby achieving layer-by-layer analysis of multiple layers of obfuscation and ensuring the reliability of the deobfuscation result.

[0088] In the embodiment of the present application, when replacing the content in the node, it can be implemented by using a dictionary. That is, the text content of each AST node can be recorded in a dictionary, and then the text content of the AST node can be replaced in the dictionary. In this way, by strictly performing in-place replacement, the semantic consistency of the deobfuscated code obtained by the deobfuscation process and the code to be processed can be guaranteed.

[0089] S105: Obtain the final deobfuscated code according to the replaced AST.

[0090] In a feasible implementation of the embodiment of the present application, the replaced AST may be converted into a second deobfuscated code, and the second deobfuscated code may be directly used as the final deobfuscated code.

[0091] However, in order to further improve readability, in another feasible implementation of the embodiment of the present application, after converting the replaced AST into the second deobfuscation code, the variable names and function names in the second deobfuscation code can also be renamed, and the second deobfuscation code can be rearranged to obtain the final deobfuscation code.

[0092] In the embodiment of the present application, the engineer can pre-set the naming rules of the variable name and the function name, so as to rename the variable name and the function name in the second deobfuscated code. Similarly, the engineer can pre-set the rearrangement rules, such as the space distance, the indentation size, etc., so as to rearrange the second deobfuscated code.

[0093] Considering that in actual application, the variable names and function names in the second deobfuscation code may not be random but have high readability, in this case, the variable names and function names in the second deobfuscation code may not be renamed, thereby reducing processing overhead.

[0094] In order to identify whether the variable names and function names in the second deobfuscation code are random, in a feasible implementation of the embodiment of the present application, the ratio between vowel characters and consonant characters in the string set composed of all non-repeating variables and functions in the second deobfuscation code can be obtained. Then it is determined whether the ratio is within a preset ratio range. If so, it can be considered that the variable names and function names in the second deobfuscation code are not currently randomly named, so there is no need to rename them. On the contrary, if the ratio is not within the preset ratio range, it can be considered that the variable names and function names in the second deobfuscation code are currently randomly named, so renaming is performed to ensure that the variable names and function names have higher readability.

[0095] It should be understood that in the embodiment of the present application, in addition to determining whether renaming is required by judging whether the ratio between vowel characters and consonant characters in the string set composed of all non-repetitive variables and functions in the second deobfuscation code is within a preset ratio range, it is also possible to determine whether renaming is required by judging the proportion of vowel characters in the total characters in the string set, or the proportion of consonant characters in the total characters in the string set. The principle is the same as the above-mentioned principle of determining whether renaming is required based on the ratio between vowel characters and consonant characters, and will not be repeated here.

[0096] The code deobfuscation method provided in the embodiment of the present application first corrects each code word of the code to be processed (i.e., correction at the Token level), and then simulates and executes the content in the target node that may be obfuscated based on AST, thereby realizing correction at the Token level and the sentence fragment level (the content in the node is the sentence fragment of the code). Compared with the existing code deobfuscation method based on regular expressions, the obfuscation correction at the code statement level is considered, which reduces the risk of missing obfuscated statements and improves the deobfuscation effect. Compared with the code deobfuscation method based on machine learning, on the one hand, the Token-level correction is considered, so that the content in the obtained AST is more readable, easier to be simulated and executed to perform correction at the sentence fragment level, and the deobfuscation reliability is higher; on the other hand, the inventor found that the nodes obfuscated in the code are usually specific types of nodes in the AST (such as "BinaryExpression" nodes, "PipeLine" nodes, etc.), so by simulating and executing these nodes to achieve the correction of the sentence fragment, compared with the existing solution that can only process the PipeLine nodes in the AST, it can also achieve more fine-grained correction, reduce the risk of missing obfuscated statements, and improve the deobfuscation effect.

[0097] In addition, in the solution in the embodiment of the present application, in-place replacement is strictly performed, and the root node can be continuously updated through the post-order traversal of the AST tree, thereby ensuring the semantic consistency between the code obtained after the deobfuscation process and the code before the deobfuscation process.

[0098] In addition, in the solution in the embodiment of the present application, variable tracking can be implemented, and the context of the obfuscated statement is improved, so that the obfuscated statement can be executed to obtain the correct restoration result, thereby improving the anti-obfuscation effect.

[0099] Embodiment 2:

[0100] Based on the first embodiment, this embodiment takes a specific example of PowerShell code as an example to further illustrate this application.

[0101] See also Figure 4 As shown, Figure 4 The left side of the figure shows the obfuscated code to be processed.

[0102] First, the code to be processed is lexically parsed, and each token in the code to be processed is deobfuscated (see Figure 4 The first deobfuscated code is obtained.

[0103] Then, see Figure 4 In the box marked with AST-based restoration, the first deobfuscated code is parsed to obtain the AST of the first deobfuscated code.

[0104] Identify the variable value and scope of the assignment type node in the AST, and record the variable value and the scope of each variable value in the record table.

[0105] Post-order traversal of each node of the AST.

[0106] When the currently traversed node does not belong to the target node of the set type and the current node has no child nodes, the next node is traversed.

[0107] When the currently traversed node does not belong to the target node of the set type, but the current node has a child node, the content of the child node of the current node is obtained, and the content corresponding to the child node in the current node is replaced with the content of the child node.

[0108] When the currently traversed node belongs to the target node of the set type and the current node has not been replaced with content, the content in the current node is executed through the Invoke function (ie, the obfuscated statement in the current node is executed), and the content in the node is replaced with the execution result.

[0109] For example, see Figure 4As shown in the dashed box marked with "Restore", the dashed box is the simulated execution process for the "BinaryExpression" node. In this process, since the execution result does not have the Invoke-Expression command or PowerShell command, the execution result is directly output and replaced.

[0110] After the AST traversal is completed, the replaced AST is converted into a second deobfuscated code, the variable names and function names in the second deobfuscated code are renamed, and the second deobfuscated code is rearranged to obtain the final deobfuscated code.

[0111] Figure 4 The lower right frame is the deobfuscated code after deobfuscation processing.

[0112] The above solution, targeting the diverse characteristics of PowerShell obfuscated code, accurately identifies obfuscated code fragments through lexical analysis and the types of each AST node. The simulated execution based on variable tracking can correctly obtain the restoration result of the obfuscated code. Strict in-place replacement can ensure that the deobfuscated code after deobfuscation is semantically consistent with the code before deobfuscation. Renaming and rearrangement make the deobfuscation result easier to read and understand.

[0113] Embodiment three:

[0114] Based on the same inventive concept, the present application also provides a code anti-obfuscation device 500. Figure 5 As shown, Figure 5 Shows the use of Figure 1 The device 500 is a device for deobfuscating the code of the method shown. It should be understood that the specific functions of the device 500 can be referred to the description above, and the detailed description is appropriately omitted here to avoid repetition. The device 500 includes at least one software function module that can be stored in a memory in the form of software or firmware or fixed in the operating system of the device 500. Specifically:

[0115] See also Figure 5 As shown, the apparatus 500 includes: an acquisition module 501 and a processing module 502. Among them:

[0116] The acquisition module 501 is used to acquire a token list of the code to be processed, where the code to be processed is the obfuscated code;

[0117] The processing module 502 is used to perform deobfuscation processing on each token in the token list to obtain a first deobfuscation code, parse the first deobfuscation code to obtain an abstract syntax tree AST of the first deobfuscation code, simulate the execution of the content in the target node of the AST to obtain a first execution result, replace the content of the target node according to the first execution result, and obtain the final deobfuscation code according to the replaced AST;

[0118] Among them, each Token in the Token list is each code word constituting the code to be processed; the target node is a node of a set type in the AST.

[0119] In a feasible implementation of the embodiment of the present application, the processing module 502 is specifically used to traverse each node of the AST in post-order, and when the currently traversed node belongs to the target node of the set type, simulate the execution of the content in the currently traversed node to obtain the first execution result of the currently traversed node, and replace the content in the currently traversed node according to the first execution result of the currently traversed node until all nodes of the AST are traversed.

[0120] In the above feasible implementation mode, in the process of post-order traversal of each node of the AST, the processing module 502 is also used to obtain the content in the child node of the currently traversed node, and replace the content corresponding to the child node in the currently traversed node with the content in the child node.

[0121] In the above-mentioned feasible implementation manner, the processing module 502 is specifically used to determine whether there is a preset obfuscation execution command in the first execution result; the obfuscation execution command is a command required for code obfuscation; if it exists, obtain the parameter part of the obfuscation execution command; perform deobfuscation processing on each code word in the parameter part to obtain the processed parameter part; simulate the execution of the processed parameter part to obtain a second execution result; replace the parameter part in the first execution result with the second execution result; when the obfuscation execution command does not exist in the second execution result, replace all the contents in the currently traversed node with the first execution result.

[0122] In a feasible implementation of an embodiment of the present application, when there are variables in the content of the target node, before simulating the execution of the content in the target node, the processing module 502 is also used to obtain the variable value of the variable in the scope to which the target node belongs from a pre-established record table; and write the variable value into the variable.

[0123] In the above-mentioned feasible implementation mode, after obtaining the AST of the first deobfuscated code and before simulating the execution of the content in the target node of the AST, the processing module 502 is also used to: obtain variable values ​​and the scope of each variable value from the assignment class node of the AST; and associate and record each variable value and the scope of each variable value in the record table.

[0124] In the embodiment of the present application, the processing module 502 is specifically used to: traverse the Token list in reverse order; modify the currently traversed Token according to the preset Token modification rule until the traversal of the Token list is completed; and obtain the first deobfuscation code according to the modified Token list.

[0125] In an embodiment of the present application, the processing module 502 is specifically used to: convert the replaced AST into a second deobfuscation code; rename the variable names and function names in the second deobfuscation code, and rearrange the second deobfuscation code to obtain the final deobfuscation code.

[0126] In an embodiment of the present application, the processing module 502 is further used to: before renaming the variable names and function names in the second deobfuscation code, determine that the ratio between vowel characters and consonant characters in a character string set consisting of all non-repeating variables and functions in the second deobfuscation code is not within a preset ratio range, so as to rename the variable names and function names in the second deobfuscation code when the ratio is not within the preset ratio range.

[0127] It should be understood that, for the sake of brevity, some of the contents described in the first embodiment will not be repeated in this embodiment.

[0128] Embodiment 4:

[0129] This embodiment provides an electronic device, see Figure 6 As shown, it includes a processor 601, a memory 602 and a communication bus 603. Among them:

[0130] The communication bus 603 is used to realize the connection and communication between the processor 601 and the memory 602 .

[0131] The processor 601 is used to execute one or more programs stored in the memory 602 to implement the code deobfuscation method in the above-mentioned embodiment 1 and / or embodiment 2.

[0132] Understandably, Figure 6 The structure shown is for illustration only. The electronic device may also include Figure 6 More or fewer components as shown, or with Figure 6Different configurations are shown. Exemplarily, the electronic device may be a server, a terminal, or other device with data processing capability.

[0133] This embodiment also provides a computer-readable storage medium, such as a floppy disk, an optical disk, a hard disk, a flash memory, a USB flash disk, an SD (Secure Digital Memory Card) card, an MMC (Multimedia Card) card, etc., in which one or more programs for implementing the above steps are stored, and the one or more programs can be executed by one or more processors to implement the code deobfuscation method in the above embodiment 1 and / or embodiment 2. No further details will be given here.

[0134] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0135] In addition, the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0136] Furthermore, the functional modules in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.

[0137] In this document, relational terms such as first and second, etc. are used merely to distinguish one entity or operation from another entity or operation, but do not necessarily require or imply any such actual relationship or order between these entities or operations.

[0138] As used herein, a plurality refers to two or more than two.

[0139] The above description is only an embodiment of the present application and is not intended to limit the protection scope of the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A code deobfuscation method, It is characterized in that include: Get the token list of the code to be processed; Each token in the token list is a code word constituting the code to be processed, and the code to be processed is an obfuscated code; Performing a deobfuscation process on each Token in the Token list to obtain a first deobfuscation code; Parsing the first deobfuscated code to obtain an abstract syntax tree AST of the first deobfuscated code; Simulate the execution of the content in the target node of the AST to obtain a first execution result, and replace the content of the target node according to the first execution result; The target node is a node of a set type in the AST; The final deobfuscated code is obtained based on the replaced AST.

2. The code deobfuscation method according to claim 1, It is characterized in that Simulating the execution of the content in the target node of the AST to obtain a first execution result, and replacing the content of the target node according to the first execution result, including: Post-order traversal of each node of the AST; When the currently traversed node belongs to the target node of the set type, simulate and execute the content in the currently traversed node to obtain a first execution result of the currently traversed node; According to the first execution result of the currently traversed node, the content in the currently traversed node is replaced until all nodes of the AST are traversed.

3. The code deobfuscation method according to claim 2, It is characterized in that In the process of post-order traversal of each node of the AST, the method further includes: Get the content of the child nodes of the currently traversed node; The content corresponding to the child node in the currently traversed node is replaced by the content in the child node.

4. The code deobfuscation method according to claim 2, It is characterized in that Replacing the content of the currently traversed node according to the first execution result of the currently traversed node includes: Determine whether there is a preset obfuscation execution command in the first execution result; the obfuscation execution command is a command required for code obfuscation; If it exists, obtain the parameter part of the obfuscated execution command; Performing a deobfuscation process on each code word in the parameter part to obtain a processed parameter part; Performing simulation execution on the processed parameter part to obtain a second execution result; Replacing the parameter portion in the first execution result with the second execution result; When the obfuscated execution command does not exist in the second execution result, all contents in the currently traversed node are replaced with the first execution result.

5. The code deobfuscation method according to claim 1, It is characterized in that When there are variables in the content of the target node, before simulating and executing the content in the target node, the method further includes: Obtaining the variable value of the variable in the scope to which the target node belongs from the pre-established record table; Write the variable value into the variable.

6. The code deobfuscation method according to claim 5, It is characterized in that After obtaining the AST of the first deobfuscated code and before simulating and executing the content in the target node of the AST, the method further includes: Obtain variable values ​​and scopes of each variable value from the assignment class node of the AST; The record table records each variable value and its scope in association with each other.

7. The code deobfuscation method according to any one of claims 1 to 6, It is characterized in that Deobfuscation processing is performed on each token in the token list to obtain a first deobfuscation code, including: Traverse the Token list in reverse order; According to the preset token modification rules, modify the currently traversed token until the token list is traversed completely; The first deobfuscation code is obtained according to the revised Token list.

8. The code deobfuscation method according to any one of claims 1 to 6, It is characterized in that The final deobfuscated code is obtained based on the replaced AST, including: Convert the replaced AST into the second deobfuscated code; The variable names and function names in the second deobfuscated code are renamed, and the second deobfuscated code is rearranged to obtain a final deobfuscated code.

9. The code deobfuscation method according to claim 8, It is characterized in that Before renaming the variable names and function names in the second deobfuscated code, the method further includes: Determine whether a ratio of vowel characters to consonant characters in a character string set consisting of all non-repeated variables and functions in the second deobfuscation code is not within a preset ratio range, so that when the ratio is not within the preset ratio range, the variable names and function names in the second deobfuscation code are renamed.

10. A code deobfuscation device, It is characterized in that include: An acquisition module is used to obtain a token list of the code to be processed, where the code to be processed is the obfuscated code; a processing module, configured to perform deobfuscation processing on each Token in the Token list to obtain a first deobfuscation code, parse the first deobfuscation code to obtain an abstract syntax tree AST of the first deobfuscation code, simulate execution of the content in the target node of the AST to obtain a first execution result, replace the content of the target node according to the first execution result, and obtain a final deobfuscation code according to the replaced AST; Wherein, each Token in the Token list is each code word constituting the code to be processed; The target node is a node of a set type in the AST.

11. An electronic device, It is characterized in that include: Processor, memory and communication bus; The communication bus is used to realize the connection and communication between the processor and the memory; The processor is used to execute the program stored in the memory to implement the code deobfuscation method according to any one of claims 1 to 9.

12. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the code deobfuscation method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Code processing method and apparatus as well as computing device

    CN106203007A

  • Anti-confusion processing method, terminal, and computer device

    CN109033764A