Method and system for generating antagonistic malicious software based on interpretability technology
By using interpretability techniques and reinforcement learning to generate adversarial malware in black box environments, the problems of low escape rates and poor migration in the existing technology are solved, and efficient and extensive malware escape capabilities are achieved.
Patent Information
- Application Number
- CN202510184820.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to efficiently generate adversarial malware with high escape rates in black box environments, and the generated adversarial malware is low in migration, so it is impossible to deal with multiple malware detectors widely.
Using an interpretability technique method, the SHAP value is calculated through the feature extractor of the EMBER model and SOREL model and the TreeExplainer module of the SHAP library, and the malware is modified by combining the UCB algorithm and reinforcement learning selection actions until the escape is successful or the maximum number of attempts is reached.
It significantly improves the escape rate and migration of adversarial malware, shortens the reinforcement learning action selection sequence, reduces the number of queries on the detection model, and accelerates the learning process.
Smart Images

Figure CN120012087A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information security technology, and in particular to a method and system for generating adversarial malware based on explainability technology. Background Art
[0002] Malware detectors are the primary defense against malware. Recently, static anti-malware detectors have made significant progress in identifying malware with the help of deep learning technology, and modern commercial malware detectors are increasingly relying on machine learning technology. Nevertheless, malware attackers use various evasion techniques to design malware that can evade machine learning detectors, making these detection systems vulnerable to adversarial attacks. In response to this situation, security companies are constantly developing new defense strategies to deal with malware attacks. At the same time, malware attackers are also constantly adopting new technologies to attack these detection systems, making the malware detection problem a continuous game between attackers and defenders. Studying malware attack strategies can not only provide new ideas for developing anti-malware solutions, but also promote security defenders to develop new defense technologies, explore and strengthen the vulnerabilities of existing detection systems, thereby enhancing the detection capabilities of unknown malware attacks and improving the overall level of network security.
[0003] In the field of adversarial malware generation, software that can evade detection is mainly generated by modifying PE (Portable Executable) files. These modifications usually include modifying certain attributes of the PE file header, appending bytes extracted from benign samples to the end of the PE file, and encrypting certain parts of the file using encryption algorithms. These modification techniques are the key to generating adversarial malware with high evasion rates. In actual malware attacks, the malware detectors faced by attackers are black-box. In a black-box environment, it is difficult to directly understand how the detector judges the maliciousness of a file because the specific architecture, parameters, or gradient information of the detection model cannot be obtained. Reinforcement learning can solve the problem of strategy selection in the process of interaction between the agent and the environment. Using reinforcement learning, the optimal modification action can be selected by the reinforcement learning algorithm to mutate the malware during the interaction with the environment, and adversarial malware can be generated by using a certain modification sequence to evade the detection of the malware detector. The selection of filling content in reinforcement learning can be regarded as a MAB (Multi-armed Bandits) problem. Faced with multiple filling contents, without prior knowledge, it is impossible to know the impact of each filling content on escaping the detection of the detector.
[0004] In view of the above analysis, the technical problems that need to be solved urgently in the prior art are:
[0005] (1) In a black-box environment, attackers cannot access key information such as the internal architecture, parameters, and gradients of the malware detection model. This results in extremely low efficiency in training the modification agent, because the modification agent cannot obtain clear guidance and useful feedback from the environment, making the process of creating adversarial software particularly difficult. In addition, the sequence of modification operations required to generate efficient adversarial software using reinforcement learning is usually long, which not only increases the query frequency of the malware detection model, but also increases the complexity of the operation and resource consumption. These factors together hinder attackers from quickly adjusting and optimizing the evasion strategies of malware, resulting in low evasion rates and limited availability of adversarial malware. For these reasons, it is difficult for attackers to quickly adjust the evasion strategies used to modify malware to bypass the detection system, resulting in low evasion rates and low availability of adversarial malware.
[0006] (2) In the process of modifying PE malware, the basic modification operation of PE malware usually randomly selects padding content, and selecting appropriate padding content is crucial to bypassing the detection system. If the selection is inappropriate, the reinforcement learning modification agent will find it difficult to obtain the maximum reward, and thus will not be able to select the most effective padding content to modify the malware, and thus will not be able to generate adversarial malware with a high escape rate.
[0007] (3) The adversarial malware generated by current research methods has low transferability. Although it is effective against a single detector, it cannot cope with multiple malware detectors on a wide scale, and its evasion effect in commercial detectors is poor. Summary of the invention
[0008] The present invention provides an adversarial malware generation method and system based on explainability technology, which are used to solve the defects of the prior art that adversarial malware with a high escape rate cannot be generated and the mobility is low, and achieve the improvement of the generation capability and mobility of adversarial malware.
[0009] The present invention provides a method for generating adversarial malware based on explainability technology, comprising:
[0010] Extract features of malware using feature extractors of the EMBER model and the SOREL model, interpret the EMBER model and the SOREL model using the TreeExplainer module of the explainability technology SHAP library, calculate SHAP values of the extracted features, and calculate SHAP priorities of actions in the action modifier of reinforcement learning according to the SHAP values of the extracted features;
[0011] Selecting filler data from a data pool using the UCB algorithm, performing the action selected by the reinforcement learning according to the filler data to modify the PE file of the malware, and performing escape evaluation on the modified malware using multiple malware detectors until the malware successfully escapes or reaches a maximum number of attempted modifications, thereby obtaining adversarial malware;
[0012] The reinforcement learning process includes: if the escape is unsuccessful, determining the reward function value according to the detection results of the malware by the multiple malware detectors, the length of the executed action and the difference before and after the PE file is modified; obtaining the external reward generated in the process of interacting with the reinforcement learning environment according to the SHAP priority, the reward function value and the TD error of the reinforcement learning agent; using the ICM module to calculate the internal reward of the reinforcement learning agent, learning according to the external reward and the internal reward, and selecting the action for each step.
[0013] According to an interpretability technology-based adversarial malware generation system provided by the present invention, a SHAP priority calculation module is used to extract features from malware using feature extractors of an EMBER model and a SOREL model, interpret the EMBER model and the SOREL model using a TreeExplainer module of an interpretability technology SHAP library, calculate SHAP values of the extracted features, and calculate the SHAP priority of an action in an action modifier of reinforcement learning according to the SHAP value of the extracted features;
[0014] an adversarial malware generation module, configured to select filler data from a data pool using a UCB algorithm, modify the PE file of the malware by executing the action selected by the reinforcement learning according to the filler data, and perform escape evaluation on the modified malware using multiple malware detectors until the malware successfully escapes the multiple malware detectors or reaches a maximum number of attempted modifications, thereby obtaining adversarial malware;
[0015] The reinforcement learning agent learning module is used to, if the escape is unsuccessful, determine the reward function value according to the detection results of the malware by the multiple malware detectors, the length of the executed action and the difference before and after the PE file is modified; obtain the external reward generated in the process of interacting with the reinforcement learning environment according to the SHAP priority, the reward function value and the TD error of the reinforcement learning agent; use the ICM module to calculate the internal reward of the reinforcement learning agent, learn according to the external reward and the internal reward, and select the action for each step.
[0016] The present invention provides an adversarial malware generation method and system based on explainability technology. Based on SHAP priority, reward function value and TD error of reinforcement learning agent, the invention realizes modification of malware by selecting actions from action space through reinforcement learning with priority experience replay, thereby effectively shortening the reinforcement learning action selection sequence and improving attack escape rate and mobility. The UCB algorithm is used to select the highest confidence modification content to solve the problem of random selection of reinforcement learning modification action operation content. The ICM module is introduced to encourage the reinforcement learning modification agent to explore those states that have not been fully explored, thereby greatly reducing the number of queries to the detection model when using the reinforcement learning agent to generate adversarial malware, improving the exploration rate of the reinforcement learning modification agent, accelerating the learning process, and improving the generation capability of adversarial malware. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 It is a flowchart of a method for generating adversarial malware based on explainability technology provided by the present invention;
[0019] Figure 2 It is a complete flow chart of the method for generating adversarial malware based on explainability technology provided by the present invention;
[0020] Figure 3 It is a schematic diagram of the framework structure of the method for generating adversarial malware based on explainability technology provided by the present invention;
[0021] Figure 4 It is a schematic diagram of the success rate of each round of malware in the adversarial malware generation method based on explainability technology provided by the present invention;
[0022] Figure 5 It is a schematic diagram of the proportion of actions performed on malware in the method for generating adversarial malware based on explainability technology provided by the present invention;
[0023] Figure 6 It is a flow chart of selecting filling content by the UCB algorithm in the adversarial malware generation method based on explainability technology provided by the present invention;
[0024] Figure 7 It is a schematic diagram of the calculation process of the priority of the SHAP value of the modification action in the method for generating adversarial malware based on the explainability technology provided by the present invention;
[0025] Figure 8 It is a flow chart of hole filling in the method for generating adversarial malware based on explainability technology provided by the present invention;
[0026] Fig. 9 It is a schematic diagram of selecting operation contents based on the UCB algorithm in the method for generating adversarial malware based on explainability technology provided by the present invention;
[0027] Fig.10 It is a schematic diagram of a detector classification pool in the adversarial malware generation method based on explainability technology provided by the present invention;
[0028] Fig.11 It is a structural schematic diagram of the adversarial malware generation system based on explainability technology provided by the present invention;
[0029] Fig.12 It is a graph showing the attack success rates of five malware families in the adversarial malware generation method based on explainability technology provided by the present invention;
[0030] Fig.13 This is a diagram showing the escaping effect of the adversarial malware generation method based on explainability technology provided by the present invention on the SOREL detector;
[0031] Fig.14 This is a diagram showing the escaping effect of the adversarial malware generation method based on explainability technology provided by the present invention on the EMBER detector;
[0032] Fig.15 This is an ablation experiment effect diagram of the adversarial malware generation method based on explainability technology provided by the present invention. DETAILED DESCRIPTION
[0033] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0034] In a black box environment, it is difficult to directly understand how the detector judges the maliciousness of a file because it is impossible to obtain the specific architecture, parameters, or gradient information of the detection model. Based on the real problems in these attacks, the SHAP value obtained using the explainable SHAP (SHapley Additive exPlanations) technology can help understand how each feature affects the detection results. Combining the SHAP value of each feature helps to select the most evasive modification operation for malware mutation.
[0035] The selection of filling content in reinforcement learning can be regarded as a MAB problem. Faced with multiple filling contents, without prior knowledge, it is impossible to know the impact of each filling content on the detection of the escape detector. The UCB (Upper Confidence Bound) algorithm selects the filling content with the largest confidence interval by calculating the confidence interval of each filling content, and combines the modification operation selected by reinforcement learning to mutate the malware.
[0036] In the face of sparse reward environments, the ICM (Intrinsic Curiosity Module) mechanism can provide intrinsic motivation for reinforcement learning and encourage agents to explore underutilized strategies. This approach not only improves the efficiency of reinforcement learning models, but also improves the ability to generate adversarial malware.
[0037] In network security defense, antivirus software and intrusion detection system (IDS) are one of the key protection methods. However, with the continuous advancement of detection technology, attackers are also constantly evolving their attack methods, generating adversarial malware to evade the recognition of detection systems, posing a huge threat to the network security of enterprises and institutions. The method of the present invention can be used to develop an adversarial malware generation tool to test and strengthen the detection capability of its enterprise-level antivirus software.
[0038] Combine the following Figure 1 A method for generating adversarial malware based on explainability technology of the present invention is described, comprising:
[0039] Step 101, extracting features from malware using feature extractors of the EMBER model and the SOREL model, interpreting the EMBER model and the SOREL model (malware detection model) using the TreeExplainer module of the explainability technology SHAP library, calculating the SHAP value of the extracted features, and calculating the SHAP priority of the action in the action modifier of reinforcement learning according to the SHAP value of the extracted features;
[0040] Step 102, using the UCB algorithm to select filler data from the data pool, performing the action selected by the reinforcement learning according to the filler data to modify the PE file of the malware, and using multiple malware detectors to perform escape evaluation on the modified malware until the malware successfully escapes the multiple malware detectors or reaches a maximum number of attempted modifications, thereby obtaining adversarial malware;
[0041] Step 103, the reinforcement learning process includes determining a reward function value according to the detection results of the malware by the multiple malware detectors, the length of the executed action, and the difference before and after the PE file is modified if the escape is not successful; obtaining an external reward generated in the process of interacting with the reinforcement learning environment according to the SHAP priority, the reward function value and the TD error of the reinforcement learning agent; using the ICM module to calculate the internal reward of the reinforcement learning agent, learning according to the external reward and the internal reward, and selecting the action for each step.
[0042] This embodiment proposes a method for generating adversarial malware based on PERD3QN (Priority Experience Replay Deep Double Q Network). The specific process is as follows: Figure 2 The corresponding framework structure is shown in Figure 3 As shown. First, the SHAP priority of the reinforcement learning action modifier is calculated, and then the adversarial malware is generated based on the reinforcement learning action modifier of the UCB algorithm. During the generation of adversarial malware, reinforcement learning agent learning based on prioritized experience replay is performed. The experience is stored with priority based on the SHAP priority and reward function priority obtained; the intrinsic reward is calculated using the ICM module and added to the external reward to perform the agent learning process; the above process is repeated until the malware escapes successfully or the maximum number of modification attempts is reached. The specific steps include:
[0043] 1. Initial preparation:
[0044] A database containing various known malware samples is selected and features are extracted from these samples using the feature extractors of the EMBER and SOREL models.
[0045] The extracted features are input into the TreeExplainer module of the SHAP library to calculate the feature SHAP value of each malware sample and determine the priority of these features in evading detection.
[0046] 2. Malware Generation:
[0047] The UCB algorithm of the present invention is used to select malware samples from the data pool for modification. By introducing multiple attack algorithms, such as modification, appending and encryption of PE files, the malware samples are automatically modified.
[0048] For each modified sample, a reinforcement learning algorithm is used to evaluate its evasion performance in multiple malware detectors, and the next modification strategy is adjusted based on the results. This process is repeated until an adversarial malware sample that can successfully evade is generated. The success rate of each round of malware is as follows: Figure 4 As shown, the actions performed on the malware are as follows: Figure 5 shown.
[0049] 3. Functional integrity test:
[0050] The generated adversarial malware samples are executed in a virtualized sandbox environment to ensure that they can fully perform their malicious functions, such as file deletion, information theft, or remote control, without being detected.
[0051] 4. Escape effect evaluation:
[0052] Use multiple antivirus software and IDS systems to perform evasion tests on the generated adversarial malware samples. By comparing the detection results, the detection capabilities of the defense system are evaluated and its vulnerabilities and blind spots are identified.
[0053] 5. Defense system optimization:
[0054] Based on the generated adversarial malware samples, update and optimize antivirus software and IDS systems to enhance their detection capabilities against the latest adversarial attacks.
[0055] Patching vulnerabilities and continuously generating new adversarial malware samples for continuous testing and improvement ensures that enterprise-level defense systems can respond quickly and effectively when facing new attacks.
[0056] Through this adversarial malware generation tool, the network security company has significantly improved the detection capabilities of its antivirus software and IDS systems, and has successfully prevented multiple malware attacks against corporate networks in actual applications, reducing the risk of network attacks and providing corporate customers with more reliable security protection.
[0057] This adversarial malware generation method is not only suitable for internal testing and defense system optimization of cybersecurity companies, but can also be extended to other companies and institutions that need to enhance their network defense capabilities for continuous detection and improvement of their network security defense capabilities.
[0058] The process of using the UCB algorithm to select fill data from the data pool to modify the PE file is as follows: Figure 6The specific operation is: each time an arm of the reinforcement learning modification action is tested, if there is an arm that has not been tested, one of the arms is selected, otherwise the UCB value of each arm is calculated according to the UCB algorithm, the data corresponding to the arm with the largest UCB value is selected, combined with the modification operation selected by reinforcement learning to modify the PE file, and the UCB algorithm is updated. The calculation formula of the UCB value is:
[0059]
[0060] in, is the average return of operation i, x is the operation, N is the total number of choices, and n i represents the number of selections of operation i.
[0061] The present invention designs a SHAP priority calculation method, uses EMBER and SOREL feature extractors to extract features from malware, then uses the TreeExplainer module of the SHAP library to interpret the EMBER and SOREL models, and calculates the SHAP value of the extracted features. Through the SHAP value, the impact of each feature on the model decision can be quantified, and then the SHAP priority of each action in the reinforcement learning action modifier can be calculated. The present invention can better understand the impact of features on detection results, so that the modification agent is more instructive when selecting modification operations, thereby improving training efficiency and escape rate.
[0062] This paper designs a reward function calculation mechanism that considers multiple factors to solve the inefficiency of generating adversarial malware in a black box environment. By combining detection rewards, distance rewards, and similarity rewards, the reinforcement learning agent can optimize the escape rate while maintaining the similarity of files.
[0063] The present invention introduces the D3QN algorithm into the ICM mechanism to provide intrinsic motivation for the reinforcement learning agent and encourage the agent to explore underutilized strategies.
[0064] The reinforcement learning content selection operation based on the UCB algorithm provided by the present invention uses the UCB algorithm to select filler data from a data pool. The UCB algorithm calculates the confidence interval of each filler content, selects the filler content with the largest confidence interval, and mutates the malware in combination with the modification operation selected by reinforcement learning. It can adjust the selection strategy based on previous trials and errors, optimize future filler content selection, solve the problem of randomization of filler content selection, make the modification operation more accurate, achieve the purpose of maximizing the reward of the reinforcement learning modification agent, and significantly improve the generation quality of adversarial malware.
[0065] This embodiment selects actions from the action space to modify malware through reinforcement learning based on SHAP priority, reward function value and TD error of reinforcement learning agent, effectively shortens the reinforcement learning action selection sequence, and can avoid the detection model after taking 4 modification actions. The escape rates for EMBER and SOREL models reach 90.45% and 93.14% respectively, which improves the attack escape rate and mobility; the use of UCB algorithm to select the highest confidence modification content solves the problem of random selection of reinforcement learning modification action operation content, and introduces the ICM module to encourage the reinforcement learning modification agent to explore those states that have not been fully explored, which greatly reduces the number of queries to the detection model when using the reinforcement learning agent to generate adversarial malware, improves the exploration rate of the reinforcement learning modification agent, accelerates the learning process, and improves the ability to generate adversarial malware.
[0066] Based on the above embodiment, in this embodiment, the SHAP priority of the action in the action modifier of reinforcement learning is calculated according to the SHAP value of the extracted feature, including:
[0067] constructing a first matrix according to the effect of each action in the action modifier on each feature;
[0068] A second matrix is formed according to the SHAP values of the features extracted by the EMBER model, a third matrix is formed according to the SHAP values of the features extracted by the SOREL model, and SHAP values less than 0 in the second matrix and the third matrix are set to zero;
[0069] Multiplying the first matrix by the second matrix after being reset to zero to obtain the SHAP priority of the action under the EMBER model;
[0070] Multiplying the first matrix by the third matrix after being reset to zero to obtain the SHAP priority of the action under the SOREL model;
[0071] The SHAP priorities of the action under the EMBER model and the SOREL model are weighted and summed to obtain the final SHAP priority of the action.
[0072] like Figure 7As shown in the figure, the feature extractors of the EMBER detection model and the SOREL detection model are used to extract the features of the selected malware samples to represent the state of the environment. The features are represented by F = [f1, f2, f3, f4, f5, f6, f7, f8, f9]. The specific features are shown in Table 1, where the extracted features of the MBER model are represented by ember_obervation_space and the extracted features of the Sorel model are represented by sorel_obervation_space.
[0073] Table 1 Feature description table
[0074]
[0075] Use the TreeExplainer module of the SHAP library to interpret the EMBER detector. The resulting interpreter ember_explainer calculates the average SHAP value S for each feature of ember_obervation_space. ember , expressed as:
[0076] S ember =[s1,s2,s3,...s7,s8,s9]
[0077] s i The calculation formula is:
[0078]
[0079] Where f(i) is the impact of feature i on model prediction when it acts alone, f(s) is the prediction output of the model based on feature subset S, N is the set of all features, and f(S∪{i}-f(S)) is the change in model prediction after feature i is added to S.
[0080] Use the TreeExplainer module of the SHAP library to interpret the Sorel detector. The interpreter sorel_explainer calculates the average SHAP value S for each feature of sorel_obervation_space. sorel , expressed as:
[0081] S sorel =[s1,s2,s3,...s7,s8,s9,]
[0082] In order to calculate the SHAP priority of the modification action in the reinforcement learning action space, we first need to quantify the impact of each modification action on the extracted features. A matrix is used to represent the impact of each modification action on the feature. The value of each row in the matrix is 1, indicating that the modification action has an impact on the feature, and the value is 0, indicating that the modification action has no impact on the feature. The matrix T is expressed as:
[0083]
[0084] Then for S ember and S sorel Each SHAP value s in i Processing to obtain S′ ember and S′ sorel , keep s greater than zero i , s less than zero i The purpose of setting it to zero is to calculate the SHAP value that retains the positive effect of each modification operation on the detection model's prediction of benignity.
[0085] Concatenate the matrix T with S′ ember , S′ sorel Multiply them together to get the operation priority P under the EMBER and SOREL models ember and P sorel , expressed as:
[0086] P ember =T*S′ ember
[0087] P sorel =T*S′ sorel
[0088] Finally, P ember and P sorel Perform weighted summation and normalization to get the SHAP priority:
[0089] P=w1*P ember +w2*P sorel
[0090] Based on the above embodiment, before calculating the SHAP priority of the action in the action modifier of reinforcement learning according to the SHAP value of the extracted feature, this embodiment further includes:
[0091] An action modifier is constructed to modify, append and encrypt the PE file.
[0092] As shown in Table 2, the main attack operations of the action modifier include the following methods:
[0093] 1. "add_imports": select an import function from the import function data pool through the UCB algorithm and add it to the added import item of the PE file.
[0094] 2. "append_benign_data_overlay": selects 1 data from a data pool containing 50 benign data through the UCB algorithm and appends it to the end of the PE file.
[0095] 3. "add_section_benign_data": select a new section from the section data pool through the UCB algorithm and add it to the PE file.
[0096] 4. "add_optional_header_dllchlist": selects a dll list from the dll list data pool through the UCB algorithm to add DLL features to the optional header of the PE file.
[0097] 5. "pert_data_directory": randomly selects a data directory and sets a new random RVA and size for it.
[0098] 6. "change_optional_header_dllch": Modify the DLLCharacteristics field in the optional header of the PE file.
[0099] 7. Modify_coff_header: Change the field value in the COFF header of the PE file, such as modifying the program's adjustment code segment and symbol information.
[0100] 8. "modify_optional_header": Modify the linker / image / operatingsystem version number information in the PE file header.
[0101] 9. "modify_timestamp": Select a timestamp from the timestamp data pool to modify the timestamp of the PE file.
[0102] 10. "dynamic_code_cave_add": select several rounds of UCB algorithm to inject some random byte sequences into the code cave, and find the bytes after the code cave injection with the lowest detector score. Figure 8As shown in the figure, the specific operations are as follows: first, the number of code hole injections is selected through the UCB algorithm, and the data of each section of the PE file is read into and stored in the memory; the prepared section data pool of benign content is traversed, and the files whose section size is larger than the section alignment value in the optional header are saved in the section candidate list; each section is traversed to add a code hole for each section, and a section containing benign data is randomly selected from the section candidate list, and data of the same size as the code hole is obtained from the selected benign data section. If the currently traversed segment is the first segment, the code hole data is directly added to the end of the current segment, if not, the code hole data is added to the end of the previous segment until the last segment, and the code hole filling is completed; then the filled binary bytes are evaluated by the EMBER detector. If the score is less than the threshold, the dynamic hole adding operation is exited, otherwise the above steps are continued to save the binary data with the lowest detection score of the EMBER detector, which is regarded as the final mutated data.
[0103] 11. "upx_pack": Select an encryption level through the UCB algorithm and use UPX to pack the PE file.
[0104] 12. "PE_pack": Perform XOR encryption on the entire PE file. The specific operation is: first generate a random key to XOR encrypt the PE file content, generate a C language program with the encrypted content, and finally use the i686w64mingw32gcc compiler to compile the generated C language program to generate an executable file. The executable file contains the decryption logic and the function of executing the decrypted file. The decryption logic will use the generated random key to XOR decrypt the encrypted data, ensuring the functional integrity of the encrypted PE file.
[0105] Table 2 Description of attack operations
[0106]
[0107]
[0108] On the basis of the above embodiment, in this embodiment, the reward function value is determined according to the detection results of the malware by the multiple malware detectors, the length of the executed action and the difference before and after the PE file is modified, including:
[0109] Calculate the detection reward according to the difference between the detection scores of each malware detector when the malware is executed and when the malware is not executed, or the reward value of the detection result after the malware is executed;
[0110] Calculating a distance reward according to the length of the action executed on the malware, wherein the longer the action length, the lower the distance reward;
[0111] Calculate similarity rewards based on the similarity between the PE files before and after modification;
[0112] The reward function value is calculated according to the detection reward, the distance reward and the similarity reward.
[0113] like Fig. 9 As shown in Figure 1, the reinforcement learning modification agent will perform modification operations based on the UCB algorithm to mutate the PE file. Fig.10 As shown in the figure, the mutated PE file is input into the academic EMBER, SOREL, Ngram detector and the commercial detector ClamAV for detection. The detection scores are S1, S2, S3, S4. The comprehensive detection score S is calculated according to a certain weight. Finally, S is compared with the threshold. If it is less than the threshold, the escape is successful, otherwise the escape fails. The calculation formula of S is:
[0114] S=w1*s1+w2*s2+w3*s3+w4*s4
[0115] If the escape is successful, the modification is stopped and adversarial malware is generated. If the escape fails, the detection reward calculation, distance reward calculation, similarity reward calculation and comprehensive reward function value calculation are performed, and reinforcement learning is performed based on the calculation results.
[0116] Based on the above embodiment, the calculation formula of the detection reward R_d in this embodiment is:
[0117] R_d=w1*ember_R+w2*Sorel_R+w3*clamav_R+w4*Ngram_R
[0118] Among them, the scores of the four malware detectors are ember_R, sorel_R, clamav_R and ngram_R, respectively. ember_R and sorel_R are the difference between the detection scores of executing the action and not executing the action on the malware. clamav_reward and ngram_reward are the rewards for the detection results after the malware detector executes the action on the malware. If the detection result is benign, the reward is 1, otherwise the reward is 0. w1, w2, w3 and w4 are the corresponding weights.
[0119] The calculation formula of the distance reward R_dis is:
[0120]
[0121] Where R is the given basic reward, t is the length of the action executed, and T is the maximum modification length;
[0122] The calculation formula of the similarity reward R_sim is:
[0123] R_sim=cos(v1,v2)
[0124] Among them, cos is the cosine similarity, v1 and v2 are the vector representations of the PE file status before and after modification;
[0125] The calculation formula of the reward function value R is:
[0126] R=w5*R_det+w6*R_dis+w7*R_sim
[0127] Among them, w5, w6 and w7 are corresponding weights.
[0128] On the basis of the above embodiment, in this embodiment, the external reward P1 generated during the interaction with the reinforcement learning environment is obtained according to the SHAP priority, the reward function value and the TD error of the reinforcement learning agent through the following formula:
[0129] P1=TD pri +SHAP pri *θ1+R pri *θ2
[0130] Among them, TD pri SHAP pri and R pri are the TD error, SHAP priority and reward function value respectively, and θ1 and θ2 are preset coefficients.
[0131] The experience is stored according to the SHAP priority and reward function priority. Figure 3 As shown, the specific operation is: according to the reward R, SHAP value and TD (Time Difference) error of the reinforcement learning agent generated during the interaction with the reinforcement learning environment, the priority P1 of the experience is designed, which is used as the index for storing the experience samples, and the experience samples are stored in SumTree.
[0132] Then, the reinforcement learning D3QN agent performs biased sampling of experiences in the experience storage pool according to the priority level for the next round of interaction with the reinforcement learning environment.
[0133] Based on the above embodiment, in this embodiment, the ICM module is used to calculate the internal reward of the reinforcement learning agent, and learning is performed according to the external reward and the internal reward, and the action of each step is selected, including:
[0134] Extract features from the current state of the input PE file using the feature extractor in the ICM module;
[0135] Using the prediction model in the ICM module to predict the features of the next state of the PE file based on the features of the current state of the PE file;
[0136] Using the reverse model in the ICM module to predict actions based on the features of the current state and the features of the next state of the PE file;
[0137] The D3QN reinforcement learning agent is used to generate an intrinsic reward according to the ICM module during the experience learning process, update the target Q value according to the intrinsic reward and the extrinsic reward, and determine the next action according to the updated target Q value.
[0138] Use the ICM module to calculate the intrinsic reward and add it to the external reward to carry out the agent's learning process. The specific operation is: first construct the ICM module, which contains three major parts: feature extractor, prediction model and reverse model. The feature extractor is used to extract features from the input state, including two layers of linear transformation and a ReLU activation function; the prediction model is used to predict the features of the next state, including two layers of linear transformation and a ReLU activation function; the reverse model predicts actions from the current state features and the next state features. Then, during the experience learning process, the D3QN reinforcement learning agent generates an intrinsic reward intrinsic_reward according to the ICM module. The reinforcement learning agent updates the target Q value according to the intrinsic reward and the external reward generated by the interaction with the environment to carry out the agent's learning process. Repeat the above process until the modified malware evades successfully or the maximum number of attempts is reached.
[0139] Based on the above embodiments, after obtaining the antagonistic malware, this embodiment further includes:
[0140] parsing the adversarial malware;
[0141] If the analysis is successful, it is known that the structure of the adversarial malware is complete, and the adversarial malware is placed in a sandbox for execution;
[0142] If the execution is successful, it is known that the functionality of the adversarial malware is complete, and a malware detector is used to perform an escape assessment and a migration assessment on the adversarial malware.
[0143] Functional integrity testing and effectiveness evaluation of adversarial malware: The generated adversarial malware samples are executed in a sandbox to ensure their functional integrity and executableness, and the generated adversarial malware is evaluated for evasion using a malware detector.
[0144] The specific operation of the anti-malware function integrity test is as follows: first, use the pefile library to parse the adversarial malware that has successfully escaped. If the parsing is successful, the surface structure is complete. Then put the structurally complete adversarial malware into the sandbox for execution, and observe whether the modified malware maintains the same malicious functions as before the modification. If the execution is successful, it means that the function is complete, otherwise the function is destroyed.
[0145] The generated adversarial malware was evaluated for evasion using academic malware detectors, and the generated malware was evaluated for migration using academic and commercial detectors. The generated adversarial software was evaluated for evasion rate using EMBER and SOREL. The generated adversarial software was tested for attack migration using EMBER, SOREL, 1gram, clamav, Tencent Manager, Comodo, Rising, etc. The evasion success rate of the generated adversarial malware for the five major malware families of backdoor, Ransom, Spyware, VirTool, and Virus was higher than all baseline methods.
[0146] The present invention introduces academic and commercial multi-detector escape attacks in the process of generating adversarial malware. By calculating a comprehensive detection score and comparing it with a threshold, adversarial malware is generated. An escape rate of more than 90% is achieved in three academic and seven commercial malware detectors, thereby improving the migration attack capability of adversarial malware.
[0147] The present invention performs a functional integrity test on the generated adversarial malware in a sandbox to ensure that the functionality is not damaged while the malware escapes detection, thereby maintaining the execution capability and functional integrity of the malware.
[0148] The adversarial malware generation system based on explainability technology provided by the present invention is described below. The adversarial malware generation system based on explainability technology described below and the adversarial malware generation method based on explainability technology described above can refer to each other.
[0149] like Fig.11 As shown, the system includes:
[0150] A SHAP priority calculation module, for extracting features from malware using feature extractors of an EMBER model and a SOREL model, interpreting the EMBER model and the SOREL model using a TreeExplainer module of an explainability technology SHAP library, calculating SHAP values of the extracted features, and calculating SHAP priorities of actions in an action modifier of reinforcement learning according to the SHAP values of the extracted features;
[0151] an adversarial malware generation module, configured to select filler data from a data pool using a UCB algorithm, modify the PE file of the malware by executing the action selected by the reinforcement learning according to the filler data, and perform escape evaluation on the modified malware using multiple malware detectors until the malware successfully escapes the multiple malware detectors or reaches a maximum number of attempted modifications, thereby obtaining adversarial malware;
[0152] The reinforcement learning agent learning module is used to, if the escape is unsuccessful, determine the reward function value according to the detection results of the malware by the multiple malware detectors, the length of the executed action and the difference before and after the PE file is modified; obtain the external reward generated in the process of interacting with the reinforcement learning environment according to the SHAP priority, the reward function value and the TD error of the reinforcement learning agent; use the ICM module to calculate the internal reward of the reinforcement learning agent, learn according to the external reward and the internal reward, and select the action for each step.
[0153] The working principle of the adversarial malware generation based on explainability technology provided by the embodiment of the present invention can be described in detail as follows:
[0154] 1.SHAP priority calculation module
[0155] Feature extraction: Use EMBER and SOREL feature extractors to perform deep feature extraction on malware. These features include static code features, dynamic behavior features, etc.
[0156] Explainability analysis: The TreeExplainer module of the SHAP library is used to interpret the EMBER and SOREL models and calculate the SHAP value of each feature. The SHAP value reflects the degree of influence of the feature on the model prediction results.
[0157] Action priority update: Based on the SHAP values of the extracted features, the SHAP priority of each action in the reinforcement learning action modifier is calculated. These priorities will guide the subsequent modification action selection to improve the efficiency of adversarial malware generation.
[0158] 2. Adversarial Malware Generation Module
[0159] Data selection: Use the UCB (Upper Confidence Bound) algorithm to select fill data from the data pool. This algorithm balances exploration and utilization and helps to discover more effective modification strategies.
[0160] Action modifier construction: Build an action modifier that contains multiple attack algorithms for modifying, appending, encrypting, etc. PE files (Portable Executable files, executable files under Windows). These actions are designed to change the characteristics of malware to evade detection.
[0161] Modification and Evaluation: Execute the modification actions selected by reinforcement learning and use multiple malware detectors to evaluate the modified malware for evasion. Repeat this process until the malware successfully evades detection or the maximum number of attempted modifications is reached.
[0162] 3. Reward function design module
[0163] Detection Reward: When malware successfully evades a detector, a certain reward value is given to encourage the generation of malware with better evasion capabilities.
[0164] Distance reward: rewards are given based on the distance between the modified malware and the original malware in the feature space to encourage the generation of samples that are significantly different from the original malware.
[0165] Similarity Reward: Encourages the generation of samples that are functionally similar to the original malware to ensure the functional integrity of the adversarial malware.
[0166] Comprehensive reward function: The above rewards are combined to form a comprehensive reward function to guide the learning process of the reinforcement learning agent.
[0167] 4. Reinforcement Learning Agent Learning Module
[0168] Priority storage: Experience is stored in priority according to SHAP priority and reward function priority, giving priority to experience that contributes more to agent learning.
[0169] Intrinsic reward calculation: Use the ICM module to calculate intrinsic rewards, which encourages the agent to explore unknown areas and discover more effective modification strategies.
[0170] Learning process: Add the intrinsic reward to the external reward to conduct the agent's learning process. Through continuous trial and error and feedback, the strategy is optimized and modified to improve the efficiency of generating adversarial malware.
[0171] 5. Functional integrity test and effect evaluation module
[0172] Sandbox execution: Execute the generated adversarial malware samples in a sandbox environment to ensure that they are fully functional and executable. The sandbox environment isolates the malware from potential harm to the real system.
[0173] Evasion Evaluation: Use malware detectors to perform evasion evaluation on the generated adversarial malware to verify its evasion capabilities.
[0174] This embodiment realizes the automatic generation of adversarial malware by integrating multiple modules such as feature extraction, explainability analysis, reinforcement learning, and reward function design. The system can not only improve the efficiency of generating adversarial malware, but also ensure its functional integrity and evasion capabilities, providing a powerful tool for research in the field of network security.
[0175] Relevant evidence of the technical effects achieved by the embodiments of the present invention.
[0176] 1. Escape rate comparison experiment: The escape rate comparison experiment generated 9387 original samples from 5 malware families and attacked the EMBER and SOREL target malware detectors. On this basis, the results of the baseline were compared and analyzed.
[0177] 2. Attack migration comparison experiment: In order to prove that the generated adversarial software has good migration attack capability, we attack three detectors from the academic community, ember, Sorel, 1gram, and the commercial community, clamav, Tencent Antivirus, Comodo, and Rising Star. The attack migration capability is compared by comparing the escape rates of all detectors.
[0178] 3. Comparison experiment of action selection sequence length: In order to prove that the present invention can shorten the action selection sequence length, the average action selection sequence length of all samples generated adversarially is shown. In the experiment, it is compared with the baseline and the action selection sequence length of each family is tested on 5 malware families.
[0179] The embodiments of the present invention have achieved some positive effects during the development or use process, and indeed have great advantages over the prior art. The following content is described in conjunction with data, charts, etc. of the test process.
[0180] 1. Dataset
[0181] The data samples of the experiment of the present invention are shown in Table 3, which include 9387 malware samples from 5 malware families.
[0182] Table 3 Dataset introduction
[0183]
[0184] 2. Target model: The target models in the escape rate experiment are two mainstream academic detectors, EMBER and SOREL detectors. The target models in the transfer attack experiment are academic detectors ember, Sorel, 1gram and commercial detectors clamav, TencentAntivirus, Comodo, and Rising Star.
[0185] 3. Baseline method: In the escape rate experiment, it is compared with six reinforcement learning-based malware adversarial generation techniques (including DQEAF, A3CMal, MAB, gymmalware, AIMED, MalFoxl) and the black box attack method GAMMA based on genetic programming. In the migration attack test, it is compared with MAB and MalFoxl. In the action selection sequence length comparison experiment, it is compared with the reinforcement learning method malware rl.
[0186] 4. Evaluation indicators: The experiment mainly compares the escape rate and the length of the reinforcement learning action selection sequence. The escape rate is the proportion of samples that escape the detector in the total samples, and the action selection sequence length is the average length of the actions taken by all samples for adversarial generation.
[0187] 5. Experimental settings: The maximum number of modifications in each round of agent training is modified to 60, the number of training batches is 32, the weight of SHAP value priority is 0.1, the priority weight of Reward is 0.0001, the weight of detection reward is 0.8, the weight of distance reward is 0.1, the weight of similarity reward is 0.1, and the model update interval is 100.
[0188] 6. Experimental results and analysis
[0189] The evasion rate detection results are shown in Table 4. The evasion success rate of adversarial malware generated for the five major malware families, backdoor, Ransom, Spyware, VirTool, and Virus, is higher than all the baseline methods.
[0190] Table 4 Escape rate detection
[0191]
[0192] The generated adversarial malware transfer attacks are shown in Table 5. The generated 9837 malwares achieved a success rate of more than 90% on ember, Sorel, 1gram, clamav, Tencent Manager, Comodo, and Rising. The generated adversarial malware can escape most commercial and academic malware detectors.
[0193] Table 5 Migration evaluation
[0194]
[0195] The average action sequence length of adversarial malware generated by the present invention is 4, which is higher than malware_rl, and the length of action selection sequence of reinforcement learning is reduced, as shown in Table 6. In the samples of five malware families, the escape rate of more than 90% can be experimentally achieved by using 7 modified actions. Fig.12 shown.
[0196] Compared with the baseline method, the present invention can greatly improve the escape rate of the generated adversarial software under different tasks, and can achieve a higher escape rate with fewer operations, such as Figures 13 to 15 shown.
[0197] Table 6 Action selection sequence length
[0198] method Avoidance success rate Average action selection sequence length The present invention 96.08% 4 malware_rl 89.20% 8.2
[0199] The expected benefits and commercial value of the technical solution of the present invention after transformation are:
[0200] Expected benefits: By simulating more advanced malware adversarial attacks, security companies can improve existing malware detection systems, improve defense capabilities against future potential threats, and enhance security defense capabilities. This invention can be directly used to develop more advanced security products, such as enhanced anti-malware tools, to meet the growing market demand for high security.
[0201] Commercial value: This invention can be used to optimize existing malware detection products and improve their market competitiveness. It can provide technical support for security companies to develop new defense strategies and products and open up new markets. This invention can promote the development of new security services and products, such as customized adversarial attack testing services.
[0202] The technical solution of the present invention fills the technical gap in the industry at home and abroad:
[0203] This paper combines the SHAP value priority calculation and the use of the UCB algorithm to provide a new perspective and method for the application of reinforcement learning in malware adversarial generation. The attack transferability of adversarial malware is improved, filling the technical gap of efficient malware evasion in actual multi-detector environments.
[0204] The technical solution of the present invention solves the technical problems that people have been eager to solve but have never been able to solve successfully:
[0205] Optimization problem in black box environment: How to efficiently generate adversarial malware in a black box environment has long been a difficult problem. The present invention provides a method that can effectively generate adversarial malware without knowing the internal structure by using SHAP value and UCB algorithm to optimize the modification operation selection.
[0206] Optimization of filling content selection: In the prior art, the randomness of filling content selection leads to poor results in generating adversarial malware. The present invention solves this problem through the UCB algorithm, making the filling content selection more accurate.
[0207] The problem of dynamically adapting to multiple detectors: Most existing adversarial malware generation methods are only effective for specific detectors. The present invention introduces the ICM module and the priority experience replay mechanism to attack the detection pool, thereby improving the migration capability of adversarial malware attacks and improving the versatility and practicality of adversarial malware.
[0208] The technical solution of the present invention overcomes technical prejudice:.
[0209] For the selection of filling content, traditional methods tend to choose randomly or experience-driven selection, but the introduction of the UCB algorithm breaks this bias and provides a method based on the confidence upper bound to optimize the selection process. Many reinforcement learning methods perform poorly in sparse reward environments, but the present invention overcomes this technical bias by introducing the ICM module and improves the learning efficiency and effect in sparse reward environments.
[0210] This paper proposes an adversarial malware generation method based on PERD3QN (Prioritized Experience Replay Deep Double Q Network), which solves multiple key problems in the prior art by combining reinforcement learning and explainability technology, and has achieved significant technological progress in industrial applications.
[0211] 1. Technical problems solved:
[0212] Traditional malware generation methods lack pertinence and flexibility: Existing malware generation techniques usually rely on manually written rules or static analysis, which is difficult to cope with the evasion requirements in multiple detector environments. This paper introduces reinforcement learning action modifiers and combines SHAP (SHapleyAdditive exPlanations) interpretive technology to achieve dynamic priority calculation of malware features, thereby guiding the adaptive generation of malware, effectively solving the limitations of traditional methods.
[0213] The efficiency and success rate of adversarial malware generation are low: Existing methods usually require a lot of attempts and modifications when generating adversarial malware, which is inefficient and has limited success rate. This invention significantly improves the efficiency of malware escape generation and reduces the computational cost of adversarial generation by combining the UCB (UpperConfidence Bound) algorithm with the priority experience replay mechanism.
[0214] Malware functional integrity is difficult to ensure: When generating adversarial malware, ensuring its functional integrity is a difficult point, and traditional methods often cause the generated malware to fail to execute normally. After generating adversarial malware, the present invention ensures that the generated malware can maintain the integrity of its original functions while escaping detection through testing and functional verification in a sandbox environment.
[0215] 2. Significant technological progress achieved:
[0216] Innovation of introducing explainability technology: This invention combines SHAP explainability technology with reinforcement learning for the first time, explains malware features through the TreeExplainer module, and uses SHAP value as the priority indicator of reinforcement learning actions, forming a dynamic modification mechanism based on data feature priority, which significantly improves the intelligence and targeting of malware generation.
[0217] Action modifier design based on UCB algorithm: By introducing the UCB algorithm, the present invention can more efficiently evaluate the potential evasion capability of each action when selecting fill data and performing modification actions in the data pool, thereby quickly generating adversarial malware that can evade multiple detectors.
[0218] Combination of priority experience replay and reward function: In the process of reinforcement learning agent learning, the present invention combines the experience replay mechanism of SHAP priority and reward function priority, and introduces intrinsic rewards through the ICM module, which further optimizes the decision-making mechanism in the learning process and improves the success rate of malware escape.
[0219] Double guarantee of functional integrity and escape effect: By executing adversarial malware in a sandbox and evaluating the effect, the present invention achieves effective escape of adversarial malware while ensuring the functional integrity of the malware, and has important practical application value.
[0220] The present invention has achieved significant technical progress in the intelligentization, efficiency improvement, functional guarantee and escape capability of malware generation, solved many difficult problems in the prior art, and provided a new technical path for malware generation and adversarial research.
[0221] The present invention solves several key problems in the generation of adversarial malware in the prior art by introducing the PERD3QN (Prioritized Experience Replay Deep Double Q Network) algorithm, combining the explainability technology SHAP and the UCB algorithm, and achieves significant technical progress in algorithm performance and generation effect.
[0222] 1. Technical problems solved:
[0223] Inefficient generation of adversarial malware: Traditional methods require a large number of attempts and modifications when generating adversarial malware, which is inefficient. This paper introduces a priority experience replay mechanism, uses the SHAP priority calculation module and the UCB algorithm to intelligently select modification actions, significantly improves the efficiency of adversarial malware generation, and reduces the number of invalid attempts.
[0224] The success rate of generating adversarial malware is limited: When faced with malware evasion evaluations from multiple detectors, traditional methods have difficulty guaranteeing a high success rate for generating adversarial malware. This paper combines multiple reward mechanisms (detection reward, distance reward, similarity reward, etc.) in the reinforcement learning process to comprehensively optimize the modification operation of malware, effectively improving the evasion success rate.
[0225] Malware functional integrity is difficult to ensure: During the generation process, adversarial malware may destroy its original functions due to multiple modification operations, affecting the detection and evasion effect. The present invention uses a functional integrity test module to ensure that the generated adversarial malware can maintain the integrity of its malicious functions while evading detection, solving the problem of functional destruction in traditional methods.
[0226] 2. Significant technological progress achieved:
[0227] Intelligent modification selection based on SHAP priority: The present invention introduces the TreeExplainer module of the SHAP library, calculates the SHAP value of malware features, and intelligently selects and updates the priority of reinforcement learning action modifiers based on this value, thereby achieving more targeted malware modification and improving the success rate of adversarial malware generation.
[0228] Application of UCB algorithm in action selection: By using the UCB algorithm to select the optimal modification action from the data pool, an efficient action modifier is constructed. The method of the present invention can quickly identify and apply modification actions with high evasion potential, thereby reducing the time and computing resources required to generate adversarial malware.
[0229] Combination of comprehensive reward function and priority experience replay: The present invention combines multiple factors such as detection reward, distance reward, similarity reward, etc. by designing a comprehensive reward function, and performs biased sampling of experience samples through the priority experience replay mechanism, thereby enhancing the stability and effectiveness of the reinforcement learning process and significantly improving the effect of adversarial malware generation.
[0230] Enhanced functional integrity verification and migration evaluation: In the functional integrity test and effect evaluation phase, the present invention ensures the functional integrity of adversarial malware by performing structural analysis and sandbox execution on the adversarial malware that has successfully escaped. In addition, the generated malware is also evaluated for migration, verifying its evasion effect on different detectors, demonstrating its wide practicality and applicability.
[0231] In summary, the present invention solves the efficiency, success rate and functional integrity issues in the prior art by introducing advanced algorithms and models in the adversarial malware generation method, achieves significant technological progress, and provides strong technical support for network security research and adversarial research on malware detection.
[0232] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for generating adversarial malware based on explainability technology, characterized in that: include: Extract features of malware using feature extractors of the EMBER model and the SOREL model, interpret the EMBER model and the SOREL model using the TreeExplainer module of the explainability technology SHAP library, calculate SHAP values of the extracted features, and calculate SHAP priorities of actions in the action modifier of reinforcement learning according to the SHAP values of the extracted features; Selecting filler data from a data pool using the UCB algorithm, performing the action selected by the reinforcement learning according to the filler data to modify the PE file of the malware, and performing escape evaluation on the modified malware using multiple malware detectors until the malware successfully escapes or reaches a maximum number of attempted modifications, thereby obtaining adversarial malware; The reinforcement learning process includes determining a reward function value according to the detection results of the malware by the multiple malware detectors, the length of the executed action, and the difference between the PE file before and after modification if the escape is unsuccessful; Obtaining an external reward generated during the interaction with the reinforcement learning environment according to the SHAP priority, the reward function value, and the TD error of the reinforcement learning agent; The ICM module is used to calculate the internal reward of the reinforcement learning agent, learn according to the external reward and the internal reward, and select the action for each step.
2. The method for generating adversarial malware based on explainability technology according to claim 1, characterized in that: Calculating the SHAP priority of the action in the action modifier of reinforcement learning according to the SHAP value of the extracted feature, including: constructing a first matrix according to the effect of each action in the action modifier on each feature; A second matrix is formed according to the SHAP values of the features extracted by the EMBER model, a third matrix is formed according to the SHAP values of the features extracted by the SOREL model, and SHAP values less than 0 in the second matrix and the third matrix are set to zero; Multiplying the first matrix by the second matrix after being reset to zero to obtain the SHAP priority of the action under the EMBER model; Multiplying the first matrix by the third matrix after being reset to zero to obtain the SHAP priority of the action under the SOREL model; The SHAP priorities of the action under the EMBER model and the SOREL model are weighted and summed to obtain the final SHAP priority of the action.
3. The method for generating adversarial malware based on explainability technology according to claim 1, characterized in that: Before calculating the SHAP priority of the action in the action modifier of reinforcement learning according to the SHAP value of the extracted feature, it also includes: An action modifier is constructed to modify, append and encrypt the PE file.
4. The method for generating adversarial malware based on explainability technology according to claim 1, characterized in that: Determining a reward function value according to the detection results of the malware by the multiple malware detectors, the length of the executed action, and the difference between the PE file before and after modification, including: Calculate the detection reward according to the difference between the detection scores of each malware detector when the malware is executed and when the malware is not executed, or the reward value of the detection result after the malware is executed; Calculating a distance reward according to the length of the action executed on the malware, wherein the longer the action length, the lower the distance reward; Calculate similarity rewards based on the similarity between the PE files before and after modification; The reward function value is calculated according to the detection reward, the distance reward and the similarity reward.
5. The method for generating adversarial malware based on explainability technology according to claim 4, characterized in that: The calculation formula of the detection reward R_d is: R_d=w1*ember_R+w2*Sorel_R+w3*clamav_R+w4*Ngram_R Among them, the scores of the four malware detectors are ember_R, sorel_R, clamav_R and ngram_R, respectively. ember_R and sorel_R are the difference between the detection scores of executing the action and not executing the action on the malware. clamav_reward and ngram_reward are the rewards for the detection results after the malware detector executes the action on the malware. If the detection result is benign, the reward is 1, otherwise the reward is 0. w1, w2, w3 and w4 are the corresponding weights. The calculation formula of the distance reward R_dis is: Where R is the given basic reward, t is the length of the action executed, and T is the maximum modification length; The calculation formula of the similarity reward R_sim is: R_sim=cos(v1,v2) Among them, cos is the cosine similarity, v1 and v2 are the vector representations of the PE file status before and after modification; The calculation formula of the reward function value R is: R=w5*R_det+w6*R_dis+w7*R_sim Among them, w5, w6 and w7 are corresponding weights.
6. The method for generating adversarial malware based on explainability technology according to claim 1, characterized in that: The external reward P generated during the interaction with the reinforcement learning environment is obtained according to the SHAP priority, the reward function value and the TD error of the reinforcement learning agent through the following formula: P=TD pri +SHAP pri *θ1+R pri *θ2 Among them, TD pri SHAP pri and R pri are the TD error, SHAP priority and reward function value respectively, and θ1 and θ2 are preset coefficients.
7. The method for generating adversarial malware based on explainability technology according to claim 1, characterized in that: The internal reward of the reinforcement learning agent is calculated using the ICM module, learning is performed based on the external reward and the internal reward, and the action of each step is selected, including: Extract features from the current state of the input PE file using the feature extractor in the ICM module; Using the prediction model in the ICM module to predict the features of the next state of the PE file based on the features of the current state of the PE file; Using the reverse model in the ICM module to predict actions based on the features of the current state and the features of the next state of the PE file; The D3QN reinforcement learning agent is used to generate an intrinsic reward according to the ICM module during the experience learning process, update the target Q value according to the intrinsic reward and the extrinsic reward, and determine the next action according to the updated target Q value.
8. The method for generating adversarial malware based on explainability technology according to any one of claims 1 to 7, characterized in that: After getting adversarial malware, it also includes: parsing the adversarial malware; If the analysis is successful, it is known that the structure of the adversarial malware is complete, and the adversarial malware is placed in a sandbox for execution; If the execution is successful, it is known that the functionality of the adversarial malware is complete, and a malware detector is used to perform an escape assessment and a migration assessment on the adversarial malware.
9. A system for generating adversarial malware based on explainability technology, characterized in that: include: A SHAP priority calculation module, for extracting features from malware using feature extractors of an EMBER model and a SOREL model, interpreting the EMBER model and the SOREL model using a TreeExplainer module of an explainability technology SHAP library, calculating SHAP values of the extracted features, and calculating SHAP priorities of actions in an action modifier of reinforcement learning according to the SHAP values of the extracted features; an adversarial malware generation module, configured to select filler data from a data pool using a UCB algorithm, modify the PE file of the malware by executing the action selected by the reinforcement learning according to the filler data, and perform escape evaluation on the modified malware using multiple malware detectors until the malware successfully escapes the multiple malware detectors or reaches a maximum number of attempted modifications, thereby obtaining adversarial malware; The reinforcement learning agent learning module is used to determine the reward function value according to the detection results of the malware by the multiple malware detectors, the length of the executed action and the difference between the PE file before and after modification if the escape is not successful; Obtaining an external reward generated during the interaction with the reinforcement learning environment according to the SHAP priority, the reward function value, and the TD error of the reinforcement learning agent; The ICM module is used to calculate the internal reward of the reinforcement learning agent, learn according to the external reward and the internal reward, and select the action for each step.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the adversarial malware generation method based on explainability technology as described in any one of claims 1 to 8 is implemented.