Confrontation method and system for GNN malicious software detection model

By inserting preprocessed benign opcode sequences into the malware and optimizing the insertion position using reinforcement learning, the adversarial attack and function retention problems of the GNN malware detection model are solved, and efficient adversarial malware generation is achieved.

CN120012088APending Publication Date: 2025-05-16ZHEJIANG NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510184821.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is difficult to carry out effective adversarial attacks on graph neural network (GNN) malware detection models, especially in implementing adversarial malware generation with functional reservations.

Method used

By preprocessing the benign opcode sequence and inserting the original malware, the reinforcement learning agent is used to select the appropriate insertion position and benign opcode sequence, so that the modified malware can bypass the detection of the GNN detection model and maintain the original function.

Benefits of technology

An effective confrontation to the GNN malware detection model is realized. The generated adversarial malware can run normally and its functions are consistent with the original samples, and the file size increases by a small increase.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012088A_ABST
    Figure CN120012088A_ABST
Patent Text Reader

Abstract

The invention provides a confrontation method and system for a GNN malicious software detection model, and the method comprises the steps: carrying out the preprocessing of a benign operation code sequence, and then inserting the benign operation code sequence into original malicious software until the modified malicious software is detected as benign software by the GNN detection model; the malicious software after the last iteration modification is used as antagonistic malicious software; the benign operation code sequence is selected from a benign operation code sequence library by a reinforcement learning agent, and the benign operation code sequence library is constructed by extracting the benign operation code sequence from a benign software sample detected by the GNN detection model; the insertion position of the benign operation code sequence is determined by the reinforcement learning agent. According to the method and the device, the CFG of the modified malicious software can be changed, and the modified malicious software can operate normally without changing functions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information security technology, and in particular to a method and system for countering a GNN malware detection model. Background Art

[0002] In recent years, the game between malware and machine learning detection models has attracted the interest of more and more research teams. There are related studies that have proposed a variety of methods, such as gradient-based attacks, randomization-based attacks, and reinforcement learning-based attacks. However, most of the existing attack methods cannot achieve adversarial attacks on the Graph Neural Network (GNN) malware model, and the adversarial attacks on GNN do not retain the function of malware. Among them:

[0003] (1) The reinforcement learning-based attack method Gym-Malware collects and compiles 10 basic modification operations of PE (Portable Executable) files with preserved functions. Using reinforcement learning, an agent is trained to learn the sequence of modification operations so that the modified malware can bypass the machine learning malware detection model based on byte sequence features. However, due to the limitations of basic modification operations, this method cannot achieve confrontation with the graph-based malware detection model.

[0004] (2) SRL is the first adversarial method for graph-based malware detection models. This method iteratively selects appropriate semantic Nops and their corresponding basic blocks, and performs semantic Nops insertion operations until the generated adversarial malware evades the detection model. However, this method can only preserve semantics and cannot generate adversarial malware with preserved functionality.

[0005] In view of the above analysis, the technical problems that need to be solved urgently in the prior art are:

[0006] (1) Most teams focus on adversarial targets such as convolutional neural networks (CNN) or recurrent neural networks (RNN), and rely too much on basic modification operations such as modifying PE file headers and inserting or appending benign byte sequences. The benign byte sequences modified or added by these operations will not be executed during the operation of the malware, and cannot change the control flow graph (CFG) of the malware, nor can they affect the results of the GNN model that relies on CFG for detection.

[0007] (2) Existing adversarial attack methods against GNN malware detection models only achieve semantic preservation and are unable to generate adversarial malware. Therefore, requiring the generated adversarial malware to be able to run and maintain functionality consistent with the original sample is one of the difficulties that needs to be solved. Summary of the invention

[0008] In response to the problems existing in the prior art, the present invention provides a method and system for countering the GNN malware detection model. Through this method, modifications can be made to the original malware to generate adversarial malware that can bypass the detection of the GNN malware detection model and achieve functional retention.

[0009] The present invention provides a countermeasure method for a GNN malware detection model, comprising:

[0010] The original malware is inserted after preprocessing the benign opcode sequence until the modified malware is detected as benign software by the GNN detection model;

[0011] The last modified iteration of the malware is used as adversarial malware;

[0012] The benign opcode sequence is selected by a reinforcement learning agent from a benign opcode sequence library, and the benign opcode sequence library is constructed by extracting benign opcode sequences from benign software samples detected by the GNN detection model;

[0013] The insertion position of the benign opcode sequence is determined by the reinforcement learning agent.

[0014] According to a countermeasure method for a GNN malware detection model provided by the present invention, a benign opcode sequence is preprocessed before being inserted into the original malware, and further includes:

[0015] Decompiling each software sample in the first software data set to obtain a CFG with node features of each software sample;

[0016] Using the GNN detection model to detect the CFG of each software sample to determine whether each software sample is a benign software sample;

[0017] Extracting a benign operation code sequence of each basic block from the benign software samples in the first software data set;

[0018] The benign operation code sequence library is constructed according to the benign operation code sequence of each basic block in the benign software sample.

[0019] According to a countermeasure method for a GNN malware detection model provided by the present invention, the original malware is inserted after preprocessing a benign opcode sequence, including:

[0020] Decompile the original malware using a reverse tool to obtain the operation code in the original malware;

[0021] The reinforcement learning agent selects an opcode from the opcodes of the original malware as the insertion position, and selects a benign opcode sequence to be inserted from the benign opcode sequence library;

[0022] After preprocessing the benign opcode sequence to be inserted, the original opcode at the insertion position and a jmp opcode for jumping to the address of the next opcode at the insertion position are added thereto to obtain the content to be inserted;

[0023] Add the content to be inserted to the end of the original malware, and calculate the insertion address of the content to be inserted in the original malware;

[0024] The original operation code at the insertion position is modified into a jmp operation code for jumping to the insertion address of the content to be inserted.

[0025] According to the benign opcode sequence of each basic block in the benign software sample, the reinforcement learning agent selects an opcode from the opcode of the original malware as the insertion position, including:

[0026] Determine the byte length of the jmp opcode according to the byte and relative offset of the jmp opcode;

[0027] The reinforcement learning agent selects a single opcode or a plurality of consecutive opcodes from the opcodes of the original malware as the insertion position;

[0028] The byte length at the insertion position is greater than or equal to the byte length of the jmp operation code, and the original operation code at the insertion position does not involve modification of the esp stack pointer register.

[0029] Preprocessing the benign operation code sequence to be inserted according to the benign operation code sequence of each basic block in the benign software sample includes:

[0030] The unfriendly opcodes, stack balancing operations, and stack reconstruction are sequentially deleted from the benign opcode sequence to be inserted, "pushf, pusha" opcodes are inserted at the front of the benign opcode sequence to be inserted, and "popa, popf" opcodes are added at the front of the benign opcode sequence to be inserted.

[0031] According to the benign opcode sequence of each basic block in the benign software sample, the reinforcement learning agent selects the benign opcode sequence and determines the insertion position of the benign opcode sequence through the following steps:

[0032] Determine the malicious probability predicted by the GNN detection model and the modified CFG of the original malware before and after each modification;

[0033] The reinforcement learning agent is used to learn the benign operation code sequence and insertion position inserted by the original malware according to the malicious probability predicted by the GNN detection model before and after each modification of the original malware and the modified CFG, so that the malicious probability of the modified malware is reduced.

[0034] According to the benign opcode sequence of each basic block in the benign software sample, the reinforcement learning agent is used to learn the benign opcode sequence and insertion position inserted by the original malware according to the malicious probability predicted by the GNN detection model before and after each modification of the original malware and the modified CFG, so that the malicious probability of the modified malware is reduced, including:

[0035] The benign opcode sequence and the insertion position inserted by the original malware are regarded as two discrete actions in a multi-dimensional discrete action space;

[0036] Determining a reward function of a reinforcement learning environment according to the maliciousness probability of the original malware before and after each modification;

[0037] The modified CFG of the original malware is used as a new state, and the modified malicious probability is used as the malicious probability before the next action is modified. The reinforcement learning agent learns according to the reward function and makes the next action.

[0038] According to the benign operation code sequence of each basic block in the benign software sample, the benign operation code sequence is preprocessed before being inserted into the original malware, further comprising:

[0039] Decompile each software sample in the second software data set to obtain a CFG with node features of each software sample;

[0040] The GNN detection model is trained using the CFG with node features and corresponding labels of each software sample in the second software data set, so that the GNN detection model distinguishes between benign software and malicious software by learning the structure and node features of the CFG.

[0041] The present invention also provides a confrontation system for the GNN malware detection model, comprising:

[0042] Benign opcode sequence insertion module, which is used to insert the original malware after preprocessing the benign opcode sequence until the modified malware is detected as benign software by the GNN detection model;

[0043] An adversarial malware acquisition module, used to use the last modified iterative malware as adversarial malware;

[0044] The benign opcode sequence is selected by a reinforcement learning agent from a benign opcode sequence library, and the benign opcode sequence library is constructed by extracting benign opcode sequences from benign software samples detected by the GNN detection model;

[0045] The insertion position of the benign opcode sequence is determined by the reinforcement learning agent.

[0046] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the countermeasure method for the GNN malware detection model as described in any one of the above-mentioned methods is implemented.

[0047] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a countermeasure method for a GNN malware detection model as described in any one of the above.

[0048] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned countermeasure methods for the GNN malware detection model.

[0049] The adversarial method and system for the GNN malware detection model provided by the present invention insert the original malware after preprocessing the benign operation code sequence, and modify the malware at the byte level so that the CFG of the modified malware can be changed, and the modified malware can run normally and the function remains unchanged; based on reinforcement learning, adversarial malware with preserved functions can be automatically generated; the file size of the generated adversarial malware is less increased than that of the original malware sample, and adversarial attacks on the GNN malware detection model can be achieved with a smaller additional load. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0051] Figure 1 It is a flowchart of the countermeasure method for the GNN malware detection model provided by the present invention;

[0052] Figure 2 It is a schematic diagram of the overall process of the confrontation method for the GNN malware detection model provided by the present invention;

[0053] Figure 3 It is a schematic diagram of live code insertion in the adversarial method for the GNN malware detection model provided by the present invention;

[0054] Figure 4 It is a schematic diagram of a reinforcement learning environment in the adversarial method for the GNN malware detection model provided by the present invention;

[0055] Figure 5 It is a grayscale schematic diagram of malware before and after modification in the adversarial method for the GNN malware detection model provided by the present invention;

[0056] Figure 6 It is a schematic diagram of the change of malware CFG before and after modification in the adversarial method for the GNN malware detection model provided by the present invention;

[0057] Figure 7 It is a bar chart comparing the average attack success rate in the adversarial method for the GNN malware detection model provided by the present invention;

[0058] Figure 8 It is a line graph of the confrontation success rate under different modification times limits in the confrontation method for the GNN malware detection model provided by the present invention;

[0059] Fig. 9 It is a structural schematic diagram of the adversarial system for the GNN malware detection model provided by the present invention. DETAILED DESCRIPTION

[0060] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0061] Combine the following Figure 1 A countermeasure method for a GNN malware detection model of the present invention is described, comprising:

[0062] Step 101, inserting the original malware after preprocessing the benign opcode sequence until the modified malware is detected as benign software by the GNN detection model;

[0063] Step 102, using the malware modified in the last iteration as adversarial malware;

[0064] The benign opcode sequence is selected by a reinforcement learning agent from a benign opcode sequence library, and the benign opcode sequence library is constructed by extracting benign opcode sequences from benign software samples detected by the GNN detection model;

[0065] The insertion position of the benign opcode sequence is determined by the reinforcement learning agent.

[0066] Live code insertion: Basic modification of the original malware input to insert live code. Live code refers to code that will be executed during actual program execution, not just filler or placeholder code.

[0067] Inserting benign opcode sequences: The extracted benign opcode sequences are inserted into the malware samples to change their CFG structure and node features. This modification is intended to mislead the GNN detection model into misclassifying the modified malware as benign.

[0068] The basic modification operation of live code insertion modifies the malware at the binary level, so that the CFG of the modified malware changes, thereby affecting the prediction results of the GNN detection model that relies on CFG for detection.

[0069] Repeated detection and modification: Repeat the steps of malware modification, and use the GNN detection model to detect the modified malware sample in each iteration. Based on the detection results, continue to adjust the inserted benign opcode sequence or modification position to gradually reduce the malicious probability of the malware.

[0070] Adversarial sample generation: When a modified malware sample can be detected as benign software by the GNN detection model, adversarial malware is considered to be generated.

[0071] In order to select appropriate insertion positions and benign opcode sequences in the original malware sample so that the modified malware can effectively reduce the malicious probability in the GNN malware detection model, this embodiment adopts a reinforcement learning approach and uses machine learning to automatically select insertion positions and benign opcode sequences that can effectively reduce the malicious probability.

[0072] This embodiment inserts the original malware after preprocessing the benign opcode sequence, and changes the bytes of the malware at the binary level, so that the CFG of the modified malware can be changed, and the modified malware can run normally and the functions remain unchanged; based on reinforcement learning, adversarial malware with preserved functions can be automatically generated; the file size of the generated adversarial malware increases less than that of the original malware sample, and adversarial attacks on the GNN malware detection model can be achieved with a smaller additional load.

[0073] Based on the above embodiment, this embodiment further includes:

[0074] Decompile each software sample in the first software data set B, and parse to obtain a CFG with node features of each software sample;

[0075] Using the GNN detection model to detect the CFG of each software sample to determine whether each software sample is a benign software sample;

[0076] Extracting a benign operation code sequence of each basic block from the benign software samples in the first software data set B;

[0077] The benign operation code sequence library is constructed according to the benign operation code sequence of each basic block in the benign software sample.

[0078] Dataset decompilation and analysis: Decompile the malware and benign software in the first software dataset B and the second software dataset A to parse their control flow graphs (CFGs). As graph structure data, CFG contains program execution flow information, and node features include opcodes, instruction types, etc.

[0079] GNN detection model training: The GNN detection model is trained using the CFG and its corresponding labels in the second software dataset A. This model distinguishes malware from benign software by learning the structure and node features of CFG.

[0080] Benign sample detection: Use the trained GNN detection model to detect the CFG in the first software dataset B, and filter out software samples that are correctly identified as benign by the model.

[0081] Benign opcode sequence extraction: Extract representative benign opcode sequences from the selected benign software samples. These sequences will be used for subsequent malware modification operations.

[0082] Figure 2This is the overall flow chart of this embodiment. First, decompile datasets A and B, parse and obtain CFG with node features, and the CFG of dataset A is used to train the GNN detection model as the adversarial target. Then, the model is used to detect the CFG of dataset B, screen benign software samples, and extract benign opcode sequences from the benign software samples. Finally, the basic modification operation of live code insertion is performed on the input original malware sample, and the extracted benign opcode sequence is inserted into the sample, and this process is repeated until the modified malware can be detected as benign by the GNN detection model, and adversarial malware can be obtained. In this repetitive process, reinforcement learning can be used to learn which modification locations and benign opcode sequences can effectively reduce the malicious probability of malware.

[0083] This embodiment can effectively generate adversarial malware samples by combining decompilation, GNN model training, benign sample screening, malware modification, and reinforcement learning optimization. These samples can mislead the GNN detection model, thereby counteracting the GNN-based malware detection system. However, it is worth noting that with the continuous development of adversarial technology, the corresponding defense technology is also constantly improving to deal with these new threats.

[0084] Based on the above embodiments, Figure 3 As shown, in this embodiment, the benign operation code sequence is preprocessed and then the original malware is inserted, including:

[0085] Decompile the original malware M using a reverse tool to obtain the operation code in the original malware;

[0086] The reinforcement learning agent selects an opcode from the opcode of the original malware as the insertion position opcode_i, selects a benign opcode sequence to be inserted from the benign opcode sequence library and records it as sequence_b, and records the next opcode address addr_i of the insertion position opcode_i;

[0087] After preprocessing the benign opcode sequence sequence_b to be inserted, add the original opcode at the insertion position and the jmp opcode that jumps to the next opcode address addr_i of the insertion position opcode_i to obtain the content sequence'_b to be inserted;

[0088] When adding the content sequence'_b to be inserted to the end of the original malware M, a new section for storing the byte information of sequence'_b needs to be created in the original malware M. At the same time, the insertion address addr_b of the content sequence'_b to be inserted in the original malware is calculated;

[0089] The original operation code at the insertion position opcode_i is modified into a jmp operation code that jumps to the insertion address addr_b of the content to be inserted. At this point, the live code insertion and modification operation is completed.

[0090] The specific process of live code insertion is as follows:

[0091] Filter insertion positions: Decompile the original malware and filter out the opcodes that meet the insertion conditions for subsequent insertion;

[0092] Extracting benign opcode sequences: Decompile the benign software sample, select a basic block, and save the bytecode sequence of this basic block for subsequent insertion;

[0093] Opcode sequence preprocessing: In order to preserve the functionality of the modified malware, benign opcode sequences need to be preprocessed, including removing unfriendly opcodes, rebuilding the stack, adding insertion point original opcodes, and adding return opcodes.

[0094] Inserting opcode sequence: constructing a new section whose content consists of the bytes of the preprocessed opcode sequence and appending the section to the end of the malware;

[0095] Establish a connection: Calculate the address of the inserted opcode sequence in the malware and modify the insertion point to a jmp opcode that jumps to that address, thereby establishing a connection from the insertion point to the inserted opcode sequence.

[0096] Based on the above embodiment, in this embodiment, the reinforcement learning agent selects an operation code from the operation code of the original malware as the insertion position, including:

[0097] Determine the byte length of the jmp opcode according to the byte and relative offset of the jmp opcode;

[0098] The reinforcement learning agent selects a single opcode or a plurality of consecutive opcodes from the opcodes of the original malware as the insertion position;

[0099] The byte length at the insertion position is greater than or equal to the byte length of the jmp operation code, and the original operation code at the insertion position does not involve modification of the esp stack pointer register.

[0100] Due to the execution logic of the PE (Portable Executable) file, the insertion position opcode_i needs to contain one or more complete opcodes, and the first requirement is that the byte length must be no less than 10. This is because at the end of the modification, pcode_i needs to be modified to an unconditional jump opcode that jumps to the inserted address addr_b, and the byte of the unconditional jump opcode is E9 plus a 32-bit relative offset, and the byte length is 10. However, in the instruction set of the X86-32 architecture, most of the opcodes are less than 10 in byte length. In order to expand as many insertion positions as possible, it is not limited to a single opcode. Multiple consecutive opcodes with a byte length of no less than 10 can also be used as insertion positions. Secondly, opcode_i cannot involve modifications to the esp stack pointer register, otherwise the stack position will be uncontrollable, causing the modified malware to fail to run or change its function.

[0101] If the byte length of opcode_i exceeds 10, the part exceeding the length needs to be replaced with a nop opcode to ensure that the part exceeding the length will not be recognized as other opcodes.

[0102] On the basis of the above embodiment, in this embodiment, the benign operation code sequence to be inserted is preprocessed, including:

[0103] The unfriendly opcodes, stack balancing operations, and stack reconstruction are sequentially deleted from the benign opcode sequence to be inserted, "pushf, pusha" opcodes are inserted at the front of the benign opcode sequence to be inserted, and "popa, popf" opcodes are added at the front of the benign opcode sequence to be inserted.

[0104] The pre-processing steps for the inserted benign opcode sequence include:

[0105] First, we need to clean up the opcodes that are not friendly to the benign opcode sequence sequence_b. Since sequence_b has been separated from the register environment of the original benign software during execution, it is impossible to determine the contents stored in the registers, which will lead to many uncontrollable situations. For example, the opcodes of the jmp and call families will cause jumps to unknown addresses, the opcodes of ret and loop will cause an infinite loop, the div opcode will have a divisor of 0, the in and out opcodes will cause communication with the I / O port of an unknown hardware device, and the cli and sti opcodes will cause unexpected interrupts, etc.

[0106] In addition, opcodes that are not friendly to insertion operations also include improper memory assignment opcodes, such as opcodes that use register indirect addressing or opcodes whose first operand is [CONST]. The addresses referenced by these opcodes will exceed the memory range allowed by the program, resulting in memory out-of-bounds errors. They may also modify the memory of stored data, causing changes in the functionality of the modified malware.

[0107] Then, the stack balance operation is performed on sequence_b. This is to ensure that after the execution of sequence_b, the esp stack pointer register remains in place and the data stored in the stack does not change, so that the register environment can be restored normally. The pop opcode is used to remove the current top element of the stack, which will cause the address in esp to decrease. The push opcode adds the element to the top of the stack, which will cause the address in esp to increase. When traversing the sequence_b opcode, the number of pop and push occurrences is counted in real time, and the redundant push opcodes are deleted to avoid stack underflow. After the traversal is completed, if the number of push occurrences is greater than that of pop, the "add esp, CONST" opcode is added to the end (the CONST is obtained by calculating the change in esp) to adjust the content in esp to change the position of the top pointer of the stack, so that the esp top pointer is restored to its original position.

[0108] Next, the stack in sequence_b needs to be rebuilt. First, clear the original opcodes involving the ebp and esp stack register operations in sequence_b, then add the "push ebp; mov ebp, esp" opcode to the head of sequence_b, and the "leave" opcode to the tail. "push ebp" is used to push the current value of the ebp register into the stack to save the base pointer of the previous function stack frame. "mov ebp, esp" is used to set the value of the ebp register to the current esp value, that is, the address of the current top of the stack. "leave" is actually a simplified form of "mov esp, ebp; pop ebp", which is used to end the stack frame of the current function and restore it to the stack frame state before calling the function. Through these operations, sequence_b can safely process local variables, and the original state of the stack can be restored after sequence_b is executed.

[0109] Finally, in order to restore the data in all registers to their original state after sequence_b is executed, it is necessary to insert the "pushf, pusha" opcode before sequence_b to push the value of the flag register and several general registers into the stack, and then add the "popa, popf" opcode after sequence_b to restore the value saved in the stack to the register. So far, the preprocessing operation of sequence_b has been completed through the above steps.

[0110] On the basis of the above embodiments, the reinforcement learning agent in this embodiment selects a benign opcode sequence and determines the insertion position of the benign opcode sequence through the following steps:

[0111] Determine the malicious probability predicted by the GNN detection model and the modified CFG of the original malware before and after each modification;

[0112] The reinforcement learning agent is used to learn the benign operation code sequence and insertion position inserted by the original malware according to the malicious probability predicted by the GNN detection model before and after each modification of the original malware and the modified CFG, so that the malicious probability of the modified malware is reduced.

[0113] Reinforcement learning agent setting: A reinforcement learning mechanism is introduced in the repeated iterative process. The reinforcement learning agent learns the most effective modification strategy and benign opcode sequence based on the CFG of the modified malware and the malicious probability predicted by the GNN detection model before and after the modification.

[0114] Learning strategy update: The reinforcement learning agent continuously optimizes its modification strategy through trial and error and feedback mechanism. Whenever an adversarial example is generated or a modification results in a decrease in malicious probability, the agent updates its internal strategy and parameters based on these positive feedback. Conversely, if the modification fails to effectively reduce the malicious probability, the agent adjusts its strategy to avoid similar ineffective operations.

[0115] On the basis of the above embodiments, in this embodiment, the reinforcement learning agent is used to learn the benign opcode sequence and insertion position inserted by the original malware according to the malicious probability predicted by the GNN detection model before and after each modification of the original malware and the modified CFG, so that the malicious probability of the modified malware is reduced, including:

[0116] The benign opcode sequence and the insertion position inserted by the original malware are regarded as two discrete actions in a multi-dimensional discrete action space;

[0117] Determining a reward function of a reinforcement learning environment according to the maliciousness probability of the original malware before and after each modification;

[0118] The modified CFG of the original malware is used as a new state, and the modified malicious probability is used as the malicious probability before the next action is modified. The reinforcement learning agent learns according to the reward function and makes the next action.

[0119] The steps of the reinforcement learning environment are as follows:

[0120] Predict the original sample score: Decompile the original malware to obtain the malware CFG with node features, and input it into the adversarial target GNN detection model to obtain the predicted malicious probability;

[0121] Modify the original sample: The reinforcement learning agent selects the appropriate insertion location and benign opcode sequence, applies the basic modification operation of live code insertion, and obtains the modified malware sample;

[0122] Predicted modified sample score: Decompile the modified malware sample to obtain the modified CFG with node features, and input it into the adversarial target GNN detection model to obtain the modified predicted malicious probability. If the modified predicted malicious probability is lower than the set threshold, it can be determined to have escaped the detection of the GNN model;

[0123] Calculate rewards and set observation states: Based on the reward function in the reinforcement learning environment and the predicted malicious probabilities before and after the modification, calculate the reward after the operation, and set the modified CFG as the observation space for the reinforcement learning agent to learn.

[0124] In order to automatically select the appropriate insertion position and the corresponding benign code sequence in the original malware sample and generate adversarial malware that can escape the detection of the GNN model, this method designs a reinforcement learning environment, such as Figure 4 shown.

[0125] When the reinforcement learning environment is initialized, it randomly selects an original malware sample M from the malware dataset and analyzes its CFG with node features. The CFG will constitute the state as the original observation space of the environment. The CFG is input into the GNN malware detection model as the adversarial target to obtain the prediction score p of M, which is used to represent the probability that the input file is malware. While decompiling the original malware sample, the opcodes that meet the requirements of the live code insertion method are selected to form a list of positions to be inserted for subsequent operations.

[0126] The malware CFG changes are different due to the differences in the insertion position's own opcode family, its position in the basic block, etc. Based on the above conditions, this embodiment divides the insertion position into 13 types, including the insertion position of a single opcode, multiple opcodes, the insertion position at the beginning, middle, and end of the basic block, and the insertion position of non-jmp family, jmp family, cjmp family, call family, and trap family opcodes.

[0127] Therefore, a multi-dimensional discrete action space is designed for the environment, which contains two discrete actions, one for selecting the insertion position according to the insertion position, and the other for selecting the benign opcode sequence. After the reinforcement learning agent takes an action, the environment selects the corresponding insertion position and benign opcode sequence according to the action, performs the live code insertion basic modification operation on the malware, and generates a modified malware M'. The modified malware is analyzed to obtain its modified CFG with node features, which is input into the GNN malware detection model to obtain the modified malware prediction score p'. By comparing the changes in p and p', it is determined whether the modification can reduce the malware prediction score. If the p' obtained after a certain action is executed is lower than the set threshold, it is determined that the modified malware has successfully escaped the detection of the GNN detection model.

[0128] The reward function of the reinforcement learning environment is calculated based on p and p'. If the reward value is a positive number, it means that this modification can lead to a decrease in the malicious probability predicted by the model. On the contrary, if the reward value is a negative number, it means that this modification leads to an increase in the malicious probability predicted by the model.

[0129] At the same time, the modified CFG is used as the new state, and p' is used as the p of the next action, and the agent learns to make the next decision action. This cycle is repeated until there is no available insertion position for the malware, or the modified malware can evade the detection of the GNN model. Then a new original malware sample is selected, the environment is initialized again, and this series of operations is repeated. Figure 5 This is a grayscale image of the malware before and after modification. Figure 6 This is the CFG change diagram of the malware before and after modification.

[0130] Based on the above embodiments, this embodiment further includes:

[0131] Decompile each software sample in the second software data set A to obtain a CFG with node features of each software sample;

[0132] The GNN detection model is trained using the CFG with node features and corresponding labels (malware or benign software) of each software sample in the second software data set A, so that the GNN detection model distinguishes benign software from malicious software by learning the structure and node features of the CFG.

[0133] The second software dataset A is a training dataset. The malware in the training dataset can be obtained from the public datasets BODMAS and SOREL-20M. The benign software can be collected from Windows system files and commonly used software, including popular applications, development tools, and security protection software. The following preprocessing work is performed on the collected samples:

[0134] 1) Sample cleaning: Some samples may fail to decompile or have only one basic block due to file damage caused by harmless processing or network transmission. Therefore, invalid samples that fail to decompile and special samples with only one basic block are removed.

[0135] 2) CFG extraction and node representation: In order to successfully train the GNN detection model for adversarial detection, a large number of real assembly code samples need to be extracted. Basic blocks are nodes in CFG and need to be represented as feature vectors. The open source reverse engineering framework Radare2 can be used to decompile binary executable files in the dataset, analyze and extract CFG of 30,000 samples, and extract node features according to different representation methods for GNN model training.

[0136] The malware dataset of the present invention can adopt two recognized public datasets, SOREL-20M and BODMAS. These two malware datasets contain many common malware families such as Ransomware, Crypto miner, Adware, Downloader, etc. The detailed information of the specific malware families is shown in Table 1. Due to copyright issues, the above two datasets do not provide binary files of benign software. The benign software dataset of the experiment of the present invention is composed of benign software samples collected from Windows system files and commonly used software, including popular applications, development tools, and security protection software.

[0137] Table 1 Dataset family details

[0138]

[0139]

[0140] The structure of the GNN detection model used in the present invention is shown in Table 2, and the SOREL-20M dataset is used for training. Table 3 shows the classification performance of the adversarial GNN malware detection model in the following experiments.

[0141] Table 2 GNN model structure

[0142]

[0143] Table 3 Classification performance of GNN malware detection model

[0144]

[0145] The baseline method includes:

[0146] Gym-malware With GNN: Based on Gym-malware, the detection model used as the adversarial target in its environment is modified to the GNN detection model, and the operation space remains unchanged.

[0147] Function-preserving Random Insertion 1 (FRI1): For the original malware sample, randomly select the insertion position in the malware, insert a fixed benign opcode sequence, and modify the malware using the live code insertion modification method.

[0148] Function-preserving Random Insertion 2 (FRI2): For the original malware sample, the insertion location in the malware and the inserted benign opcode sequence are randomly selected, and the malware is modified using the live code insertion modification method.

[0149] Function-preserving Accumulated Insertion (FAI): This method is based on the hill climbing method. Each time a modification is made, an insertion position and a benign opcode sequence are randomly selected. Then, the prediction scores before and after the modification are compared. If the prediction score decreases after the modification, the modification is retained. If the prediction score increases, the modification is rolled back to the previous modification.

[0150] The experiment mainly compares the present invention with other baseline methods from the attack success rate. The attack success rate is the number of generated adversarial malware that can escape the detection model divided by the number of original malware samples, indicating the proportion of adversarial malware that can be successfully generated in the input malware samples. The experiment also introduces a functional retention indicator to evaluate the functional retention of the generated adversarial malware, indicating the proportion of adversarial malware that can run normally and keep its functions unchanged in the sampled samples.

[0151] The attack success rates of the present invention, gym-malware with GNN, FRI, FAI and other methods on different GNN detection models are shown in Table 4. Figure 7 According to the experimental results, the present invention shows good adversarial effect, and achieves an average adversarial success rate of 96.08% for various GNN detection models, which is much higher than other baseline methods.

[0152] Table 4 Attack success rate

[0153]

[0154]

[0155] At the same time, the present invention also randomly selects a GNN malware detection model, and the maximum number of modifications is limited to 50, 100, 250, 500, 750 and 1000 respectively. The attack success rates in these six cases are compared, and the comparison results are as follows Figure 8 As shown in Figure 2, it can be seen that, in extreme cases, this method also shows a good attack success rate.

[0156] Table 5 shows the function retention ratio of randomly selected adversarial malware generated by the modification of each family of malware by the present invention. The experimental results show that 78.43% of the adversarial malware generated by the present invention can run normally and the functions are consistent with the original malware samples.

[0157] Table 5 Function retention rate

[0158]

[0159] The effectiveness of the attack method was evaluated on a variety of data sets. The experimental results showed that the method not only had a high success rate in the attack, but also achieved good results in the tests of function retention rate, anti-migration and other aspects.

[0160] The expected benefits and commercial value of the technical solution of the present invention after transformation are as follows: the present invention helps to reveal new technologies that malicious authors may use to evade existing security protection measures, prompting researchers to improve detection models and conduct corresponding adversarial training. In addition, in the field of security protection, it helps to predict the development trend of malware, explore and discover attack paths and vulnerabilities that may be ignored by existing malware detection methods, deploy effective protection strategies in advance, achieve the advancement of security technology, and respond to ever-changing security challenges.

[0161] The technical solution of the present invention fills the technical gap in the industry at home and abroad: the present invention is the first function-preserving black-box adversarial attack method for the GNN malware detection model. The adversarial malware generated by the present invention can run normally and its functions are consistent with the original malware samples.

[0162] The technical solution of the present invention solves a technical problem that people have been eager to solve but have never been successful: the present invention proposes a method for inserting and modifying live code with preserved functions. During the operation of the executable file obtained by the modification method, the inserted operation code sequence can be executed by the processor, and the CFG obtained by decompiling it is also changed, and it is guaranteed that the modified executable file can run normally and the function remains unchanged.

[0163] The adversarial system for the GNN malware detection model provided by the present invention is described below. The adversarial system for the GNN malware detection model described below and the adversarial method for the GNN malware detection model described above can refer to each other.

[0164] like Fig. 9 As shown, the system includes:

[0165] Benign opcode sequence insertion module, which is used to insert the original malware after preprocessing the benign opcode sequence until the modified malware is detected as benign software by the GNN detection model;

[0166] An adversarial malware acquisition module, used to use the last modified iterative malware as adversarial malware;

[0167] The benign opcode sequence is selected by a reinforcement learning agent from a benign opcode sequence library, and the benign opcode sequence library is constructed by extracting benign opcode sequences from benign software samples detected by the GNN detection model;

[0168] The insertion position of the benign opcode sequence is determined by the reinforcement learning agent.

[0169] An application embodiment of the present invention provides a computer device, which includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of the countermeasure method for the GNN malware detection model.

[0170] An application embodiment of the present invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the steps of a countermeasure method for a GNN malware detection model.

[0171] An application embodiment of the present invention provides an information data processing terminal, which includes an adversarial system for a GNN malware detection model.

[0172] In modern computing environments, the threat of malware is increasing, and traditional detection methods are becoming insufficient. Malware detection models based on graph neural networks (GNNs) have been widely used to detect complex malware attacks. However, adversarial attack techniques against these detection models are also developing. This paper proposes an adversarial method for GNN malware detection models, which can effectively generate adversarial malware for testing and improving the robustness of malware detection models.

[0173] 1. Decompilation and analysis:

[0174] Hardware equipment: high-performance decompilation server

[0175] Implementation principle: Decompile dataset A and dataset B and parse them to obtain a control flow graph (CFG) with node features. The CFG of dataset A is used to train the GNN detection model as the adversarial target.

[0176] 2. Testing and screening:

[0177] Hardware equipment: High performance computing server

[0178] Implementation principle: Use the trained GNN detection model to detect the CFG of data set B, screen benign software samples, and extract benign opcode sequences from them.

[0179] 3. Code insertion and modification:

[0180] Hardware equipment: Automatic code insertion device

[0181] Implementation principle: Perform basic modification operations of live code insertion on the input original malware sample and insert the extracted benign opcode sequence into the sample.

[0182] 4. Repeated and reinforced learning:

[0183] Hardware equipment: reinforcement learning training server

[0184] Implementation principle: Repeat the steps of decompilation, detection, code insertion, etc., and use the reinforcement learning agent to learn which modification locations and benign opcode sequences can effectively reduce the malicious probability of malware based on the CFG of the modified malware and the malicious probability predicted by the GNN detection model before and after the modification, until adversarial malware that can be detected as benign is generated.

[0185] Through the above-mentioned mechanical structure and intelligent control system, the malware detection and protection system achieves the ability to efficiently generate adversarial malware, significantly improves the robustness and protection capabilities of the detection model, and can more effectively respond to complex malware attacks.

[0186] In the process of software development and maintenance, security testing is an important part of ensuring software security. Traditional manual testing methods are inefficient and prone to missing potential security vulnerabilities. The adversarial method against the GNN malware detection model can be used in an automated security testing platform to improve the efficiency and coverage of security testing.

[0187] 1. Decompile and analyze module:

[0188] Hardware equipment: high-performance decompilation server

[0189] Implementation principle: Decompile the software sample data set and parse it to obtain the control flow graph (CFG) with node features to provide basic data for subsequent security testing.

[0190] 2. Malware Detection Module:

[0191] Hardware equipment: High performance computing server

[0192] Implementation principle: Use the GNN malware detection model to detect the decompiled CFG and screen out potential malware samples.

[0193] 3. Code modification and module insertion:

[0194] Hardware equipment: Automatic code insertion device

[0195] How it works: Live code insertion and base modification operations are performed on detected potential malware samples to insert benign opcode sequences to generate adversarial malware.

[0196] 4. Reinforcement Learning Module:

[0197] Hardware equipment: reinforcement learning training server

[0198] Implementation principle: In the process of generating adversarial malware, a reinforcement learning agent is used to optimize the modification locations and benign opcode sequences according to the CFG of the modified malware and the prediction results of the GNN detection model of the modified malware until effective adversarial malware is generated.

[0199] 5. Safety testing and feedback module:

[0200] Hardware equipment: integrated testing and feedback system

[0201] Implementation principle: The generated adversarial malware is used for automated security testing. Through the integrated testing and feedback system, potential security vulnerabilities can be discovered and fixed in a timely manner to improve software security.

[0202] Through the above-mentioned mechanical structure and intelligent control system, the automated security testing platform realizes an efficient security testing process and can automatically generate adversarial malware for testing, which significantly improves the test coverage and efficiency and effectively ensures the security of the software.

[0203] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A countermeasure method for GNN malware detection model, characterized in that: include: The original malware is inserted after preprocessing the benign opcode sequence until the modified malware is detected as benign software by the GNN detection model; The last modified iteration of the malware is used as adversarial malware; The benign opcode sequence is selected by a reinforcement learning agent from a benign opcode sequence library, and the benign opcode sequence library is constructed by extracting benign opcode sequences from benign software samples detected by the GNN detection model; The insertion position of the benign opcode sequence is determined by the reinforcement learning agent.

2. The method for combating the GNN malware detection model according to claim 1, characterized in that: The benign opcode sequence is preprocessed before being inserted into the original malware, and also includes: Decompiling each software sample in the first software data set to obtain a CFG with node features of each software sample; Using the GNN detection model to detect the CFG of each software sample to determine whether each software sample is a benign software sample; Extracting a benign operation code sequence of each basic block from the benign software samples in the first software data set; The benign operation code sequence library is constructed according to the benign operation code sequence of each basic block in the benign software sample.

3. The method for combating the GNN malware detection model according to claim 1, characterized in that: The original malware is inserted after preprocessing the benign opcode sequence, including: Decompile the original malware using a reverse tool to obtain the operation code in the original malware; The reinforcement learning agent selects an opcode from the opcodes of the original malware as the insertion position, and selects a benign opcode sequence to be inserted from the benign opcode sequence library; After preprocessing the benign opcode sequence to be inserted, the original opcode at the insertion position and a jmp opcode for jumping to the address of the next opcode at the insertion position are added thereto to obtain the content to be inserted; Add the content to be inserted to the end of the original malware, and calculate the insertion address of the content to be inserted in the original malware; The original operation code at the insertion position is modified into a jmp operation code for jumping to the insertion address of the content to be inserted.

4. The method for combating the GNN malware detection model according to claim 3, characterized in that: Selecting, by the reinforcement learning agent, an opcode from the opcode of the original malware as the insertion position, comprises: Determine the byte length of the jmp opcode according to the byte and relative offset of the jmp opcode; The reinforcement learning agent selects a single opcode or a plurality of consecutive opcodes from the opcodes of the original malware as the insertion position; The byte length at the insertion position is greater than or equal to the byte length of the jmp operation code, and the original operation code at the insertion position does not involve modification of the esp stack pointer register.

5. The method for combating the GNN malware detection model according to claim 3, characterized in that: Preprocessing the benign operation code sequence to be inserted includes: The unfriendly opcodes are deleted from the benign opcode sequence to be inserted, the stack is balanced, the stack is rebuilt, the "pushf, pusha" opcode is inserted at the front of the benign opcode sequence to be inserted, and the "popa, popf" opcode is added at the front of the benign opcode sequence to be inserted.

6. The method for combating the GNN malware detection model according to any one of claims 1 to 5, characterized in that: The reinforcement learning agent selects a benign opcode sequence and determines an insertion position of the benign opcode sequence by the following steps: Determine the malicious probability predicted by the GNN detection model and the modified CFG of the original malware before and after each modification; The reinforcement learning agent is used to learn the benign operation code sequence and insertion position inserted by the original malware according to the malicious probability and CFG changes predicted by the GNN detection model before and after each modification of the original malware, so that the malicious probability of the modified malware is reduced.

7. The method for combating the GNN malware detection model according to claim 6, characterized in that: Using the reinforcement learning agent to learn the benign opcode sequence and insertion position inserted by the original malware according to the malicious probability predicted by the GNN detection model before and after each modification of the original malware and the modified CFG, so that the malicious probability of the modified malware is reduced, including: The benign opcode sequence and the insertion position inserted by the original malware are regarded as two discrete actions in a multi-dimensional discrete action space; Determining a reward function of a reinforcement learning environment according to the maliciousness probability of the original malware before and after each modification; The modified CFG of the original malware is used as a new state, and the modified malicious probability is used as the malicious probability before the next action is modified. The reinforcement learning agent learns according to the reward function and makes the next action.

8. The method for combating the GNN malware detection model according to any one of claims 1 to 5, characterized in that: The benign opcode sequence is preprocessed before being inserted into the original malware, and also includes: Decompile each software sample in the second software data set to obtain a CFG with node features of each software sample; The GNN detection model is trained using the CFG with node features and corresponding labels of each software sample in the second software data set, so that the GNN detection model distinguishes between benign software and malicious software by learning the structure and node features of the CFG.

9. An adversarial system for a GNN malware detection model, characterized in that: include: Benign opcode sequence insertion module, which is used to insert the original malware after preprocessing the benign opcode sequence until the modified malware is detected as benign software by the GNN detection model; An adversarial malware acquisition module, used to use the last modified iterative malware as adversarial malware; The benign opcode sequence is selected by a reinforcement learning agent from a benign opcode sequence library, and the benign opcode sequence library is constructed by extracting benign opcode sequences from benign software samples detected by the GNN detection model; The insertion position of the benign opcode sequence is determined by the reinforcement learning agent.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the adversarial method for the GNN malware detection model as described in any one of claims 1 to 8 is implemented.