Malicious code self-adaptive detection and repair method and device based on Transform model

Through the Transformer model, the code is subject to global semantic understanding and malicious behavior detection, and the repair code is automatically generated, which solves the problems of insufficient recognition capabilities and inefficient repair efficiency of traditional malicious code detection methods, and realizes efficient and flexible malicious code repair.

CN120408606APending Publication Date: 2025-08-01XIAMEN UNIV OF TECH +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510493227.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing malicious code detection methods have weak ability to identify new or encrypted malicious codes, and traditional detection methods lack a comprehensive analysis of the dynamic behavior of the code when running, resulting in frequent missed and false positives, and malicious code repair is inefficient in human intervention.

Method used

The Transformer model is used to understand the code syntax and semantics. By converting the code into a token sequence, the pre-trained code embedding model is used to generate local semantic vectors, combining the Transformer encoder and decoder to generate global semantic vectors, malicious behavior detection and automatic generation of repair codes, introducing the diversity reward mechanism optimization and repair process.

Benefits of technology

It realizes automated detection and repair of malicious code, improves detection accuracy and coverage, reduces manual intervention, and improves repair efficiency and flexibility and diversity of code repair.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408606A_ABST
    Figure CN120408606A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of computer security, and discloses a malicious code self-adaptive detection and repair method and device based on a Transform model, and the method comprises the steps: obtaining a to-be-detected code, and converting the to-be-detected code into a token sequence; inputting the token sequence into a code embedding model to obtain a local semantic vector of each token; the local semantic vectors of all the tokens in the token sequence are input into a Transform model, and an encoder of the Transform model generates a global semantic vector of the code; a decoder of the Transform model generates a suspected malicious behavior sequence based on the global semantic vector, and matches each suspected malicious behavior with a malicious behavior rule to obtain a real malicious behavior sequence; the method comprises the steps of generating a real malicious behavior sequence, generating a repair target based on the real malicious behavior sequence and a context sensing semantic vector of a target token, repairing the target token based on the repair target and the context sensing semantic vector of the target token, and converting the repaired token into a readable code to obtain a repaired code. The malicious codes can be automatically detected and repaired accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer security, and particularly relates to a method and device for malicious code adaptive detection and repair based on a Transformer model. Background Art

[0002] With the wide application of information technology, computer security issues have become increasingly serious, especially the threat of malicious code. Malicious code (such as viruses, Trojans, ransomware, backdoor programs, etc.) attacks computer systems, network infrastructures, and user data through various channels, resulting in serious consequences such as system crashes, data leaks, and property losses. Despite the continuous update of existing protection means, the complexity and variety of malicious code attacks pose huge challenges to traditional detection and defense methods.

[0003] Traditional malicious code detection methods mainly rely on signature matching technology, which uses the characteristics or signatures of known malicious code for matching. This method shows weak recognition ability for new malicious code or malicious code that has been encrypted or mutated, and is prone to false negatives and false positives. That is, the technical solutions of the prior art have a low detection accuracy for malicious code. Summary of the Invention

[0004] The purpose of the present invention is to achieve automated detection and repair of malicious code, improve the accuracy and coverage rate of malicious code detection, automatically generate repair code based on the discovery of malicious behaviors, reduce manual intervention, and improve the repair efficiency.

[0005] In a first aspect, an embodiment of the present invention provides a method for malicious code adaptive detection and repair based on a Transformer model, the method comprising:

[0006] Obtain the code to be detected, and convert the code to be detected into a token sequence, where each token in the token sequence is a basic component unit of the code to be detected at the syntactic and semantic levels;

[0007] Input the token sequence into a pre-trained code embedding model to obtain the local semantic vector of each token in the token sequence;

[0008] Input the local semantic vectors of all tokens in the token sequence into a Transformer model to perform the following steps by the Transformer model:

[0009] The encoder of the Transformer model generates a context-aware semantic vector for each token based on the semantic local vectors of all tokens, and generates a global semantic vector of the code to be detected based on the context-aware semantic vectors of all tokens;

[0010] The decoder of the Transformer model generates a sequence of suspected malicious behaviors based on the global semantic vector, and matches each suspected malicious behavior in the sequence of suspected malicious behaviors with a preset malicious behavior rule to obtain a sequence of actual malicious behaviors;

[0011] The decoder of the Transformer model generates a repair target based on the sequence of actual malicious behaviors and the context-aware semantic vector of the target token, repairs the target token based on the repair target and the context-aware semantic vector of the target token to generate a repaired token, and converts the repaired token into readable code to obtain repaired code, where the target token is a malicious token located based on the sequence of actual malicious behaviors.

[0012] Optionally, the matching each suspected malicious behavior in the sequence of suspected malicious behaviors with a preset malicious behavior rule to obtain a sequence of actual malicious behaviors includes:

[0013] Obtain a preset malicious behavior rule library, where the malicious behavior rule library includes multiple malicious behavior rules;

[0014] For each suspected malicious behavior in the sequence of suspected malicious behaviors, match the suspected malicious behavior with the malicious behavior rules included in the malicious behavior rule library to obtain a matching result;

[0015] If the matching result is that the suspected malicious behavior does not match the malicious behavior rule, determine that the suspected malicious behavior does not belong to the actual malicious behavior;

[0016] If the matching result is that the suspected malicious behavior matches the malicious behavior rule, determine that the suspected malicious behavior belongs to the actual malicious behavior, and use the determined actual malicious behavior as the sequence of actual malicious behaviors.

[0017] Optionally, the method further includes:

[0018] The decoder of the Transformer model determines the harm degree of each suspected malicious behavior in the sequence of suspected malicious behaviors based on a preset rule, and determines the weight of each suspected malicious behavior based on the determined harm degree, where the weight of a suspected malicious behavior is proportional to the harm degree of the suspected malicious behavior;

[0019] The decoder of the Transformer model determines the indication values corresponding to the respective suspected malicious behaviors through a preset indication function. If a suspected malicious behavior belongs to a real malicious behavior, the indication value corresponding to the suspected malicious behavior is a first value. If a suspected malicious behavior does not belong to a real malicious behavior, the indication value corresponding to the suspected malicious behavior is a second value, and the first value is greater than the second value;

[0020] The decoder of the Transformer model calculates the probability of malicious behavior in the code to be detected based on the indication values corresponding to the respective suspected malicious behaviors and the weights of the respective suspected malicious behaviors.

[0021] Optionally, the encoder of the Transformer model includes a repair target generation module;

[0022] The encoder of the Transformer model generates a repair target based on the real malicious behavior sequence and the context-aware semantic vector of the target token, including:

[0023] The repair target generation module generates a repair target based on the real malicious behavior sequence and the context-aware semantic vector of the target token;

[0024] Among them, the total loss function of the repair target generation module includes a first loss function and a second loss function. The first loss function is used to make the similarity between the global semantic vector of the repaired code and the global semantic vector of the code to be detected greater than a preset similarity, and the second loss function is used to make the matching degree between the repaired code and the repair target greater than a preset matching degree.

[0025] Optionally, the method further includes:

[0026] The decoder of the Transformer model calculates the uniqueness score of each repair target based on the code repair method of each repair target, where the uniqueness score of a repair target is inversely proportional to the occurrence frequency of the code repair method of the repair target;

[0027] The decoder of the Transformer model calculates the diversity score of the repair targets based on the uniqueness scores of multiple repair targets;

[0028] The decoder of the Transformer model generates multiple new repair targets based on the diversity score.

[0029] Optionally, the method further includes:

[0030] The decoder of the Transformer model extracts new malicious behavior rules based on the matching results between each suspected malicious behavior in the suspected malicious behavior sequence and the malicious behavior rules.

[0031] The decoder of the Transformer model updates the malicious behavior rule library based on the new malicious behavior rules.

[0032] In a second aspect, an embodiment of the present invention provides a malicious code adaptive detection and repair device based on a Transformer model, the device includes:

[0033] A code conversion module, configured to obtain the code to be detected and convert the code to be detected into a token sequence, where each token in the token sequence is a basic component unit of the code to be detected at the syntactic and semantic levels;

[0034] A local semantic vector determination module, inputs the token sequence into a pre-trained code embedding model to obtain the local semantic vector of each token in the token sequence;

[0035] A local semantic vector input module, inputs the local semantic vectors of all tokens in the token sequence into a Transformer model to perform the following steps through the Transformer model:

[0036] The encoder of the Transformer model generates a context-aware semantic vector for each token based on the semantic local vectors of all tokens, and generates a global semantic vector of the code to be detected based on the context-aware semantic vectors of all tokens;

[0037] The decoder of the Transformer model generates a suspected malicious behavior sequence based on the global semantic vector, and matches each suspected malicious behavior in the suspected malicious behavior sequence with a preset malicious behavior rule to obtain a true malicious behavior sequence;

[0038] The decoder of the Transformer model generates a repair target based on the true malicious behavior sequence and the context-aware semantic vector of the target token, repairs the target token based on the repair target and the context-aware semantic vector of the target token to generate a repaired token, and converts the repaired token into a readable code to obtain a repaired code, where the target token is a malicious token located based on the true malicious behavior sequence.

[0039] In a third aspect, an embodiment of the present invention provides an electronic device, including:

[0040] At least one processor;

[0041] A memory for storing instructions executable by the at least one processor;

[0042] Wherein, the at least one processor is configured to execute the instructions to implement the method described in the first aspect.

[0043] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the method described in the first aspect.

[0044] In a fifth aspect, an embodiment of the present invention provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method described in the first aspect.

[0045] The present invention converts the code to be detected into a token sequence, inputs the token sequence into a pre-trained code embedding model to obtain the local semantic vectors of each token in the token sequence, inputs the local semantic vectors of all tokens in the token sequence into a Transformer model, and the encoder of the Transformer model generates the context-aware semantic vectors of each token; generates the global semantic vector of the code to be detected based on the context-aware semantic vectors of all tokens; generates a suspected malicious behavior sequence based on the global semantic vector, matches the suspected malicious behavior sequence with a preset malicious behavior rule to obtain a real malicious behavior sequence, locates the malicious token through the real malicious behavior sequence, and finally generates a repair target according to the real malicious behavior sequence and the context-aware semantic vector of the malicious token, and repairs the malicious token according to the repair target and the context-aware semantic vector of the malicious token, and converts the repaired token into readable code.

[0046] As can be seen from the above description, the present invention realizes the automatic detection and repair of malicious code, not only improves the accuracy and coverage rate of malicious code detection, but also can automatically generate repair code on the basis of discovering malicious behaviors, reduces manual intervention, and improves the repair efficiency. Description of the Drawings

[0047] Figure 1 It is a flowchart of a method for adaptive detection and repair of malicious code based on a Transformer model provided by an embodiment of the present invention;

[0048] Figure 2Schematic diagram of a malicious code adaptive detection and repair device based on the Transformer model provided by an embodiment of the present invention;

[0049] Figure 3 Schematic diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0050] The present invention will be described in detail below through embodiments.

[0051] With the wide application of information technology, computer security issues have become increasingly serious, especially the threat of malicious code. Malicious code (such as viruses, Trojans, ransomware, backdoor programs, etc.) attacks computer systems, network infrastructures, and user data through various channels, resulting in serious consequences such as system crashes, data leaks, and property losses. Although existing protection means are constantly updated, the complexity and variety of malicious code attacks pose huge challenges to traditional detection and defense methods.

[0052] Traditional malicious code detection methods mainly rely on signature matching technology, which uses the characteristics or signatures of known malicious code for matching. This method shows weak recognition ability for new malicious code or malicious code that has been encrypted or mutated, and is prone to false negatives and false positives. At the same time, traditional detection methods often rely on static features and lack a comprehensive analysis of the dynamic behavior of the code during runtime. In addition, the repair of malicious code often requires manual intervention, and the repair process is cumbersome and inefficient.

[0053] In recent years, deep learning technology, especially the application of the Transformer model, has achieved remarkable results in the fields of natural language processing, image recognition, etc. Malicious code detection methods based on deep learning can effectively identify complex malicious code by automatically learning the characteristics and behavior patterns of the code, especially being able to discover some hidden malicious behaviors that cannot be captured by traditional methods. However, existing deep learning-based malicious code detection methods still have some deficiencies, especially in aspects such as the understanding of code syntax and semantics, the accurate positioning of malicious behaviors, and automatic repair. In particular, existing methods often ignore the modeling of the global semantics of the code, resulting in difficulties in accurate positioning and repair in complex or highly mutated malicious code scenarios.

[0054] To solve the above technical problems existing in the prior art, the present invention proposes a malicious code automatic detection and repair method based on the Transformer, which can be roughly divided into the following steps:

[0055] The first step is to convert the input source code or binary code into a token sequence, where each token in the token sequence is a basic unit of the code at the syntactic and semantic levels.

[0056] The second step is to use a pre-trained code embedding model to map each token in the token sequence to a semantic vector, which can be called a local semantic vector.

[0057] The third step is to use a Transformer encoder to capture the context dependencies in the code and generate a global semantic vector of the code through an attention mechanism.

[0058] The fourth step is to generate a potential malicious behavior sequence based on the global semantic vector using a Transformer decoder, match it with a predefined malicious behavior rule library, calculate the probability of malicious code, and locate the position of the malicious code. By combining the behavior matching and the weight calculation of the harm degree, the accuracy of malicious behavior detection is further improved.

[0059] The fifth step is to generate a repair target based on the detected malicious behavior, and use a Transformer decoder to generate a repaired code sequence in combination with the code semantic vector. By optimizing the loss function of semantic similarity and repair target correctness, and introducing a repair code diversity reward mechanism, the diversity and flexibility of the repair are enhanced.

[0060] The malicious code automatic detection and repair method provided by the embodiments of the present invention can not only accurately detect and locate malicious behaviors, but also automatically generate repair codes, improving the security and stability of software, and is applicable to various scenarios of malicious code repair.

[0061] After elaborating on the overall technical solution of the embodiments of the present invention, the following emphasizes several key points of the embodiments of the present invention:

[0062] The first key point is code semantic understanding. By converting the source code or binary code into a token sequence and using a pre-trained code embedding model (CodeBERT) to generate a semantic vector of the code. Then, use a Transformer encoder to perform context-aware processing on the code to obtain a context-aware semantic vector, capture the syntactic and semantic relationships between different tokens, and generate a global semantic vector based on the context-aware semantic vector, so as to more deeply understand the potential malicious behaviors in the code.

[0063] The second key point is malicious behavior detection and localization. The Transformer decoder can generate a potential malicious behavior sequence based on the global semantic vector. By matching with a predefined malicious behavior rule library, the probability of malicious code is calculated, and the location of the malicious code is accurately located. As the number of code identifications increases, the Transformer model will self-learn to update the malicious behavior rule library, increasing the number of matchable malicious behaviors in the malicious behavior rule library, thereby improving the accuracy and repair effect of malicious code identification next time. Compared with traditional feature matching methods, this method can identify complex logic and latent threats in the code that are not apparent.

[0064] The third key point is automatic repair. Based on the detected malicious behavior, corresponding repair targets are generated, and the repaired code sequence is generated through the Transformer decoder combined with the semantic vector of the code. In the process of generating repair targets, not only the semantic similarity of the code is considered, but also a diversity reward mechanism is used to promote the generation of diverse repair solutions, thereby avoiding a single repair method and improving the flexibility and adaptability of code repair.

[0065] The present invention can achieve automatic detection and repair of malicious code, not only improving the accuracy and coverage of malicious code detection, but also automatically generating repair code based on the discovery of malicious behavior, reducing manual intervention and improving repair efficiency. The innovation of this method lies in making full use of the semantic modeling ability of the Transformer model to identify potential threats from the global semantic level of the code, and combining a diversity reward mechanism to enhance the flexibility and diversity of the repair process, significantly improving the effect of malicious code repair.

[0066] In summary, the present invention has an important technical breakthrough in the field of malicious code detection and repair, can effectively improve the security and stability of computer systems, has a wide range of application prospects, especially in defending against new attacks such as advanced persistent threats (APTs) and zero-day attacks, and can provide strong guarantees.

[0067] After expounding the overall technical solution of the embodiments of the present invention, a method for adaptive detection and repair of malicious code based on the Transformer model provided by the present invention will be elaborated in detail below.

[0068] As Figure 1 shown, a method for adaptive detection and repair of malicious code based on the Transformer model provided by the embodiments of the present invention may include the following steps:

[0069] S110, obtain the code to be detected and convert the code to be detected into a token sequence.

[0070] Among them, each token in the token sequence is a basic component unit of the code to be detected at the syntactic and semantic levels.

[0071] Specifically, the code to be detected can be source code or binary code. After obtaining the code C to be detected, the code to be detected can be converted into a token sequence T = {t1, t2,..., t n}, and each token t i represents the basic component unit of the code at the syntactic and semantic levels, such as operators, API calls, variable names, etc. After this step, the syntactic and semantic information of the code is represented as a token sequence. To further enhance the understanding of the code, syntactic and semantic information (such as control flow, data flow, etc.) is added.

[0072] S120. Input the token sequence into the pre-trained code embedding model to obtain the local semantic vector of each token in the token sequence.

[0073] Specifically, use the pre-trained code embedding model (which can be called CodeBERT) to map each token into a vector representation e i , and this vector representation e i is the local semantic vector, and the local semantic vector corresponding to each token represents the semantic information of the token. Compared with traditional machine learning methods or methods that do not use CodeBERT, CodeBERT has the following advantages: First, in terms of semantic understanding, it can capture the deep semantics and context information in the code. Second, it has the ability to automatically learn features, which can avoid cumbersome manual feature engineering. Third, it supports cross-languages and can process the code of multiple programming languages. Fourth, it has better generalization ability. It can learn general patterns from a large number of codes and has better performance in different tasks and code libraries. Fifth, it has the ability to generate code. It can not only understand the code but also generate repaired and completed code.

[0074] S130. Input the local semantic vectors of all tokens in the token sequence into the Transformer model to execute the following steps S140 to S160 through the Transformer model.

[0075] Specifically, after obtaining the local semantic vectors of each token in the token sequence, the local semantic vectors of all tokens in the token sequence can be input into the Transformer model. The Transformer model may include a Transformer encoder (described below as the encoder of the Transformer model) and a Transformer decoder (described below as the decoder of the Transformer model), and the encoder and decoder of the Transformer model perform S140 to S170 on the local semantic vectors of the tokens.

[0076] S140, the encoder of the Transformer model generates a context-aware semantic vector for each token based on the semantic local vectors of all tokens, and generates a global semantic vector of the code to be detected based on the context-aware semantic vectors of all tokens.

[0077] Specifically, input the token vector sequence E = {e1, e2,..., e n} into the encoder of the Transformer model to generate a context-aware semantic vector H = {h1, h2,..., h n}. The Transformer model captures the dependencies between different tokens through a multi-layer self-attention mechanism, thus effectively modeling the syntax structure and semantic information of the code. h i represents the context-aware semantic vector of the i-th token, which can effectively capture the potential logic and malicious behavior in the code.

[0078] Moreover, in order to capture the global semantic information of the entire code snippet, the attention mechanism is used to aggregate the vectors in H to generate the global semantic vector h global of the code to be detected. The attention mechanism enables the Transformer model to dynamically adjust its contribution value according to the importance of each token, so as to more accurately obtain the global semantics of the code to be detected.

[0079]

[0080] where α i is the attention weight calculated according to the semantic importance of each token, and w is a learnable weight vector. Through this global semantic vector, the Transformer model can obtain the context information of the entire code to be detected.

[0081] S150: The decoder of the Transformer model generates a suspected malicious behavior sequence based on the global semantic vector, and matches each suspected malicious behavior in the suspected malicious behavior sequence with a preset malicious behavior rule to obtain a true malicious behavior sequence.

[0082] Specifically, the Transformer decoder is based on the global semantic vector h global Generate a potential malicious behavior sequence, which can be called a suspected malicious behavior sequence, that is, the behavior in the suspected malicious behavior sequence may be malicious behavior. The suspected malicious behavior sequence can be called A = {a1, a2, ..., a m}, a j It represents the jth suspected malicious behavior in the suspected malicious behavior sequence. In actual applications, the suspected malicious behavior can be file deletion or network connection, etc.

[0083] The decoder gradually generates a sequence of suspected malicious behaviors through autoregressive generation:

[0084] A=TransformerDecoder(h global )

[0085] A new malicious behavior is generated each time until all sequences are generated.

[0086] Next, the generated suspected malicious behavior sequence A is compared with the predefined malicious behavior rule library R = {r1, r2, ..., r k} to obtain the real malicious behavior sequence. i represents the i-th malicious behavior rule.

[0087] As an implementation of an embodiment of the present invention, S150, matching each suspected malicious behavior in the suspected malicious behavior sequence with a preset malicious behavior rule to obtain a true malicious behavior sequence, may include the following steps, namely, steps a1 to a4:

[0088] Step a1: Obtain a preset malicious behavior rule library, which includes multiple malicious behavior rules.

[0089] Step a2: For each suspected malicious behavior in the suspected malicious behavior sequence, the suspected malicious behavior is matched with the malicious behavior rules included in the malicious behavior rule library to obtain a matching result.

[0090] In step a3, if the matching result is that the suspected malicious behavior does not match the malicious behavior rule, it is determined that the suspected malicious behavior is not a real malicious behavior.

[0091] Step a4, if the matching result is that the suspected malicious behavior matches the malicious behavior rule, determine that the suspected malicious behavior belongs to a real malicious behavior, and use the determined real malicious behavior as the real malicious behavior sequence.

[0092] Specifically, the preset malicious behavior rule library can be a predefined malicious behavior rule library, and this malicious behavior rule library can be generated according to the occurred malicious behaviors. Each suspected malicious behavior in the suspected malicious behavior sequence can be compared with the malicious behavior rules in the malicious behavior rule library one by one. If a suspected malicious behavior does not hit the malicious behavior rules in the malicious behavior rule library, that is, the suspected malicious behavior does not match the malicious behavior rules, then it is determined that this suspected malicious behavior does not belong to a real malicious behavior, and 0 is returned through the indicator function. If a suspected malicious behavior hits the malicious behavior rules in the malicious behavior rule library, that is, the suspected malicious behavior matches the malicious behavior rules, then it is determined that this suspected malicious behavior belongs to a real malicious behavior, and 1 is returned through the indicator function. Finally, the sequence composed of the determined real malicious behaviors is determined as the real malicious behavior sequence.

[0093] Moreover, as an implementation manner of the embodiment of the present invention, this malicious code adaptive detection and repair method based on the Transformer model may further include the following steps, namely step b1 and step b2:

[0094] Step b1, the decoder of the Transformer model extracts new malicious behavior rules based on the matching results of each suspected malicious behavior in the suspected malicious behavior sequence and the malicious behavior rules.

[0095] Step b2, update the malicious behavior rule library based on the new malicious behavior rules.

[0096] Specifically, as the number of code identifications increases, the decoder of the Transformer model will self-learn to update the malicious behavior rule library, so that the number of matching malicious behaviors in the malicious behavior rule library increases, thereby improving the accuracy and repair effect of code identification next time. Compared with the traditional feature matching method, this method can identify complex logics and potential threats that are not yet apparent in the code.

[0097] For example, if the matching result of a suspected malicious behavior and the malicious behavior rules is a mismatch, but the decoder of the Transformer model determines that this suspected malicious behavior belongs to a real malicious behavior, then the decoder of the Transformer model can extract the malicious behavior rule of this suspected malicious behavior as a new malicious behavior rule, and update the new malicious behavior rule to the malicious behavior rule library.

[0098] Based on the above embodiments, in one implementation, the malicious code adaptive detection and repair method based on the Transformer model may further include the following steps, namely steps c1 to c3 respectively:

[0099] Step c1, determine the harm degree of each suspected malicious behavior in the suspected malicious behavior sequence based on a preset rule, and determine the weight of each suspected malicious behavior based on the determined harm degree, where the weight of a suspected malicious behavior is proportional to the harm degree of the suspected malicious behavior;

[0100] Step c2, determine the indication value corresponding to each suspected malicious behavior through a preset indication function. If a suspected malicious behavior belongs to a real malicious behavior, the indication value corresponding to the suspected malicious behavior is the first value. If a suspected malicious behavior does not belong to a real malicious behavior, the indication value corresponding to the suspected malicious behavior is the second value, and the first value is greater than the second value.

[0101] Step c3, calculate the probability of the existence of malicious behavior in the code to be detected based on the indication value corresponding to each suspected malicious behavior and the weight of each suspected malicious behavior.

[0102] To determine the probability of the existence of malicious behavior in the code to be detected, this embodiment provides a calculation method for determining the probability of the existence of malicious behavior in the code to be detected:

[0103]

[0104] Among them, P malicious is the malicious behavior probability, m is the number of suspected malicious behaviors included in the suspected malicious behavior sequence, w j is the weight of the j-th suspected malicious behavior determined based on the harm degree of the j-th suspected malicious behavior. The greater the harm degree of the j-th suspected malicious behavior, the greater w j ; the smaller the harm degree of the j-th suspected malicious behavior, the smaller w j . is the indication function. If a j belongs to a real malicious behavior, then the value returned by the indication function can be 1; if a j does not belong to a real malicious behavior, then the value returned by the indication function can be 0. Finally, the malicious behavior probability is calculated based on the above formula. The traditional calculation of malicious behavior probability is only based on the number of matches and does not consider the harm degree of different malicious behaviors. The malicious behavior probability calculation method provided in this embodiment can improve the calculation accuracy of malicious behavior probability.

[0105] S160. The decoder of the Transformer model generates a repair target based on the true malicious behavior sequence and the local semantic vector of the target token. Based on the repair target and the context-aware semantic vector of the target token, the target token is repaired to generate a repaired token, and the repaired token is converted into readable code to obtain the repaired code.

[0106] Among them, the target token is a malicious token located based on the true malicious behavior sequence.

[0107] Specifically, according to the true malicious behavior sequence, the corresponding malicious code segment, that is, the malicious token, can be located, which is called the target token.

[0108] L malicious ={l j |a j ∈R, l j is the code position corresponding to a j}

[0109] In this step, by associating the malicious behavior with the code position, the specific position of the malicious token can be determined and repaired. For example, if the located position is the third line, then the token on the third line can be determined as the malicious token.

[0110] According to the true malicious behavior sequence and the context-aware semantic vector H of the target token, a repair target G={g1, g2,... g p} is generated, where g k represents the code segment after repairing the target token. The generation of the repair target is obtained based on the reasoning process of code semantics.

[0111] For example, assume the true malicious behavior is to delete a file, and the repair target is to remove this malicious operation. A repair target G={g1} is generated, where g1 is the repair target code segment:

[0112] Next, the Transformer decoder combines the repair target G and the context-aware semantic vector H to generate a repaired code sequence C repaired , that is, to generate a repaired token:

[0113] C repaired =TransformerDecoder(H, G)

[0114] The Transformer decoder repairs malicious behaviors based on the original code semantics through conditional generation, generates repaired tokens, and converts the repaired tokens into readable code to obtain the repaired code.

[0115] In the present invention, the code to be detected is converted into a token sequence, the token sequence is input into a pre-trained code embedding model to obtain the local semantic vectors of each token in the token sequence, and the local semantic vectors of all tokens in the token sequence are input into the Transformer model. The encoder of the Transformer model generates the context-aware semantic vectors of each token; a global semantic vector of the code to be detected is generated based on the context-aware semantic vectors of all tokens; a suspected malicious behavior sequence is generated based on the global semantic vector, the suspected malicious behavior sequence is matched with a preset malicious behavior rule to obtain a real malicious behavior sequence, and the malicious tokens are located through the real malicious behavior sequence. Finally, a repair target is generated according to the real malicious behavior sequence and the context-aware semantic vectors of the malicious tokens, and the malicious tokens are repaired according to the repair target and the context-aware semantic vectors of the malicious tokens. The repaired tokens are converted into readable code.

[0116] As can be seen from the above description, the present invention realizes the automatic detection and repair of malicious code, which not only improves the accuracy and coverage rate of malicious code detection, but also can automatically generate repair code based on the discovery of malicious behaviors, reduce manual intervention, and improve the repair efficiency.

[0117] On the basis of the above embodiments, in one implementation manner, the encoder of the Transformer model may include a repair target generation module, and the repair target generation module is used to generate a repair target.

[0118] At this time, the encoder of the Transformer model generates a repair target based on the real malicious behavior sequence and the context-aware semantic vectors of the target tokens, which may include the following steps:

[0119] The repair target generation module generates a repair target based on the real malicious behavior sequence and the context-aware semantic vectors of the target tokens.

[0120] Among them, the total loss function of the repair target generation module includes a first loss function and a second loss function. The first loss function is used to make the similarity between the global semantic vector of the repaired code and the global semantic vector of the code to be detected greater than a preset similarity, and the second loss function is used to make the matching degree between the repaired code and the repair target greater than a preset matching degree.

[0121] Specifically, the loss function of the repair target generation module combines code semantic similarity and repair target correctness. The code semantic similarity loss (the first loss function) measures the global semantic difference between the repaired code and the original code through cosine similarity. The repair target correctness loss (the second loss function) measures the generation quality of the repaired code snippet through cross-entropy.

[0122]

[0123] Among them, is the total loss function of the repair target generation module, and λ1 and λ2 are weight coefficients, which can be determined according to the actual situation, and the embodiments of the present invention do not make specific limitations on this.

[0124] is the code semantic similarity, calculated using cosine similarity:

[0125]

[0126] where h repaired is the global semantic vector of the repaired code.

[0127] For example, assume that the global semantic vector of the original code to be detected is: h global = [0.5, 0.2, 0.3], and the global semantic vector of the repaired code is: h repaired = [0.6, 0.1, 0.3]. By calculating the cosine similarity:

[0128]

[0129] It is calculated that:

[0130]

[0131] is the repair target correctness loss, calculated using cross-entropy.

[0132] Given the repair target G = {g1, g2,..., g p} and the repaired code sequence C repaired = {c′1, c′2,..., c′ p}, use the attention mechanism to aggregate the vectors in C repaired to obtain the global semantic vector h repaired of the repaired code.

[0133]

[0134] where w is a learnable attention weight vector.

[0135] Cross - entropy loss The formula is as follows:

[0136]

[0137] Among them, P(c′ k |g k ) is the conditional probability of the repaired code snippet c′ k under the given repair goal g k , and usually a probability model (such as the softmax function) is used to calculate it.

[0138] For example, assume that the expected repair goal is to replace the code for deleting a file with "skip operation".

[0139] Use cross - entropy to measure the quality of the generated repair goal:

[0140]

[0141] Assume that the generation probability P(c′ k |g k ) of the repaired code snippet c′ k = 0.8, then

[0142]

[0143] Calculate the total loss function of the repair goal generation module:

[0144] Assume that λ1 = 1, λ2 = 1, then

[0145]

[0146] It can be seen that the loss function of the repair goal generation module combines code semantic similarity and repair goal correctness, making the accuracy of the repaired code higher.

[0147] Based on the above - mentioned embodiments, in one implementation manner, the malicious code adaptive detection and repair method based on the Transformer model may further include the following steps:

[0148] The first step: The decoder of the Transformer model calculates the uniqueness score of each repair goal based on the code repair method of each repair goal.

[0149] Among them, the uniqueness score of a repair goal is inversely proportional to the occurrence frequency of the code repair method of the repair goal.

[0150] The second step: The decoder of the Transformer model calculates the diversity score of the repair goals based on the uniqueness scores of multiple repair goals.

[0151] In the third step, the decoder of the Transformer model generates multiple repair targets based on the diversity score.

[0152] Specifically, to improve the diversity of the repaired code, a diversity reward mechanism (which can be called the diversity score) is added By encouraging the generation of diverse repair results, a single repair method is avoided.

[0153]

[0154] Among them, Unique(g k ) represents the uniqueness score of the repair target g k , aiming to increase the flexibility of the system through diverse repair paths.

[0155] For example, assume that two repair targets G = {g1, g2} are generated.

[0156] The repair target 1 is to directly remove the file deletion operation.

[0157] The repair target 2 is to replace the deletion operation with a logging operation.

[0158] Calculate the uniqueness scores Unique(g1) and Unique(g2) of each repair target.

[0159] Suppose Unique(g1) = 1 (the repair target 1 is a common repair method)

[0160] Unique(g2) = 1.5 (the repair target 2 is a relatively innovative repair method)

[0161] Then, the diversity reward score is:

[0162]

[0163] This reward mechanism encourages the Transformer model to generate more diverse repair targets to avoid fixed code repair patterns. Specifically, in practical applications, a diversity score threshold can be set. If the calculated diversity score is less than the diversity score threshold, it means that the diversity score is low and the Transformer model generates fewer repair targets. In this case, the Transformer model can continue to generate multiple new repair targets until the calculated diversity score is greater than the diversity score threshold. By encouraging the Transformer model to generate diverse repair targets, the flexibility and adaptability of code repair can be improved.

[0164] As can be seen from the above description, the main advantages of the embodiments of the present invention include:

[0165] 1. End-to-end learning. Achieve end-to-end learning of code semantic understanding, malicious behavior detection, and code repair through a unified Transformer framework.

[0166] 2. Dynamic weight mechanism. Improve the accuracy of malicious behavior detection by introducing the weight of the harm degree of malicious behavior.

[0167] 3. Diversity repair of malicious code. Generate multiple repair solutions through a diversity reward mechanism to improve the flexibility of code repair.

[0168] Moreover, train and test using publicly available malicious code datasets and open-source code libraries. Evaluation metrics include: the detection accuracy of malicious code, the repair success rate of malicious code, the functional correctness of the repaired code, and the performance overhead. And compare with traditional rule-based and machine learning-based malicious code detection and repair methods. The comparison results show that the technical solution of the embodiment of the present invention has a higher detection accuracy of malicious code, a higher repair success rate of malicious code, a higher functional correctness of the repaired code, and a smaller performance overhead.

[0169] For a clearer description of the solution, the embodiments of the present invention will be elaborated in detail below with specific examples.

[0170] The present invention proposes a code repair model based on Transformer, aiming to automatically repair potential errors in the given code while maintaining its syntactic and semantic consistency. The core of this solution lies in utilizing the self-attention mechanism of the Transformer model and optimizing the effect of code repair through a multi-task loss function.

[0171] 1. Word embedding layer. The role of the word embedding layer is to convert the input code tokens into low-dimensional vector representations for subsequent processing. Given a vocabulary V, each token t in the vocabulary i will obtain the corresponding embedding vector by looking up the matrix W embed Assume the dimension of the embedding vector is d embed , then for each token t i the embedding vector e i is given by the following formula:

[0172] e i = W embed [t i where t i ∈ {1, 2,..., V}

[0173] In this process, the matrix is obtained through model training and continuously updated through an optimization process.

[0174] Suppose there is a piece of code:

[0175] def add(a, b):

[0176] return a + b

[0177] Suppose the vocabulary contains 10 tokens, namely: ‘def’: 1, ‘add’: 2, ‘(‘: 3, ‘a’: 4, ‘b’: 5, ‘)’: 6, ‘:’: 7, ‘return’: 8, ‘+’: 9, ‘end’: 10. Then, for the input tokens ‘[‘def’, ‘add(‘, ‘a’, ‘,‘, ‘b’, ‘]’, ‘:’, ‘return’, ‘a’, ‘+’, ‘b’]‘, their embedding representations are respectively:

[0178] e1 = W embed [1], e2 = W embed [2],... e 12 = W embed

[12]

[0179] Each token will be converted into the corresponding embedding vector, Tokenization -> Vocabulary ID Embedding Lookup -> Use the token ID to look up the corresponding embedding vector position encoding -> Add the position encoding to ensure that the sequence information is retained with context information (in the pre-trained model) -> Further adjust the embedding of the token through the context.

[0180] 2. Position Encoding Layer. In order to enable the Transformer to handle the position information in the sequence, the present invention adopts the position encoding technology. The position encoding matrix provides an additional vector information for each position, where L is the length of the input sequence. The final input representation corresponding to each token t i is: x i = e i + P[i, :]

[0181] This position encoding matrix is obtained by encoding each position of the input, and the update method of the position encoding is similar to that of the word embedding layer.

[0182] The following is an example for illustration. Suppose the length of the input sequence is 12, and the dimension of the embedding vector of each token is 3, and the dimension of the position encoding matrix is 12×3. Suppose the position encoding matrix P is as follows (for simplicity, random values are used):

[0183]

[0184] Then, the positional encoding is added to the word embedding vector to obtain the final input representation:

[0185] x1 = e1 + P[1, :], x2 = e2 + P[2, :],...

[0186] For example, the final input representation of the first token is:

[0187] x1 = e1 + [0.1, 0.2, 0.3]

[0188] 3. Transformer Encoder Layer. The present invention uses a Transformer encoder to process the input sequence, which mainly consists of a self-attention mechanism and a feed-forward neural network. The calculation process of each Transformer encoder layer can be described by the following steps:

[0189] Self-attention mechanism: Each input token t i calculates the attention scores through the query Q i , key K i and value V i . The formula for self-attention is as follows:

[0190]

[0191] where, A i is the attention weight of the i-th token, are the vectors of the query, key, and value.

[0192] Feed-forward neural network:

[0193] After the output A i of the self-attention mechanism, it is further processed through a feed-forward neural network. The calculation method of the feed-forward neural network is:

[0194] H i = ReLU(W2(ReLU(W1A i )) + b2)

[0195] where, W1, W2 are the weight matrices of the linear transformation, and b2 is the bias term.

[0196] Suppose the embedding representations of two input tokens are:

[0197] e1 = [0.5, 0.3, 0.7], e2 = [0.2, 0.4, 0.6]

[0198] When calculating self-attention, first generate the query, key, and value:

[0199] Q1 = [0.1, 0.3, 0.5], K1 = [0.2, 0.4, 0.6], V1 = [0.6, 0.8, 0.1]

[0200] Then, calculate the attention weights:

[0201]

[0202] Next, input the obtained attention values into the feed-forward neural network, and after ReLU activation, the final output H1 is obtained.

[0203] 4. The final output H of the Transformer encoder in the output layer i will be mapped through a fully connected layer to generate predictions for the next token at each position. Specifically, for each token t i , the prediction of the output layer is:

[0204]

[0205] where is the weight matrix of the output layer, is the probability distribution predicted by the model. For example, assume the output after encoding is:

[0206] H1 = [0.1, 0.2, 0.7]

[0207] And the weight matrix of the output layer is:

[0208]

[0209] Then the predicted probability output by the model is:

[0210]

[0211] 5. Loss function calculation

[0212] The present invention designs two loss functions to train the model, namely cross-entropy loss and semantic loss.

[0213] Cross-entropy loss: used to calculate the difference between the model output and the target distribution. The calculation formula of the cross-entropy loss is:

[0214]

[0215] where is the predicted probability of the model for the target token y i .

[0216] Semantic loss: It is used to measure the similarity between the semantics of the model output and the target code. The semantic loss is defined by calculating the cosine similarity between the model output and the target representation, and the formula is as follows:

[0217]

[0218] where and are the normalized vectors of the model output and the target code representation.

[0219] Total loss. The total loss combines the cross-entropy loss and the semantic loss, and the weights of the two are adjusted by the hyperparameter λ. The calculation formula of the total loss is:

[0220]

[0221] 6. Optimizer and training process.

[0222] Optimizer: The present invention uses the Adam optimizer to minimize the total loss, and the update formula of the Adam optimizer is:

[0223]

[0224] where η is the learning rate, m t and υ t are the estimates of the first moment and the second moment of the gradient respectively, and ∈ is a small constant.

[0225] Gradient clipping: To avoid gradient explosion, gradient clipping is applied during the training process to ensure that the gradient norm does not exceed the set maximum value.

[0226] The code repair model based on the Transformer self-attention mechanism proposed by the present invention repairs the syntax and semantics of the input code through the combination of an embedding layer, a position encoding layer, a self-attention mechanism, a feed-forward neural network, and a fully connected output layer. At the same time, by introducing a multi-task loss function, the Adam optimizer, and gradient clipping technology, the performance in the code repair task is effectively optimized.

[0227] In a second aspect, an embodiment of the present invention provides a malicious code adaptive detection and repair device 20 based on a Transformer model, as Figure 2 shown, the device includes:

[0228] A code conversion module 210, configured to obtain the code to be detected and convert the code to be detected into a token sequence, and each token of the token sequence is a basic unit of the code to be detected at the syntax and semantic levels;

[0229] The local semantic vector determination module 220 inputs the token sequence into a pre-trained code embedding model to obtain the local semantic vector of each token in the token sequence;

[0230] The local semantic vector input module 230 inputs the local semantic vectors of all tokens in the token sequence into a Transformer model to perform the following steps through the Transformer model:

[0231] The encoder of the Transformer model generates the context-aware semantic vector of each token based on the semantic local vectors of all tokens, and generates the global semantic vector of the code to be detected based on the context-aware semantic vectors of all tokens;

[0232] The decoder of the Transformer model generates a suspected malicious behavior sequence based on the global semantic vector, and matches each suspected malicious behavior in the suspected malicious behavior sequence with a preset malicious behavior rule to obtain a real malicious behavior sequence;

[0233] The decoder of the Transformer model generates a repair target based on the real malicious behavior sequence and the context-aware semantic vector of the target token, repairs the target token based on the repair target and the context-aware semantic vector of the target token to generate a repaired token, and converts the repaired token into readable code to obtain the repaired code, where the target token is a malicious token located based on the real malicious behavior sequence.

[0234] In a third aspect, an embodiment of the present invention provides an electronic device, including:

[0235] At least one processor;

[0236] A memory for storing instructions executable by the at least one processor;

[0237] Wherein, the at least one processor is configured to execute the instructions to implement the method described in the first aspect.

[0238] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the method described in the first aspect.

[0239] In a fifth aspect, an embodiment of the present invention provides a computer program product, including a computer program, where the computer program implements the method described in the first aspect when executed by a processor.

[0240] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention without departing from the principles and spirit of the present invention.

Claims

1. An adaptive detection and repair method for malicious code based on the Transformer model, characterized in that The method includes: Obtain the code to be detected and convert the code to be detected into a token sequence, where each token in the token sequence is a basic component unit of the code to be detected at the syntactic and semantic levels; Input the token sequence into a pre-trained code embedding model to obtain the local semantic vector of each token in the token sequence; Input the local semantic vectors of all tokens in the token sequence into a Transformer model to perform the following steps through the Transformer model: The encoder of the Transformer model generates the context-aware semantic vector of each token based on the semantic local vectors of all tokens, and generates the global semantic vector of the code to be detected based on the context-aware semantic vectors of all tokens; The decoder of the Transformer model generates a suspected malicious behavior sequence based on the global semantic vector, and matches each suspected malicious behavior in the suspected malicious behavior sequence with a preset malicious behavior rule to obtain a real malicious behavior sequence; The decoder of the Transformer model generates a repair target based on the real malicious behavior sequence and the context-aware semantic vector of the target token, repairs the target token based on the repair target and the context-aware semantic vector of the target token to generate a repaired token, and converts the repaired token into readable code to obtain the repaired code, where the target token is a malicious token located based on the real malicious behavior sequence; 2. The method according to claim 1, wherein The step of matching each suspected malicious behavior in the suspected malicious behavior sequence with a preset malicious behavior rule to obtain a real malicious behavior sequence includes: Obtain a preset malicious behavior rule library, where the malicious behavior rule library includes multiple malicious behavior rules; For each suspected malicious behavior in the suspected malicious behavior sequence, match the suspected malicious behavior with the malicious behavior rules included in the malicious behavior rule library to obtain a matching result; If the matching result is that the suspected malicious behavior does not match the malicious behavior rule, determine that the suspected malicious behavior does not belong to the real malicious behavior; If the matching result is that the suspected malicious behavior matches the malicious behavior rule, determine that the suspected malicious behavior belongs to the real malicious behavior, and use the determined real malicious behavior as the real malicious behavior sequence; 3. The method according to claim 1, wherein The method further includes: The decoder of the Transformer model determines the harm degree of each suspected malicious behavior in the suspected malicious behavior sequence based on a preset rule, and determines the weight of each suspected malicious behavior based on the determined harm degree, where the weight of a suspected malicious behavior is proportional to the harm degree of the suspected malicious behavior; The decoder of the Transformer model determines the indication values corresponding to the respective suspected malicious behaviors through a preset indication function. If a suspected malicious behavior belongs to a true malicious behavior, the indication value corresponding to the suspected malicious behavior is a first value. If a suspected malicious behavior does not belong to a true malicious behavior, the indication value corresponding to the suspected malicious behavior is a second value, and the first value is greater than the second value; The decoder of the Transformer model calculates the probability of malicious behavior in the code to be detected based on the indication values corresponding to the respective suspected malicious behaviors and the weights of the respective suspected malicious behaviors.

4. The method according to any one of claims 1 to 3, characterized in that, The encoder of the Transformer model includes a repair target generation module; The encoder of the Transformer model generates a repair target based on the true malicious behavior sequence and the context-aware semantic vector of the target token, including: The repair target generation module generates a repair target based on the true malicious behavior sequence and the context-aware semantic vector of the target token; Wherein, the total loss function of the repair target generation module includes a first loss function and a second loss function. The first loss function is used to make the similarity between the global semantic vector of the repaired code and the global semantic vector of the code to be detected greater than a preset similarity. The second loss function is used to make the matching degree between the repaired code and the repair target greater than a preset matching degree.

5. The method according to any one of claims 1 to 3, characterized in that The method further includes: The decoder of the Transformer model calculates the uniqueness score of each repair target based on the code repair method of each repair target, wherein the uniqueness score of a repair target is inversely proportional to the occurrence frequency of the code repair method of the repair target; The decoder of the Transformer model calculates the diversity score of the repair targets based on the uniqueness scores of multiple repair targets; The decoder of the Transformer model generates multiple new repair targets based on the diversity score.

6. The method according to claim 2, characterized in that, The method further includes: The decoder of the Transformer model extracts new malicious behavior rules based on the matching results of each suspected malicious behavior in the suspected malicious behavior sequence and the malicious behavior rules; The decoder of the Transformer model updates the malicious behavior rule library based on the new malicious behavior rules.

7. An adaptive malicious code detection and repair device based on the Transformer model, characterized in that, The device includes: A code conversion module, configured to obtain the code to be detected and convert the code to be detected into a token sequence, where each token of the token sequence is a basic component unit of the code to be detected at the syntax and semantic levels; A local semantic vector determination module, which inputs the token sequence into a pre-trained code embedding model to obtain the local semantic vector of each token in the token sequence; The local semantic vector input module inputs the local semantic vectors of all tokens in the token sequence into the Transformer model to perform the following steps through the Transformer model: The encoder of the Transformer model generates a context-aware semantic vector for each token based on the semantic local vectors of all tokens, and generates a global semantic vector of the code to be detected based on the context-aware semantic vectors of all tokens; The decoder of the Transformer model generates a sequence of suspected malicious behaviors based on the global semantic vector, and matches each suspected malicious behavior in the sequence of suspected malicious behaviors with a preset malicious behavior rule to obtain a sequence of actual malicious behaviors; The decoder of the Transformer model generates a repair target based on the sequence of actual malicious behaviors and the context-aware semantic vector of the target token, repairs the target token based on the repair target and the context-aware semantic vector of the target token to generate a repaired token, and converts the repaired token into readable code to obtain repaired code, where the target token is a malicious token located based on the sequence of actual malicious behaviors.

8. An electronic device, characterized in that, Comprising: At least one processor; A memory for storing instructions executable by the at least one processor; Wherein, the at least one processor is configured to execute the instructions to implement the method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method according to any one of claims 1-6.

10. A computer program product, characterized in that, Comprising a computer program, which when executed by a processor implements the method according to any one of claims 1-6.

Citation Information

Cited By

  • Generative artificial intelligence-based malicious code detection method and system

    CN121009547A

  • Automatic program repairing method and device based on non-autoregression parallel generation

    CN122411910A