A method for generating adversarial samples of smart contracts

By employing word vector classification and gradient descent optimization with balance factors and masks, the method generates smart contract adversarial samples that evade detection by minimizing semantic differences and perturbation density, ensuring high stealthiness.

CN115659334BActive Publication Date: 2025-07-15HUAZHONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211271076.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-17
Publication Date
2025-07-15
Estimated Expiration
2042-10-17

AI Technical Summary

Technical Problem

When existing code-adversarial sample generation technology generates malicious smart contracts, disturbances are easily manually identified, resulting in insufficient concealment and inability to effectively evade detection.

Method used

By building a vocabulary library, use word vectors and support vector machines to divide the hyperplane of variable names, calculate the distance between variable names and hyperplanes, generate malicious weight vectors, guide the gradient descent optimization function to select new variable names similar to the nouns of the original variable, and use balance factors and masks to optimize perturbation sparseness to generate highly concealed adversarial samples.

Benefits of technology

The generated smart contract adversarial samples not only circumvent the detection model, but also significantly improve the concealment of the human eye and reduce the possibility of being discovered.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115659334B_ABST
    Figure CN115659334B_ABST
Patent Text Reader

Abstract

The present invention belongs to the fields of deep learning adversarial samples and smart contracts, and specifically relates to a method for generating smart contract adversarial samples. In the process of generating smart contract adversarial samples by modifying variable names, the concept of the nominal concealment of variable names is introduced, and a method for generating smart contract adversarial samples with higher concealment is proposed. It includes: for the new scenario where malicious smart contract creators evade detection through adversarial samples, systematically studying the method for generating smart contract adversarial samples, and proposing a weight-guided high-concealment smart contract adversarial sample generation technology, so that when generating adversarial samples by replacing variable names, new variable names with similar nominal properties to the original variable names are selected to increase the concealment of smart contract adversarial samples; using a balance factor and a mask, introducing the average value, variance of the nominal property gap between the old and new variable names when multiple variable names are modified, and the sparsity of perturbation addition into the optimization function. When finding the solution that minimizes the optimization function, the concealment of the adversarial samples is balanced at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of deep learning adversarial samples and smart contracts, and more specifically, relates to a method for generating adversarial samples of smart contracts. Background Art

[0002] Smart contracts allow for the execution of trustworthy transactions without a third party on the blockchain, and these transactions are traceable and irreversible. Compared with exploiting the vulnerabilities of existing smart contracts, attackers can adopt a more proactive approach to lure others into traps. Common methods include "honeypot smart contracts" and "Ponzi scheme smart contracts". A "honeypot" is a smart contract that seems to have obvious defects in its design, luring others to transfer funds to it. The Ponzi scheme promises users high profits and ultimately defrauds users of their principal. As the security issues of smart contracts have received increasing attention, many deep learning-based malicious smart contract detection tools have been proposed in the academic community. However, attackers deploying malicious smart contracts are very likely to design malicious smart contracts that can evade detection based on the characteristics of smart contracts and corresponding deep learning detection tools.

[0003] Existing methods for generating adversarial samples mainly stem from the research on generating adversarial samples for image data. When facing discrete fields such as natural language processing and programs, the technical methods for generating adversarial samples for images are no longer applicable. This is because image data is continuous, while code data is symbolic and discrete, making it difficult to define perturbations on code. At the same time, compared with image data, code data also has the characteristics that small changes are easily recognized by humans and changing symbols in the code easily changes the code semantics and may fail to pass compilation.

[0004] There were some early attempts to generate adversarial samples for machine learning-based malware detection tools on the Android platform. However, the practical feasibility of these tools often has certain problems. Chen et al. [Sen Chen, Minhui Xue, Lingling Fan, Shuang Hao, Lihua Xu, Haojin Zhu, et al. Automated poisoning attacks and defenses in malware detection systems: An adversarial machine learning approach. Computers & Security, 2018, 73: 326 - 344] proposed injecting carefully crafted adversarial examples into the training dataset to reduce the detection accuracy, but this method is difficult to apply in reality because in most cases, it is difficult for attackers to access the training dataset.

[0005] Liu et al. [Xinbo Liu, Jiliang Zhang, Yaping Lin, He Li. Adversarial examples: Attacks on machine learning-based malware visualization detection methods. In: 17th IEEE International Conference On Trust, Security And Privacy In Computing And Communications (TrustCom 2018), New York, NY, USA, August 1-3, 2018: 1354-1359] created adversarial samples based on grayscale images and successfully deceived the classifier at the image level. However, they did not consider how to convert the grayscale image samples into executable files, and their research was still limited to the image level. During the generation process of adversarial samples, there will be quite a lot of pixel changes, which will cause the generated adversarial images to not be restored to the executable file level. It can only deceive the image-based classifier and cannot generate code that can actually run.

[0006] The existing Discrete Adversarial Manipulation of Programs (DAMP) technology for generating code adversarial samples, in order to keep the code semantically unchanged, chooses to add perturbations by changing variable names. A vector is used to represent the probabilities of each variable name in the vocabulary being selected. The discrete problem of selecting a new variable name to replace the original variable name is made continuous. Then, while keeping the weights of the detection model unchanged, the probability vector is iteratively updated according to the gradient calculated based on the current probability vector. Finally, a variable name with the highest probability is obtained and used to replace the existing variable name to generate adversarial samples. This method only selects one variable name for replacement and has greatly improved the concealment of perturbations. However, code is different from images. When attacking images, single-pixel attacks are highly concealed in images and are hardly detectable. However, the concealment of the adversarial sample generation method in code cannot be simply judged from the number of perturbation sites. Every perturbation in the code can definitely be identified manually. If the difference between the new variable name and the context and the old variable name is too large, the perturbation is very likely to be discovered manually.

[0007] Shashank Srikant et al. [Shashank Srikant, Sijia Liu, Tamara Mitrovska, Shiyu Chang, Quanfu Fan, Gaoyuan Zhang, et al. Generating adversarial computer programs using optimized obfuscations. In: the 9th International Conference on Learning Representations (ICLR 2021), Virtual Event, Austria, May 3 - 7, 2021] transformed the process of generating code adversarial samples into two problems: which parts of the program to transform and which ways to add perturbations, represented this problem using mathematical formulas, and proposed a set of first-order optimization algorithms, using projected gradient descent to find a suitable adversarial sample generation method. Similar to DAMP, this method did not consider the problem that the part-of-speech gap between each group of old and new variable names might lead to a decrease in concealment, and this method modified multiple variable names to generate adversarial samples, and the overly dense perturbation addition might further reduce the concealment of adversarial samples.

[0008] Existing code adversarial sample generation techniques add perturbations by modifying one or more variable names. The problem is that they have not deeply studied the problem that the difference between old and new variable names might cause adversarial samples to be discovered by the human eye. Especially for malicious smart contracts, the process of defrauding victims requires victims to carefully read the source code of the smart contract. At this time, the concealment of the added perturbations for the human eye is crucial. Summary of the Invention

[0009] In view of the defects of the existing technology and the improvement requirements, the present invention provides a method for generating smart contract adversarial samples, aiming to propose a method for generating smart contract adversarial samples with high concealment, so that the generated adversarial samples are not easily discovered by the human eye.

[0010] To achieve the above object, according to one aspect of the present invention, a method for generating smart contract adversarial samples is provided, including:

[0011] According to the labels of whether all variable names in the vocabulary belong to a legitimate smart contract or not, classify the word vectors corresponding to all variable names, obtain a hyperplane that divides the word vectors into two categories: those from malicious smart contracts and those from non-malicious smart contracts, and calculate the distances from the word vectors of each variable name to the hyperplane; map the distances from all variable names whose categories belong to malicious smart contracts to the hyperplane within the range of [0, 1], and represent them with the vector D′ = {d1′, d2′,...}, where, if the word vector of a certain variable name whose category belongs to a malicious smart contract is on the non-malicious smart contract side, the corresponding distance value in the vector D′ is 0, and add 1 to each value in the vector D′ to obtain the maliciousness weight vector;

[0012] Keep the weights of the attacked detection model unchanged, and fix the output result as a wrong result that misleads the attacked detection model. Use the maliciousness weight vector to guide the iterative process of the gradient descent attack optimization function, so that the probability of selecting a variable name with a similar part of speech to the variable name to be replaced from the vocabulary is increased during the iteration, and the generation of adversarial samples for smart contracts is completed.

[0013] Furthermore, the construction method of the vocabulary is as follows:

[0014] Preprocess multiple smart contract source codes, extract all variable names in the multiple smart contract source codes, and form a vocabulary, including variable names and their corresponding labels of whether the smart contracts they belong to are legitimate or not; among them, the multiple smart contract source codes include malicious and non-malicious smart contract source codes.

[0015] Furthermore, use the cosine similarity between the word vectors corresponding to the variable name to be replaced and the new variable name to evaluate the semantic gap between the two variable names. The closer the cosine similarity is to 1, the more similar the parts of speech are, and the better the concealment of the generated adversarial samples.

[0016] Furthermore, during the iteration process, the optimization function used is: the new optimization function obtained by introducing the average value and variance of the part-of-speech gap between the old and new variable names when multiple variable names are modified into the original optimization function of the attacked detection model by using a balance factor.

[0017] Furthermore, during the iteration process, the optimization function used is: also introduce the sparsity of the perturbation added when multiple variable names are modified into the original optimization function of the attacked detection model by using a mask, and the new optimization function obtained.

[0018] Furthermore, during the iteration process, the optimization function used is:

[0019]

[0020] Wherein, D represents the sum of the cosine vector values of the word vectors before and after perturbation at each perturbation site, P' and P respectively represent the smart contract source code after and before adding perturbation, 1 represents a mask set to prevent too many perturbation addition sites, k represents the maximum number of allowable perturbation addition sites, Ω represents the total number of variable names in the vocabulary, n represents the total number of sites to be perturbed, σ 2 represents the variance of the perturbation degree, the balance factor λ1 and the balance factor λ2 are artificially set constants, θ is the detection model to be attacked, l attack represents the loss function of the attack. Generating a perturbation for the symbol at the i-th position using the j-th variable in the vocabulary is represented as [u i '] j = 1 and z i = 1.

[0021] Furthermore, in the process of iteration, an alternating iteration method is adopted. First, the vector u i is fixed, and the vector z is iterated. Then, the value of the new vector z is regarded as a fixed value to iterate the vector u i . Among them, the vector z ∈ {0, 1} n marks whether each site is selected to add perturbation. When z i = 1, it means that the position of P i is selected to add perturbation, and n represents the total number of sites to be perturbed; the one-hot vector u i = {0, 1} |Ω| is used to represent the variable name to be selected for adding perturbation. Each position where z i = 1 corresponds to a vector u i , and Ω represents the total number of variable names in the vocabulary.

[0022] The present invention also provides a computer-readable storage medium. The computer-readable storage medium includes a stored computer program. Among them, when the computer program is run by a processor, it controls the device where the storage medium is located to execute a method for generating adversarial samples of a smart contract as described above.

[0023] Generally speaking, through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:

[0024] (1) The method for generating adversarial samples of a smart contract with high concealment provided by the present invention, aiming at the brand-new scenario where malicious smart contract creators avoid detection through adversarial samples, by proposing a weight-guided high-concealment smart contract adversarial sample generation technology, when generating adversarial samples by replacing variable names, new variable names with similar parts of speech to the original variable names are selected, increasing the concealment of smart contract adversarial samples.

[0025] (2) The highly concealed smart contract adversarial sample generation method provided by the present invention uses a balancing factor and a mask to introduce the average value, variance, and sparsity of the part-of-speech difference between the new and old variable names when multiple variable names are modified into the optimization function. When finding the solution that minimizes the optimization function, the concealment of the adversarial sample is balanced at the same time, so that the generated smart contract adversarial sample can avoid model detection while being difficult to be discovered by the human eye. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 A structural diagram of a site selection vector and a variable name selection vector provided in an embodiment of the present invention;

[0027] Figure 2 A process diagram of generating a malicious weight vector provided by an embodiment of the present invention;

[0028] Figure 3 The overall structure of using malicious weight vectors to generate adversarial samples provided by the embodiments of the present invention. DETAILED DESCRIPTION

[0029] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0030] Embodiment 1

[0031] A method for generating adversarial samples of smart contracts, comprising:

[0032] According to the labels of whether all variable names in the vocabulary belong to good-faith or not smart contracts, the word vectors corresponding to all variable names are classified to obtain a hyperplane that divides the word vectors into two categories: those from malicious smart contracts and those from non-malicious smart contracts. The distance from the word vector of each variable name to the hyperplane is calculated. The distance from all variable names belonging to malicious smart contracts to the hyperplane is mapped to the range of [0,1] and represented by a vector D′={d1′,d2′,...}, where if the word vector of a variable name belonging to a malicious smart contract is on the side of a non-malicious smart contract, the corresponding distance value in the vector D′ is 0. Each value in the vector D′ is added with 1 to obtain a malicious weight vector.

[0033] The weight of the attacked detection model is kept unchanged, and the output result is fixed as an erroneous result that misleads the attacked detection model. The malicious weight vector is used to guide the iterative process of the gradient descent attack optimization function, so that the probability of selecting a variable name that is similar to the noun of the variable to be replaced from the vocabulary library is increased during the iteration process, thereby completing the generation of smart contract adversarial samples.

[0034] It should be noted that the iterative process of guiding the gradient descent to attack the optimization function is actually to guide the gradient descent to find the disturbance addition vector [u i ] j = 1, the vector [u i ] j =1 means that the jth variable name in the vocabulary library is used to generate perturbations for the i-th position symbol in the target smart contract source code.

[0035] When selecting new variable names from the vocabulary to replace the original variable names, it is proposed to increase the probability of selecting new variable names that are closer to the noun nature of the original variable through weight guidance. Word2vec is used to generate word vectors corresponding to the variable names in all vocabularies. SVM is used to find the hyperplane that divides the word vectors into two categories: those from malicious smart contracts and those from non-malicious smart contracts. The distance from each word vector to this hyperplane is calculated to generate a malicious weight vector, which is used for guidance in the process of finding a perturbation adding method by gradient descent, thereby increasing the concealment of the new variable name and reducing the possibility of being discovered.

[0036] As a preferred embodiment, the above vocabulary library is constructed in the following manner:

[0037] Preprocessing is performed on multiple smart contract source codes, and all variable names in the multiple smart contract source codes are extracted to form a vocabulary library, including variable names and labels of whether their corresponding smart contracts are benign or not; wherein the multiple smart contract source codes include malicious and non-malicious smart contract source codes.

[0038] Specifically, when the above method is executed in actual application, it may include the following steps in sequence:

[0039] (1) Take the preprocessed smart contract source code P as the initial input;

[0040] (2) Use vectors to represent the perturbation addition sites and perturbation addition methods used to generate adversarial samples;

[0041] (21) Define a vector z∈{0,1} n To mark whether each site is selected to add perturbation, when z i =1 indicates P i The positions are selected to add perturbations, and n represents the total number of sites to be perturbed.

[0042] (22) For the variable name sites in the code, the variable name to be perturbed is selected, and a one-hot vector u i ={0, 1} |Ω| is used to represent it. For each z i = 1, there is a corresponding vector u i ; Ω represents the total number of variable names in the vocabulary;

[0043] (23) For the variable name at the P i position, the perturbation is generated using the j-th variable name in the vocabulary, which can be expressed as [u i j = 1 and z i = 1;

[0044] (3) Construct a weight generation module to generate a malicious weight vector;

[0045] (31) Preprocess all the smart contract code samples in step (1), extract all the variable names therein, and obtain word vectors for these variable names using word2vec;

[0046] (32) Use TFIDF (preferred) for weighting, and then classify using SVM according to the labels (malicious smart contract or non-malicious smart contract) of the smart contracts to which these variable names belong recorded in the vocabulary, and calculate the distance from each variable name word vector to the hyperplane;

[0047] (33) Map the distances from all variable names whose categories belong to malicious smart contracts to the hyperplane to the range of [0, 1]. The values corresponding to each variable name are represented by the vector D' = {d1', d2',...}. Among them, if the word vector is on the non-malicious smart contract side, its corresponding value is 0. Adding 1 to each value in this vector can obtain the malicious degree vector E = {e1, e2,...}, and use it as the weight to guide the process of gradient descent attack, so as to increase the probability of selecting variable names with similar parts of speech during the iteration process.

[0048] Specifically, Algorithm 1.1 can be used to generate a malicious weight vector, and the optimal hyperplane a T ​x+b=0 is obtained by the support vector machine (SVM); by finding a dividing hyperplane in the sample space, samples of different categories are separated, and at the same time, the minimum distance between the two point sets and this plane is maximized. Such an optimal hyperplane maximizes the distance between the edge points of the two point sets and this plane; the word vector for each variable name in the vocabulary library is obtained in the previous article, and the optimal hyperplane that divides the two vector sets from malicious smart contracts and non-malicious smart contracts is obtained. After the optimal hyperplane is obtained, the distance from each variable name in the vocabulary library to the optimal hyperplane can be calculated, and the distance from all variable names in the malicious smart contract half area to the hyperplane is calculated. When mapped to the range of [0,1], a new vector corresponding to the variable name in the vocabulary library is the required maliciousness weight vector E; Algorithm 1.1 is as follows:

[0049]

[0050] (4) Use the site selection vector and variable name selection vector in step (21) and step (22) to represent the generation process of smart contract adversarial samples.

[0051] (5) While keeping the weight of the smart contract benevolence detection model unchanged, the probability vector is iteratively updated according to the gradient calculated by the current representation probability vector, and the malicious weight vector generated in step (3) is used to guide the gradient descent attack process to increase the probability of selecting a variable name with a smaller part of speech difference. The construction of the weight generation module and the guiding process are as follows: Figure 3 As shown;

[0052] (6) using a projected gradient descent algorithm based on a balancing factor to iteratively search for the values of the site selection vector and the variable name selection vector in step (21) and step (22);

[0053] (61) The trained parameters in the attacked detection model are regarded as fixed values, the output results are fixed as erroneous results that mislead the attacked model, and the loss function value of the input data is calculated;

[0054] (62) Use the balancing factor to introduce the mean and variance of the part-of-speech distance of the variable name into the optimization function, and then perform a projected gradient descent attack on the new optimization function;

[0055] (63) Define the concept of sparsity of adding perturbations to smart contract code, using a balancing factor to introduce sparsity, i.e. the overall perturbation density, into the optimization function;

[0056] (64) To prevent the situation of excessive disturbance addition sites locally, a mask is set, and z is processed with this mask during each alternating iteration, so that some sites in the smart contract cannot be used as disturbance addition sites, ensuring that there will be no excessive disturbance addition density in any local area;

[0057] Steps (63) and (64) are to introduce the average value and variance of the part-of-speech distance of variable names into the optimization function using the balance factor, and define the concept of sparsity of adding disturbances to the smart contract code. The balance factor is used to introduce sparsity, that is, the overall disturbance density, into the optimization function; to prevent the situation of excessive disturbance addition sites locally, a mask is set, and z is processed with this mask during each alternating iteration, so that some sites in the smart contract cannot be used as disturbance addition sites, ensuring that there will be no excessive disturbance addition density in any local area. The new optimization function is expressed as:

[0058]

[0059] where σ 2 represents the variance of the disturbance degree. The balance factor λ1 and the balance factor λ2 are artificially set constants. θ is the detection model to be attacked, and l attack represents the loss function of the attack. Generating a disturbance for the symbol at the i-th position using the j-th variable in the vocabulary can be expressed as [u i ′] j = 1 and z i = 1. D represents the sum of the cosine vector values of the word vectors before and after the disturbance at each disturbance site. P′ and P respectively represent the source code of the smart contract after and before adding the disturbance. 1 represents the mask set to prevent the situation of excessive disturbance addition sites, and k represents the maximum value of the number of allowable disturbance addition sites.

[0060] (65) Along the direction of gradient descent, continuously iterate the vectors z and u i . Because there are two vectors that need to be iterated, an alternating iteration method is adopted during the iteration. First, fix the vector u i , iterate the vector z, and then use the value of the new vector z as a fixed value to iterate the vector u i ;

[0061] That is, on this basis, perform a projected gradient descent attack on the new optimization function, along the direction of gradient descent, continuously iterate the vectors z and u i . Because there are two vectors that need to be iterated, an alternating iteration method is adopted during the iteration. First, fix the vector u i , iterate the vector z, and then use the value of the new vector z as a fixed value to iterate the vector u iIteration is performed, and the specific iteration process is shown in Algorithm 1.2:

[0062]

[0063] In Algorithm 1.2, the difference value between each group of old and new variable names introduced by the balance factor into the optimization function is obtained by calculating the cosine similarity of the word vectors corresponding to the old and new variable names obtained from the word2vec model. The cosine similarity between two variable names is used to evaluate the semantic gap between these two variable names. The cosine similarity between the original variable name and the new variable name represents the concealment of the added perturbation. When the cosine similarity is closer to 1, the concealment of the adversarial sample is better.

[0064] (7) After iterating to the pre-set number of iterations, the site selection method and variable name selection method represented by the vectors z and u i are used as the perturbation addition methods to generate adversarial samples of the original smart contract.

[0065] The construction of the site selection vector and variable name selection vector provided in this embodiment is as Figure 1 shown, the process of generating the maliciousness weight vector is as Figure 2 shown, and the overall structure of using the maliciousness weight vector to generate adversarial samples is as Figure 3 shown.

[0066] Embodiment 2

[0067] A computer-readable storage medium, the computer-readable storage medium includes a stored computer program, wherein when the computer program is run by a processor, it controls the device where the storage medium is located to execute the method for generating adversarial samples of a smart contract as described above.

[0068] The related technical solutions are the same as those in Embodiment 1 and will not be elaborated here.

[0069] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for generating adversarial samples of smart contracts, characterized in that, Including: Classify the word vectors corresponding to all variable names according to the labels of whether the smart contracts to which all variable names in the vocabulary belong are bona fide or not, obtain a hyperplane that divides the word vectors into two categories, namely, those from malicious smart contracts and those from non-malicious smart contracts, and calculate the distances from the word vectors of each variable name to the hyperplane; map the distances from the variable names whose categories belong to malicious smart contracts to the hyperplane to the range of [0, 1], and represent them with the vector D′ = {d′1, d′2,...}, where, if the word vector of a certain variable name whose category belongs to a malicious smart contract is on the non-malicious smart contract side, the corresponding distance value in the vector D′ is 0, and add 1 to each value in the vector D′ to obtain the maliciousness weight vector; Keep the weights of the attacked detection model unchanged, and fix the output result to a wrong result that misleads the attacked detection model. Use the maliciousness weight vector to guide the iterative process of the gradient descent attack optimization function, so that the probability of selecting a variable name with a similar part of speech to the variable name to be replaced from the vocabulary is increased during the iteration, and the generation of adversarial samples for smart contracts is completed; Wherein, during the iteration process, the optimization function used is: a new optimization function obtained by introducing the average value and variance of the part-of-speech gaps between the old and new variable names when multiple variable names are modified into the original optimization function of the attacked detection model by using a balance factor; During the iteration process, the optimization function used is: a new optimization function obtained by introducing the sparsity of the perturbation added when multiple variable names are modified into the original optimization function of the attacked detection model by using a mask; During the iteration process, the optimization function used is: Where D represents the sum of the cosine vector values of the word vectors before and after perturbation at each perturbation site, P' and P represent the smart contract source code after and before adding perturbation respectively, 1 represents a mask set to prevent too many perturbation addition sites, k represents the maximum number of allowed perturbation addition sites, Ω represents the total number of variable names in the vocabulary, n represents the total number of sites to be perturbed, σ 2 represents the variance of the perturbation degree, the balance factor λ1 and the balance factor λ2 are constants set artificially, θ is the detection model to be attacked, l attack represents the loss function of the attack. Generating a perturbation for the symbol at the i-th position using the j-th variable in the vocabulary is represented as [u' i j = 1 and z i = 1.​ 2. The method for generating adversarial samples of smart contracts according to claim 1, wherein, The construction method of the vocabulary is: Preprocess multiple smart contract source codes, extract all variable names in the multiple smart contract source codes to form a vocabulary, including variable names and their corresponding labels of whether the smart contracts to which they belong are bona fide or not; wherein, the multiple smart contract source codes include malicious and non-malicious smart contract source codes.

3. The method for generating adversarial samples of smart contracts according to claim 1, wherein Use the cosine similarity between the word vectors corresponding to the variable name to be replaced and the new variable name to evaluate the semantic gap between the two variable names. The closer the cosine similarity is to 1, the more similar the parts of speech are, and the better the concealment of the generated adversarial sample is.

4. The method for generating an intelligent contract adversarial sample according to claim 1, wherein In the iterative process, an alternating iteration method is adopted. First, the vector u i is fixed, and the vector z is iterated. Then, the value of the new vector z is regarded as a fixed value to iterate the vector u i . Among them, the vector z ∈ {0, 1} n marks whether each site is selected to add perturbations. When z i = 1, it means that the P i position is selected to add perturbations, and n represents the total number of sites to be perturbed; the one-hot vector u i = {0, 1} |Ω| is used to represent the variable name to be selected for adding perturbations. Each position where z i = 1 corresponds to a vector u i,Ω indicating the total number of variable names in the vocabulary.

5. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein, when the computer program is run by a processor, it controls the device where the storage medium is located to execute a method for generating adversarial samples for smart contracts according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Feature extraction method, model generation method and malicious code detection method

    CN109308413A

  • Secure messaging in a machine learning blockchain network

    US11081219B1