Intelligent contract vulnerability detection method based on adaptive data partitioning and GPT-4o fusion preprocessing technology
By using adaptive data chunking and GPT-4o methods in smart contract vulnerability detection, data augmentation and segmentation of the smart contract source code, combined with the CodeBERT model for feature extraction and classification, the existing detection methods are solved, and efficient and accurate vulnerability detection is achieved.
Patent Information
- Application Number
- CN202510009689.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
AI Technical Summary
The existing smart contract vulnerability detection methods have low detection efficiency and poor accuracy, and there are problems such as high environmental configuration time cost, low detection efficiency, and uninterpretation of deep learning and semantic breakage.
The smart contract vulnerability detection method based on adaptive data chunking and GPT-4o is adopted, and the smart contract source code is enhanced and token sequence mapping is carried out through a large model. Combined with data segmentation methods of adaptive length and semantics, multiple parallel-set CodeBERT models are used for feature extraction and text classification to detect the probability of vulnerabilities.
It improves the accuracy and speed of smart contract vulnerability detection, and can efficiently perform feature extraction and detection, especially to provide better solutions for the problem of excessive data, improving the efficiency and accuracy of detection.
Smart Images

Figure CN119939599A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of blockchain smart contract security vulnerability analysis, and in particular to a smart contract vulnerability detection method based on adaptive data segmentation and GPT-4o. Background Art
[0002] With the booming development of blockchain technology, smart contracts, as one of its core technologies, are gradually penetrating into all walks of life and becoming a key force in promoting digital transformation. Smart contracts define transaction rules in the form of code and can be automatically executed without the need for third-party trust, greatly improving the transparency and efficiency of transactions. However, this innovation also brings new security challenges. Since smart contracts are difficult to modify once deployed, their potential vulnerabilities, if maliciously exploited, may cause huge economic losses and a crisis of trust.
[0003] Ethereum is one of the most well-known deployment platforms for smart contracts. The number of smart contracts running on it has exceeded millions, managing a huge amount of digital assets. However, the code quality of these smart contracts varies. Many Solidity contract codes written by developers are potentially risky due to the lack of sufficient security audits. Historically, events such as The DAO incident, the multi-signature vulnerability of the Parity wallet, and the integer overflow vulnerability of the BEC (Beauty Chain) contract have revealed the fragility of smart contract security and sounded the alarm for the industry. Therefore, vulnerability detection of smart contracts has become an important research topic in the field of blockchain security.
[0004] At present, the vulnerability detection methods based on Solidity code are developing rapidly, but the detection methods have the following problems in actual use: 1. The environment configurations required by the existing detection methods are different. For actual vulnerability detection, huge debugging time costs are required, resulting in too long project completion time; 2. The existing traditional methods exist independently as various tools. After the development work is completed, developers are required to manually input the contracts into various detection tools for testing. The detection efficiency is too low and the detection results are too subjective; 3. The existing deep learning methods are unexplainable and often ignore the detection difficulties caused by the length of the smart contract source code. The length segmentation uses a fixed-length segmentation mode, which is easy to cause semantic breaks.
[0005] The invention patent with application number 202410927303.1 discloses a multi-modal smart contract vulnerability detection method and system, which includes the following steps: Step 1: Obtain the source code of the smart contract to be detected; Step 2: Compile the source code to obtain the operation code and abstract syntax tree of the smart contract to be detected; Step 3: Obtain the word vectors of the three modal data of the source code, operation code and abstract syntax tree of the smart contract to be detected respectively; Step 4: Input the word vectors of the three modal data into the preset smart contract vulnerability detection model to obtain the vulnerability status of the smart contract to be detected. The above invention effectively improves the accuracy of the smart contract vulnerability detection method based on deep learning, ensures that the legal assets of Ethereum smart contract users are not infringed, and maintains the stability of the Ethereum blockchain network. However, the above patent has some smart contract vulnerabilities with too long data length, and is not converted into the form of operation code and abstract syntax tree. The feature extraction of the above patent for vulnerabilities relies too much on the three modalities of data, and requires source code, operation code and abstract syntax tree at the same time. Summary of the invention
[0006] In view of the technical problems of low detection efficiency and poor accuracy of existing smart contract vulnerability detection methods, the present invention proposes a smart contract vulnerability detection method based on adaptive data segmentation and GPT-4o to help users quickly and accurately detect vulnerabilities in smart contract codes.
[0007] In order to achieve the above object, the technical solution of the present invention is implemented as follows: a smart contract vulnerability detection method based on adaptive data segmentation and GPT-4o fusion preprocessing technology, the steps of which are as follows:
[0008] Step 1: Use the large model GPT-4o to perform data enhancement on the input data contract source code, and map the enhanced code into a Token sequence; use the adaptive length and semantic data segmentation method to process the length of the Token sequence to obtain multiple Token fragments;
[0009] Step 2: Use multiple CodeBERT models set up in parallel to build a mcCodeBERT network model, and use the mcCodeBERT network model to extract features from multiple Token fragments to obtain embedding vectors;
[0010] Step 3: Use text classification methods to classify the embedded vectors and obtain the probability of vulnerability.
[0011] Preferably, the three strategy forms of the large model GPT-4o include: (1) adding English comments to the smart contract source code; (2) adding smart contract code for normal execution functions to the smart contract source code to expand the source code data; (3) copying the smart contract source code to increase the data length.
[0012] Preferably, the large model GPT-4o builds a prompt structure through the CO-STAR framework, integrates the required context on the basis of the CO-STAR framework to complete the construction of the prompt structure, and then inputs it into the large model GPT-4o; the CO-STAR framework includes a Context module, an Objective module, a Style module, a Tone module, an Audience module and a Response module. The Context module explains that data enhancement is to be performed on the data set of the smart contract source code, and the Objective module informs that data enhancement is to be achieved by adding English comments to the source code and attaching the source code data to be processed in response. The Style module generates the corresponding style of a smart contract coder, and the Tone module explains that the tone of the response must be professional; the Audience module is set as the audience group to be researchers related to smart contract vulnerability detection, and the Response module requires that the response content must maintain the source code with annotations.
[0013] Preferably, the data-enhanced code is optimized and mapped to a Token sequence, and the optimization rules are: the line breaks in the data-enhanced source code are uniformly replaced with spaces; the complex closed brackets that may appear in the source code are uniformly replaced with spaces followed by right brackets; the complex closed brackets include double right brackets, semicolons followed by right brackets, and right brackets, to obtain preliminary optimized code.
[0014] Preferably, the preliminary optimized code is converted into a Token sequence in which each code character is represented by a Token value to obtain the contract code Code T .
[0015] Preferably, the method of processing the length of the Token sequence by using the adaptive length and semantic data segmentation method is: implementing the contract code Code by using the adaptive length and semantic slider segmentation method T The length of the source code is processed based on the token position. The source code is segmented by intelligently identifying the left bracket and the space followed by the right bracket as the semantic boundary.
[0016] Preferably, the method for adaptive length and semantics data segmentation performs adaptive semantics and length code segmentation as follows: input the tokenized contract code Code T , the output is the segmented Token fragment set Code IT , the steps are as follows:
[0017] Step 1: Initialize the string sequence set Code IT ; i = 1; position set The contract code T The first Token bit is marked as the initial position p0;
[0018] Step 2: Determine the remaining contract code T Is the length greater than 512? If so, continue executing; otherwise, jump to step 5.
[0019] Step 3: Traverse the contract code from the initial position p0 T , find the 512th bit of the Token and record its position p m and Token value ids pm ; Let the target Token be located at position p l =LEFT(Code T ,p m ,24303); if p l The value is in the position set P, let the initial position p0 = p l , repeat step 3; otherwise, proceed to step 4; LEFT() is a left-rounding positioning function;
[0020] Step 4: Move the initial position p0 to the target Token position p l The fragments between are represented as string sequences Code iT ; Let P = P∪{p l}; Code T =Code T -Code iT ; i = i + 1; p m =p l ; Initial position p0 = LEFT (Code T ,p m ,45152);
[0021] Step 5: Let i = i + 1; Code iT =CodeT;Code IT =Code IT ∪{Code iT};
[0022] Step 6: Output the Token fragment Code IT .
[0023] Preferably, the left-rounding positioning function LEFT (Code T ,p0,ids p ) is implemented as follows:
[0024] Step 11: Initialize the target Token location p l, specify the Token position input p, target Token value ids pl , let p = p0; p0 is the initial position;
[0025] Step 12: Calculation and Contract Code T The token value ids corresponding to the character at position p in p , when ids p The value is not equal to the target Token value ids pl When p--, repeat step 12; otherwise, jump out of the loop and execute step 13;
[0026] Step 13: Let p l =p, output the target Token location p l ;
[0027] and the contract code T The token value corresponding to the character at position p is ids p =VALUE T (Code T ,p); p is the position input of the specified Token.
[0028] Preferably, according to the number of Token fragments n, n independent CodeBERT model channels are deployed to form an mcCodeBERT network model. The CodeBERT model consists of 12 layers of Transformer encoders, each layer of Transformer encoder has 12 self-attention heads, each head size is 64, and the hidden layer dimension is 768; in each layer of Transformer encoder, the attention mechanism automatically adjusts the weight distribution between words through the correlation between words in the code to obtain the final representation of each word; the 12 self-attention heads divide the input sequence into multiple sets of independent queries, keys and values through a multi-head attention mechanism through linear transformation; each set of queries, keys and values is independently calculated through the attention mechanism, and the calculation results are spliced and integrated to generate the final output;
[0029] The output of each CodeBERT model is concatenated into the final embedding vector.
[0030] Preferably, the method of classifying the embedded vector using the text classification method is: after the embedded vector is input into the vulnerability detection module, in the vulnerability detection module, classification is performed through a linear layer, and the Softmax function is used as a classifier. After the embedded vector is input into the Softmax layer for normalization processing, the probability of the final existence of the vulnerability is obtained; the implementation steps are:
[0031] F. Each embedding vector containing data features is normalized and then input into the vulnerability detection module in turn;
[0032] G. If there is a vulnerability, output 1 as the identifier; if there is no vulnerability, output 0 as the identifier;
[0033] H. Repeat steps A and B until all embedded vectors EV are identified;
[0034] I. Perform mathematical operations on all 0 and 1 identifiers after detection, and calculate the detection index based on the label and predicted value;
[0035] J. Calculate the detection index based on the true positive TP, true negative TF, false positive FP and false negative FN to determine the vulnerability rate.
[0036] Compared with the prior art, the present invention has the following beneficial effects: in the preprocessing process, the semantics of the segmented smart contract code fragments are enhanced to make up for the insufficient coverage of the smart contract code set in the vulnerability features; before the smart contract code is written and deployed on the chain, the smart contract code is automatically detected for vulnerabilities, and the preprocessed smart contract code fragments are input into the vulnerability detection model for detection, thereby performing security analysis on the smart contract code, and improving the accuracy and speed of code vulnerability detection. The present invention only needs the smart contract source code to perform efficient feature extraction and detection, and has a better solution to the problem of too long data to help feature extraction, which can more efficiently help detect vulnerabilities with too long data length. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0038] Figure 1 It is a process flow chart of the present invention.
[0039] Figure 2 The figure is a flow chart of data preprocessing of the present invention.
[0040] Figure 3 A schematic diagram of a code example for performing data enhancement in the present invention.
[0041] Figure 4 The present invention discloses a processing flow of the method for adaptive length and semantics data segmentation.
[0042] Figure 5 Schematic diagram of model training of the present invention.
[0043] Figure 6 This is the processing flow of vulnerability detection of the present invention. DETAILED DESCRIPTION
[0044] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0045] like Figure 1 As shown in the figure, a smart contract vulnerability detection method based on adaptive data segmentation and GPT-4o is shown, and the steps are as follows:
[0046] Step 1: Data preprocessing: Use the large model GPT-4o to perform data enhancement on the input data contract source code, and map the enhanced code into a Token sequence; use the adaptive length and semantic data segmentation method to process the length of the Token sequence to obtain multiple Token fragments.
[0047] Data preprocessing is responsible for preliminary processing of the user's contract data to improve the accuracy and efficiency of subsequent testing. Figure 2 As shown in the figure, the data preprocessing module includes adaptive data segmentation and GPT-4o fusion preprocessing technology to perform initial processing on the contracts uploaded by users to improve the efficiency and accuracy of subsequent detection. Users can input the data contract source code into the large model GPT-4o for data enhancement in three strategies. The process of uploading the contract source code is very simple. The calling interface of the large model GPT-4o has been built. Just click the upload button and select the local data contract source code file. The three strategies of the large model GPT-4o include: (1) adding English comments to the smart contract source code; (2) adding the smart contract code of the normal execution function to the smart contract source code to expand the source code data; (3) copying the smart contract source code to increase the data length. Through the three strategies, the sample and capacity of the contract training process are expanded to enhance the robustness of model detection and expand the coverage of the smart contract data set.
[0048] Building the prompt structure through the CO-STAR framework can obtain better generation effects for the large language model GPT-4o. Specifically, the required context is integrated into the CO-STAR framework to complete the construction of the prompt structure, and then input into the large model GPT-4o to complete the set response work, ensuring that data enhancement is completed through three strategy forms. The CO-STAR framework mainly consists of six setting modules: Context, Objective, Style, Tone, Audience, and Response. The six modules have standardized input requirements and cooperate with each other, which can comprehensively consider the effectiveness and relevance of all aspects of the GPT-4o response.
[0049] Figure 3 The example of the first strategy "adding English comments to the smart contract source code" is shown. The prompt structure constructed by the CO-STAR framework is input to the large model GPT-4o. The prompt structure constructed here has 6 modules of the CO-STAR framework. In the Context module, it is explained that data enhancement should be done on the dataset of the smart contract source code to indicate the goal for the large model; in the Objective module, it is clearly stated that data enhancement needs to be achieved by adding English comments to the source code, and the source code data to be processed for the response is attached, indicating the operation task object for the large model; in the Style module, it is emphasized that the corresponding generation should be done in the style of a smart contract coder, which can make the contract code more professional; at the same time, in the Tone module, it is also explained that the tone of the response should be professional; in the Audience module, the audience is set to be researchers related to smart contract vulnerability detection; finally, in the Response module, it is required that the response content must maintain the source code with comments, which is easy to understand and can run normally. The above processing methods can make the generated new smart contracts more professional and cover more features. Figure 2 The right side shows that the new smart contract code generated according to the above framework contains semantic comments that are easy to understand, and the contract context logic is more concise and clear.
[0050] In addition, the present invention also adopts a series of rules to optimize the code structure and improve the subsequent code segmentation effect. In the optimized contract code structure, the line breaks in the source code are uniformly replaced with spaces. At the same time, the complex closed brackets that may appear in the source code, such as double right brackets ("}}"), semicolons followed by right brackets (";}"), and right brackets ("}"), are uniformly replaced with spaces followed by right brackets ("□}", where □ represents a space). These measures provide segmentation reference points for subsequent semantic-based code snippets, ensuring that each segmentation point can accurately cut out a complete code logic unit for subsequent compilation and testing. After obtaining the preliminary optimized code, the processed code is accurately converted into a Token sequence in which each code character is represented by a Token value, that is, the contract code Code T .
[0051] Subsequently, the contract source code is segmented by the designed adaptive length and semantic data segmentation method (ALSDC), which is mainly implemented by the adaptive length and semantic slider segmentation method. T Length processing of data source code. Based on the special token position, the efficient segmentation of source code is achieved by intelligently identifying "{" and "□}" as semantic boundaries. Figure 4 Demonstrates the segmentation process through ALSDC.
[0052] First, the relevant functions are given. Formula (1) is the calculation function of the unique identifier of p (the token value of p), where p is the position input of the specified token. VALUE T Function used to return the contract code T The token value corresponding to the character at position p in the final ids p express.
[0053] ids p =VALUE T (Code T ,p) (1)
[0054] In the slicing process, it is necessary to trace back to the left from a specific position to locate the location of the target token, so a left-rounding positioning method is given, named LEFT (Code T ,p0,ids p ), the input of this algorithm is the tokenized contract code Code T , initial position p0, target Token value ids pl , the output is the location p of the target Token l . LEFT(Code T ,p0,ids p)The steps are as follows:
[0055] Step 1: Initialize pl, p, ids pl , let p=p0.
[0056] Step 2: Calculate ids p =VALUE T (Code T ,p); when ids p The value is not equal to ids pl When p--, repeat step 2; otherwise, jump out of the loop and execute step 3;
[0057] Step 3: Let p l =p, output p l ,Finish.
[0058] Based on Algorithm 1, Algorithm 2 gives the process of obtaining code slices with adaptive semantics and length. The algorithm inputs the tokenized contract code Code T , the output is the segmented Token fragment set Code IT The algorithm steps are as follows:
[0059] Step 1: Initialize the string sequence set Code IT ; i = 1; position set The contract code T The first Token bit is marked as p0;
[0060] Step 2: Determine the remaining contract code T Is the length greater than 512? If so, continue executing; otherwise, jump to step 5.
[0061] Step 3: Traverse the contract code from position p0 T , find the 512th bit of the Token and record its position p m and Token value ids pm ; Let p l =LEFT(Code T ,p m ,24303); if p l The value is in the position set P, let the initial position p0 = p l , repeat step 3; otherwise, proceed to step 4; Step 4: Move the initial position p0 to p l The fragments between are represented as string sequences Code iT ; Let P = P∪{p l}; Code T =Code T -Code iT; i = i + 1; p m =p l ; Initial position p0 = LEFT (Code T ,p m ,45152);
[0062] Step 5: Let i = i + 1; Code iT =CodeT;Code IT =Code IT ∪{Code iT};
[0063] Step 6: Output Code IT , the algorithm ends.
[0064] In the segmentation process of ALSDC, the ids corresponding to the 512th token (ids are numbers mapped to different characters in the CodeBERT dictionary) are first accurately located and used as the initial segmentation point. Then, from this position, backtrack to the left to accurately find the first position with ids of 24303 (the ids of "□}" corresponds to 24303) to achieve the first effective segmentation. This step ensures the complete extraction of key semantic fragments while avoiding unnecessary data truncation. Next, the length of the remaining code fragments is evaluated. If it has naturally shrunk to less than 512 bits, it is directly regarded as processed, which effectively reduces the computational overhead. Otherwise, continue to search from the left side of 24303 until a new segmentation point with ids of 45152 is found (the ids of "{" corresponds to 45152), and repeat the above steps until all code fragments meet the input requirements.
[0065] The adaptive length and semantic data segmentation (ALSDC) method fundamentally solves the problem of ignoring important semantic information due to data truncation, and opens up a new path for the security analysis of smart contracts.
[0066] Step 2: Use multiple CodeBERT models set up in parallel to build a mcCodeBERT network model, and use the mcCodeBERT network model to extract features of multiple Token fragments to obtain embedding vectors.
[0067] The model training module is a core component in smart contract vulnerability detection, and is mainly used to conduct comprehensive training and learning on the segmented contract source code. After a refined upstream preprocessing phase, the data contract source code is successfully decomposed into i carefully selected token fragments, which not only meet the input specifications of CodeBERT, but also retain the relatively coherent semantic information between the codes. Figure 5As shown in Figure 1, in order to cope with the challenge of very long smart contracts, a parallel strategy is adopted in the downstream network architecture, and i independent CodeBERT model channels are deployed to form mcCodeBERT, ensuring the integrity and completeness of the training and detection process. T To mark special delimiters to facilitate further data processing, [CLS] is usually used as a special marker at the beginning of the input sequence, and [SEP] is usually used as the end of a single input sequence, or to divide natural language and code segments.
[0068] CodeBERT uses the same network structure as RoBERTa, which is composed of 12 layers of Transformer encoder parts. Each layer of Transformer encoder has 12 self-attention heads, each head size is 64, the hidden layer dimension is unified to 768, and the total number of model parameters is 125M.
[0069] In each layer of Transformer encoder, the attention mechanism automatically adjusts the weight distribution between words according to the correlation between them in the code, so as to obtain the final representation of each word. After weight distribution, the model's ability to understand sequence data can be significantly improved. The so-called multi-head attention mechanism in the encoder is to divide the input sequence into multiple independent sets of queries (Q), keys (K), and values (V) through linear transformation; then, each set of Q, K, and V is calculated independently by the attention mechanism, and finally the calculation results are spliced and integrated to generate the final output. The weight distribution is realized through the self-attention mechanism, which mainly includes calculating the dot product of the query Query and the key Key to obtain the relevance score, and then weighted summing the value Value through scaling and Softmax to generate the final contextual representation.
[0070] The data processing of a single-layer Transformer encoder is completed. After processing by a 12-layer Transformer encoder, the data feature extraction in a single CodeBERT model can be completed. The i Token sequences generated by the upstream module enter the i CodeBERT models for parallel training. The data processed by the upstream is input into the corresponding CodeBERT model for learning and training. Finally, the output of each CodeBERT is spliced into the final Embedding Vector (EV). At this point, the data training and feature extraction of the mcCodeBERT network model are completed. Figure 5V1~Vn in the figure represent the data features in Embedding form output after learning and training of the mcCodeBERT model. Training is to input the training set data into the network model for training.
[0071] Step 3: Use text classification methods to classify the embedded vectors and obtain the probability of vulnerability.
[0072] The vulnerability detection module is the final module for detecting smart contracts. After the contract data is extracted by the upstream mcCodeBERT network model, an Embedding Vector containing data feature representation is generated, which is then input into the vulnerability detection module to obtain the probability of the final vulnerability. Figure 6 As shown, in the vulnerability detection module, the data contract processed by the model training module will be input into the module in the form of Embedding Vector for subsequent detection and analysis.
[0073] The Embedding Vector contains rich data feature representation after training and learning. In the vulnerability detection module, classification is performed through the linear layer, and the Softmax function is used as the classifier. After the Embedding Vector (EV) is input into the Softmax layer for normalization, the probability of the final vulnerability will be obtained. The steps are as follows:
[0074] K. Each embedding vector EV containing data features will be input into the vulnerability detection module in turn after normalization.
[0075] L. If there is a vulnerability, 1 will be output as an identifier; if there is no vulnerability, 0 will be output as an identifier. The model will test the validation set based on the data features learned during training. According to the learned features, the data in the validation set will be divided into features. If it is determined that there is a vulnerability, 1 will be output. If it is determined that it does not meet the vulnerability characteristics, 0 will be output. Finally, the integration calculation will be performed to obtain the final result.
[0076] M. Repeat the above steps until all embedded vectors EV are identified.
[0077] N. Perform mathematical operations on all 0 and 1 identifiers after detection. Make judgments based on the validation set labels and model predictions, and then calculate the detection indicators.
[0078] O. Determine the vulnerability rate and calculate it based on four detection indicators: TP (true positive), TF (true negative), FP (false positive) and FN (false negative).
[0079] During the final testing process, the vulnerability detection module will eventually generate a test report and store the test results in the file server for subsequent viewing of the test results.
[0080] Specific examples:
[0081] (1) Dataset: We used two self-built datasets, SBHD and ESBHD, which contain multiple vulnerabilities, for experiments. We integrated parts of the SmartWild dataset and the RSC dataset into the SBHD dataset. These data sources provide a rich basis for smart contract samples. In order to further enhance the randomness and universality of the dataset, we crawled some real smart contract source codes from the Ethereum official website. These crawled original contract codes were then preliminarily screened by automated tools such as Oyente and Smartcheck, and supplemented by manual fine annotation to ensure the accuracy and diversity of the contract data. This part was then merged into the SBHD dataset and randomly mixed to form the ESBHD dataset. Reentrancy and Timestamp Dependency vulnerabilities, and Integer Overflow vulnerabilities exist in each dataset.
[0082] (2) Model benchmark: In order to more intuitively demonstrate the effectiveness of the present invention, a total of 6 models were learned and trained in this specific example. They are divided into two categories, deep learning model comparison and ablation model comparison. The deep learning models mainly include 3 currently mainstream models (BERT, LSTM and GRU), and the ablation models include 2 baseline models of the present invention (CodeBERT and GCodeBERT). Specifically as follows: BERT is a pre-trained language representation model based on the Transformer architecture. LSTM is a special recurrent neural network structure. GRU is another recurrent neural network variant that can process sequence data. CodeBERT is the baseline model, and GCodeBERT is the network model without the ALSDC module and is named GCodeBERT.
[0083] (3) Parameter setting: The learning rate of the model of the present invention is set to 0.00005, the dropout is set to 0.4, the model has undergone 80 iterations, and the optimizer is AdamW. The operating system in the computer is Windows 11, and the vulnerability detection model uses Python3 and Pytorch2 framework.
[0084] (4) Evaluation indicators: The four most common evaluation indicators are used to reasonably evaluate the model checking effect of the present invention. These four indicators are: accuracy (Acc), precision (Pre), recall (Rec) and F1. These four indicators are calculated from true positive examples (TP), true negative examples (FP), false positive examples (TN) and false negative examples (FN).
[0085] (5) Experimental results:
[0086] Comparison of deep learning models: Under the operating system Windows 11, the model Python3 and Pytorch2 framework, the model learning rate was set to 0.00005, the dropout was set to 0.4, the model went through 80 iterations, the optimizer was AdamW, and then experiments were conducted on the self-built SBHD and ESBHD data sets containing multiple vulnerabilities, and the detection results in Table 1 were visualized, which can draw conclusions more intuitively. Compared with other advanced deep learning models, the method of the present invention has better detection results. The method of the present invention shows excellent performance in these three vulnerability detections, which are significantly better than the comparison models. Among them, the F1 index of the two vulnerabilities of Reentrancy and Timestamp Dependency exceeded 95%, which is at least 3% higher than other advanced models. The F1 index of the Integer Overflow vulnerability also reached 83.44%, which is 7.11% higher than other advanced models.
[0087] Ablation model comparison: Under the operating system Windows 11, the model Python3 and Pytorch2 framework, the model learning rate was set to 0.00005, the dropout was set to 0.4, the model went through 80 iterations, and the optimizer was AdamW. Then, experiments were conducted on two self-built datasets containing multiple vulnerabilities, SBHD and ESBHD, and the following results were obtained: Figure 2 The visualization of the detection results shows that the baseline model has F1 detection results of 92.61%, 92.90% and 80.00% for the three vulnerabilities. After adding the large model GPT-4o data enhancement module, the F1 value of the GCodeBERT model increased by 1.5%, 0.97% and 1.76% respectively. After adding the ALSDC module, the F1 value of vulnerability detection in this example method is further improved. The addition of the data enhancement module and the ALSDC module has improved the detection performance of this example method to varying degrees.
[0088] Table 1 Data comparison between the present invention and existing deep learning models
[0089]
[0090] Table 2 Data comparison between the present invention and the ablation model
[0091]
[0092] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A smart contract vulnerability detection method based on adaptive data segmentation and GPT-4o fusion preprocessing technology, characterized in that: The steps are as follows: Step 1: Use the large model GPT-4o to perform data enhancement on the input data contract source code, and map the enhanced code into a Token sequence; use the adaptive length and semantic data segmentation method to process the length of the Token sequence to obtain multiple Token fragments; Step 2: Use multiple CodeBERT models set up in parallel to build a mcCodeBERT network model, and use the mcCodeBERT network model to extract features from multiple Token fragments to obtain embedding vectors; Step 3: Use text classification methods to classify the embedded vectors and obtain the probability of the existence of vulnerabilities.
2. According to claim 1, the smart contract vulnerability detection method based on adaptive data segmentation and GPT-4o fusion preprocessing technology is characterized in that: The three strategies of the large model GPT-4o include: (1) adding English comments to the smart contract source code; (2) adding smart contract code for normal execution functions to the smart contract source code to expand the source code data; (3) copying the smart contract source code to increase the data length.
3. According to claim 2, the smart contract vulnerability detection method based on adaptive data segmentation and GPT-4o fusion preprocessing technology is characterized in that: The large model GPT-4o builds a prompt structure through the CO-STAR framework, integrates the required context on the basis of the CO-STAR framework to complete the construction of the prompt structure and then inputs it into the large model GPT-4o; the CO-STAR framework includes a Context module, an Objective module, a Style module, a Tone module, an Audience module and a Response module. The Context module explains that data enhancement is to be performed on the data set of the smart contract source code, and the Objective module informs that data enhancement is to be achieved by adding English comments to the source code and attaching the source code data to be processed in response. The Style module generates the corresponding style of a smart contract coder, and the Tone module explains that the tone of the response must be professional; the Audience module is set as the audience group to be researchers related to smart contract vulnerability detection, and the Response module requires that the response content must maintain the source code with annotations.
4. The smart contract vulnerability detection method based on adaptive data segmentation and GPT-4o fusion preprocessing technology according to claim 2 or 3 is characterized in that: The data-enhanced code is optimized and mapped to a Token sequence. The optimization rules are as follows: the line breaks in the data-enhanced source code are uniformly replaced with spaces; the complex closed brackets that may appear in the source code are uniformly replaced with spaces followed by right brackets; the complex closed brackets include double right brackets, semicolons followed by right brackets, and right brackets, to obtain preliminary optimized code.
5. According to claim 4, the smart contract vulnerability detection method based on adaptive data segmentation and GPT-4o fusion preprocessing technology is characterized in that: Convert the preliminary optimized code into a Token sequence where each code character is represented by a Token value to obtain the contract code Code T .
6. The smart contract vulnerability detection method based on adaptive data segmentation and GPT-4o fusion preprocessing technology according to claim 1 or 5 is characterized in that: The method of processing the length of the Token sequence by using the adaptive length and semantic data segmentation method is: implementing the contract code Code by using the adaptive length and semantic slider segmentation method T The length of the source code is processed based on the token position. The source code is segmented by intelligently identifying the left bracket and the space followed by the right bracket as the semantic boundary.
7. The smart contract vulnerability detection method based on adaptive data segmentation and GPT-4o fusion preprocessing technology according to claim 6 is characterized in that: The method for adaptive length and semantics data segmentation performs adaptive semantics and length code segmentation as follows: Input the tokenized contract code Code T , the output is the segmented Token fragment set Code IT , the steps are as follows: Step 1: Initialize the string sequence set Code IT ; i = 1; position set The contract code T The first Token bit is marked as the initial position p0; Step 2: Determine the remaining contract code T Is the length greater than 512? If so, continue executing; otherwise, jump to step 5. Step 3: Traverse the contract code from the initial position p0 T , find the 512th bit of the Token and record its position p m and Token value ids pm ; Let the target Token be located at position p l =LEFT(Code T ,p m ,24303); if p l The value is in the position set P, let the initial position p0 = p l , repeat step 3; otherwise , proceed to step 4; LEFT() is a left-rounding positioning function; Step 4: Move the initial position p0 to the target Token position p l The fragments between are represented as string sequences Code iT ; Let P = P∪{p l }; Code T =Code T -Code iT ; i = i + 1; p m =p l ; Initial position p0 = LEFT (Code T ,p m ,45152); Step 5: Let i = i + 1; Code iT = CodeT; Code IT = Code IT ∪{Code iT}; Step 6: Output the Token fragment Code IT .
8. The smart contract vulnerability detection method based on adaptive data segmentation and GPT-4o fusion preprocessing technology according to claim 7 is characterized in that: The left rounding positioning function LEFT (Code T ,p0,ids p ) is implemented as follows: Step 11: Initialize the target Token location p l , specify the Token position input p, target Token value ids pl , let p = p0; p0 is the initial position; Step 12: Calculation and Contract Code T The token value ids corresponding to the character at position p in p , when ids p The value is not equal to the target Token value ids pl When p--, repeat step 12; otherwise, jump out of the loop and execute step 13; Step 13: Let p l =p, output the target Token location p l ; and the contract code T The token value corresponding to the character at position p is ids p =VALUE T (Code T ,p); p is the position input of the specified Token.
9. The smart contract vulnerability detection method based on adaptive data segmentation and GPT-4o fusion preprocessing technology according to claim 6 or 7 is characterized in that: According to the number of token fragments n, n independent CodeBERT model channels are deployed to form the mcCodeBERT network model. The CodeBERT model consists of 12 layers of Transformer encoders. Each layer of Transformer encoder has 12 self-attention heads, each head size is 64, and the hidden dimension is 768. In each layer of Transformer encoder, the attention mechanism automatically adjusts the weight distribution between words according to the correlation between words in the code to obtain the final representation of each word. The 12 self-attention heads divide the input sequence into multiple sets of independent queries, keys and values through linear transformation through the multi-head attention mechanism. Each set of queries, keys and values is independently calculated through the attention mechanism, and the calculation results are spliced and integrated to generate the final output. The output of each CodeBERT model is concatenated into the final embedding vector.
10. The smart contract vulnerability detection method based on adaptive data segmentation and GPT-4o fusion preprocessing technology according to claim 9 is characterized in that: The method of classifying the embedded vector using the text classification method is as follows: after the embedded vector is input into the vulnerability detection module, in the vulnerability detection module, classification is performed through a linear layer, and the Softmax function is used as a classifier. After the embedded vector is input into the Softmax layer for normalization processing, the probability of the final vulnerability is obtained; the implementation steps are as follows: A. Each embedding vector containing data features is normalized and then input into the vulnerability detection module in turn; B. If there is a vulnerability, output 1 as the identifier; if there is no vulnerability, output 0 as the identifier; C. Repeat steps A and B until all embedded vectors EV are identified; D. Perform mathematical operations on all 0 and 1 identifiers after detection, and calculate the detection index based on the label and predicted value; E. Calculate the detection index based on the true positive TP, true negative TF, false positive FP and false negative FN to determine the vulnerability rate.
Citation Information
Patent Citations
Intelligent contract vulnerability detection method and system based on multiple modes
CN118940272A
Cited By
Intelligent contract vulnerability detection method based on hierarchical multi-granularity coding
CN121706106A