Software vulnerability assessment optimization method based on code simplification
By introducing code simplified steps in software vulnerability evaluation, the problem of high computing resource requirements for fine-tuning pre-trained models is solved, and more efficient vulnerability evaluation is achieved, suitable for resource-constrained scenarios.
Patent Information
- Application Number
- CN202510091992.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-27
AI Technical Summary
In existing software vulnerability assessment methods, fine-tuning pre-trained models requires strong hardware support, resulting in high computing resource requirements, affecting the accuracy and efficiency of assessment.
Code simplification is introduced as a preprocessing step, through code segmentation and attention weight acquisition, code simplification is performed based on greedy algorithms, and the simplified code is input into the Transformer model for fine-tuning for comparison and evaluation.
It significantly reduces the training and inference time overhead of the model, improves the computing efficiency, and can quickly complete vulnerability assessment in resource-constrained scenarios, optimizing the applicability and scalability of the model.
Smart Images

Figure CN120046157A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of software engineering, and particularly to the application of code simplification in the field of vulnerability assessment. Background Art
[0002] Software vulnerabilities refer to design flaws, implementation errors, or configuration errors existing in software systems. Attackers can exploit these flaws to obtain unauthorized access, steal data, or disrupt services. With the continuous progress of digital and network technologies, the number and complexity of software vulnerabilities are increasing continuously, posing a major threat to the security of information systems. In recent years, a large number of serious security incidents have shown that vulnerability exploitation may lead to catastrophic consequences, including sensitive data leakage, financial losses, and paralysis of critical infrastructure. To address these threats, timely discovery and repair of vulnerabilities have become a key priority in maintaining network security. To cope with these challenges, software vulnerability assessment has emerged as a systematic approach to identify and quantify potential security risks in software. Software vulnerability assessment not only focuses on detecting vulnerabilities but also provides scientific support for determining the priority of vulnerability repair.
[0003] Currently, machine learning techniques and artificial intelligence have been widely applied in software vulnerability assessment, especially for vulnerability assessment by fine-tuning pre-trained models, which greatly improves the accuracy and efficiency of software vulnerability assessment. Although existing methods have achieved certain results, fine-tuning usually requires powerful hardware support, such as GPUs or TPUs, to optimize the parameters of large models like BERT, CodeBERT, or CodeT5. These models usually contain hundreds of millions or even billions of parameters and require huge memory and computing power during training and inference. Therefore, the high computational resource requirements of pre-trained models using the fine-tuning paradigm pose a challenge in software vulnerability assessment.
[0004] It can be seen that considering reducing the high computational requirements of fine-tuning is also a technical problem to be solved that affects the accuracy and efficiency of software vulnerability assessment. Summary of the Invention
[0005] The purpose of the present invention is to provide an optimized method for software vulnerability assessment based on code simplification. By introducing code simplification as a preprocessing step in software vulnerability assessment, it can not only alleviate the computational resource bottleneck but also further optimize the effect of existing software vulnerability assessment methods, providing a new idea for the analysis of complex vulnerability scenarios.
[0006] The inventive concept of the present invention is as follows: The present invention comprises four stages: In the data preprocessing stage, for the source code in the dataset, code splitting is performed according to the code splitting criteria, and each piece of code retains complete semantic information. This meets the requirements of the vulnerability assessment scenario based on code simplification. In the attention weight acquisition stage, the source code is used as the input of the Transformer model, and the attention weights of the tokens and statements are extracted by the preprocessed Transformer model. This helps with the implementation of the subsequent code simplification strategy. In the code simplification stage, based on the obtained token and statement weights and the greedy algorithm, a series of segmented code snippets are simplified, and the simplified code is obtained. In the model fine-tuning stage, the simplified code is input into the Transformer model for fine-tuning, and the test results are compared and evaluated with the test results obtained by inputting the source code into the pre-trained model.
[0007] The present invention is implemented by the following measures: An optimized method for software vulnerability assessment based on code simplification, which includes the following steps:
[0008] (1) Use the MegaVul dataset as the basis for the experiment, focusing on vulnerable functions in C / C++. Since there are vulnerability scores that do not conform to the CVSS 3.0 standard and redundant code in the dataset, preprocessing operations are taken to delete low-quality data. Finally, 14,235 vulnerable functions are selected as the experimental objects. Before performing code simplification, the code needs to be split into a series of code snippets. The specific preprocessing operations include the following steps:
[0009] (1-1) First, according to the CVSS 3.0 vulnerability scoring standard, delete the data that does not conform to this standard;
[0010] (1-2) To ensure the prediction accuracy of the model, delete the code comments to avoid them being misinterpreted as executable code;
[0011] (1-3) Given the length limitation of the pre-trained language model, we delete many redundant elements, such as blank lines and spaces, to enhance the attention to relevant code details;
[0012] (1-4) Through the statistical analysis of the vulnerability locations in the dataset, it is found that vulnerabilities usually appear in control statements and preprocessing statements. Therefore, a code splitting criterion for vulnerable functions is formulated to perform program slicing of the source code;
[0013] (2) Code has many levels of granularity, such as tokens, statements, and functions. First, study the atomic unit of the target code, that is, tokens. Next, study the statement-level knowledge learned by the Transformer model, which contains basic structures and semantic units. Finally, explore the function-level knowledge learned by the Transformer model through downstream tasks. The specific steps are as follows:
[0014] (2-1) Since the core of the Transformer model is the self-attention network, where the hidden state of each token is calculated layer by layer according to the self-attention weights, the attention weights in the Transformer layer of the Transformer model are used to measure the importance of each token after pre-training;
[0015] (2-2) After obtaining the token weights, the attention weights of each statement can be calculated. The weighted average attention weights of all tokens in a statement are summed to obtain the attention weight a(S) of the statement S:
[0016]
[0017] where a(t) is the attention weight of the token in the statement; w(t) is the normalized attention weight of the token t in the entire dataset. The normalization process is implemented by the Softmax function, that is:
[0018] w(t) = Softmax t∈S (a(t))
[0019] (2-3) Different from tokens, each statement is basically unique in the dataset. It is unrealistic to write all statements into the attention weight dictionary. Therefore, all statements are classified according to the C / C++ language specification. The statements are classified into the following categories: expression statements, variable and function declaration statements, conditional statements, loop statements, jump statements, and exception handling statements. Obtain the attention weights assigned to each type of C / C++ statement;
[0020] (3) The Transformer model will pay attention to specific types of tokens and statements. The tokens that receive high attention are not the same as the tokens with high frequencies. In order to be able to more reasonably select important tokens and statements, a code simplification strategy based on attention is selected. The tokens and statements to be retained are selected based on their attention weights, and code simplification is carried out in two stages, namely selecting statements and pruning tokens. The specific steps are as follows:
[0021] (3-1) In the selection statement stage, according to the obtained statement attention weight dictionary, the key statement with the highest attention weight is selected. At the same time, we want to ensure that the selected statement does not exceed the length limit. This can be formulated as a 0-1 knapsack problem, where the statements are regarded as items to be selected into the knapsack, the attention weight of the statement is the value, the statement length is the weight, and the length limit can be regarded as the capacity of the knapsack.
[0022] (3-2) In the pruning token stage, through a greedy strategy, tokens with the lowest attention weight in the statement with the lowest weight are iteratively removed until the requirement of the maximum number of tokens is met. To avoid the problem that shorter sequences with lower attention weights are more likely to be selected by the algorithm because their value ratio (attention weight divided by sequence length) is much larger than that of longer sequences, min-max normalization is used to amplify the attention weight and multiply it by the number of tokens in the statement, that is:
[0023]
[0024] where N represents the number of tokens in the statement;
[0025] (4) Fine-tuning the pre-trained model for the vulnerability assessment task can be carried out through a classification method, which is specifically designed to ensure the consistency between the pre-trained model and the task. Therefore, it is modeled as a multi-classification task and a classification head is customized. Classes provided by the Transformer library are used for classification fine-tuning, but since it is a multi-classification task, this class is adjusted. The specific steps are as follows:
[0026] (4-1) The customized classification head is optimized for the multi-classification task and is designed to consist of three fully connected networks, gradually reducing the input feature dimension from 768 to 512, 256, and finally mapping to the output classes (i.e., 4 severity levels) to achieve more efficient feature extraction;
[0027] (4-2) By introducing layer normalization, the classification head normalizes the features after the first fully connected layer, which helps to stabilize the training process and alleviate the problem of gradient vanishing or explosion;
[0028] (4-3) In addition, the GELU activation function is used instead of traditional activation functions (such as tanh), which enhances the non-linear expression ability of the model. At the same time, the Dropout ratio is set to 0.3 to effectively prevent overfitting. The overall architecture, through hierarchical feature compression and regularization design, not only improves the training stability of the model but also significantly enhances the generalization ability of the task. For computational efficiency, an approximate form of GELU is often used:
[0029]
[0030] Among them, x is the input.
[0031] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention proposes an optimization method for vulnerability assessment based on code simplification, which combines code slicing and simplification strategies. On the premise of ensuring the performance of the vulnerability assessment model, it significantly reduces the complexity and scale of the input code. By defining code slicing criteria, the semantic integrity of code segments is guaranteed; the relevance of elements in the code is captured during the attention weight generation stage; through the simplification strategy, code segments with low contribution to prediction are effectively removed while retaining key vulnerability-related features. This not only reduces the training and inference time overhead of the model, but also significantly improves the computational efficiency of the model. In addition, this method can still quickly complete vulnerability assessment in resource-constrained scenarios, optimizing the applicability and scalability of the model, and providing a more efficient solution for handling complex vulnerability scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a system framework diagram of the optimization method for vulnerability assessment based on code simplification provided by the present invention.
[0033] Figure 2 It is a specific example of the simplification strategy of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] In order to enable developers in this field to better understand the specific implementation process of the present invention, the following will be described in detail in combination with the attached drawing information. The description in this part is only for the explanation and demonstration of the invention, and does not limit the scope of action of the invention.
[0035] Embodiment 1
[0036] Please refer to Figure 1 , this embodiment is divided into four stages: data preprocessing stage, attention weight acquisition stage, code simplification stage, and model fine-tuning stage. The specific implementation includes the following contents:
[0037] (1) The original data of the experimental object is to construct a high-quality vulnerability assessment data set from vulnerable code, including the source code and severity information of 17,380 vulnerabilities. The following preprocessing is performed on the original data:
[0038] (1-1) First, according to the CVSS 3.0 standard, delete the data rows where the vulnerability scoring vectors do not meet the standard.
[0039] (1-2) Since there is redundant information (such as blank lines, comments, spaces) in the source code, in order to improve the quality of the data set, functions such as strip are applied to delete the redundant information. Finally, 14,235 pairs of vulnerable functions are selected as the experimental object.
[0040] (1-3) Table 1 shows the statistics of vulnerability occurrence locations. According to the vulnerability occurrence locations, code slicing criteria for vulnerable functions are formulated: statements that match the control keyword set or the preprocessing keyword set are grouped into complete code segments, while statements that do not meet this condition are split into separate code segments. According to this criterion, the source code is sliced.
[0041] (1-4) The dataset is divided into a training set (80%), a validation set (10%), and a test set (10%). After the division, the dataset is divided into 11,388 training sets, 1,423 validation sets, and 1,424 test sets.
[0042] Statistics of vulnerability occurrence locations in Table 1
[0043]
[0044] (2) In the attention weight acquisition stage, since the core of the Transformer model is the self-attention network, where the hidden state of each token is calculated layer by layer according to the self-attention weights, we use the attention weights in the Transformer layer of the Transformer model after pre-training to measure the importance of each token. The process of measuring the importance of tokens and statements is as follows:
[0045] (2-1) The Transformer model has multiple self-attention layers and heads. Each layer generates an attention weight for the same token, and the attention weights of all layers and heads for each token are averaged. Each vulnerable function in the dataset is input into the Transformer model, and then, the average attention weight assigned by the Transformer model to each token is calculated. Tables 2 and 3 show the key tokens that receive the most attention from the CodeBERT and CodeT5 models.
[0046] Table 2 Key tokens that receive the most attention from the CodeBERT model
[0047]
[0048] Table 3 Key tokens that receive the most attention from the CodeT5 model
[0049]
[0050]
[0051] After obtaining the token weights, the attention weights of each statement can be calculated. If we simply sum the average attention weights of all the obtained tokens as the attention weight of a statement, then we do not consider that different tokens play different roles in a statement. Therefore, we consider using the weighted average attention weights of all tokens in a statement S for summation to obtain the attention weight a(S) of this statement:
[0052]
[0053] (2-3) Classify all statements according to the C / C++ language specifications. The statements are classified into the following categories: expression statements, variable and function declaration statements, conditional statements, loop statements, jump statements, and exception handling statements. Tables 4 and 5 show the attention weights assigned by CodeBERT and CodeT5 to each type of C / C++ statement.
[0054] Table 4 Attention weights assigned by the CodeBERT model to each type of C / C++ statement
[0055]
[0056]
[0057] Table 5 Attention weights assigned by the CodeT5 model to each type of C / C++ statement
[0058]
[0059] (3) In the code simplification stage, in order to more reasonably select important tokens and statements, an attention-based code pruning strategy is adopted to select the tokens and statements to be retained based on their attention weights. For a given C / C++ language, the classification of statements is obtained. Then, an attention weight dictionary for statements and tokens is created. This code pruning strategy simplifies the code in two steps, namely selecting statements and pruning tokens. Figure 2 Figure is the code example diagram before and after using the code simplification method. The specific steps are as follows:
[0060] (3-1) When selecting statements, according to the obtained statement attention weight dictionary, select the key statements with the highest attention weights. At the same time, it is necessary to ensure that the selected statements do not exceed the length limit. The problem of selecting statements is formulated as a 0-1 knapsack problem. The statement s is regarded as an item to be selected into the knapsack. The attention weight a(s) of the statement is the value, and the statement length |s| is the weight. The length limit Max Length can be regarded as the capacity of the knapsack. The mathematical expression is as follows:
[0061]
[0062] x i ∈ {0, 1} for i = 1, 2, ..., |s|;
[0063] where x i indicates whether the i-th statement is selected.
[0064] (3-2) When pruning tokens, the token with the lowest attention weight in the statement with the lowest weight is removed by the greedy strategy until the maximum token number requirement is met. At the same time, min-max normalization is used to amplify the attention weights and multiply them by the number of tokens in the statement, i.e.:
[0065]
[0066] (4) In the model fine-tuning stage, the fine-tuning of the pre-trained model for the vulnerability assessment task can be carried out by a classification method, which is specially designed to ensure the consistency between the pre-trained model and the task. For the vulnerability assessment task, it is modeled as a multi-classification task and a classification head is customized for vulnerability assessment. The classes provided by the Transformer library are used for classification fine-tuning. However, since the vulnerability assessment is a multi-classification task, this class is adjusted. This customized classification head is optimized for multi-classification tasks and is designed to consist of three fully-connected networks, gradually reducing the input feature dimension from 768 to 512, 256, and finally mapping to 4 output severity categories to achieve more efficient feature extraction. By introducing layer normalization, this classification head standardizes the features after the first fully-connected layer, which helps to stabilize the training process and alleviate the problem of gradient vanishing or explosion. In addition, the GELU activation function is used instead of the traditional activation function (such as tanh), which enhances the non-linear expression ability of the model. At the same time, the Dropout ratio is set to 0.3 to effectively prevent overfitting. The overall architecture, through hierarchical feature compression and regularization design, not only improves the training stability of the model but also significantly enhances the generalization ability of the task. Table 6 shows the hyperparameter configuration obtained according to specific experiments in this embodiment.
[0067] Table 6 Hyperparameter configuration obtained according to specific experiments
[0068]
[0069] (4) Consider performance evaluation metrics such as F1-score and MCC. To comprehensively evaluate the performance of the vulnerability assessment model, the macro-average versions of F1-score and MCC are used, which are calculated by taking the average of the metrics independently calculated for each category in the multi-classification task. This choice is due to the multi-class nature and class imbalance problem of the vulnerability assessment task. To ensure fairness, the same dataset and partitioning method are used for experiments in this embodiment and all baseline methods. The results are shown in Table 7:
[0070] Table 7 Comparison results of this embodiment and baseline methods in terms of two performance metrics
[0071]
[0072]
[0073] In this table, we emphasize the best performance of different performance metrics in bold. As can be seen from the table, the performance of this embodiment reaches 0.48 and 0.32 in terms of F1 and MCC respectively. Compared with the baseline methods, the performance of the two metrics has increased by 29.73% and 30.08% on average. It is worth noting that in terms of the F1 metric, this embodiment has improved the performance by at least 14.29%. This shows that this embodiment is more effective than the five mainstream methods and has higher practical application potential in the vulnerability assessment task.
[0074] Experiments show that generating simplified code based on the code simplification strategy helps improve the performance of the vulnerability assessment task. Compared with the current mainstream vulnerability assessment methods, this embodiment can not only obtain more accurate prediction results, but also make up for the problems of the mainstream methods such as too long code input and large time overhead in model construction and inference, further indicating the rationality and competitiveness of the method of this embodiment.
[0075] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A software vulnerability assessment optimization method based on code simplification, characterized in that: The steps include: (1) Select vulnerability codes that meet the CVSS 3.0 standard to build a high-quality vulnerability assessment dataset, including the source code and severity information of the vulnerability, and perform program slicing on the source code in the dataset; (2) Taking the source code as input, the code pre-training Transformer model is pre-trained. During the pre-training process, a dictionary of attention weight values of the model's tags and statements is generated and stored to capture and record the relationship between the elements in the code and their contribution to the model prediction; (3) Simplify the source code. The simplification process includes removing code segments that have low contribution to model prediction and reorganizing the remaining code to improve processing efficiency and model prediction accuracy. (4) Take the simplified code as input, fine-tune the code pre-trained Transformer model, and train a vulnerability assessment model.
2. The software vulnerability assessment optimization method based on code simplification according to claim 1 is characterized in that: The code slicing criteria defined in step (2) include the following steps: (2-1) For the existing vulnerable source code, first count the types of statements that are prone to vulnerabilities; (2-2) Then, according to the programming language standard of the target code, the target programming language is defined for statement type keywords; (2-3) Finally, the keywords of the final code segmentation criteria are determined by combining the statement types that are prone to vulnerabilities and the defined statement type keywords. If the statements contain the control body keywords, the control body statements are treated as a code fragment. Otherwise, the individual statements are treated as a code fragment.
3. The software vulnerability assessment optimization method based on code simplification according to claim 1 is characterized in that: The code simplification strategy defined in step (3) comprises the following steps: (3-1) First, according to the specified code simplification rate, the maximum length of the simplified statement is calculated; (3-2) Next, based on the sentence attention weight dictionary obtained in step (2), sentence selection is performed on the code snippet processed in step (1), and sentences with high weights in the sentence attention weight dictionary are retained first. These sentences are considered to play a key role in vulnerability triggering and propagation; (3-3) Finally, if the selected statement sequence exceeds the maximum length of the statement, a pruning mark is set for the selected statement sequence, and statements with low weight marks are pruned first.
4. The software vulnerability assessment optimization method based on code simplification according to claim 1 is characterized in that: In step (4), the model is fine-tuned. For the vulnerability assessment task, it is modeled as a multi-classification task and a customized classification head is set, which consists of a three-layer fully connected network. The input feature dimension is gradually reduced from 768 to 512 and 256, and finally mapped to the output category. Layer normalization and GELU activation function are introduced, and the Dropout ratio is set to 0.3.