Lightweight software vulnerability assessment method based on particle swarm optimization and knowledge distillation
Through particle swarm optimization and knowledge distillation methods, students' model architecture is optimized and teachers' model knowledge is migrated, which solves the application bottleneck of large code models in resource-constrained environments, and realizes efficient and real-time evaluation of lightweight software vulnerability assessment.
Patent Information
- Application Number
- CN202510357372.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-04
AI Technical Summary
Large code models consume high computing resources in software vulnerability assessment, resulting in slow inference speed and difficulty in deploying and applying in resource-constrained environments.
Combining particle swarm optimization algorithm and knowledge distillation technology, the compression and performance optimization of the model are achieved by searching for the best student model architecture and migrating teacher model knowledge.
Significantly reduce model size and inference time while maintaining high accuracy, suitable for real-time vulnerability assessments for resource-constrained devices.
Smart Images

Figure CN120258089A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a lightweight software vulnerability assessment method based on particle swarm optimization and knowledge distillation. Background Art
[0002] With the rapid development of pre-trained models, large code models have been widely used in software engineering tasks. However, the high number of parameters and complex network structures of these large models lead to slow inference speed and high storage occupancy, making it difficult to directly deploy them in resource-constrained environments. For example, the CodeBERT model for code generation and vulnerability detection has more than 120 million parameters and a model size of more than 460MB. The computational overhead of such large-scale models is huge, which not only affects the inference performance but also brings significant energy consumption and carbon emission problems, having a negative impact on environmental sustainability.
[0003] Traditional software vulnerability assessment methods mainly rely on pre-trained large-scale models to evaluate potential risks in software systems. Although these methods improve the automation of vulnerability detection and reduce the need for manual review, their high computational cost and resource consumption limit their application in real-time or low-latency environments. Especially in integrated development environments (IDEs) or embedded devices, the high storage and computational requirements of the models become a bottleneck for their large-scale application.
[0004] To address the above problems, researchers have proposed various model compression techniques in recent years, attempting to reduce the model size and computational resource consumption while maintaining the model performance. Among them, knowledge distillation is an effective model compression method that transfers the knowledge of a large teacher model to a smaller student model, thereby maintaining a high prediction accuracy while significantly reducing the model scale. However, knowledge distillation also faces some challenges. Especially when selecting an appropriate student model architecture to carry the knowledge of the teacher model, it may lead to performance loss. In addition, the selection of the student model architecture involves a complex parameter search space, which makes it extremely time-consuming and difficult to find the optimal architecture. Summary of the Invention
[0005] The purpose of the present invention is to provide a lightweight software vulnerability assessment method based on particle swarm optimization and knowledge distillation to address the challenges of existing vulnerability assessment models in terms of efficiency and resource occupancy, especially the application bottleneck of large code models in software vulnerability assessment tasks. Through the method of the present invention, it is possible to significantly reduce the model size and inference time while retaining high performance, thereby achieving real-time or near-real-time software vulnerability assessment on resource-constrained devices.
[0006] The idea of the present invention is as follows: The present invention proposes a lightweight software vulnerability assessment method based on particle swarm optimization and knowledge distillation. The technical solution of the present invention includes the following core innovations: First, the particle swarm optimization (PSO) algorithm is used to search in the model architecture space to determine the optimal student model architecture, so that the student model can exhibit good performance with a limited number of parameters. Second, knowledge distillation technology is used to transfer the knowledge in the teacher model to the student model to achieve a balance between model compression and performance. Finally, the performance and time cost of the compressed model in the vulnerability assessment task are evaluated. Specifically, in the implementation process of the present invention, an optimized student model is constructed to complete the vulnerability assessment task with a smaller computational cost, effectively reducing memory occupancy and inference latency.
[0007] The present invention is implemented by the following measures: A lightweight software vulnerability assessment method based on particle swarm optimization and knowledge distillation, which includes the following steps:
[0008] 1.1: Collect code snippets of C++ vulnerabilities, CVSS (Common Vulnerability Scoring System) severity scores and classification information of vulnerabilities from the National Vulnerability Database (NVD). Clean and preprocess the collected data, including removing comments, blank lines, irrelevant symbols and elements in the code snippets, to form a structured vulnerability assessment dataset D. The severity levels in dataset D include four levels: "Critical" (9.0 - 10.0 points), "High" (7.0 - 8.9 points), "Medium" (4.0 - 6.9 points) and "Low" (0.1 - 3.9 points);
[0009] 1.2: Use the particle swarm optimization (PSO) algorithm to search in the architecture parameter space of the student model, design a fitness function to evaluate the performance of each student model architecture, and find the optimal lightweight student model architecture;
[0010] 1.3: Use the knowledge distillation method to take the output of the teacher model as "soft labels" and transfer knowledge to the student model to achieve model compression and performance optimization;
[0011] 1.4: Use the obtained lightweight student model to perform software vulnerability assessment tasks, calculate the evaluation metrics of the model, including accuracy (Accuracy), F1 score (F1 Score) and time cost (Time cost), to verify the compression effect and performance of the student model.
[0012] The particle swarm optimization algorithm in step 1.2 searches the configuration space, specifically including:
[0013] 2.1: Set the initial parameters of the PSO algorithm, including the particle swarm size, inertia weight (set to 0.5), individual learning factor c1, and swarm learning factor c2 (set to 2.0) to control the speed adjustment and global search behavior of the particles;
[0014] 2.2: Randomly initialize the positions and velocities of the particles in the model architecture parameter space. Each particle represents a candidate student model architecture. The position parameters include structural parameters such as the number of hidden layers, number of attention heads, embedding dimension, and hidden dimension. For example, a possible student model architecture is ["Byte-Pair Encoding", 3000, 12, 96, "GELU", 0.1, 3072, 12, 0.1, 512, "absolute", 5e-5, 32]. Each value in the vector represents a configuration parameter to ensure the diversity and global exploration of the search process;
[0015] 2.3: Design a fitness function to evaluate the performance of each student model architecture. The fitness function comprehensively considers the performance of the model in terms of the accuracy of vulnerability assessment while maintaining a relatively small number of parameters, as well as the model size, to achieve a balance between model compression and performance. The formula for the fitness function is Fitness(i) = GFLOPs - |S - s i |, where GFLOPs represents the computing power of the model, S represents the capacity of the teacher model, and s i represents the capacity of the i-th possible student model;
[0016] 2.4: Update the velocity of each particle using the weighted sum of the inertia term, cognitive term, and social term. The formula is: v = w * v + c1 * r1 * (pBest - position) + c2 * r2 * (gBest - position), where w represents the inertia weight, and r1 and r2 are random numbers used to enhance the randomness of the search;
[0017] 2.5: Update the individual best position (pBest) and global best position (gBest) of the particles in each generation of iteration, and select the particle with the highest fitness after multiple rounds of iteration to determine the optimal lightweight student model architecture.
[0018] The use of the knowledge distillation technique in step 1.3 is as follows:
[0019] 3.1: Use the teacher model to predict the unlabeled training data to generate "soft labels" as the target labels for the student model to learn. These labels contain the subtle classification information of the teacher model under different inputs;
[0020] 3.2: When training the student model, the parameters of the student model are gradually optimized by minimizing the KL (Kullback-Leibler Divergence) loss between the output of the student model and the soft labels of the teacher model. The loss function is loss = KL(softmax(T_output), softmax(S_output)), where T_output and S_output are the outputs of the teacher and student models respectively;
[0021] 3.3: Introduce a temperature adjustment parameter T (set to 3) to smooth the output distribution of the teacher model, making it easier for the student model to learn the knowledge of the teacher model. The adjustment of temperature T improves the learning ability of the student model for complex tasks in knowledge distillation;
[0022] 3.4: After multiple rounds of training iterations, the student model can retain the vulnerability assessment prediction performance of the teacher model as much as possible while significantly reducing the number of parameters, achieving efficient model compression and knowledge transfer.
[0023] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0024] 1. The present invention proposes a lightweight software vulnerability assessment method based on particle swarm optimization and knowledge distillation, innovatively combining an optimization algorithm with model compression technology to address the application limitations of large-scale pre-trained code models in resource-constrained environments. By using the particle swarm optimization algorithm to search for the optimal student model architecture, automatic adjustment of the model structure is achieved, enabling the student model to maintain high performance in the vulnerability assessment task while reducing the number of model parameters. This method can operate efficiently on devices with limited resources, broadening the application scenarios of software vulnerability assessment models.
[0025] 2. Compared with traditional vulnerability assessment models, the present invention uses knowledge distillation technology to transfer the knowledge of a large teacher model to a smaller student model, significantly reducing the storage requirements and computational costs of the model. Compared with traditional models that require a large amount of computing resources, the method of the present invention not only maintains high prediction accuracy but also significantly reduces the model inference time and memory occupancy. In addition, the method of the present invention demonstrates strong robustness and stability in performance evaluation experiments, can adapt to the evaluation requirements under different datasets and environments, and provides an efficient and lightweight solution for software security assurance in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a system framework diagram of a lightweight software vulnerability assessment method based on particle swarm optimization and knowledge distillation provided by the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0027] To more clearly elaborate the purpose, technical details, and advantages of the present invention, the following will elaborate on the present invention in detail through the accompanying drawings and specific implementation cases. It should be clear that these implementation cases are only for explanation and illustration purposes and do not limit the scope of the present invention.
[0028] Example 1
[0029] See Figure 1 As shown, this example provides a lightweight software vulnerability assessment method based on particle swarm optimization and knowledge distillation, specifically including the following:
[0030] (1) Construction and preprocessing of the vulnerability assessment dataset, including the following steps:
[0031] (1-1) Obtain code snippets containing C++ vulnerabilities and their CVSS (Common Vulnerability Scoring System) scores and grading information from the vulnerability database (NVD) as the basic dataset for vulnerability assessment.
[0032] (1-2) Process each vulnerability record, including removing comments, blank lines, and irrelevant elements in the code, unifying the code format, and generating a structured vulnerability assessment dataset D. Dataset D contains four severity levels: Critical (9.0 - 10.0), High (7.0 - 8.9), Medium (4.0 - 6.9), and Low (0.1 - 3.9), which are used to support the training and testing of subsequent models. Table 1 shows the detailed information of the training set (80%), validation set (10%), and test set (10%).
[0033] Table 1 Statistical information of the experimental objects
[0034]
[0035] (2) Search and optimization of the student model architecture, specifically including the following steps:
[0036] (2-1) Search for the best architecture in the architecture parameter space of the student model through the particle swarm optimization (PSO) algorithm. The pseudo-code of this algorithm is shown in Table 2. Each particle represents a candidate student model architecture, which contains specific model parameters such as the number of layers, embedding dimension, and number of attention heads when initialized.
[0037] (2-2) Design a fitness function to comprehensively consider the vulnerability assessment accuracy and resource consumption of the model while maintaining a small number of parameters to optimize the architecture of the student model. The formula of the fitness function is Fitness(i) = GFLOPs - |S - s i |.
[0038] (2-3) Adjust the formula according to the velocity and position of the particles, and gradually update the particle positions through weighted calculations of the inertia term, cognitive term, and social term, so that the particle swarm gradually converges to the optimal solution, thereby obtaining the optimal lightweight student model architecture. The formula is: v = w*v + c1*r1*(pBest - position) + c2*r2*(gBest - position).
[0039] (3) The process of enabling the small model to learn from the large model through knowledge distillation specifically includes the following steps:
[0040] (3-1) Use the teacher model to generate "soft labels", and transfer knowledge to the student model through the probability distribution information predicted from unlabeled data.
[0041] Table 2 PSO algorithm pseudocode
[0042]
[0043] (3-2) When training the student model, gradually optimize the parameters of the student model by minimizing the difference between the output of the student model and the soft labels of the teacher model to ensure performance retention with a smaller model capacity. The loss function is loss = softmax(p / T) * log(softmax(q / T)) * T 2 , and the pseudocode is shown in Table 3.
[0044] Table 3 Knowledge distillation algorithm pseudocode
[0045]
[0046] (3-3) Introduce a temperature adjustment parameter during the knowledge distillation process to smooth the output distribution of the teacher model and enhance the adaptability of the student model to complex tasks. The time cost required for model training is shown in Table 4.
[0047] Table 4 Comparison of time costs between the teacher model (CodeBERT) and the student models (BiLSTMsoft and the method of this embodiment)
[0048]
[0049] (4) The vulnerability assessment process of the lightweight model specifically includes the following steps:
[0050] (4-1) Apply the optimized lightweight student model to the software vulnerability assessment task. After inputting the vulnerability code snippet to be evaluated, the model outputs the vulnerability severity level (Critical, High, Medium, Low) of the code.
[0051] (4-2) Evaluate the compression effect and performance of the model using metrics such as the accuracy, F1 score, and model size of the model prediction to verify the practicability and performance of the model on resource-constrained devices. The results are shown in Table 5.
[0052] Table 5 Comparison of Compression Effect and Performance between the Teacher Model (CodeBERT) and the Student Models (BiLSTMsoft and the Method of this Embodiment)
[0053]
[0054]
[0055] The experimental results show that a lightweight software vulnerability assessment method based on particle swarm optimization and knowledge distillation proposed in this embodiment is significantly superior to the existing baseline methods in terms of the key evaluation metrics of accuracy and F1 score.
[0056] Specifically, the method of this embodiment reaches 60.89% in terms of accuracy and 53.01% in terms of F1 score. Compared with the baseline method BiLSTM soft in this embodiment, the accuracy is improved by 5.41% and the F1 score is improved by 4.5%. Secondly, as shown in Table 4, the training time of the method of this embodiment is only 19 minutes. Compared with the 68 minutes required for the fine-tuning of CodeBERT, the method of this embodiment reduces the time cost by 72.06%. Finally, compared with the size of the original teacher model, which is 476MB, the capacity of the model compressed by the method of this embodiment is only 3MB, which is only 0.6% of the original. This shows that the method of this embodiment can significantly reduce the model size.
[0057] These experimental results not only verify the effectiveness of the method of this embodiment in model compression and performance optimization, but also demonstrate its practical value in resource-constrained environments. By combining the architecture search ability of the particle swarm optimization algorithm and the model compression advantage of the knowledge distillation technology, this embodiment can significantly reduce the model size and inference time while maintaining a high vulnerability assessment accuracy, providing an efficient and lightweight solution for the field of software vulnerability assessment.
[0058] Embodiment 2
[0059] These experimental results not only verify the effectiveness of the method in this embodiment for model compression and performance optimization, but also demonstrate its practical value in resource-constrained environments. By combining the architecture search ability of the particle swarm optimization algorithm with the model compression advantage of knowledge distillation technology, this embodiment can significantly reduce the model size and inference time while maintaining a high vulnerability assessment accuracy, providing an efficient and lightweight solution for the field of software vulnerability assessment.
[0060] This embodiment tests the lightweight software vulnerability assessment method based on particle swarm optimization and knowledge distillation on different datasets to further verify its effectiveness.
[0061] Collect vulnerability code snippets in Java language and their corresponding vulnerability severity information from another authoritative vulnerability database (such as CVE Details). Similarly, clean and preprocess the data, removing comments, blank lines, and irrelevant elements to generate a new structured vulnerability assessment dataset. This dataset is also divided into a training set, a validation set, and a test set in the ratio of 80%, 10%, and 10%. The training set contains 15,000 data, the validation set contains 1,875 data, and the test set contains 1,875 data. Use the same particle swarm optimization and knowledge distillation methods as in Embodiment 1 for model training and optimization. In the student model architecture search stage, use the same particle swarm optimization algorithm.
[0062] In the knowledge distillation stage, select a large Java vulnerability assessment model based on the Transformer architecture as the teacher model. After training, use the optimized lightweight student model for the software vulnerability assessment task and compare it with traditional rule-based vulnerability assessment methods and assessment methods based on simple neural networks (such as shallow multi-layer perceptron MLP). The results are shown in Table 6:
[0063] Table 6 Comparison of compression effects and performance between the teacher model and the student model
[0064]
[0065] Embodiment 3
[0066] To verify the performance of the technical solution of the present invention in different hardware environments, experiments are carried out on an embedded device (Raspberry Pi 4B). The experimental dataset uses the same C++ vulnerability data collected from NVD as in Embodiment 1.
[0067] On the Raspberry Pi 4B, due to its limited resources, the parameters in the particle swarm optimization algorithm and the knowledge distillation process are adjusted appropriately. The particle swarm size is set to 30, and the inertia weight (set to 0.5), the individual learning factor c1, and the swarm learning factor c2 (set to 2.0) are used to reduce the computational load. The temperature adjustment parameter T in the knowledge distillation stage is set to 2.5 to enable the student model to converge faster.
[0068] Similarly, it is also compared with traditional static analysis-based vulnerability assessment tools (such as Cppcheck) and lightweight assessment models based on small convolutional neural networks (CNNs). The experimental results are shown in Table 7.
[0069] Table 7 Comparison of the compression effects and performance between the teacher model and the student model
[0070]
[0071] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A lightweight software vulnerability assessment method based on particle swarm optimization and knowledge distillation, characterized in that It includes the following steps: 1.1: Collect C++ vulnerability code snippets, vulnerability severity scores, and their classification basic information from the vulnerability database NVD. Preprocess the collected data to remove comments, blank lines, and other irrelevant elements in the code, and generate a vulnerability assessment dataset D in a unified format. The vulnerability severity in the dataset is divided into four levels: "Critical", "High", "Medium", and "Low". 1.2: Use the particle swarm optimization algorithm to search in the architecture parameter space of the student model to determine the final lightweight student model architecture. This step includes initializing the particle swarm, updating the position and velocity of the particles based on the fitness function, and gradually guiding the particle swarm to converge to the optimal architecture. 1.3: Through the knowledge distillation method, use the output of the teacher model as soft labels to transfer knowledge to the student model with the determined lightweight architecture. This step minimizes the difference between the output of the student model and the teacher model, enabling the student model to maintain the high performance of the teacher model while reducing the number of parameters. 1.4: Use the optimized lightweight student model to perform the software vulnerability assessment task. Input the source code snippets of the vulnerabilities to be evaluated into the trained lightweight student model, predict the severity of the vulnerabilities, and divide the severity into four levels: "Critical", "High", "Medium", and "Low". Calculate the evaluation metrics of precision, F1 score, and time cost to comprehensively evaluate the prediction performance and resource consumption of the software vulnerability assessment model after compression.
2. The lightweight software vulnerability assessment method based on particle swarm optimization and knowledge distillation according to claim 1, wherein The particle swarm optimization algorithm in step 1.2 includes the following steps: 2.1: Set the initial values of the parameters of the particle swarm optimization algorithm, including the particle swarm size, inertia weight, individual learning factor, and swarm learning factor, to control the search behavior of the particles. Each particle represents a student model architecture, and the ability to regulate the search process through parameter configuration. 2.2: Randomly initialize the architecture parameter positions and velocities of each particle in the search space to ensure the diversity of the search process. 2.3: Design the fitness function, and the calculation formula is Fitness(i) = GFLOPs - |S - s i |, where GFLOPs in the formula represents the computational cost of the model, S represents the capacity of the teacher model, and s i represents the capacity of the i-th possible student model. The fitness value of each particle is calculated through this fitness function, which is used to evaluate the effect of the architecture; 2.4: Adjust the velocity and position of each particle according to the velocity update formula, including the inertia weight, individual learning factor, and swarm learning factor, so that the particles gradually move towards the optimal solution. 2.5: Adjust the velocity and position of each particle according to the velocity update formula, including the weighted calculation of the inertia weight, individual learning factor, and swarm learning factor, so that the particles gradually approach the optimal solution.
3. A lightweight software vulnerability assessment method based on particle swarm optimization and knowledge distillation according to claim 1, characterized in that The knowledge distillation process in step 1.3 includes the following steps: 3.1: Use the teacher model to generate prediction results for the unlabeled software vulnerability assessment dataset, and use these predictions as the "soft labels" of the student model to provide rich knowledge for the student model to learn. 3.2: When training the student model, minimize the difference between the output of the student model and the soft labels of the teacher model through KL divergence. The loss function is set as loss = KL(softmax(T_output), softmax(S_output)), and gradually optimize the parameters of the student model to achieve knowledge transfer. 3.3: By introducing temperature adjustment parameters, the smoothness of the teacher model output is controlled to improve the effect of knowledge distillation and help the student model absorb the software vulnerability assessment knowledge in the teacher model; 3.4: By iteratively updating the parameters of the student model, the prediction performance of the teacher model is retained while the number of parameters is reduced, thus achieving compression and knowledge retention of the software vulnerability assessment model.
4. A lightweight software vulnerability assessment method based on particle swarm optimization and knowledge distillation according to claim 1, characterized in that, In step 1.1, C++ vulnerability code snippets, CVSS severity scores and classification information of the vulnerabilities are collected from the vulnerability database NVD. During the data collection process, web crawler technology or the API interface provided by NVD is used to filter out vulnerability data related to the C++ language according to the set search rules; The collected data is cleaned and preprocessed. Regular expressions are used to match and delete comments in code snippets. Blank lines are removed by judging whether the code lines are empty or contain only blank characters. At the same time, irrelevant symbols and elements in the code are removed. The processed code snippets are unified in format to form a structured vulnerability assessment dataset D. In the dataset, the severity of vulnerabilities is divided into four levels: Critical, High, Medium, and Low–. The dataset is divided into training set, validation set, and test set in a ratio of 80%, 10%, and 10%, providing data support for subsequent model training, parameter adjustment, and performance evaluation.