Code large model security reinforcement method based on knowledge distillation and co-decoding
Through knowledge distillation and co-decoding technologies, and the functional correctness of the basic model and the best model are aligned, and the safety reinforcement loss function is designed in combination with oversampling and low-rank adaptation methods, the problem of high cost and poor effect of safety reinforcement of code large models in the existing technology is solved, and more efficient and safe code generation is achieved.
Patent Information
- Application Number
- CN202510528116.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art has problems such as high data cleaning cost, poor model adaptability and limited security reinforcement effect in the security reinforcement process of code large models. Especially when using smaller security models to reinforce larger basic models, the generation results are unstable and potential security vulnerabilities may be introduced.
The knowledge distillation and co-decoding methods are used to align the functional correctness of the basic model and the best model through knowledge distillation, combine the oversampling strategy and the low-rank adaptation method to design the safety reinforcement loss function, and use the CoSec strategy for collaborative decoding, and embed the security model to reinforce the target model.
It significantly improves the security and functional accuracy of code generation, improves the model's protection against unlearned vulnerabilities, reduces development costs and resource consumption, and achieves a more efficient security reinforcement effect.
Smart Images

Figure CN120387484A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and more specifically, to a method for securely strengthening a large code model based on knowledge distillation and co-decoding. Background Art
[0002] In recent years, with the exponential growth of computing power and the accumulation of massive data resources, large language models (LLMs) based on the Transformer architecture have made breakthroughs in the field of code generation. Through pre-training with massive open-source code libraries (such as GitHub) and their associated natural language data of documents and annotations, code LLMs not only demonstrate powerful context-aware code completion capabilities but can also complete complex software engineering tasks such as code summarization, defect repair, and cross-language conversion. Such models are gradually reconstructing the software development paradigm, and intelligent programming assistants such as GitHub Copilot and Codeium have penetrated into the daily work processes of more than 76% of developers globally. However, there are serious security risks hidden behind this technological revolution, mainly reflected in two major dimensions: First, at the data level, the several terabytes of code data used for pre-training are mixed with a large number of code samples containing known vulnerabilities. Related research has found that approximately 12.3% of the code snippets in the mainstream code LLM training sets (such as The Stack) exhibit CWE Top25 vulnerability patterns, including but not limited to high-risk vulnerabilities such as buffer overflow (CWE-119), SQL injection (CWE-89), and cross-site scripting (CWE-79). These defective code patterns are implicitly absorbed by the model through self-supervised learning, resulting in the possible recurrence of similar vulnerabilities when generating code. Second, at the generation mechanism level, causal language models based on maximum likelihood estimation tend to generate token sequences with the highest statistical probabilities. Empirical studies have shown that when the model faces multiple possible code implementation schemes, the probability of choosing a secure implementation method is on average 28.7 percentage points lower than that of a dangerous implementation method. This "statistical priority" rather than "security priority" generation strategy makes the code generated by LLMs perform excellently in terms of syntactic correctness and context adaptability but may introduce potential security vulnerabilities.
[0003] Current mainstream solutions focus on constructing a purified dataset for model retraining and then retraining or fine-tuning the model. However, under this method, the required amount of data is huge and a large amount of manual supervision is needed. At the same time, if the model parameters are too large or multiple models with the same parameter scale have the same optimization requirements, the above method will become infeasible. For example, more than 300 million lines of secure code are screened out through static analysis tools (such as CodeQL, SonarQube) to form the CleanCode dataset, and models such as CodeGen and StarCoder are fine-tuned. But this method has significant drawbacks: First, data cleaning requires more than 2,000 person-days of expert review; second, the parameter scale of the fine-tuned model generally exceeds 30B, resulting in a single training cost of more than $1.5 million; finally, under this method, the multi-model adaptability is poor, and each model needs to be fine-tuned independently.
[0004] Although the existing co-decoding strategy CoSec performs well, in terms of functional correctness, CoSec can still become unstable in some scenarios, especially when using a smaller security model to strengthen a larger base model. In terms of security reinforcement, when the training is insufficient, the security model lacks sufficient confidence to offset the security tendency of the base model for sparse representation of CWE types, limiting the effect of security reinforcement. Summary of the Invention
[0005] Aiming at the above problems existing in the prior art, the technical problem to be solved by the present invention is: how to improve the capabilities of the model in terms of security reinforcement and functional correctness.
[0006] To solve the above technical problems, the present invention adopts the following technical solutions: A method for securing a large code model based on knowledge distillation and co-decoding, including the following steps:
[0007] Step S1: Select a base model and an optimal model from any model family;
[0008] Step S2: Select any code snippet dataset as the training set C, use the knowledge distillation method to align the functional correctness of the base model and the optimal model, and then use the Adam optimizer to train the base model after the functional correctness alignment. After the training is completed, the trained base model Q with aligned functional correctness is obtained;
[0009] Step S3: Select several vulnerability data pairs to form a vulnerability dataset. Each vulnerability data pair in the vulnerability dataset includes the vulnerability type of the code snippet and the programming language type of the edited code snippet. The oversampling strategy is used to preprocess the vulnerability dataset to obtain a security reinforcement training set D;
[0010] Step S4: Design a security reinforcement loss function Taking D as the input of Q, the low-rank adaptation method is used to train Q to obtain the final security-enhanced basic model M;
[0011] Step S5: Select the target model X from the model family, and embed M into X to obtain the security-enhanced target model X';
[0012] Step S6: Arbitrarily select the target model Y to be security-enhanced, and repeat Step S2-Step S5 to obtain the security-enhanced target model Y'.
[0013] Preferably, the database used in the knowledge distillation method in Step S2 is Magicoder.
[0014] Preferably, the steps to obtain the trained basic model Q after functional correctness alignment in Step S2 are as follows:
[0015] Design the rKLD loss function and combine it with the cross-loss function Design the loss function through weighting At the same time, set the maximum number of training rounds. Take C as the input data of the basic model and the best model, and use the Adam optimizer to train the basic model. When the training reaches the maximum number of training rounds or the loss function no longer changes, the training ends, and the trained basic model Q after functional correctness alignment is obtained;
[0016] The calculation expression of
[0017]
[0018] The calculation formula of
[0019]
[0020] The calculation expression of
[0021]
[0022] Among them, N represents the total number of training data in C, q θ represents the output of the basic model, p is the output of the target model, x i represents the i-th data in C, X 1:i-1 represents the set of the first to the i-1-th data in C, and λ is the ratio of, with a value range of [0,1].
[0023] Preferably, in Step S3, the content of preprocessing the vulnerability data set by oversampling strategy to obtain the security-enhanced training set D is as follows:
[0024] S3-1: Count the number of code snippets for each vulnerability type in the vulnerability dataset, and calculate the average value n of all vulnerability types in the vulnerability dataset; at the same time, count the number of programming languages used in the edited code snippets for each same vulnerability type, and set the programming language quantity threshold k;
[0025] S3-2: Compare the number of code snippets of each vulnerability type in the vulnerability dataset with n. If the number of code snippets of a vulnerability type is less than n, copy the existing code snippets of this vulnerability type so that the total number of code snippets of this vulnerability type is equal to n; otherwise, keep the existing state unchanged;
[0026] If the number of the d-th programming language in a same vulnerability type is less than k, randomly copy the number of code snippets so that the number of code snippets represented by the d-th programming language is equal to k; among them, the randomly copied code snippets are consistent with the vulnerability type and programming language of the code snippets represented by the d-th programming language.
[0027] Preferably, in the step S4, design a security reinforcement loss function as follows:
[0028] S4-1: Use the cross-entropy loss function to construct The calculation formula is as follows:
[0029]
[0030] Among them, Q θ represents the output of Q, and m i represents the security parameter corresponding to x i ;
[0031] Use the acceptance algorithm to calculate the value of m i , set the acceptance threshold a, and the value-taking process of m i is as follows:
[0032]
[0033] Among them, y n+1 represents the token value output by X when y1,..., y n are jointly used as the input of X; y' n+1 represents the token value output by X' when y1,..., y n are jointly used as the input of X'; represents the probability of generating the next token by y n+1 in X, represents the probability of generating the next token by y' n+1 in X';
[0034] If The calculated value of m is greater than a i then take 0; otherwise, for m i then take 1;
[0035] S4-2: Modify the rKLD function designed in the functional correctness alignment phase to obtain the function required in the secure training phase The calculation expression is as follows:
[0036]
[0037] wherein represents the predicted distribution of the secure model at the current time step represents the predicted distribution of the target model
[0038] S4-3: Use and to obtain the loss function finally required for secure training The calculation expression is as follows:
[0039]
[0040] Preferably, in the step S4, the collaborative decoding strategy adopted by the secure reinforcement base model M is CoSec
[0041] Compared with the prior art, the present invention has at least the following advantages:
[0042] 1. Comprehensively improved the security of code generation: CoSec+ introduces a small independent security model that, through collaborative decoding with the target model, guides code large models with different parameters to generate more secure code. This method can strengthen security without modifying the internal parameters of the target model, thus avoiding the problems of a large amount of manual intervention and resource consumption in traditional methods. As a result, it can effectively guide code large models with different parameters to generate more secure and effective code in actual application scenarios. Specifically, when comparing with the current state-of-the-art CoSec in three code large models, for the CodeGen-350M model, CoSec+ achieved a 27.9% relative security improvement, 7.2% higher than CoSec; for the larger CodeGen-2.7B model, CoSec+ achieved a significant improvement of 37.7, far exceeding CoSec's 19.7% improvement; for the CodeGen-6.1 model with a relatively high security rate baseline of 68.2%, CoSec+ achieved an additional 12.3% relative security improvement, almost twice that of CoSec. Furthermore, CoSec+ also has strong generalization ability in models of the StartCoder-Base series and DeepSeek-Coder series. For the StartCoder-Base-1B model, CoSec+ improved security by 7.7% more than CoSec, reaching 22.8%. For the StartCoder-Base-7B model, CoSec+ increased the security performance by 18.2%, showing an improvement of 15.8% beyond CoSec's 2.4%; for the DeepSeek-Coder-1.3B model, CoSec+ achieved a 16.8% relative security improvement, exceeding CoSec. For the DeepSeek-Cpder-6.7B model, CoSec+ significantly improved by 14.4%, exceeding CoSec's 4.5%. This shows that CoSec+ can effectively improve code security in a variety of code models.
[0043] 2. While improving the security of code generation, it can better maintain and even further improve the functional correctness of code generation: In the model functional alignment training stage, CoSec+ adopts the knowledge distillation technique to align the functional correctness of the base model. This process not only ensures that the finally trained secure model can generate high-quality outputs, but also reduces the functional performance loss caused by secure training. By using the Pass@1, Pass@5, and Pass@10 metrics of HumanEval as the evaluation of code generation correctness in actual application scenarios, among which, the improvement of Pass@1 indicates a significant increase in the probability that the model can output both secure and correct code at the first generation, which has important value for reducing debugging costs in actual development scenarios. The continuous optimization of Pass@5 and Pass@10 verifies the ability of the model to maintain stable output quality in multiple iterations. Compared with CoSec, after securing the CodeGen, StarCoderBase, and DeepSeek-Coder series of code large models, CoSec+ increased Pass@1 by 5.2% to reach 78.2%, and Pass@5 and Pass@10 increased by 1.6% to reach 82.8%. This shows that compared with the traditional security hardening method that causes an average functional performance loss of 3-8%, CoSec+ can effectively secure the code large model while taking into account functional correctness to a certain extent.
[0044] 3. It can effectively protect against security vulnerability types that the model has not learned before: The design of the CoSec+ framework enables it to effectively deal with security vulnerability types that have not been learned. By training and evaluating on different code models, CoSec+ demonstrates its effectiveness in a wide range of application scenarios. By evaluating the security of CoSec+ against four vulnerability types that do not appear in the security training dataset, it increased by 24.8% in CodeGen2.7B, by 1.3% in StarCoderBase-6B, and by 11.0% in DeepSeek-Coder. These results all show the excellent generalization ability of CoSec+ in security hardening, and it can still have a certain effective security hardening effect on vulnerability types that have never been trained. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 Schematic diagram of the method architecture of the present invention;
[0046] Figure 2 Schematic diagram of supervised collaborative decoding;
[0047] Figure 3 Comparison table of security hardening effects;
[0048] Figure 4Schematic diagram of the code before and after security enhancement. Detailed implementation mode
[0049] The present invention will be further described in detail below.
[0050] The present invention belongs to the field of artificial intelligence. The method is mainly based on the CoSec technology, and at the same time uses a new three-stage optimization process constructed based on knowledge distillation to solve and make up for the deficiencies of the CoSec technology, solves the deficiencies of the original CoSec in terms of security enhancement, and further improves the functional correctness of CoSec.
[0051] See Figures 1 - 4 , a method for enhancing the security of a code large model based on knowledge distillation and co-decoding, comprising the following steps:
[0052] Step S1: Select a basic model and an optimal model in any model family; the family model refers to a model with the same theoretical basic architecture and the same design principle. The basic model refers to the model with the simplest function and the fewest parameters in a family; the optimal model refers to the model with the most parameters and the most complete functions in the family; they share some core features and technologies, but are different in scale, function, and optimization direction, and will have certain application focuses.
[0053] Step S2: Select any code snippet dataset as the training set C, use the knowledge distillation method to align the functional correctness of the basic model and the optimal model, and then use the Adam optimizer to train the basic model after the functional correctness alignment. After the training is completed, the trained basic model Q after the functional correctness alignment is obtained; the knowledge distillation method is an existing technology;
[0054] The database used in the knowledge distillation method in step S2 is Magicoder, and Magicoder is an existing publicly available dataset.
[0055] The steps to obtain the trained basic model Q after the functional correctness alignment in step S2 are as follows:
[0056] Design the rKLD loss function And combine the cross loss function Design the loss function through weighting At the same time, set the maximum number of training rounds, use C as the input data of the basic model and the optimal model, and use the Adam optimizer to train the basic model. When the training reaches the maximum number of training rounds or the loss function no longer changes, the training ends, and the trained basic model Q after the functional correctness alignment is obtained. The calculation expression of is as follows:
[0057]
[0058] The calculation formula is as follows:
[0059]
[0060] The calculation expression is as follows:
[0061]
[0062] Among them, N represents the total number of training data in C, q θ represents the output of the base model, p is the output of the target model, and x i represents the i-th data in C, and X 1:i-1 represents the set of the first to the (i - 1)-th data in C, and λ is the ratio of, and its value range is [0, 1].
[0063] Step S3: Select several pairs of vulnerability data to form a vulnerability data set. Each pair of vulnerability data in the vulnerability data set includes the vulnerability type of the code snippet and the programming language type for editing the code snippet. The pair of vulnerability data refers to a pair of code snippets with the same attribute, one is correct and the other is vulnerable. In actual use, 9 key vulnerabilities highlighted in the MITRE Top25 report are used. Each key vulnerability contains a pair of functions before and after repair, as well as the corresponding editing points for the modification. The oversampling strategy is used to preprocess the vulnerability data set to obtain the security reinforcement training set D. The oversampling strategy is a prior art. The purpose of using this technology here is to solve the problem of unbalanced distribution of training data, which can improve the training effectiveness and also make the calculation of the trained model more accurate;
[0064] In the said step S3, the content of preprocessing the vulnerability data set with the oversampling strategy to obtain the security reinforcement training set D is as follows:
[0065] S3-1: Count the number of code snippets of each vulnerability type in the vulnerability data set, and calculate the average value n of all vulnerability types in the vulnerability data set. At the same time, count the number of programming languages used for editing the code snippets in each same vulnerability type, and set the programming language number threshold k;
[0066] S3-2: Compare the number of code snippets of each vulnerability type in the vulnerability data set with n. If the number of code snippets of a vulnerability type is less than n, then copy the existing code snippets of this vulnerability type to make the total number of code snippets of this vulnerability type equal to n; otherwise, keep the existing state unchanged;
[0067] If the number of code snippets in the d-th programming language in the same vulnerability type is less than k, randomly duplicate the number of code snippets so that the number of code snippets represented by the d-th programming language is equal to k; where the randomly duplicated code snippets are consistent with the vulnerability type and programming language of the code snippets represented by the d-th programming language; generally, the value of k is 10.
[0068] Step S4: Design the security reinforcement loss function Use D as the input of Q and adopt the low-rank adaptation method to Train Q to obtain the final security reinforcement basic model M; the low-rank adaptation method is a prior art;
[0069] In the said step S4, design the security reinforcement loss function The content is as follows:
[0070] S4-1: Construct using the cross-entropy loss function The calculation formula is as follows:
[0071]
[0072] Among them, Q θ represents the output of Q, and m i represents the security parameter corresponding to x i ;
[0073] Use the acceptance algorithm to calculate the value of m i The acceptance algorithm is a prior art. Set the acceptance threshold a. The value-taking process of m i is as follows:
[0074]
[0075] Among them, y n+1 represents the token value output by X when y1,..., y n are jointly used as the input of X; y' n+1 represents the token value output by X' when y1,..., y n are jointly used as the input of X'; represents the probability of generating the next token by y n+1 in X, represents the probability of generating the next token by y' n+1 in X';
[0076] If the calculated value of is greater than a, m i then takes 0; otherwise, m i then takes 1; its specific meaning is to specify whether the token at the current time step is classified as a security change: if the value is 1, it means it is a safe change, and if the value is 0, it means it is an unsafe change;
[0077] S4-2: Modify the rKLD function designed in the functional correctness alignment stage to obtain the function required in the secure training stage The calculation expression is as follows:
[0078]
[0079] where represents the predicted distribution of the security model at the current time step, represents the predicted distribution of the target model.
[0080] S4-3: Use and to obtain the final loss function required for secure training The calculation expression is as follows:
[0081]
[0082] In the step S4, the collaborative decoding strategy adopted by the security-enhanced basic model M is CoSec; CoSec belongs to the prior art.
[0083] Step S5: Select the target model X from the model family, and embed M into X to obtain the security-enhanced target model X';
[0084] Step S6: Arbitrarily select the target model Y to be security-enhanced, and repeat steps S2 - S5 to obtain the security-enhanced target model Y'.
[0085] Experiments and analysis:
[0086] In the experiments, the current basic code large models CodeGen, StarCoderBase, and DeepSeekCoder were selected for code security enhancement. They all have a complete open-source ecosystem and a large download volume, and at the same time have versions with different parameter sizes available for use, which are suitable for researching and testing the CoSec+ security enhancement effect. In terms of the dataset, the evol-codealpaca-v1 dataset was used as the training dataset for knowledge distillation. It was composed of adding code snippets to 80,000 documents screened from StarCoderData using the evolution-directive method and 75,000 data synthesized by GPT-4-0613, which can make the code snippets generated by the code large model more diverse, more realistic, and more controllable. 10k entries were selected as the test set in the experiment, and the remaining data was used as the training set. All experiments were completed using the PyTorch framework on an Ubuntu server with an NVIDIA A800 GPU.
[0087] During the functional correctness alignment phase of the experiment, for three different model families, the model with the smallest number of parameters was selected as the student model, namely CodeGen350M, StarCoderBase-1B, and DeepSeek-Coder-1.3b, while the model with the largest number of parameters was used as the teacher model, namely CodeGen-6.1B, StarCoderBase-7B, and DeepSeek-Code-6.7b. The experiment set the hyperparameters to a maximum of 3 epochs, a batch size of 24, a learning rate of 1e-5, gradient clipping of 1.0, λ of 0.5, a temperature of 1.0, and a maximum input text length of 1024.
[0088] In the subsequent secure training phase, it was necessary to perform secure training on the three student models that had already completed functional correctness alignment to serve as the basis for the secure model. The hyperparameters for this phase of training were set with the LoRA rank of 8, alpha of 16, and a dropout rate of 0.1 to achieve better training efficiency. The maximum input length remained 1024, the maximum number of epochs was limited to 8, the learning rate was set to 5e-5, and the round with the smallest loss value during training was used as the best checkpoint.
[0089] In the final inference phase using the collaborative decoding strategy, the number of sampling iterations was set to a fixed 25, the maximum number of newly generated tokens was 256, TOP-P was 0.95, the minimum prediction probability was 0, the acceptance threshold was 0.3, and the sampling temperature for both the secure model and the target base model was 0.4.
[0090] Finally, the evaluation and comparison of model performance found that CoSec+ can effectively guide large code models with different parameters to generate more secure and effective code in actual application scenarios. Specifically, when comparing with the current state-of-the-art CoSec in three large code models, for the CodeGen-350M model, CoSec+ achieved a 27.9% relative security improvement, 7.2% higher than CoSec; for the larger CodeGen-2.7B model, CoSec+ achieved a significant improvement of 37.7%, far exceeding CoSec's 19.7% improvement; for the CodeGen-6.1 model, which already has a relatively high security rate baseline of 68.2%, CoSec+ achieved an additional 12.3% relative security improvement, almost twice that of CoSec. Further, CoSec+ also has strong generalization ability in models of the StartCoder-Base series and DeepSeek-Coder series. For the StartCoder-Base-1B model, CoSec+ improved security by 7.7% more than CoSec, reaching 22.8%. For the StartCoder-Base-7B model, CoSec+ increased the security performance by 18.2%, showing an improvement of 15.8% beyond CoSec's 2.4%; for the DeepSeek-Coder-1.3B model, CoSec+ achieved a 16.8% relative security improvement, exceeding CoSec. For the DeepSeek-Cpder-6.7B model, CoSec+ significantly improved by 14.4%, exceeding CoSec's 4.5%. This indicates that CoSec+ can achieve good results in improving code security in a variety of code models. While improving the security of code generation, CoSec+ can maintain and even further improve the functional correctness of code generation: By using the Pass@1, Pass@5, and Pass@10 metrics of HumanEval as the evaluation of code generation correctness in actual application scenarios. Among them, the improvement of Pass@1 indicates a significant increase in the probability that the model can output both secure and correct code in the first generation, which has important value for reducing debugging costs in actual development scenarios. The continuous optimization of Pass@5 and Pass@10 verifies the model's ability to maintain stable output quality in multiple iterations. Compared with CoSec, after strengthening the security of the CodeGen, StarCoderBase, and DeepSeek-Coder series of large code models, CoSec+ increased Pass@1 by 5.2% to reach 78.2%, and Pass@5 and Pass@10 increased by 1.6% to reach 82.8%.This indicates that compared with traditional security hardening methods that average 3-8% loss of functional performance, CoSec+ can effectively secure large language models while taking into account functional correctness to a certain extent. Finally, CoSec+ can also effectively protect against security vulnerability types that the model has not learned before: by evaluating the security of CoSec+ against four vulnerability types not present in the security training dataset, it improved by 24.8% in CodeGen2.7B, by 1.3% in StarCoderBase-6B, and by 11.0% in DeepSeek-Coder. These results all demonstrate the excellent generalization ability of CoSec+ in security hardening, and it can still have a certain effective security hardening effect on vulnerability types that have never been trained.
[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A code model security reinforcement method based on knowledge distillation and co-decoding, characterized in that: It includes the following steps: Step S1: Select the base model and the best model in any model family; Step S2: Select any code snippet dataset as the training set C, use the knowledge distillation method to align the functional correctness of the base model and the best model, and then use the Adam optimizer to train the base model after the functional correctness alignment. After the training is completed, obtain the trained base model Q with the functional correctness alignment; Step S3: Select several vulnerability data pairs to form a vulnerability dataset. Each vulnerability data pair in the vulnerability dataset includes the vulnerability type of the code snippet and the programming language type of the edited code snippet. Use the oversampling strategy to preprocess the vulnerability dataset to obtain the security reinforcement training set D; Step S4: Design a security reinforcement loss function Take D as the input of Q and use the low-rank adaptation method Train Q to obtain the final security reinforcement basic model M; Step S5: Select the target model X from the model family, and embed M into X to obtain the security-reinforced target model X'; Step S6: Optionally select the target model Y to be security-reinforced, and repeat Step S2 - Step S5 to obtain the security-reinforced target model Y'.
2. The method for securely strengthening a code large model based on knowledge distillation and co-decoding according to claim 1, characterized in that: The database used in the knowledge distillation method in Step S2 is Magicoder.
3. The method for securely strengthening a large code model based on knowledge distillation and co-decoding according to claim 2, wherein: The steps to obtain the trained base model Q with the functional correctness alignment in Step S2 are as follows: Design the rKLD loss function and combine it with the cross-entropy loss function Design the loss function through weighting At the same time, set the maximum number of epochs. Use C as the input data for the base model and the best model. Train the base model using the Adam optimizer. When the training reaches the maximum number of epochs or the loss function no longer changes, the training ends, and the trained base model Q after functional correctness alignment is obtained; The calculation expression is as follows: The calculation formula is as follows: The calculation expression is as follows: Among them, N represents the total number of training data in C, and q θ represents the output of the base model, p is the output of the target model, and x i represents the i-th data in C, and X 1:i-1 represents the set of the first to the (i - 1)-th data in C, and λ is the ratio of, with a value range of [0, 1].
4. The method for securely strengthening a large code model based on knowledge distillation and co-decoding according to claim 3, wherein: In Step S3, the content of preprocessing the vulnerability dataset with the oversampling strategy to obtain the security reinforcement training set D is as follows: S3-1: Count the number of code snippets of each vulnerability type in the vulnerability dataset, and calculate the average value n of all vulnerability types in the vulnerability dataset; at the same time, count the number of programming languages used in the edited code snippets of each same vulnerability type, and set the programming language number threshold k; S3-2: Compare the number of code snippets of each vulnerability type in the vulnerability dataset with n. If the number of code snippets of a vulnerability type is less than n, then copy the existing code snippets of this vulnerability type so that the total number of code snippets of this vulnerability type is equal to n; otherwise, keep the existing state unchanged; If the number of the d-th programming language in a same vulnerability type is less than k, then randomly copy the number of code snippets so that the number of code snippets represented by the d-th programming language is equal to k; among them, the randomly copied code snippets are consistent with the vulnerability type and programming language of the code snippets represented by the d-th programming language.
5. The method for securely strengthening a code large model based on knowledge distillation and co-decoding according to claim 4, wherein: In step S4, a security reinforcement loss function is designed. The content is as follows: S4-1: Construct using the cross-entropy loss function The calculation formula is as follows: Among them, Q θ represents the output of Q, and m i represents the safety parameter corresponding to x i ; Calculate m using the acceptance algorithm i The value of m, set the acceptance threshold a, m i The value acquisition process is as follows: where y n+1 represents the token value output by X when y1, …, y n are jointly used as the input of X; y′ n+1 represents the token value output by X′ when y1, …, y n are jointly used as the input of X′; represents the probability that y n+1 generates the next token in X, represents the probability that y′ n+1 generates the next token in X′; If the calculated value of is greater than a, m i then take 0; otherwise, m i then take 1; S4-2: Modify the rKLD function designed in the functional correctness alignment phase to obtain the function required for the security training phase The calculation expression is as follows: in, represents the predicted distribution of the security model at the current time step, represents the prediction distribution of the target model; S4-3: Using and to obtain the loss function finally required for secure training The calculation expression is as follows:
6. The method for strengthening the security of a large code model based on knowledge distillation and co-decoding according to claim 5, characterized in that: In Step S4, the co-decoding strategy adopted for the security reinforcement of the base model M is CoSec.
Citation Information
Patent Citations
Lightweight source code vulnerability detection method based on knowledge distillation
CN115809464A
Self-training iterative artificial intelligence model training method
CN116628510A
Large model-oriented security reinforcement and credible generation method
CN118296610A
Large model generation content information security enhancement-oriented system and method
CN119558403A
Self-distillation training method and device for convolutional neural network, and scalable dynamic prediction method
WO2021023202A1
Cited By
A Secure Code Generation Method Based on Edit-Aware Large Language Model
CN122569901A