An automated code security optimization method based on large language models

Through an automated code security optimization method based on a large language model, combined with static analysis and semantic discriminator, vulnerabilities in code are identified and repaired, and an efficient closed-loop system is built, which solves the security and efficiency problems in code generation and realizes automated vulnerability detection and repair.

CN120046149BActive Publication Date: 2025-07-25SHANGHAI QITONG INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510528338.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-25
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

The existing technology fails to effectively combine security considerations when generating code, resulting in the generated code that poses security risks, and traditional vulnerability detection and repair methods are inefficient, making it difficult to meet the requirements of rapid iteration and high security of modern software.

Method used

A primary training model is built based on a large language model, combined with static analysis and semantic discriminator for code evaluation, identify and classify vulnerabilities through predefined vulnerabilities pattern collections, capture potential vulnerabilities using real-time monitoring, establish vulnerabilities and repair strategy libraries, optimize vulnerabilities time and repair algorithms, and build an efficient closed-loop code security optimization system.

Benefits of technology

It realizes vulnerability detection, location and repair in automated code generation, improves the security and efficiency of code generation, ensures that the generated code avoids security vulnerabilities during compilation and execution, and reduces manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046149B_ABST
    Figure CN120046149B_ABST
Patent Text Reader

Abstract

The present invention discloses an automated code security optimization method based on large language models, belonging to the technical field of code security optimization. Obtain large-scale programming data to establish a programming data set, build a primary training model, generate initial code and establish an initial code set; scan the initial code set with a predefined vulnerability pattern set, identify the vulnerabilities in the initial code set as initial vulnerabilities, and locate and classify the initial vulnerabilities; use static analysis in cooperation with a semantic discriminator to evaluate code snippets to obtain newly generated strategies, adjust the primary generation strategies to obtain an optimized training model; capture potential vulnerabilities as training signals, optimize the optimized training model to obtain a continuously optimized training model; establish a vulnerability library and a repair strategy library, and optimize the vulnerability time and repair algorithm through data analysis of the vulnerability library and the repair strategy library. The present invention can automatically identify and repair potential vulnerabilities in the generated code, without a large amount of manual intervention, and improve the security of code generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of code security optimization, and particularly relates to an automated code security optimization method based on a large language model. Background Art

[0002] In recent years, with the continuous expansion of software development scale and the increasingly strict requirements for network security, automated code generation technology has gradually become an important research direction in the industry. Currently, large-scale pre-trained language models (such as CodeT5, GPT series, etc.) show powerful capabilities in generating natural language and code. However, the current technological trend usually treats code generation separately from security considerations such as vulnerability detection and repair. This separation means that although the automatically generated code may be impeccable at the syntax level, it often implies security risks, such as buffer overflows, SQL injections, cross-site scripting attacks, type mismatches, and uninitialized variables. In addition, traditional vulnerability detection and repair methods mostly rely on static rules and manual intervention, making it difficult to efficiently handle complex and changing vulnerability scenarios and unable to meet the requirements of rapid iteration and high security of modern software. Summary of the Invention

[0003] Object of the Invention: To provide an automated code security optimization method based on a large language model, which solves the above problems existing in the prior art.

[0004] Technical Solution: An automated code security optimization method based on a large language model includes the following steps:

[0005] Obtain large-scale programming data to establish a programming data set, and a large language model f based on the Transformer architecture θ Build a primary training model, generate initial code, and establish an initial code set;

[0006] Scan the initial code set through a predefined set of vulnerability patterns, identify the vulnerabilities in the initial code set as initial vulnerabilities, and locate the initial vulnerabilities through symbolic execution and data flow analysis. According to the positioning results, combined with the predefined vulnerability categories in the predefined pattern set, classify the initial vulnerabilities;

[0007] During the process of generating the initial code, use static analysis combined with a semantic discriminator to evaluate the code fragments during the process of generating the initial code to obtain a new generation strategy, and adjust the primary generation strategy in the primary training model through the new generation strategy to obtain an optimized training model;

[0008] By monitoring the generated initial code in real time, capture potential runtime vulnerabilities as training signals, and optimize the optimized training model using the training signals to obtain a continuously optimized training model;

[0009] Record the initial vulnerabilities, runtime vulnerabilities, and the corresponding location and repair processes, establish a vulnerability database and a repair strategy database, and optimize the vulnerability time and repair algorithms through data analysis of the vulnerability database and the repair strategy database.

[0010] Preferably, the programming dataset is constructed as follows:

[0011] Obtain a large amount of programming data as the initial dataset D, where the initial dataset D = {(x i , y i )}, where x i represents the given input code, and y i represents the expected output code. Clean the input code x i and the output code y i , remove the invalid code segments in the input code x i and the output code y i , and use the tokenizer to split the input code x i and the output code y i into tokens. Tokenize each input code x i and the output code y i to convert the code into a token sequence, that is, generate the sequence dataset S → [token1, token2,..., token N ;

[0012] That is, the token sequence converted from the input code x i is: x i → [x i1 , x i2 ,..., x iTX , and the token sequence converted from the output code y i is: y i → [y i1 , y i2 ,..., y iTY , where TX and TY are the lengths of the input and output sequences respectively.

[0013] Preferably, in the architecture large language model f θ of the Transformer, θ is the weight parameter of the model. In the primary training model, the weight parameter θ represents the pre-trained weight in transfer learning. Among them, the calculation formula of the architecture large language model f θ is as follows:

[0014] f θ (x, y<t) = Softmax(W o h t );

[0015] Where: x represents the input problem description or partial code, and y<t represents the tokens generated before position t N-1 sequence, and t represents the position of generating the current token N in the sequence, and h t represents the hidden state at the t-th position calculated by the primary training model, and W o represents the weight matrix that maps the hidden state to the output vocabulary dimension, and Softmax represents the activation function to convert the output to the probability distribution of the next token N+1 sequence.

[0016] Preferably, the architecture of the Transformer large language model f θ The internal calculation process of building the primary training model is as follows:

[0017] Input the initial information a and the token sequence corresponding to the generated y<t into the primary training model, and calculate the dependency of each token sequence on other token N sequences through the multi-head self-attention mechanism, which is the self-attention layer. After non-linearly transforming the output of the self-attention with the feed-forward neural network, through residual connection and layer normalization, repeat L times to obtain the hidden state h corresponding to the position t t , map the hidden state to the dimension of the output vocabulary through the weight matrix, and then convert the output to the probability distribution of the next token through the Softmax activation function N+1 ;

[0018] Among them, the multi-head self-attention calculation formula is as follows:

[0019]

[0020] Where: Q represents the query vector of the current token N sequence, which is obtained by linearly transforming the input vector z i ; K represents the key vectors of all token sequences, which are obtained by linearly transforming the input vector z i ; V represents the value vectors of all token sequences, which are obtained from the input vector z through the corresponding linear transformation i ; d k represents the dimension of each attention head; T represents the length of a single token sequence;

[0021] Among them, the weight matrix formula is as follows:

[0022]

[0023] Where: n represents the corresponding sequence length calculated in formula (1), and d modelFor the dimension of the primary training model, through the weight matrix, a linear transformation is performed on the input X to obtain the query vector Q, the key vector K, and the value vector V, that is, the query vector Q = XW Q , the key vector K = XW K , the value vector V = XW V , where W represents the learning parameter matrix for linear transformation in the primary training model. After calculating the query vector Q, the key vector K, and the value vector V, substitute the query vector Q, the key vector K, and the value vector V into the Softmax activation function to generate the final attention output.

[0024] Preferably, by giving a real number vector z to the Softmax activation function, the Softmax activation function maps the real number vector z to obtain the corresponding probability vector, and the components are calculated through the following calculation formula:

[0025]

[0026] Where: is the exponential transformation of the input vector z i ; represents the sum of the exponential values of all elements of the real number vector z, and the sum of all output probabilities is 1; the real number vector z = z1, z2,..., z K ; i represents the i-th element in the real number vector z.

[0027] Preferably, during the process of converting the code into a token N sequence, given the previous token N-1 , the predicted current token N value is obtained through formula (2), where formula (2) is as follows:

[0028] P(Y t |x, Y<t) = f θ (x, Y<t) (2)

[0029] Where: P represents the probability distribution of the token N +1 sequence, Y t represents predicting the current token N sequence at time t, x represents the input code, Y represents the output token sequence, f θ represents the primary training model, and t represents the time of the current generated token N sequence;

[0030] The difference between the predicted value and the true value is calculated through the cross-entropy loss function, where the calculation formula of the cross-entropy loss function is as follows:

[0031]

[0032] Where: T Y represents the length of the output sequence tokens.

[0033] Preferably, the optimizer AdamW is used to update the parameters θ of the primary training model to increase the stable convergence during the training process of the primary training model. The formula for updating the parameters θ of the primary training model is as follows:

[0034]

[0035] where ηt is the learning rate, is the gradient of the cross-entropy loss with respect to the parameter θ, and during the training process, the training data is divided into N training batches, and forward propagation and backward propagation are performed on each training batch until an optimized training model is obtained. The formula for forward propagation is as follows:

[0036]

[0037] Where: represents the token sequence corresponding to y<t generated by the primary training model according to the input code x at time t;

[0038] The formula for backward propagation is as follows:

[0039]

[0040] Through the perplexity Perplexity, the generation effect of the model is evaluated on the validation set D val ={(x i ,y i )}, where the formula for the perplexity Perplexity is as follows:

[0041]

[0042] Where: N represents the total number of token sequences, l represents the index for traversing each token sequence, and P represents the probability generated for the l-th token sequence;

[0043] When the perplexity of the optimized training model on the validation set reaches the expected standard, save the current parameter θ * , and deploy it to the value reinforcement learning and vulnerability repair tasks as the basis for subsequent vulnerability detection and repair. The formula for the parameter θ * is as follows:

[0044]

[0045] Where: θ represents the weight parameter in the primary training model.

[0046] Preferably, the process of vulnerability detection is as follows:

[0047] Obtain a set of predefined vulnerability patterns, denoted as p, i.e., the set of predefined vulnerability patterns p = {p1, p2, …… p M} The severity weights corresponding to the set of predefined vulnerability patterns p are {ω1, ω2, …… ω M}, and a predefined vulnerability detection function V(x; φ), where the parameter φ in the vulnerability detection function includes at least the set of vulnerability patterns and the detection threshold. Use the detection function to detect each vulnerability p in the set of vulnerability patterns p M The detection results are as follows:

[0048]

[0049] During the process of detecting vulnerabilities in the set of vulnerability patterns P, a predefined severity function vulnerability classification function c M and vulnerability location function Use the predefined severity function to detect the severity of vulnerabilities in the set of vulnerability patterns p and output a value between 0 and 1 to reflect the severity of the vulnerability. Use the vulnerability location function π i (x) to determine the location of the vulnerability in the code x, and then use formula (3) to obtain the vulnerability detection report R(x). Formula (3) is as follows;

[0050]

[0051] In the formula: represents the set of precise vulnerability locations determined after symbolic execution and data flow analysis; c M ∈C represents the type of the vulnerability, and C is a predefined set of vulnerability categories;

[0052] At the same time, perform a weighted sum of the severity of each vulnerability to calculate the risk score corresponding to the vulnerability, which is used as the basis for automated vulnerability repair. Among them, the formula for the weighted sum of the severity of each vulnerability is as follows:

[0053]

[0054] Preferably, the static analysis process is as follows:

[0055] Define a basic error detection function, and use the error detection function to count the basic errors that occur in the generated code. Among them, the basic errors include at least syntax errors, type mismatches, and uninitialized variables. The calculation formula of the detection function is as follows:

[0056]

[0057] In the formula: yF The output code to be detected for errors is denoted as, and n represents the total number of output codes to be detected for errors, where I(e j ) represents the indication function for judging the output code, that is, when the j-th basic error e j is detected, I(e j ) = 1; conversely, I(e j ) = 0. The vulnerability pattern set p is obtained by collecting the detected basic errors, that is, the vulnerability pattern set p = {p1, p2,..., p i}}. For each vulnerability pattern p i the indication function Ip i (y) is defined. The overall vulnerability detection metric can be calculated by the following formula:

[0058]

[0059] A static analysis and compiler feedback function R static (y) is preliminarily determined, and its calculation formula is as follows:

[0060] R static (y) = -α·E basic (y) - β·E vuln ();

[0061] In the formula: α is the penalty weight for basic errors, α > 0; β is the penalty weight for vulnerability detection,

[0062] The number of incorrect output codes in the output code is obtained, and the error ratio R of the number of incorrect output codes to the total number of output codes is calculated;

[0063] When the error ratio R ≥ 5%, at this time, the result of the detection function E basic (y) increases, triggering a negative reward;

[0064] When the generated code has the same vulnerability as the known vulnerability pattern, E vuln (y) increases, generating a negative reward.

[0065] Preferably, the discrimination process of the semantic discriminator is as follows:

[0066] Pre-train a semantic model. The input code x i is the input context, y i is the output code, and the concatenation of the context and the production code is represented by ".". The semantic score is obtained through mapping by a multi-layer perceptron (MLP). The semantic score reflects the matching degree of the generated code with the expected function semantically. A high semantic score indicates that the generated code meets the requirements functionally; conversely, it indicates that the generated code does not meet the requirements functionally. The calculation formula of the semantic score is as follows:

[0067] S semantic (y) = tanh(MLP(z y ))), where

[0068] The score given by the preset vulnerability detection module is denoted as:

[0069]

[0070] In the formula: Indicates the indicator function for judging the output code content;

[0071] Takes the value 1 when a vulnerability pattern is detected in the generated code y, otherwise 0; Is the severity score of vulnerability p M In the generated code y; ω M Is the weight of vulnerability p M Reflects the importance of its risk, and when the score given by the preset vulnerability detection module is greater than the threshold, triggers the automated repair process;

[0072] Among them, the automated repair process is as follows:

[0073] The predefined repair function F takes the input and output code y and the vulnerability detection report R′(y), and after automatically adjusting the structure of the code according to the type of the vulnerability using the repair function F, outputs the repaired code y fix , where the repaired code y fix The calculation formula of is y fix = F(y, R′(y));

[0074] The predefined semantic discriminator feedback function is substituted with the semantic score and the score given by the vulnerability detection module to obtain the overall score. When the overall score is high,

[0075] Then the code has no security problems, otherwise, the code has security problems. Among them, the calculation formula of the predefined semantic discriminator feedback function is as follows:

[0076] R semantic (y) = S semantic (y) - λ vuln ·Risk(y);

[0077] In the formula: λ vuln Represents the adjustment parameter used to balance the weights of the semantic score and the penalty, λ vuln > 0;

[0078] The semantic score extracts the semantic features of the generated code through the pre-trained semantic model and the multi-layer perceptron, and is used to evaluate the consistency between its function and the expected function;

[0079] Penalty λ vuln ·Risk(y) applies a negative reward to the generated code with vulnerabilities according to the output of the vulnerability detection module;

[0080] Using the repaired code y fix It will be re-analyzed statically and semantically judged to obtain a comprehensive feedback. The calculation formula is as follows:

[0081] R total (y fix ) = R static (y fix ) + R semantic (y fix );

[0082] In the formula: R static (y fix ) represents the basic errors in the repaired code y fix , and R semantic (y fix ) represents the semantic score in the repaired code y fix ;

[0083] When R total (y fix ) reaches the preset security and function standards, it is regarded as a successful repair; otherwise, further repair or adjustment is triggered.

[0084] Beneficial effects: The present invention relates to an automated code security optimization method based on a large language model. Using a programming dataset to build a primary training model for a large language model f based on a Transformer architecture, enabling the primary training model to master the basic grammar and logical structure of the programming language, generating initial code that meets the syntax requirements, and through the collaborative scanning of the initial code by static analysis, compiler feedback, and semantic discriminator, using a predefined set of vulnerability patterns and corresponding severity weights to identify potential vulnerabilities in the initial code, generating a vulnerability library and a repair strategy library, and through data analysis of the vulnerability library and the repair strategy library, optimizing the vulnerability time and repair algorithm, realizing vulnerability detection, location, automatic repair, and re-verification in automated code generation, constructing an efficient and closed-loop code security optimization system, that is, it can automatically identify and repair potential vulnerabilities in the generated code without a large amount of manual intervention, improving the security and efficiency of code generation. θ

[0085] ​Secondly, static analysis is used in conjunction with a semantic discriminator to evaluate code snippets during the process of generating initial code, obtaining a newly generated strategy. The primary generation strategy in the primary training model is adjusted through the newly generated strategy, and then the initial code is regenerated for real-time monitoring to capture potential runtime vulnerabilities as training signals. The training signals are used to optimize the optimized training model to obtain a continuously optimized training model. Through reinforcement learning techniques, the generated code can adjust and optimize the strategy according to feedback at each step, gradually reducing vulnerabilities and errors, ensuring that the code meets security requirements during the generation process. The code generated using static analysis and the semantic discriminator can not only pass compilation checks but also ensure the avoidance of security vulnerabilities during actual execution, improving the security and reliability of the code. BRIEF DESCRIPTION OF THE DRAWINGS

[0086] Figure 1 It is a system block diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0087] As Figure 1 shown, the present invention provides a technical solution: an automated code security optimization method based on a large language model, including the following steps:

[0088] Obtain large-scale programming data to establish a programming data set. The process of building the programming data is as follows:

[0089] Obtain large-scale programming data as the initial data set D, where the initial data set D = {(x i , y i )}, where x i represents the given input code, y i represents the expected output code. Clean the input code x i and the output code y i , remove invalid code snippets from the input code x i and the output code y i , and use a tokenizer to split the input code x i and the output code y i into tokens. Tokenize each input code x i and output code y i , convert the code into a token sequence, that is, generate a sequence data set S → [token1, token2,..., token N ;

[0090] That is, the token sequence converted from the input code x i is: x i → [x i1 , x i2 ,..., x iTX , and the output code yi The transformed token sequence is: y i → [y i1 , y i2 , …, y iTY , where TX and TY are the lengths of the input and output sequences respectively. After the establishment of the programming dataset, based on the Transformer architecture large language model f θ Build a primary training model, generate initial code and establish an initial code set. In the large language model f with the Transformer architecture θ , θ is the weight parameter of the model. In the primary training model, the weight parameter θ represents the pre-trained weight in transfer learning. Among them, the calculation formula of the large language model f with the architecture θ is as follows:

[0091] f θ (x, y < t) = Softmax(W o h t )

[0092] In the formula: x represents the input problem description or partial code, y < t represents the token sequence generated before position t N-1 , t represents the position of generating the current token N sequence, h t represents the hidden state at the t-th position calculated by the primary training model, W o represents the weight matrix that maps the hidden state to the output vocabulary dimension, and Softmax represents the probability distribution of converting the activation function output to the next token N+1 sequence. Among them, the calculation process inside the large language model f with the Transformer architecture for building the primary training model is as follows: θ Input the initial information a and the token sequence corresponding to the generated y < t into the primary training model. Calculate the dependency relationship of each token sequence on other token

[0093] sequences through the multi-head self-attention mechanism, which is the self-attention layer. After non-linearly transforming the output of the self-attention with the feed-forward neural network, through residual connection and layer normalization, repeat L times to obtain the hidden state h N corresponding to the position t. Map the hidden state to the dimension of the output vocabulary through the weight matrix, and then convert the output to the probability distribution of the next token t through the Softmax activation function; N+1 Among them, the calculation formula of the multi-head self-attention is as follows:

[0094]

[0095] ​

[0096] Where: Q represents the current token N The query vector of the sequence, obtained by linearly transforming the input vector z i is obtained; K represents the key vectors of all token sequences, obtained by linearly transforming the input vector z i is obtained; V represents the value vectors of all token sequences, obtained from the input vector z through corresponding linear transformations i is obtained; d k represents the dimension of each attention head; T represents the length of a single token sequence;

[0097] Among them, the weight matrix formula is as follows:

[0098]

[0099] Where: n represents the corresponding sequence length calculated in formula (1), d model is the dimension of the primary training model. Through the weight matrix, the input X is linearly transformed to obtain the query vector Q, the key vector K, and the value vector V, that is, the query vector Q = XW Q , the key vector K = XW E , the value vector V = XW V , where W represents the learning parameter matrix for linear transformation in the primary training model. After calculating the query vector Q, the key vector K, and the value vector V, the query vector Q, the key vector K, and the value vector V are substituted into the Softmax activation function to generate the final attention output.

[0100] The initial code set is scanned through a predefined vulnerability pattern set, and the vulnerabilities in the initial code set are identified as initial vulnerabilities. Through symbolic execution and data flow analysis, the initial vulnerabilities are located. Based on the positioning results, combined with the predefined vulnerability categories in the predefined pattern set, the initial vulnerabilities are classified. Among them, by giving a real number vector z to the Softmax activation function, the Softmax activation function is used to map the real number vector z to obtain the corresponding probability vector, and the components are calculated through the following calculation formula:

[0101]

[0102] Where: is the exponential transformation of the input vector z i is performed; represents the sum of the exponential values of all elements of the real number vector z, and the sum of all output probabilities is 1; the real number vector z = z1, z2,..., z K ; i represents the i-th element in the real number vector z.

[0103] Preferably, the code is converted to tokensN During the sequence process, given the previous token N-1 , the predicted current token is obtained through formula (2) N , where formula (2) is as follows:

[0104] P(Y t |x, Y<t) = f θ (x, Y<t) (2)

[0105] In the formula: P represents the probability distribution of the token N +1 sequence, Y t represents predicting the current token N sequence at time t, x represents the input code, Y represents the output token sequence, f θ represents the primary training model, and t represents the time of the current generated token N sequence;

[0106] The difference between the predicted value and the true value is calculated through the cross-entropy loss function. The formula for the cross-entropy loss function is as follows:

[0107]

[0108] In the formula: T Y represents the length of the output sequence token. After calculating the difference between the predicted value and the true value, the optimizer AdamW is used to update the parameters θ of the primary training model to increase the stable convergence during the training process of the primary training model. The formula for updating the parameters θ of the primary training model is as follows:

[0109]

[0110] where ηt is the learning rate, is the gradient of the cross-entropy loss with respect to the parameter θ. During the training process, the training data is divided into N training batches, and forward propagation and backward propagation are performed on each training batch until an optimized training model is obtained. The formula for forward propagation is as follows:

[0111]

[0112] In the formula: represents at time t, the token sequence corresponding to the input code x and the already generated Y<t by the primary training model;

[0113] The formula for backward propagation is as follows:

[0114]

[0115] Evaluate the generation effect of the model through Perplexity on the validation set D val ={(x t , y t )}, where the calculation formula of Perplexity is as follows:

[0116]

[0117] In the formula: N represents the total number of token sequences, l represents the index for traversing each token sequence, and P represents the probability generated for the l-th token sequence;

[0118] When the Perplexity of the optimized training model on the validation set reaches the expected standard, save the current parameter θ * , and deploy it to the value reinforcement learning and vulnerability repair tasks as the basis for subsequent vulnerability detection and repair. The calculation formula of the parameter θ * is as follows:

[0119]

[0120] In the formula: θ represents the weight parameter in the primary training model. After calculating the parameter θ * , perform vulnerability detection. The process of vulnerability detection is as follows:

[0121] Obtain the predefined vulnerability pattern set, denoted as p, that is, the predefined vulnerability pattern set p = {p1, p2,... p M}, and the corresponding severity weights of the predefined vulnerability pattern set p are {ω1, ω2,... ω M}, and the predefined vulnerability detection function V(x; φ), where the parameters φ in the vulnerability detection function include at least the vulnerability pattern set and the detection threshold. Use the detection function to detect each vulnerability p M in the vulnerability pattern set p. The detection results are as follows:

[0122]

[0123] During the process of detecting vulnerabilities in the vulnerability pattern set p, simultaneously predefined the severity function the vulnerability classification function c M and the vulnerability location function Use the predefined severity function to detect the severity of vulnerabilities in the vulnerability pattern set P and output a value between 0 and 1 to reflect the severity of the vulnerability. Use the vulnerability location function π M (x) to determine the location of the vulnerability in the code x, and then use formula (3) to obtain the vulnerability detection report R(x). Formula (3) is as follows;

[0124]

[0125] In the formula: represents the set of precise vulnerability locations determined after symbolic execution and data flow analysis; c M ∈C represents the type of vulnerability, and C is a predefined set of vulnerability categories;

[0126] Meanwhile, the severity of each vulnerability is weighted and summed to calculate the risk score corresponding to the vulnerability, which is used as the basis for automated vulnerability repair. Among them, the formula for weighted summation of the severity of each vulnerability is as follows:

[0127] After completing the weighted summation of the severity of vulnerabilities, the predefined set of vulnerability patterns and the corresponding severity weights are used to identify potential vulnerabilities in the initial code, generating a vulnerability library and a repair strategy library. By performing data analysis on the vulnerability library and the repair strategy library, the vulnerability time and repair algorithm are optimized, realizing vulnerability detection, location, automatic repair, and re-verification in automated code generation, and constructing an efficient and closed-loop code security optimization system, which can automatically identify and repair potential vulnerabilities in the generated code without a large amount of manual intervention, improving the security and efficiency of code generation.

[0128] Among them, during the process of generating the initial code, static analysis is used in combination with a semantic discriminator to evaluate code fragments during the process of generating the initial code to obtain a newly generated strategy, and the primary generation strategy in the primary training model is adjusted through the newly generated strategy to obtain an optimized training model. Among them, the static analysis process is as follows:

[0129] Define a basic error detection function, and count the basic errors that occur in the generated code through the error detection function. Among them, the basic errors at least include syntax errors, type mismatches, and uninitialized variables. The calculation formula of the detection function is as follows:

[0130]

[0131] In the formula: y F represents the output code for which errors are to be detected, n represents the total number of output codes for which errors are to be detected, I(e j ) represents an indicator function for judging the output code, that is, when the jth basic error e j is detected, I(e j ) = 1, otherwise, I(e j ) = 0. Obtain the detected basic errors to establish a set of vulnerability patterns p, that is, the set of vulnerability patterns p = {p1, p2,..., p M}}, and define an indicator function Ip for each vulnerability pattern p M M ​(y), that is, the overall vulnerability detection index can be calculated through the following formula:

[0132]

[0133] Pre-determine the static analysis and compiler feedback function R static (y), and its calculation formula is as follows:

[0134] R static (y) -- α·E basic (y) - β·E vuln (y)

[0135] In the formula: α is the penalty weight for basic errors, is the penalty weight for vulnerability detection,

[0136] Obtain the number of incorrect output codes in the output code, and calculate the error ratio R of the number of incorrect output codes to the total number of output codes;

[0137] When the error ratio R ≥ 5%, at this time, the result of the detection function E basic (y) increases, triggering a negative reward;

[0138] When a vulnerability identical to the known vulnerability pattern appears in the generated code, E vuln (y) increases, generating a negative reward.

[0139] The discrimination process of the semantic discriminator is as follows:

[0140] Pre-train the semantic model, input the code x i is the input context, y i is the output code, and the concatenation of the context and the production code is represented by ".", and the semantic score is obtained through mapping by a multi-layer perceptron (MLP). The semantic score reflects the matching degree of the generated code with the expected function semantically. A high semantic score indicates that the generated code meets the requirements functionally, otherwise, it indicates that the generated code does not meet the requirements functionally. The calculation formula of the semantic score is as follows:

[0141] S semantic (y) = tanh(MLP(z y ), where

[0142] The score given by the preset vulnerability detection module is denoted as:

[0143]

[0144] In the formula: represents the indicator function for judging the output code content;

[0145] It takes the value of 1 when a vulnerability pattern is detected in the generated code y, otherwise it is 0; For vulnerability p M The severity score in the generated code y; w M For vulnerability p M The weight of, reflecting the importance of its risk, and when the score given by the preset vulnerability detection module is greater than the threshold, it triggers an automated repair process to repair the detected vulnerability. Among them, the automated repair process is as follows:

[0146] A predetermined repair function F, which takes the input and output code y and the vulnerability detection report R′(y), and uses the repair function F to automatically adjust the structure of the code according to the type of the vulnerability, and then outputs the repaired code y fix , where the repaired code y fix The calculation formula of is y fix =P(y, R t (y));

[0147] A pre-determined semantic discriminator feedback function is substituted with the semantic score and the score given by the vulnerability detection module to obtain an overall score. When the overall score is high,

[0148] Then the code has no security issues. Otherwise, the code has security issues. Among them, the calculation formula of the pre-determined semantic discriminator feedback function is as follows:

[0149] R semantic (y)=S semantic (y)λ vuln ·Risk(y)

[0150] In the formula: λ vuln Represents an adjustment parameter used to balance the weights of the semantic score and the penalty. λ vuln >0;

[0151] The semantic score is used to extract the semantic features of the generated code through a pre-trained semantic model and a multi-layer perceptron, and is used to evaluate the consistency between its function and the expected function;

[0152] The penalty λ vuln ·Risk(y) applies a negative reward to the generated code with vulnerabilities according to the output of the vulnerability detection module;

[0153] Using the repaired code y fix Will be re-analyzed statically and semantically judged to obtain a comprehensive feedback. The calculation formula is as follows:

[0154] R total (y fix )=R static (y fix )+R semantic (yfix )

[0155] Where: R static (y fix ) represents the basic error in the repaired code y fix ; R semantic (y fix ) represents the semantic score in the repaired code y fix ;

[0156] When R total (y fix ) reaches the preset safety and function standards, it is regarded as successful repair; otherwise, further repair or adjustment is triggered. By monitoring the generated initial code in real time, potential runtime vulnerabilities are captured as training signals. The optimized training model is optimized using the training signals to obtain a continuously optimized training model. The initial vulnerabilities, runtime vulnerabilities, and the corresponding location and repair processes during the optimization process are recorded to establish a vulnerability library and a repair strategy library. By analyzing the data in the vulnerability library and the repair strategy library, the vulnerability time and repair algorithm are optimized, that is, static analysis is combined with a semantic discriminator to evaluate the code fragments during the generation of the initial code to obtain a newly generated strategy. The primary generation strategy in the primary training model is adjusted using the newly generated strategy, and then the initial code is regenerated for real-time monitoring. Potential runtime vulnerabilities are captured as training signals, and the optimized training model is optimized using the training signals to obtain a continuously optimized training model. Through reinforcement learning technology, the generated code can adjust and optimize the strategy according to the feedback at each step, gradually reducing vulnerabilities and errors, ensuring that the code meets the safety requirements during the generation process. The code generated using static analysis and the semantic discriminator can not only pass the compilation check but also ensure the avoidance of security vulnerabilities during actual execution, improving the security and reliability of the code.

[0157] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all belong to the protection scope of the present invention.

Claims

1. An automated code security optimization method based on large language models, characterized in that, It includes the following steps: Obtain large-scale programming data to establish a programming dataset, and a large language model f based on the Transformer architecture θ Build a primary training model, generate initial code and establish an initial code set; Scan the initial code set through a predefined vulnerability pattern set, identify the vulnerabilities in the initial code set as initial vulnerabilities, locate the initial vulnerabilities through symbolic execution and data flow analysis, and classify the initial vulnerabilities according to the located results and the predefined vulnerability categories in the predefined pattern set; During the process of generating the initial code, use static analysis in cooperation with a semantic discriminator to evaluate the code fragments during the process of generating the initial code to obtain a newly generated strategy, and adjust the primary generation strategy in the primary training model through the newly generated strategy to obtain an optimized training model; Through real-time monitoring of the generated initial code, capture potential runtime vulnerabilities as training signals, and optimize the optimized training model with the training signals to obtain a continuously optimized training model; Record the initial vulnerabilities, runtime vulnerabilities, and the corresponding location and repair processes, establish a vulnerability library and a repair strategy library, and optimize the vulnerability time and repair algorithm through data analysis of the vulnerability library and the repair strategy library.

2. The automated code security optimization method based on a large language model according to claim 1, wherein The process of building a programming data set is as follows: Obtain large-scale programming data as the initial dataset D, where the initial dataset D = {(x i , y i )}, where x i represents the given input code, and y i represents the expected output code. Clean the input code x i and the output code y i , remove the invalid code segments in the input code x i and the output code y i , and use the tokenizer to split the input code x i and the output code y i into tokens. Tokenize each input code x i and the output code y i , convert the code into a token sequence, that is, generate the sequence dataset S → [token1, token2, …, token N ; That is, the input code is x i The transformed token sequence is: x i → [x i1 , x i2 , …, x iTX , and the output code is y i The transformed token sequence is: y i → [y i1 , y i2 , …, y iTY , where TX and TY are the lengths of the input and output sequences respectively.

3. The automated code security optimization method based on a large language model according to claim 2, wherein, The architecture large language model f of the Transformer θ where θ is the weight parameter of the model. In the primary training model, the weight parameter θ represents the pre-trained weight in transfer learning. Among them, the architecture large language model f θ has the following calculation formula: f θ (x,y<t) = Softmax(W o h t ); Where: x represents the input problem description or partial code, and y<t represents the tokens that have been generated before position t N-1 sequence, and t represents the generation of the current token N position in the sequence, h t represents the hidden state at the t-th position calculated by the primary training model, W o represents the weight matrix that maps the hidden state to the output vocabulary dimension, and Softmax represents the activation function output converted to the next token N+1 probability distribution of the sequence.

4. An automated code security optimization method based on a large language model according to claim 3, characterized in that, Transformer-based architecture large language model f θ The internal calculation process of building the primary training model is as follows: The initial information a and the token sequence corresponding to the generated y<t> are input into the primary training model. The dependency of each token sequence on other token sequences is calculated through the multi-head self-attention mechanism, which is the self-attention layer. After non-linearly transforming the output of the self-attention with the feed-forward neural network, through residual connection and layer normalization, after repeating L times, the hidden state h corresponding to the position t is obtained. N The dependency of each token sequence on other token sequences is calculated through the multi-head self-attention mechanism, which is the self-attention layer. After non-linearly transforming the output of the self-attention with the feed-forward neural network, through residual connection and layer normalization, after repeating L times, the hidden state h corresponding to the position t is obtained. t The hidden state is mapped to the dimension of the output vocabulary through the weight matrix, and then the output is converted into the probability distribution of the next token through the Softmax activation function. N+1 ; Among them, the calculation formula of the multi-head self-attention is as follows: Where: Q represents the current token N The query vector of the sequence, obtained by linearly transforming the input vector z i ; K represents the key vectors of all token sequences, obtained by linearly transforming the input vector z i ; V represents the value vectors of all token sequences, obtained from the input vector z through the corresponding linear transformation i ; d k represents the dimension of each attention head; T represents the length of a single token sequence Among them, the formula of the weight matrix is as follows: Where: n represents the corresponding sequence length calculated in formula (1), d model is the dimension of the primary training model. Through the weight matrix, a linear transformation is performed on the input X to obtain the query vector Q, the key vector K, and the value vector V, that is, the query vector Q = XW Q , the key vector K = XW K , the value vector V = XW V . Wherein, W represents the learning parameter matrix for linear transformation in the primary training model. After calculating the query vector Q, the key vector K, and the value vector V, the query vector Q, the key vector K, and the value vector V are substituted into the Softmax activation function to generate the final attention output.

5. An automated code security optimization method based on a large language model according to claim 4, characterized in that, Among them, by giving a real number vector z to the Softmax activation function, the real number vector z is mapped by the Softmax activation function to obtain a corresponding probability vector, and the components are calculated through the following calculation formula: Wherein: performs an exponential transformation on the input vector z i ; represents the sum of the exponential values of all elements of the real vector z, and the sum of all output probabilities is 1; the real vector z = z1, z2,..., z K ; i represents the i-th element in the real vector z.

6. An automated code security optimization method based on a large language model according to claim 3, characterized in that, Convert the code into tokens N During the sequence process, given the previous token N-1 , the predicted current token is obtained through formula (2) N The value of which, formula (2) is as follows: P(Y t | x, Y < t) = f g (x, Y < t) (2); Where: P represents the token N Probability distribution of the +1 sequence, Y t represents predicting the current token N sequence at time t, x represents the input code, Y represents the output token sequence, f θ represents the primary training model, t represents the time of the currently generated token N sequence; Calculate the difference between the predicted value and the true value through the cross-entropy loss function, where the calculation formula of the cross-entropy loss function is as follows: where: T Y represents the length of the output sequence tokens.

7. An automated code security optimization method based on a large language model according to claim 3, characterized in that Use the optimizer AdamW to update the parameters θ of the primary training model to increase the stable convergence during the training process of the primary training model, where the formula for updating the parameters θ of the primary training model is as follows: where ηt is the learning rate, is the gradient of the cross-entropy loss with respect to the parameter θ, and during the training process, the training data is divided into N training batches, and forward propagation and backward propagation are performed on each training batch until an optimized training model is obtained. The calculation formula for forward propagation is as follows: Wherein: represents, at time t, the token sequence corresponding to Y<t generated by the primary training model according to the input code x; The calculation formula of backpropagation is as follows: Evaluate the generation effect of the model through Perplexity on the validation set D val ={(x i , y i )}, where the calculation formula of Perplexity is as follows: In the formula: N represents the total number of token sequences, l represents the index for traversing each token sequence, and P represents the probability generated for the l-th token sequence; When the perplexity of the optimized training model on the validation set reaches the expected standard, save the current parameter θ * , and deploy it in value reinforcement learning and vulnerability repair tasks as the basis for subsequent vulnerability detection and repair, where the parameter θ * is calculated as follows: In the formula: θ represents the weight parameters in the primary training model.

8. An automated code security optimization method based on a large language model according to claim 6, characterized in that, The process of vulnerability detection is as follows: Obtain a set of predefined vulnerability patterns, denoted as p, i.e., the set of predefined vulnerability patterns p = {p1, p2, …… p M}, and the severity weights corresponding to the set of predefined vulnerability patterns p are {ω1, ω2, …… ω M}. There is a predefined vulnerability detection function V(x; φ), where the parameter φ in the vulnerability detection function includes at least the set of vulnerability patterns and the detection threshold. Use the detection function to detect each vulnerability p in the set of vulnerability patterns p M The detection results are as follows: During the vulnerability detection process for the vulnerability pattern set p, a severity function is predefined simultaneously vulnerability classification function c M and the vulnerability location function Using the predefined severity function perform severity detection on the vulnerabilities in the vulnerability pattern set p and output a value between 0 and 1, reflecting the severity of the vulnerability. Use the vulnerability location function π i (x) to determine the location of the vulnerability in the code x, and then use formula (3) to obtain the vulnerability detection report R(x). Formula (3) is as follows; In the formula: represents the set of precise vulnerability locations determined after symbolic execution and data flow analysis; c M ∈C represents the type of vulnerability, where C is a pre-defined set of vulnerability categories; At the same time, perform a weighted sum of the severity of each vulnerability to calculate the risk score corresponding to the vulnerability, which is used as the basis for automated vulnerability repair. Among them, the formula for performing a weighted sum of the severity of each vulnerability is as follows:

9. An automated code security optimization method based on a large language model according to claim 1, characterized in that, The process of static analysis is as follows: Define a basic error detection function, and count the basic errors that appear in the generated code through the error detection function. Among them, the basic errors at least include syntax errors, type mismatches, and uninitialized variables. The calculation formula of the detection function is as follows: where: y F represents the output code to be detected for errors, n represents the total number of output codes to be detected for errors, I(e j ) represents the indication function for judging the output code, that is, when the j-th basic error e j is detected, I(e j ) = 1, otherwise, I(e j ) = 0, obtain the detected basic error to establish the vulnerability pattern set p, that is, the vulnerability pattern set p = {p1, p2,..., p M}}, for each vulnerability pattern p M define the indication function Ip M (y), that is, the overall vulnerability detection index can be calculated through the following formula: Pre-determined static analysis and compiler feedback function R static (y), and its calculation formula is as follows: R static f(y) = -α·E basic f(y) - β·E vuln f(y); Where: α is the penalty weight for basic errors, α > 0; β is the penalty weight for vulnerability detection, Obtain the number of output codes with errors in the output code, and calculate the error ratio R of the number of output codes with errors to the total number of output codes; When the error ratio R ≥ 5%, at this time, the result of the detection function E basic (y) increases, triggering a negative reward; When a vulnerability identical to a known vulnerability pattern appears in the generated code, E vuln (y) increases, generating a negative reward.

10. An automated code security optimization method based on a large language model according to claim 8, characterized in that, The discrimination process of the semantic discriminator is as follows: Pre-trained semantic model, input code x i is the input context, y i is the output code, and the concatenation of the context and the production code is represented by ".", and the semantic score is obtained through mapping by a multi-layer perceptron (MLP). The semantic score reflects the matching degree of the generated code with the expected function semantically. A high semantic score indicates that the generated code meets the requirements functionally, otherwise, it indicates that the generated code does not meet the requirements functionally. The calculation formula of the semantic score is as follows: S semantic (y) = tanh(MLP(z y ))), where The score given by the preset vulnerability detection module is denoted as: In the formula: Indicates an indicator function for judging the output code content; Take the value 1 when a vulnerability pattern is detected in the generated code y, otherwise 0; for vulnerability p M the severity score in the generated code y; ω M for vulnerability p M the weight, reflecting the importance of its risk, and triggering an automated repair process when the score given by the preset vulnerability detection module is greater than the threshold; Among them, the automatic repair process is as follows: Given a predefined repair function F, with input-output code y and vulnerability detection report R′(y), after automatically adjusting the code structure according to the type of vulnerability using the repair function F, the repaired code y is output. fix , where the repaired code y fix has the calculation formula y fix = F(y, R′(y)); Pre-determine the semantic discriminator feedback function, substitute the semantic score and the score given by the vulnerability detection module to obtain the overall score. When the overall score is high, The code has no security issues. Otherwise, the code has security issues. Among them, the calculation formula of the pre-determined semantic discriminator feedback function is as follows: R semantic (y) = S semantic (y) - λ vuln ·Risk(y); Where: λ vuln represents a tuning parameter for balancing the weights of semantic scores and penalties, and λ vuln > 0; Semantic scoring extracts the semantic features of the generated code through a pre-trained semantic model and a multi-layer perceptron to evaluate the consistency between its function and the expected function; Penalty λ vuln ·Risk(y) applies a negative reward to the generated code with vulnerabilities according to the output of the vulnerability detection module; Using the repaired code y fix It will be statically analyzed and semantically evaluated again to obtain comprehensive feedback. The calculation formula is as follows: R total (y fix ) = R static (y fix ) + R semantic (y fix ); Where: R static (y fix ) represents the basic error in the repaired code y fix , and R semantic (y fix ) represents the semantic score in the repaired code y fix ; When R total (y fix ) meets the preset safety and functional standards, it is considered a successful repair; otherwise, further repair or adjustment is triggered.

Citation Information

Patent Citations

  • Code vulnerability detection large model construction method and device and electronic equipment

    CN118171291A

  • Software vulnerability automatic repair method based on thinking chain and storage medium

    CN119128893A