Automatic code security optimization method based on large language model

Through an automated code security optimization method based on a large language model, combined with static analysis and semantic discrimination, vulnerabilities in the code are identified and fixed, and the problem of difficult code generation in the existing technology is to ensure syntax accuracy and security, and efficient and automatic code security optimization is achieved.

CN120046149AActive Publication Date: 2025-05-27SHANGHAI QITONG INFORMATION TECH CO LTD +1

Patent Information

Application Number
CN202510528338.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-27
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

The prior art is difficult to ensure the syntax accuracy and security of the code at the same time in automated code generation, especially when facing complex and changeable vulnerability scenarios, it cannot meet the requirements of rapid iteration and high security of modern software.

Method used

Using automated code security optimization methods based on large language models, we use large language models of Transformer architecture to generate initial code and perform static analysis and semantic discrimination, identify and repair potential vulnerabilities, establish vulnerability libraries and repair policy libraries, and optimize vulnerability detection and repair algorithms.

Benefits of technology

It realizes vulnerability detection, location, automatic repair and re-verification in automated code generation, and builds an efficient, closed-loop code security optimization system that can automatically identify and repair potential vulnerabilities in generated code, improving the security and efficiency of code generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046149A_ABST
    Figure CN120046149A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic code security optimization method based on a large language model, and belongs to the technical field of code security optimization. Acquiring large-scale programming data to establish a programming data set, establishing a primary training model, generating an initial code and establishing an initial code set; the predefined vulnerability mode set scans the initial code set, vulnerabilities in the initial code set are identified and defined as initial vulnerabilities, and the initial vulnerabilities are positioned and classified; static analysis is matched with a semantic discriminator to evaluate the code snippets to obtain a new generation strategy, and the primary generation strategy is adjusted to obtain an optimized training model; capturing potential vulnerabilities as training signals, and optimizing the optimization training model to obtain a continuous optimization training model; and establishing a vulnerability library and a repair strategy library, and optimizing vulnerability time and a repair algorithm by analyzing data of the vulnerability library and the repair strategy library. Potential vulnerabilities in generated codes can be automatically recognized and repaired, a large amount of manual intervention is not needed, and code generation safety is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of code security optimization, and in particular relates to an automatic code security optimization method based on a large language model. Background Art

[0002] In recent years, with the continuous expansion of software development scale and increasingly stringent network security requirements, automatic code generation technology has gradually become an important research direction in the industry. CodeT 5. GPT series, etc.) have shown strong capabilities in generating natural language and code. However, the current technical trend is usually to separate code generation from security considerations such as vulnerability detection and repair. This separation means that although the automatically generated code may be impeccable at the syntactic level, it often contains security risks such as buffer overflows, SQL Injection, cross-site scripting attacks, type mismatches, and uninitialized variables. In addition, traditional vulnerability detection and repair methods mostly rely on static rules and manual intervention, which are difficult to efficiently deal with complex and changing vulnerability scenarios and cannot meet the needs of modern software rapid iteration and high security requirements. Summary of the invention

[0003] Purpose of the invention: To provide an automated code security optimization method based on a large language model to solve the above-mentioned problems existing in the prior art.

[0004] Technical solution: An automated code security optimization method based on a large language model, comprising the following steps: Acquire large-scale programming data to build a programming dataset based on Transformer Architecture of Large Language Model f θ Build a primary training model, generate initial code and establish an initial code set; Scan the initial code set through the predefined vulnerability pattern set, identify the vulnerabilities in the initial code set as initial vulnerabilities, locate the initial vulnerabilities through symbolic execution and data flow analysis, and classify the initial vulnerabilities based on the positioning results and the predetermined vulnerability categories in the predefined pattern set; In the process of generating the initial code, static analysis is used in conjunction with a semantic discriminator to evaluate the code snippets in the process of generating the initial code to obtain a new generation strategy, and the primary generation strategy in the primary training model is adjusted through the new generation strategy to obtain an optimized training model; By real-time monitoring of the generated initial code, potential runtime vulnerabilities are captured as training signals, and the training signals are used to optimize the optimized training model to obtain a continuously optimized training model; Record the initial vulnerabilities, runtime vulnerabilities, and the corresponding location and repair processes, establish a vulnerability library and a repair strategy library, and optimize the vulnerability time and repair algorithms by analyzing the data in the vulnerability library and the repair strategy library.

[0005] Preferably, the programming dataset is built as follows: Obtain large-scale programming data as the initial dataset D , where the initial dataset D = {( x i , y i )}, where, x i represents the given input code, y i represents the expected output code. Clean the input code x i and the output code y i , remove the invalid code segments in the input code x i and the output code y i , and use tokenizer to split the input code x i and the output code y i into tokens . For each input code x i and the output code y i perform token conversion to convert the code into token sequences, that is, generate the sequence dataset S → token 1 , token 2 , …, token N ; That is, the x i sequence converted from the input code is: token x i → x i1 , x i2 , …, x iTX , and the y i sequence converted from the output code is: token y i→ y i1 , y i2 ,…, y iTY , where TX and TY are the lengths of the input and output sequences respectively. and TY are the lengths of the input and output sequences respectively.

[0006] Preferably, in the architecture large language model Transformer , f θ is the weight parameter of the model. In the primary training model, the weight parameter θ represents the pre-trained weight in transfer learning. Among them, the calculation formula of the architecture large language model f θ is as follows: Transformer the architecture large language model f θ in θ is the weight parameter of the model. In the primary training model, the weight parameter θ represents the pre-trained weight in transfer learning. Among them, the calculation formula of the architecture large language model f θ is as follows: θ represents the pre-trained weight in transfer learning. Among them, the calculation formula of the architecture large language model f θ is as follows: f θ The calculation formula is as follows: ; In the formula: x represents the input problem description or partial code, y < t represents at position t the token N-1 sequence that has been generated before, t represents the position where the current token N sequence is generated, h t represents the hidden state at the t th position calculated by the primary training model, W o represents the weight matrix that maps the hidden state to the output vocabulary dimension, Softmax represents the conversion of the activation function output to the probability distribution of the next token N+1 sequence.

[0007] Preferably, Transformer the architecture large language model f θ The internal calculation process of building the primary training model is as follows: Input the initial information a and the y < t corresponding token sequence to the primary training model. Calculate the dependency relationship of each token sequence on other token N sequences through the multi-head self-attention mechanism, which is the self-attention layer. After non-linearly transforming the output of the self-attention with the feed-forward neural network, through residual connection and layer normalization, repeat LAfter that, the hidden state corresponding to the position t is obtained h t , and the hidden state is mapped to the dimension of the output vocabulary through a weight matrix, and then through Softmax the activation function, the output is converted into the probability distribution of the next token N+1 ; Among them, the formula for multi-head self-attention is as follows: (1); In the formula: Q represents the query vector of the current token N sequence, which is obtained by linearly transforming the input vector z i ; K represents the key vectors of all token sequences, which are obtained by linearly transforming the input vector z i ; V represents the value vectors of all token sequences, which are obtained from the input vector z i through the corresponding linear transformation; represents the dimension of each attention head; T represents a single token sequence length; Among them, the formula for the weight matrix is as follows: ; In the formula: n represents the corresponding sequence length calculated in formula (1), is the dimension of the primary training model. Through the weight matrix, the input x is linearly transformed to obtain the query vector Q , the key vector K and the value vector V , that is, the query vector , the key vector , the value vector . In the formula, W represents the learning parameter matrix for linear transformation in the primary training model. After calculating the query vector Q , the key vector K and the value vector V , substitute the query vector Q , the key vector K and the value vector V into Softmax the activation function to generate the final attention output.

[0008] Preferably, by applying Softmax an activation function to a real number vector z , using Softmax the activation function to map the real number vector z to obtain a corresponding probability vector, and calculating the components through the following calculation formula: ; In the formula: is the exponential transformation of the input vector z i ; represents the sum of the exponential values of all elements of the input vector z, and the sum of all output probabilities is 1; the real number vector z = z 1 , z 2 ,..., z K ; i represents the z th element in the input vector i .

[0009] Preferably, in the process of converting the code into token N sequence, given the previous token N-1 , the predicted current token N value is obtained through formula (2), where formula (2) is as follows: (2) In the formula: P represents the token N + 1 sequence probability distribution, y t represents predicting the current t at time token N sequence, x represents the input code, y represents the output token sequence, f θ represents the primary training model, t represents the current generation token N sequence time; The difference between the predicted value and the true value is calculated through the cross-entropy loss function, where the cross-entropy loss function calculation formula is as follows: ; In the formula: T y represents the length of the output sequence token , y represents the number of output sequences token .

[0010] Preferably, an optimizer AdamW is used to update the parameters θ of the primary training model, and the stability and convergence during the training process of the primary training model are increased. The formula for updating the parameters θ of the primary training model is as follows: ; where η t is the learning rate, is the gradient of the cross-entropy loss with respect to the parameter θ, and during the training process, the training data is divided into N training batches, and forward propagation and backward propagation are performed on each training batch until an optimized training model is obtained. The calculation formula for forward propagation is as follows: ; In the formula: represents at time t , the primary training model generates a sequence x corresponding to the input code y < t ; token sequence; The calculation formula for backward propagation is as follows: ; The generation effect of the model is evaluated in the validation set Perplexity through the perplexity . The calculation formula for the perplexity Perplexity is as follows: ; In the formula: N represents token the total number of sequences, l represents traversing each token sequence index, P represents the probability of generating the l th token sequence; When the perplexity of the optimized training model on the validation set reaches the expected standard, the current parameters are saved and deployed to the value reinforcement learning and vulnerability repair tasks as the basis for subsequent vulnerability detection and repair. The calculation formula for the parameter is as follows: ; In the formula: θ represents the weight parameter in the primary training model.

[0011] Preferably, the process of vulnerability detection is as follows: Obtain a set of predefined vulnerability patterns, denoted by p , that is, the set of predefined vulnerability patterns p = { p 1 , p 2 ,... p M}}, and the corresponding severity weights of the set of predefined vulnerability patterns are {ω p , ω 1 , ω 2 ,... ω M}, and a predefined vulnerability detection function , where the parameter ϕ in the vulnerability detection function includes at least the set of vulnerability patterns and the detection threshold. Use the detection function to detect each vulnerability p in the set of vulnerability patterns p M , and the detection results are as follows: ; During the process of detecting vulnerabilities in the set of vulnerability patterns P , simultaneously define a severity function , a vulnerability classification function , and a vulnerability location function . Use the predefined severity function to detect the severity of vulnerabilities in the set of vulnerability patterns p , and output a value between 0 and 1 to reflect the severity of the vulnerability. Use the vulnerability location function to determine the location of the vulnerability in the code x , and then use formula (3) to obtain the vulnerability detection report R ( x ), and formula (3) is as follows; (3); In the formula: represents the set of precise vulnerability locations determined after symbolic execution and data flow analysis; represents the type of the vulnerability, is the predefined set of vulnerability categories; Meanwhile, perform a weighted sum of the severity of each vulnerability to calculate the risk score corresponding to the vulnerability, as the basis for automated vulnerability repair. Among them, the formula for performing a weighted sum of the severity of each vulnerability is as follows: .

[0012] Preferably, the static analysis process is as follows: Define the basic error detection function. The basic errors in the generated code are counted through the error detection function. Among them, the basic errors at least include syntax errors, type mismatches, and uninitialized variables. The calculation formula of the detection function is as follows: ; In the formula: y F represents the output code to be detected for errors, n represents the total number of output codes to be detected for errors, I ( e j ) represents the indicator function for judging the output code, that is, the j th basic error e j when detected, I ( e j ) = 1. Conversely, I ( e j ) = 0. Obtain the detected basic errors to establish a vulnerability pattern set p , that is, the vulnerability pattern set p = { p 1 , p 2 , ……, p i}. For each vulnerability pattern p i define the indicator function Ip i ( y ). That is, the overall vulnerability detection index can be calculated through the following formula: ; Pre-determine the static analysis and compiler feedback function , and its calculation formula is as follows: ; In the formula: is the penalty weight for basic errors, ; is the penalty weight for vulnerability detection, Obtain the number of error output codes in the output code, and calculate the error ratio of the number of error output codes to the total number of output codes R ; When the error ratio R ≥ 5%, at this time, the result of the detection function increases, and a negative reward is triggered; When there are vulnerabilities in the generated code that are the same as the known vulnerability patterns, Increases, resulting in a negative reward.

[0013] Preferably, the discrimination process of the semantic discriminator is as follows: Pre-train the semantic model and input the code x i as the input context, y i as the output code, and use "." to represent the splicing of the context and the generated code. Map through a multi-layer perceptron ( MLP ) to obtain a semantic score, and reflect the matching degree of the generated code with the expected function semantically through the semantic score. A high semantic score indicates that the generated code meets the requirements functionally, otherwise, it indicates that the generated code does not meet the requirements functionally. The calculation formula of the semantic score is as follows: , where ; The score given by the preset vulnerability detection module is denoted as: ; In the formula: represents the indicator function for judging the output code content; When the generated code y detects a vulnerability pattern, it takes the value of 1, otherwise it is 0; is the vulnerability p M in the generated code y severity score; ω M is the weight of the vulnerability p M , reflecting the importance of its risk, and when the score given by the preset vulnerability detection module is greater than the threshold, trigger the automated repair process; Among them, the automated repair process is as follows: The predetermined repair function F , input the input and output code y and the vulnerability detection report , and use the repair function F to automatically adjust the structure of the code according to the type of the vulnerability, and then output the repaired code , where the repaired code The calculation formula of is ; Pre-determine the semantic discriminator feedback function, substitute the semantic score and the score given by the vulnerability detection module to obtain the overall score. When the overall score is high, then the code has no security issues, otherwise, the code has security issues. Among them, the calculation formula of the pre-determined semantic discriminator feedback function is as follows: ; In the formula: Represents a tuning parameter used to balance the weights of semantic scoring and penalty. ; The semantic score is used to extract the semantic features of the generated code through a pre-trained semantic model and a multi-layer perceptron, and is used to evaluate the consistency between its function and the expected function. Penalty According to the output of the vulnerability detection module, a negative reward is imposed on the generated code with vulnerabilities. Using the repaired code The repaired code will be re-analyzed statically and semantically judged to obtain a comprehensive feedback. The calculation formula is as follows: ; In the formula: Represents the basic error in the repaired code in Represents the semantic score in the repaired code ; When meets the preset security and function standards, it is regarded as a successful repair; otherwise, further repair or adjustment is triggered.

[0014] Beneficial effects: The present invention relates to an automated code security optimization method based on a large language model. By using a programming dataset, a primary training model is built for the large language model Transformer with the f θ architecture, enabling the primary training model to master the basic syntax and logical structure of the programming language, generating initial code that meets the syntax requirements. Through the collaborative scanning of the initial code by static analysis, compiler feedback, and a semantic discriminator, using a predefined set of vulnerability patterns and corresponding severity weights to identify potential vulnerabilities in the initial code, generating a vulnerability library and a repair strategy library. By analyzing the data in the vulnerability library and the repair strategy library, optimizing the vulnerability time and repair algorithm, it realizes vulnerability detection, location, automatic repair, and re-verification in automated code generation, constructs an efficient and closed-loop code security optimization system, which can automatically identify and repair potential vulnerabilities in the generated code without a large amount of manual intervention, improving the security and efficiency of code generation.

[0015] Secondly, static analysis is used in conjunction with a semantic discriminator to evaluate code snippets during the process of generating initial code to obtain a newly generated strategy. The primary generation strategy in the primary training model is adjusted through the newly generated strategy, and then the initial code is generated for real-time monitoring to capture potential runtime vulnerabilities as training signals. The training signals are used to optimize the optimized training model to obtain a continuously optimized training model. Through reinforcement learning techniques, the generated code can adjust and optimize the strategy at each step, gradually reducing vulnerabilities and errors, ensuring that the code meets security requirements during the generation process. The code generated using static analysis and the semantic discriminator can not only pass compilation checks but also ensure the avoidance of security vulnerabilities during actual execution, improving the security and reliability of the code. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 The system block diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] As Figure 1 shown, the present invention provides a technical solution: an automated code security optimization method based on a large language model, comprising the following steps: Obtain large-scale programming data to establish a programming data set. Among them, the process of building the programming data is as follows: Obtain large-scale programming data as the initial data set D , where the initial data set D = {( x i , y i )}, where x i represents the given input code, y i represents the expected output code. Clean the input code x i and the output code y i , remove the invalid code snippets in the input code x i and the output code y i , and use tokenizer to split the input code x i and the output code y i into tokens . For each input code x i and the output code y i , perform token transformation to convert the code into tokenSequence, that is, generate a sequence dataset S → token 1 , token 2 ,…, token N ; That is, the input code x i Converted token The sequence is: x i → x i1 , x i2 ,…, x iTX , the output code y i Converted token The sequence is: y i → y i1 , y i2 ,…, y iTY , where TX and TY are the lengths corresponding to the input and output sequences respectively. After the establishment of the programming dataset, based on Transformer The architecture of the large language model f θ Build a primary training model, generate initial code and establish an initial code set. The Transformer The architecture of the large language model f θ In θ Are the weight parameters of the model. In the primary training model, the weight parameter θ Represents the pre-trained weight in transfer learning, where the architecture of the large language model f θ The calculation formula is as follows: ; In the formula: x Represents the input problem description or part of the code, y < t Represents at the position t The token N-1 Sequence that has been generated before, t Represents the generation of the current token N Sequence position, h t Represents the hidden state of the t th position calculated by the primary training model,W o The weight matrix representing the mapping of the hidden state to the output vocabulary dimension, Softmax indicating the conversion of the activation function output to the next token N+1 probability distribution of the sequence, where, Transformer architecture of the large language model f θ The internal calculation process of building the primary training model is as follows: Input the initial information into the primary training model a and the already generated y < t corresponding token sequence, calculate the dependency of each token sequence on other token N sequences through the multi-head self-attention mechanism, which is the self-attention layer. After non-linearly transforming the output of the self-attention with the feed-forward neural network, through residual connection and layer normalization, repeat L times, and obtain the hidden state corresponding to the position in t ; h t , map the hidden state to the dimension of the output vocabulary through the weight matrix, and then through Softmax the activation function convert the output to the probability distribution of the next token N+1 sequence; Among them, the multi-head self-attention calculation formula is as follows: (1) In the formula: Q represents the query vector of the current token N sequence, obtained by linearly transforming the input vector z i ; K represents the key vectors of all token sequences, obtained by linearly transforming the input vector z i ; V represents the value vectors of all token sequences, obtained from the input vector z i through the corresponding linear transformation; represents the dimension of each attention head; T represents a single token sequence length; Among them, the weight matrix formula is as follows: ; In the formula: nRepresents the corresponding sequence length calculated in formula (1). Is the dimension of the primary training model. Through the weight matrix, the input x Is linearly transformed to obtain the query vector Q , the key vector K And the value vector V , that is, the query vector , the key vector , the value vector , where W Represents the learning parameter matrix for linear transformation in the primary training model. After calculating the query vector Q , the key vector K And the value vector V , substitute the query vector Q , the key vector K And the value vector V Into Softmax The activation function to generate the final attention output.

[0018] The initial code set is scanned through a predefined vulnerability pattern set, and the vulnerabilities in the initial code set are identified as initial vulnerabilities. Through symbolic execution and data flow analysis, the initial vulnerabilities are located. Based on the positioning results, combined with the predefined vulnerability categories in the predefined pattern set, the initial vulnerabilities are classified, where by Softmax Given a real number vector z To the activation function, use Softmax The activation function maps the real number vector z To obtain the corresponding probability vector, and calculate the components through the following calculation formula: ; Where: Is the exponential transformation of the input vector z i ; Represents the sum of the exponential values of all elements of the input vector z, and the sum of all output probabilities is 1; the real number vector z = z 1 , z 2 ,..., z K ; i Represents the z In the input vector i Th element.

[0019] During the process of converting the code to token N Sequence, given the previous token N-1 , the predicted current is obtained through formula (2).token N The value, where formula (2) is as follows: (2) In the formula: P represents token N + 1 the probability distribution of the sequence, y t represents predicting the current t at time token N sequence, x represents the input code, y represents the output token sequence, f θ represents the primary training model, t represents the time of the currently generated token N sequence; Calculating the difference between the predicted value and the true value through the cross - entropy loss function, where the calculation formula of the cross - entropy loss function is as follows: ; In the formula: T y represents the length of the output sequence token and y represents the number of output sequences token . After calculating the difference between the predicted value and the true value, use the optimizer AdamW to update the parameters θ of the primary training model, increasing the stable convergence during the training process of the primary training model. The formula for updating the parameters θ of the primary training model is as follows: ; where η t is the learning rate, ∇θ is the gradient of the cross - entropy loss with respect to the parameter θ, and during the training process, divide the training data into N training batches, and perform forward propagation and backward propagation on each training batch until an optimized training model is obtained. The calculation formula for forward propagation is as follows: ; In the formula: represents at time t , the sequence generated by the primary training model according to the input code x and the already generated y < t corresponding token sequence; The calculation formula for backward propagation is as follows: ; Evaluate the generation effect of the model on the validation set by perplexity, where the calculation formula of perplexity is as follows: Perplexity in the validation set wherein, the perplexity Perplexity has the following calculation formula: ; In the formula: N represents token the total number of sequences, l represents traversing each token sequence index, P represents the probability generated for the l th token sequence; When the perplexity of the optimized training model on the validation set reaches the expected standard, save the current parameters and deploy them to the value reinforcement learning and vulnerability repair tasks as the basis for subsequent vulnerability detection and repair, where the calculation formula of the parameter is as follows: ; In the formula: θ represents the weight parameter in the primary training model. After calculating the parameter , perform vulnerability detection, and the process of vulnerability detection is as follows: Obtain the predefined vulnerability pattern set, denoted by p , that is, the predefined vulnerability pattern set p ={ p 1 , p 2 ,... p M}, and the corresponding severity weights of the predefined vulnerability pattern set p are {ω 1 , ω 2 ,... ω M}, and the predefined vulnerability detection function , where the parameters in the vulnerability detection function ϕ include at least the vulnerability pattern set and the detection threshold. Use the detection function to detect each vulnerability p in the vulnerability pattern set p M , and the detection results are as follows: ; During the process of detecting vulnerabilities in the vulnerability pattern set p , simultaneously predefined severity function , vulnerability classification function and vulnerability location function , using a predefined severity function , perform severity detection on the vulnerability pattern set P and output a value between 0 and 1, reflecting the severity of the vulnerability. Using the vulnerability location function , determine the location of the vulnerability in the code x , and then use formula (3) to obtain the vulnerability detection report R ( x ), formula (3) is as follows; (3) In the formula: represents the set of precise vulnerability locations determined after symbolic execution and data flow analysis; represents the type of the vulnerability, is the predefined set of vulnerability categories; At the same time, perform a weighted sum of the severity of each vulnerability to calculate the risk score corresponding to the vulnerability, which is used as the basis for automated vulnerability repair. Among them, the formula for the weighted sum of the severity of each vulnerability is as follows: , after completing the weighted sum of the severity of the vulnerabilities, use the predefined vulnerability pattern set and the corresponding severity weights to identify the possible vulnerabilities in the initial code, generate a vulnerability library and a repair strategy library, and optimize the vulnerability time and repair algorithm through data analysis of the vulnerability library and the repair strategy library, realizing vulnerability detection, location, automatic repair and re-verification in automated code generation, and constructing an efficient and closed-loop code security optimization system, which can automatically identify and repair potential vulnerabilities in the generated code without a large amount of manual intervention, improving the security and efficiency of code generation.

[0020] Among them, during the process of generating the initial code, static analysis is used in combination with a semantic discriminator to evaluate the code fragments during the process of generating the initial code to obtain a newly generated strategy, and the initial generation strategy in the primary training model is adjusted through the newly generated strategy to obtain an optimized training model. Among them, the static analysis process is as follows: Define a basic error detection function, and count the basic errors that occur in the generated code through the error detection function. Among them, the basic errors include at least syntax errors, type mismatches, and uninitialized variables. The calculation formula of the detection function is as follows: ; In the formula: y F represents the output code for which errors are to be detected, n represents the total number of output codes for which errors are to be detected, I ( e j)The indicator function for judging the output code, that is, the j basic error e j is detected, I ( e j ) = 1, conversely, I ( e j ) = 0, obtain the set of vulnerability patterns established from the detected basic errors p , that is, the set of vulnerability patterns p = { p 1 , p 2 , ……, p M}, for each vulnerability pattern p M define the indicator function Ip M ( y ), that is, the overall vulnerability detection index can be calculated by the following formula: ; Pre-determine the static analysis and compiler feedback function , and its calculation formula is as follows: ; In the formula: is the penalty weight for basic errors, is the penalty weight for vulnerability detection, Obtain the number of error output codes in the output code, and calculate the error ratio of the number of error output codes to the total number of output codes R ; When the error ratio R ≥ 5%, at this time, the result of the detection function increases, and a negative reward is triggered; When there are vulnerabilities in the generated code that are the same as the known vulnerability patterns, increases, and a negative reward is generated.

[0021] The discrimination process of the semantic discriminator is as follows: Pre-train the semantic model, input the code x i is the input context, y i is the output code, and the splicing of the context and the production code is represented by ".", through a multi-layer perceptron ( MLP)Mapping obtains a semantic score, which reflects the matching degree of the generated code with the expected function semantically. A high semantic score indicates that the generated code meets the requirements functionally, otherwise, it means that the generated code does not meet the requirements functionally. The calculation formula of the semantic score is as follows: , where ; The score given by the preset vulnerability detection module is denoted as: ; In the formula: Indicates the indicator function for judging the output code content; When the generated code y detects a vulnerability pattern, it takes the value of 1, otherwise it is 0; Is the vulnerability p M In the generated code y The severity score; Is the vulnerability p M The weight of reflects the importance of its risk, and when the score given by the preset vulnerability detection module is greater than the threshold, it triggers an automated repair process to repair the detected vulnerability. Among them, the automatic repair process is as follows: The predetermined repair function F , input and output the code y and the vulnerability detection report , use the repair function F According to the type of the vulnerability, automatically adjust the structure of the code, and then output the repaired code , where the repaired code The calculation formula of is ; Pre-determine the semantic discriminator feedback function, substitute the semantic score and the score given by the vulnerability detection module to obtain the overall score. When the overall score is high, Then the code has no security problems, otherwise, the code has security problems. Among them, the calculation formula of the pre-determined semantic discriminator feedback function is as follows: ; In the formula: Indicates the adjustment parameter, which is used to balance the weights of the semantic score and the penalty, > 0; The semantic score is to extract the semantic features of the generated code through a pre-trained semantic model and a multi-layer perceptron, and is used to evaluate the consistency between its function and the expected function; Penalty According to the output of the vulnerability detection module, apply a negative reward to the generated code with vulnerabilities; Using the repaired code It will be re-analyzed statically and semantically evaluated to obtain comprehensive feedback. The calculation formula is as follows: ; In the formula: represents the basic errors in the repaired code ; represents the semantic score in the repaired code ; When meets the preset safety and functional standards, it is regarded as a successful repair; otherwise, further repair or adjustment is triggered. By real-time monitoring of the generated initial code, potential runtime vulnerabilities are captured as training signals, and the optimization training model is optimized using the training signals to obtain a continuously optimized training model. The initial vulnerabilities, runtime vulnerabilities, and their corresponding location and repair processes during the optimization process are recorded to establish a vulnerability library and a repair strategy library. Through data analysis of the vulnerability library and the repair strategy library, the vulnerability time and repair algorithm are optimized, that is, static analysis is used in conjunction with a semantic discriminator to evaluate code segments during the generation of the initial code to obtain a newly generated strategy. The primary generation strategy in the primary training model is adjusted using the newly generated strategy, and then the initial code is generated for real-time monitoring. Potential runtime vulnerabilities are captured as training signals, and the optimization training model is optimized using the training signals to obtain a continuously optimized training model. Through reinforcement learning technology, the generated code can adjust and optimize the strategy at each step according to the feedback, gradually reducing vulnerabilities and errors, ensuring that the code meets safety requirements during the generation process. The code generated using static analysis and the semantic discriminator can not only pass the compilation check but also ensure the avoidance of security vulnerabilities during actual execution, improving the security and reliability of the code.

[0022] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the protection scope of the present invention.

Claims

1. An automated code security optimization method based on a large language model, characterized in that: The following steps are involved: Acquire large-scale programming data to build a programming dataset based on Transformer Architecture of Large Language Model f θ Build a primary training model, generate initial code and establish an initial code set; Scan the initial code set through the predefined vulnerability pattern set, identify the vulnerabilities in the initial code set as initial vulnerabilities, locate the initial vulnerabilities through symbolic execution and data flow analysis, and classify the initial vulnerabilities based on the positioning results and the predetermined vulnerability categories in the predefined pattern set; In the process of generating the initial code, static analysis is used in conjunction with a semantic discriminator to evaluate the code snippets in the process of generating the initial code to obtain a new generation strategy, and the primary generation strategy in the primary training model is adjusted through the new generation strategy to obtain an optimized training model; By real-time monitoring of the generated initial code, potential runtime vulnerabilities are captured as training signals, and the training signals are used to optimize the optimized training model to obtain a continuously optimized training model; Record initial vulnerabilities, runtime vulnerabilities, and the corresponding positioning and repair processes, establish a vulnerability library and a repair strategy library, and optimize vulnerability time and repair algorithms by performing data analysis on the vulnerability library and repair strategy library.

2. The method for automatic code security optimization based on a large language model according to claim 1, characterized in that: The process of building a programming data set is as follows: Obtaining large-scale programming data as an initial dataset D , where the initial data set D ={( x i , y i )},in, x i Represents a given input code, y i Indicates the expected output code, for the input code x i and the output code y i To clean, enter the code x i and the output code y i Remove invalid code snippets from the code and use Tokenizer Enter the code x i and the output code y i Split into Tokens , for each input code x i and the output code y i conduct token , converting the code into token Sequence, that is, generate a sequence dataset S →[ token 1 , token 2 , …, token N ]; Enter the code x i Transformed token The sequence is: x i →[ x i1 , x i2 , …, x iTX ], output code y i Transformed token The sequence is: y i →[ y i1 , y i2 , …, y iTY ],in, TX and TY are the lengths of the input and output sequences respectively.

3. The method for automatic code security optimization based on a large language model according to claim 2, characterized in that: Said Transformer Architecture of Large Language Model f θ middle θ is the weight parameter of the model. In the primary training model, the weight parameter θ Represents the pre-trained weights in transfer learning, where the architecture is a large language model f θ The calculation formula is as follows: ; Where: x Represents the input problem description or part of the code. y < t Indicates at location t Previously generated token N-1 sequence, t Indicates the generation of the current token N The position of the sequence, h t Indicates the first t The hidden state of the position, W o The weight matrix representing the mapping of hidden states to the output vocabulary dimensions, Softmax Indicates that the activation function output is transformed into the next token N+1 Probability distribution of a sequence.

4. The method for automatic code security optimization based on a large language model according to claim 3, characterized in that: Transformer Architecture of Large Language Model f θ The internal calculation process of building a primary training model is as follows: Input initial information to the primary training model a and Generated y < t Corresponding token Sequence, calculate each through the multi-head self-attention mechanism token Sequence to other token N The dependency relationship of the sequence is the self-attention layer. After the output of the self-attention is nonlinearly transformed with the feedforward neural network, it is repeated through residual connection and layer normalization. L After that, get t The hidden state corresponding to the position h t , mapping the hidden state to the dimension of the output vocabulary through the weight matrix, and then through Softmax The activation function transforms the output into the next token N+1 The probability distribution of Among them, the calculation formula of multi-head self-attention is as follows: (1); Where: Q Indicates the current token N The query vector of the sequence is obtained by inputting the vector z i Perform a linear transformation to obtain; K Indicates all token The key vector of the sequence, through the input vector z i Perform a linear transformation to obtain; V Indicates all token The value vector of the sequence is transformed from the input vector by the corresponding linear transformation z i get; Represents the dimension of each attention head; T Indicates a single token The length of the sequence; Among them, the weight matrix formula is as follows: ; Where: n represents the length of the corresponding sequence calculated in formula (1), is the dimension of the primary training model, through the weight matrix, the input x Perform linear transformation to obtain the query vector Q , key vector K Sum value vector V , which is the query vector , the key vector , value vector , where W Represents the learning parameter matrix for linear transformation in the primary training model. Q , key vector K Sum value vector V After the calculation of Q , key vector K Sum value vector V Substitution Softmax In the activation function, the final attention output is generated.

5. The method for automatic code security optimization based on a large language model according to claim 4, characterized in that: Among them, through Softmax The activation function is given a real vector z ,use Softmax The activation function converts the real vector z The corresponding probability vector is obtained by mapping, and the components are calculated by the following calculation formula: ; Where: is the input vector z i Perform exponential transformation; represents the sum of the exponential values ​​of all elements of the input vector z, and the sum of all its output probabilities is 1; real number vector z = z 1, z 2, ..., z K ; i Represents the input vector z The i elements.

6. The method for automatic code security optimization based on a large language model according to claim 3, characterized in that: Convert the code to token N In the sequence process, given the previous token N-1 , the predicted current token N The value of , where formula (2) is as follows: (2); Where: P express token N + 1 The probability distribution of the sequence, y t Indicates at time t Next prediction current token N sequence, x Indicates the input code. y Representation output token sequence, f θ represents the primary training model, t Indicates the current generation token N The timing of the sequence; The difference between the predicted value and the true value is calculated by the cross entropy loss function, where the cross entropy loss function calculation formula is as follows: ; Where: T y Represents the output sequence token Length, y Represents the output sequence token The number of .

7. The method for automatic code security optimization based on a large language model according to claim 1, characterized in that: Using the Optimizer AdamW To update the primary training model parameters θ, increase the stable convergence of the primary training model during training, where the formula for updating the primary training model parameters θ is as follows: ; Among them, η t is the learning rate, is the gradient of the cross entropy loss with respect to the parameter θ, and during the training process, the training data is divided into N training batches, and perform forward propagation and back propagation on each training batch until the optimized training model is obtained. The forward propagation calculation formula is as follows: ; Where: Indicates at time t At the place where the primary training model is trained according to the input code x And generated y < t Corresponding token sequence; The back propagation calculation formula is as follows: ; Through perplexity Perplexity , in the validation set The generation effect of the model is evaluated in Perplexity The calculation formula is as follows: ; Where: N express token The total number of sequences, l Indicates traversing each token the index of the sequence, P Indicates l indivual token The probability of sequence generation; When the perplexity of the optimized training model on the validation set reaches the expected standard, save the current parameters , and deployed in value reinforcement learning and vulnerability repair tasks as the basis for subsequent vulnerability detection and repair, where the parameters The calculation formula is as follows: ; Where: θ Represents the weight parameters in the primary training model.

8. The method for automatic code security optimization based on a large language model according to claim 6, characterized in that: The vulnerability detection process is as follows: Get a set of predefined vulnerability patterns, using p Represents a set of predefined vulnerability patterns p = p 1, p 2,…… p M }, a set of predefined vulnerability patterns p The corresponding severity weights are {ω1, ω2, ... ω M }, scheduled vulnerability detection function , where the parameters in the vulnerability detection function are ϕ At least including a vulnerability pattern set and a detection threshold, using the detection function to detect the vulnerability pattern set p Each vulnerability in p M The test results are as follows: ; In the vulnerability pattern collection p During the vulnerability detection process in , Vulnerability classification function and vulnerability location function , using the predefined severity function , for the vulnerability pattern set p The vulnerability in the vulnerability is detected and a value between 0 and 1 is output to reflect the severity of the vulnerability. , determine the vulnerability in the code x Then use formula (3) to obtain the vulnerability detection report R ( x ), formula (3) is as follows; (3); Where: It means that after symbolic execution and data flow analysis, the exact set of vulnerability locations is determined; Indicates the type of vulnerability. is a predefined set of vulnerability categories; At the same time, the severity of each vulnerability is weighted and summed to calculate the risk score corresponding to the vulnerability, which serves as the basis for automated vulnerability repair. The formula for weighted summation of the severity of each vulnerability is as follows: 。 9. The method for automatic code security optimization based on a large language model according to claim 1, characterized in that: The static analysis process is as follows: Define a basic error detection function, and use the error detection function to count the basic errors that occur in the generated code. The basic errors at least include syntax errors, type mismatches, and uninitialized variables. The calculation formula of the detection function is as follows: ; Where: y F Indicates the output code of the error to be detected, n Indicates the total number of output codes to be detected as errors, I ( e j ) represents the indicator function for judging the output code, i.e. j Basic Error e j When detected, I ( e j )=1, otherwise, I ( e j )=0, get the basic error detected to establish a vulnerability pattern set p , that is, the set of vulnerability patterns p = p 1 , p 2 ,……, p M }, for each vulnerability pattern p M Defining indicator functions Ip M ( y ), that is, the overall vulnerability detection index can be calculated by the following formula: ; Pre-planned static analysis and compiler feedback functions , and its calculation formula is as follows: ; Where: is the penalty weight for the basic error, ; is the penalty weight for vulnerability detection, Get the number of wrong output codes in the output code, and calculate the error ratio of the number of wrong output codes to the total number of output codes R ; When the error ratio R ≥5%, at this time, the detection function If the result of increases, a negative reward is triggered; When a vulnerability with the same pattern as a known vulnerability appears in the generated code, Increase, resulting in negative rewards.

10. The method for automatic code security optimization based on a large language model according to claim 8, characterized in that: The discriminant process of the semantic discriminator is as follows: Pre-trained semantic model, input code x i is the input context, y i is the output code, and "." is used to represent the concatenation of the context and the production code, through a multi-layer perceptron ( MLP ) is mapped to obtain a semantic score, which reflects the degree of match between the generated code semantically and the expected function. A high semantic score indicates that the generated code meets the requirements in terms of function. On the contrary, a low semantic score indicates that the generated code does not meet the requirements in terms of function. The calculation formula of the semantic score is as follows: , among which ; The score given by the preset vulnerability detection module is recorded as: ; Where: Indicates the indicator function for judging the content of the output code; When generating code y The value is 1 when a vulnerability mode is detected, otherwise it is 0; For loopholes p M In the generated code y The severity score in ω M For loopholes p M The weight of the vulnerability detection module reflects the importance of the risk, and when the score given by the preset vulnerability detection module is greater than the threshold, the automated repair process is triggered; The automatic repair process is as follows: Scheduled repair function F , input and output code y and vulnerability detection reports , using the repair function F Automatically adjust the code structure according to the type of vulnerability and output the repaired code , where the fixed code The calculation formula is ; The semantic discriminator feedback function is pre-planned, and the semantic score and the score given by the vulnerability detection module are substituted to obtain the overall score. When the overall score is high, If the code has no security issues, then the code has security issues. Otherwise, the code has security issues. The calculation formula of the pre-planned semantic discriminator feedback function is as follows: ; Where: represents the adjustment parameter used to balance the weight of semantic score and penalty, ; Semantic scoring is to extract the semantic features of the generated code through pre-trained semantic models and multi-layer perceptrons to evaluate the consistency between its functions and the expected functions; punish Based on the output of the vulnerability detection module, negative rewards are imposed on the generated code with vulnerabilities; Utilizing the fixed code After static analysis and semantic evaluation, comprehensive feedback is obtained. The calculation formula is as follows: ; Where: Indicates the repaired code Basic errors in Indicates the repaired code mid-semantic score; when If the preset safety and functional standards are met, the repair is considered successful; otherwise, further repairs or adjustments are triggered.

Citation Information

Patent Citations

  • Code vulnerability detection large model construction method and device and electronic equipment

    CN118171291A

  • Source code vulnerability detection method and system based on Transform language model

    CN118886016A

  • Software vulnerability automatic repair method based on thinking chain and storage medium

    CN119128893A

  • Code review and optimization method driven by large language model

    CN119512556A

  • Source code vulnerability detection and repair through machine learning

    WO2024019848A1

Cited By

  • Software package source code optimization method of cross-instruction-set architecture

    CN120371387A

  • A cross-instruction set architecture software package source code optimization method

    CN120371387B

  • Large visual language model vulnerability detection method and device based on security sensitive layer activation guidance and storage medium

    CN121302380A

  • Method, device and storage medium for detecting vulnerabilities of large visual language model based on security-sensitive layer activation guidance

    CN121302380B

  • Code development system and method with self-debugging capability

    CN121349416A