Prompt for code repair based on code changes

By generating small sample examples based on code differences and clustering of large language models, source code vulnerabilities are automatically repaired, and inefficient problems in the existing technology are solved and efficient vulnerability repair is achieved.

CN120476389APending Publication Date: 2025-08-12MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480006273.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-06-02
Filing Date
2024-05-23
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing technology is difficult to efficiently detect and repair vulnerabilities in source code, especially software vulnerabilities, and manual repair methods are inefficient.

Method used

Build hints to fix source code vulnerabilities by generating small sample examples based on code differences, using large language model clustering and contour analysis.

Benefits of technology

It realizes efficient and automated source code vulnerability repair, reduces the dependence of manual repair, and improves repair efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120476389A_ABST
    Figure CN120476389A_ABST
Patent Text Reader

Abstract

The source code repair system generates a prompt that includes a less sample example of code changes previously made to correct a particular source code vulnerability. The prompt is used for the large language model to generate repaired code of the source code fragments with the same vulnerability. Code changes made to specific source code vulnerabilities in a code difference format are clustered into closely related code change embedding groups. The small number of code differences selected in each group with the closest mean value of the cluster is used as a small sample example of the respective source code vulnerabilities.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Source code vulnerabilities are errors in a program's source code that cause it to behave in unintended ways, such as producing incorrect results. There are various types of source code vulnerabilities. Functional vulnerabilities are vulnerabilities in which a program fails to perform according to its functional description or specification. Compiler errors are software vulnerabilities that fail to conform to the syntax of the program's programming language. Runtime errors occur during runtime and include logic errors, I / O errors, undefined object errors, and divide-by-zero errors.

[0002] Software vulnerabilities differ from source code vulnerabilities, such as functional bugs, compiler errors, and runtime errors, because they do not produce incorrect results. In contrast, software vulnerabilities are programming flaws that lead to significant performance degradation, such as excessive resource usage, increased latency, reduced throughput, and overall degraded performance, or can be exploited for malicious purposes. Software vulnerabilities are difficult to detect due to the lack of fail-stop symptoms. As software system complexity increases, emphasis is placed on the efficient use of resources and system security, leading to improvements in detecting and remediating software vulnerabilities. Summary of the Invention

[0003] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

[0004] The source code remediation system generates hints that include few-shot examples of code changes previously made to correct a specific source code vulnerability. These hints are used with a large language model to generate fixed code for source code snippets with the same vulnerability. Code changes made to address a specific source code vulnerability are clustered into groups based on closely related code changes in the form of code diffs. Within each group, a subset of code diffs with the closest average intra-cluster distance and average closest-cluster distance for each vector of the group are selected as few-shot examples for the hint.

[0005] These and other features and advantages will be apparent from a reading of the following detailed description and an overview of the associated drawings.It is to be understood that both the foregoing general description and the following detailed description are explanatory only and are not restrictive of the aspects as claimed. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Figure 1 is a diagram illustrating the generation of a few-shot example of similar code changes for a software vulnerability type.

[0007] Figure 2 is a diagram showing code for generating hints to a large language model to obtain repairs.

[0008] Figures 3A to 3B is a schematic diagram illustrating an exemplary application of a source code repair system.

[0009] Figure 4 is a flow chart illustrating an exemplary method of a source code repair system.

[0010] Figure 5 is a flow chart illustrating an exemplary method for generating hints to a large language model of code for repair.

[0011] Figure 6 is a block diagram illustrating an exemplary operating environment. DETAILED DESCRIPTION

[0012] Overview

[0013] This disclosure relates to generating hints for a large language model to predict the presence of fixed source code correcting a software vulnerability of a known vulnerability type. The hints include a few-shot examples of code changes used to fix previous software vulnerabilities of the same vulnerability type. The few-shot examples guide the large language model to predict the correct output.

[0014] Code changes made to correct known software vulnerabilities are clustered into groups. The code change encoding and the corresponding natural language text description of the code change are used to cluster similar code changes into groups. Each group is associated with a specific type of vulnerability. Based on profile analysis, the most representative changes from each group are selected and used as a few-shot example of the software vulnerability in the alert.

[0015] Prompts are generated for a large language model to predict code that fixes a software vulnerability. In one aspect, the large language model is trained on source code and natural language text. Prompts are generated in a format that includes instructions, a few-shot example, and a conversation about the software vulnerability. Given the prompts, the large language model generates fixed code that fixes the software vulnerability. In one aspect, the fixed code is in the form of a code diff that shows the changes required to correct the original source code snippet.

[0016] The systems, methods, and components used in the source code repair system will now be described in more detail.

[0017] system

[0018] Figure 1 Illustrated is an exemplary system for generating groups of similar code changes for a software vulnerability 100. In one aspect, the system 100 utilizes a source code repository 102, a static analyzer 104, a pull request engine 106, a neural encoder 108, a neural decoder 110, and a clustering engine 112.

[0019] The source code repository 102 is a file archive and network hosting facility that stores a large number of artifacts, such as source code files and code libraries. Programmers (i.e., developers, users, end users, etc.) typically utilize the shared source code repository 102 to store source code and other programming artifacts that can be shared between different programmers. Programming artifacts are files generated by programming activities, such as source code, program configuration data, documents, etc. The shared source code repository 102 can be configured as a source control system or version control system, which stores each version of an artifact, such as a source code file, and tracks changes or differences between different versions. The repository managed by the source control system is distributed so that each user of the repository has a working copy of the repository. The source control system coordinates the distribution of changes made to the contents of the repository to different users.

[0020] Static analyzer 104 discovers software vulnerabilities through a code library or source code repository. Static analyzer 104 does not execute the source code to discover software vulnerabilities, but instead relies on static analysis. Examples of static analyzers include, but are not limited to, Infer, CodeQL, and source code security analyzers (i.e., BASH, dotTEST, etc.). Compilers differ from static analyzers in that they detect syntax errors, which are different from software vulnerabilities.

[0021] There are various types of vulnerabilities, such as network security vulnerabilities listed in the National Institute of Standards and Technology's Public Vulnerabilities and Exposures Database, SQL injection, and cross-site scripting vulnerabilities, network vulnerabilities that manage the flow of data workloads, user traffic, and computing requests.

[0022] Infer is an interprocedural static code analyzer based on separation logic that performs Hoare logic reasoning about programs that mutate data structures. Infer uses an analysis language, Smallfoot Intermediate Language (SIL), to represent programs in a simpler instruction set that describes the program's actions on a symbolic heap. Infer symbolically executes SIL commands on the symbolic heap according to a set of separation logic verification rules in order to discover program paths with symbolic heaps that violate heap-based properties. Infer discovers software vulnerabilities consisting of null pointer exceptions, resource leaks, annotated reachability, lost lock protection, and concurrency race conditions.

[0023] It should be noted that SIL is different from an intermediate language such as CIL, which represents instructions that can be converted into native machine code. SIL instructions are used as symbolic execution for logic-based verification analysis. SIL instructions are not structured to execute on a processor or CPU like CIL instructions.

[0024] CodeQL is a static analysis tool that runs on a codebase or repository. CodeQL creates a CodeQL database for the codebase and includes information about the grammatical structure, data flow, and control flow of each code. CodeQL monitors the compilation process of the codebase and extracts information about the code, such as syntactic data from the abstract syntax tree of each code and semantic data about name bindings and type information. A set of queries is run on the CodeQL database to search for these artifacts from the database. Each query or group of queries is written to identify patterns for a specific vulnerability type. A list of code errors or vulnerabilities with the codebase is generated from the query. CodeQL identifies software vulnerabilities such as null pointer dereferences, uninitialized variables, hardcoded credentials, and SQL injections.

[0025] Static analyzer 104 generates a list of software vulnerabilities 114 found in each source code file of source code repository 102. Pull request engine 106 retrieves pull requests that include changes to the source code program that correct the vulnerabilities identified by the static analyzer. Pull requests include code changes 116 in code diff format. Code diff 116 is a comparison of data between two files or two versions of the same file, output by a diff tool. The differences are displayed in a standard format called a code diff or diff format.

[0026] The code diff format begins with a two-line header that includes the modification timestamps of the "from file" and the "to file." It then includes difference blocks, each of which includes a first line indicating the line numbers that changed as follows: @@@from-file-line-numbers to-file-line-numbers@@. Lines common to both files are then prepended with a space character, lines added to the first file are prepended with a '+' character, and lines deleted from the first file are prepended with a '-' character.

[0027] The code difference 118 is encoded into an embedding 120 using a deep learning encoder 108. A description of the code change 122 is generated by a deep learning decoder 110 of the code difference given the code change 118. In one aspect, the deep learning encoder 108 is an encoder neural transformer model with attention, and the deep learning decoder 110 is a decoder neural transformer model with attention. Examples of encoder neural transformer models with attention include Bidirectional Encoder Representations from Transformers (BERT), and examples of decoder neural transformer models with attention include Generative Pretrained Transformers (GPT) models. Other exemplary encoder models include OpenAI's neural encoder model, Huggingface's neural encoder model, or Microsoft's neural encoder model. Other exemplary decoder models include OpenAI's Codex model, Meta's Large Language Model MetaAI (LLaMA), and Microsoft's neural decoder model.

[0028] The Encoder Neural Transformer model with Attention is pre-trained on an unsupervised dataset of natural language text and source code snippets using a Masked Language Model objective. The Masked Language Model randomly masks some tokens from the input and aims to predict the original vocabulary of the masked token based on its context. The Masked Language Model objective enables the embedding representation to incorporate both left and right context, thereby incorporating the bidirectional nature of the embedding representation.

[0029] Deep learning machine learning models differ from traditional machine learning models that do not use neural networks. Machine learning involves the use and development of computer systems that can learn and adapt without explicit instructions, using algorithms and statistical models to analyze and draw inferences from patterns in data. Machine learning uses different types of statistical methods to learn from data and predict future decisions. Traditional machine learning includes statistical techniques, data mining, Bayesian networks, Markov models, clustering, support vector machines, and visual data mapping.

[0030] Deep learning differs from traditional machine learning because it uses multi-level data processing through many hidden layers of a neural network to learn and interpret features and the relationships between them. Deep learning embodies neural networks, unlike traditional machine learning techniques that do not use neural networks. There are various types of deep learning models that generate source code, such as recurrent neural network (RNN) models, convolutional neural network (CNN) models, long short-term memory (LSTM) models, and neural transformers.

[0031] The Neural Transformer with Attention model utilizes an attention mechanism. Attention is used to determine which part of the input sequence is important for each token / subtoken, particularly since the encoder is limited to encoding fixed-size vectors. The attention mechanism collects information about the relevant context of a given token / subtoken and then encodes this context into a vector representing the token / subtoken. It is used to identify relationships between subtokens in long sequences while ignoring other subtokens that do not have a significant impact on a given prediction.

[0032] The attention mechanism indicates how much attention a particular input should pay to other elements in a given input sequence. The attention mechanism can be implemented in the self-attention layer of the model. In the self-attention layer, each token in the input sequence is transformed into a query (Q), a key (K), and a value (V), which is used to calculate a score indicating the degree to which a particular token should pay attention to other tokens in the input sequence. The self-attention layer is integrated into the encoder neural transformer model with attention and the decoder neural transformer model with attention.

[0033] The attention mechanism used in the encoder neural transformer model with attention is a self-attention layer preceding the neural network layer. The self-attention layer focuses on both the right and left sides of the token being computed. The decoder neural transformer model with attention has a masked self-attention layer preceding the neural network layer. The masked self-attention layer masks the token to the right of the token being computed.

[0034] Clustering engine 112 uses code difference embeddings 120 and code change descriptions 122 to generate vectors representing code changes. Each code change vector is then clustered into groups representing similar code changes for each vulnerability type. There can be several groups for a particular vulnerability type. In one aspect, k-means clustering is used to form groups 124A-124N.

[0035] Profile analysis is then used to select the top k code differences that represent the group 126A-126N. Profile analysis determines how similar a vector is to the other vectors in the group. Each vector is given a profile value that measures how similar it is to the group. A high value indicates that the vector matches the other vectors in the group well. The profile value can be based on any distance measure, such as Euclidean distance or Manhattan distance. The vectors with the top k profile values are selected as representative data points for the group, where k is a predetermined value.

[0036] like Figure 1As shown, there are several groups 124A-124N, each associated with a specific software vulnerability. There can be several groups for a specific software vulnerability. Each group includes several data points, where each data point includes a code difference, an associated code difference embedding (embedding), a code change description (change description), and the number of lines changed. For each group 126A-126N, the top k code differences are identified.

[0037] We now turn our attention to further discussion of the source code repair system. Figure 2 , Figure 2 An exemplary system 200 for generating hints for a large language model to produce fixed code 217 is shown. System 200 includes a static analyzer 202, a hint generator 204, a set of groups 206, and a large language model 208. As described above, static analyzer 202 analyzes source code for software vulnerabilities. Static analyzer 202 identifies software vulnerabilities of a specific vulnerability type and generates a warning message. Hint generator 204 receives a vulnerable code snippet 210 and a vulnerability type 212 and, using the vulnerability type, retrieves the top k code differences 214 from each group 206 associated with the vulnerability type. Hint generator 204 generates hints 216, which include the top k code differences or few-shot examples 214, the vulnerable code snippet 210, and the instruction. Hints 216 are transmitted to large language model 208 for predicted fixed code 218.

[0038] In one aspect, prompt 216 may include a first instruction 218 describing a task to be performed by the large language model. In this exemplary prompt 216, first instruction 218 is as follows: You are fixing a security vulnerability. The trusted tool has flagged the following vulnerability in the code: <Vulnerability Type> with the following description: <Warning Message>. <Vulnerability Type> is a software vulnerability identified by a static analyzer, and <Warning Message> is a message output by the static analyzer.

[0039] Subsequently, the top k code differences 220 selected from the cluster group 206 are then input into the hint along with the source code snippet with the vulnerability and its context 224. The context is the source code surrounding the source code snippet and may include the source code immediately preceding, the source code immediately following the source code snippet, or a combination thereof.

[0040] Next, a second instruction 222 describes a task for the large language model to perform, including outputs 226 through 232 formatted as code differences. The output format includes the average number of lines that the repaired code should include as a variable k_lines 230. The value k_line 230 is the median number of changed lines across all changed lines in the group. The hint includes the number of lines so that the model knows how many lines of change it has typically seen in previous examples.

[0041] A large language model is a deep machine learning model that includes billions of parameters. The parameters are the part of the model that is learned from a training dataset that defines the model's skill in generating predictions for a target task. In one aspect, the large language model is a unified cross-modal neural transformer model with attention. The unified cross-modal neural transformer model with attention is a neural transformer model that is pre-trained on multimodal content, such as natural language text and source code, to support a variety of code-related tasks. The large language model can be implemented as a neural transformer model with attention in either an encoder-decoder configuration or a decoder-only configuration.

[0042] In one aspect, the large language model can be hosted in a remote server, access to which is provided as a service, wherein access to the large language model can be provided through an application programming interface (API). Examples of large language models include OpenAI's Chat GPT model or other GPT models provided as a service as described above.

[0043] Attention now turns to a more detailed discussion of the application of the source code repair system.

[0044] Figure 3A and 3B An exemplary system utilizing a source code repair system is shown. Figure 3AFIG3 illustrates a system 300 in which a source code repair system 308 operates within an integrated development environment (IDE) 302. IDE 302 is a software development tool that provides tools for software development, such as, but not limited to, a source code editor, a compiler, a debugger, and build automation tools. IDE 302 may include a source code editor 304, which interacts with developers to generate or edit source code. Source code repair system 308 may operate as a background process that monitors source code input into source code editor 304. At various times during development, source code repair system 308 may receive source code snippets as expressions 306 extracted from source code editor 304. As described above, expressions 306 are analyzed by a static analyzer to identify any software vulnerabilities in the source code being developed. Upon detecting such a software vulnerability, source code repair system 308 constructs a hint, as described above, which is sent to large language model 312 for use in remediating code. Source code repair system 308 notifies source code editor 304 that the expression may be a software vulnerability and provides remediating code 310 without the vulnerability.

[0045] Figure 3B System 312 is shown, in which a source code repair system 318 operates within a source code version control system 314 to identify vulnerabilities in the source code of a pull request 324 submitted by a developer 326. Source code repository 316 detects pull request 324 and initiates a request 320 for source code repair system 318 to analyze the source code subject to pull request 324 for software vulnerabilities. Source code repair system 318 utilizes a static analyzer, as described above, to detect software vulnerabilities. If a software vulnerability is found, source code repair system 318 generates a hint, as described above, which is transmitted to large language model 328 for use in the repaired source code. System 318 notifies source code repository 316 of the vulnerability and returns repaired source code 322.

[0046] method

[0047] Attention will now be directed to the description of various exemplary methods utilizing the systems and devices disclosed herein. The operation of these aspects can be further described with reference to the various exemplary methods. It will be understood that, unless otherwise indicated, the representative methods do not necessarily have to be performed in the order presented or in any particular order. In addition, the various activities described with respect to the methods can be performed in a serial or parallel manner, or in any combination of serial and parallel operations. In one or more aspects, the methods illustrate the operation of the systems and devices disclosed herein.

[0048] Figure 4 An exemplary method of a source code repair system 400 is shown. Initially, a group is formed before operating a target vulnerability repair system.

[0049] These groups are formed by finding code changes made to correct specific vulnerabilities. A static analyzer is used to detect vulnerabilities in the source code. The static analyzer analyzes the program and indicates the vulnerability type of any detected vulnerability. A pull request engine extracts code diffs of the changes made to correct the vulnerabilities. A neural encoder is applied to the code diffs and generates corresponding embeddings. The code changes are fed to a neural decoder that generates a natural language description of the code changes. (Generally, block 402).

[0050] Each data point including a code difference embedding and a corresponding natural language description of the code change is clustered to form a group of similar code changes for a specific vulnerability type. On the one hand, a k-means clustering method is used. The k-means clustering method forms groups based on the mean or center centroid value of the code difference embedding and the code change description. The k-means clustering method divides n data points into k<=n sets S={S1, S2,...Sk} given a set of data points (x1, x2,...xn), where each data point is represented by a vector including a code difference embedding and a code change description so as to minimize the difference between the vectors. The k-means clustering method minimizes the Euclidean distance between a data point and the cluster center to which the data point belongs. It should be noted that other clustering techniques may be used, such as Gaussian mixture models, density-based clustering algorithms (DBSCAN), etc. (in general, block 404).

[0051] For each group, select the top k code differences that represent the group. Use profile analysis to determine the code differences that are closest to the group's average. A profile is based on how close or far a vector is from the group's average. Calculate a profile value for each vector in the group, where the vector includes an embedding of the code difference and a description of the code change. The profile value is calculated using the average intra-cluster distance a and the average nearest cluster distance b for each vector, which is the distance between the vector and the nearest cluster of which the vector is not a part. This calculation is mathematically represented as (b a) / max(a, b). Profile values range from -1 to 1. Select the k vectors with the highest scores to represent the group, where k is a user-defined value. (Overall, block 406).

[0052] For each group associated with a particular vulnerability type, the median number of lines changed is calculated from all code differences across all groups associated with the vulnerability type. This median is then used in the hint to guide the large language model toward the desired length of the output. (Generally, block 408).

[0053] Once the group is formed, it is deployed to the target source code repair system (block 410).

[0054] We now turn our attention to the discussion of the inference phase, where the source code repair system generates hints for the large language model to predict the repaired code. Figure 5, an exemplary method 500 for constructing hints for a large language model to generate repaired code is shown.

[0055] The static analyzer identifies vulnerabilities in code snippets of a particular vulnerability type with warning messages (block 502). The hint generator obtains the top k code differences associated with each vulnerability type from the clustered groups (block 504). The hint generator then generates a hint that includes the instruction, the top k code differences, the vulnerable code, and associated context (block 506). The hint is applied to the large language model, which returns the predicted fixed code in the form of code differences (block 508).

[0056] Technical effects / improvements

[0057] Aspects of the subject matter disclosed herein relate to the technical problem of generating hints to guide a large, pre-trained language model to predict fixes for source code excerpts with known vulnerability types. The pre-trained large language model is trained on natural language text and source code, rather than on the task of generating fixes. The hints include a few-shot example of a code diff format that illustrates the fix generation task for guiding the large language model to produce correct output.

[0058] The technical feature associated with solving this problem is to include in the prompts a small number of examples based on code changes made to correct the same vulnerability type. The technical effect achieved is to avoid fine-tuning a large language model on the code repair task without losing accuracy.

[0059] Creating a small set of examples for each vulnerability type consumes significant computing resources and time. The inference phase of the source code remediation system in the target system must execute within strict timing requirements in order to be feasible in the target system. For at least these reasons, the set creation and inference phases performed by the source code remediation system need to be performed on a computing device. Therefore, the operations performed are inherently digital. The human mind cannot directly interface with a CPU, or a network interface card, or other processor, or with RAM or digital storage devices, to read and write the necessary data and perform the necessary operations and processing steps taught herein.

[0060] It is still assumed that the embodiments are capable of operating "at scale," ie, capable of handling larger volumes, in a production environment or in a test lab for a production environment, rather than being considered merely experimental.

[0061] The technique described in this paper is a technical improvement over existing solutions where example selection is either manually hand-picked or uses a hard-coded list of examples. The technique described in this paper can be used in dynamic settings without the need for hand-picked examples.

[0062] Exemplary Operating Environment

[0063] Attention now turns to a discussion of exemplary operating environment 600 . Figure 6 An exemplary operating environment 600 is shown having one or more computing devices 602 and 604 communicatively coupled to a network 606. In one aspect, group construction and prompting can be processed on one computing device 602, and a large language model can be hosted as a service on a second computing device 604. In another aspect, group construction, prompt creation, and the large language model can be hosted on the same computing device. Aspects of the operating environment are not limited to a particular configuration.

[0064] Computing devices 602 and 604 can be any type of electronic device, such as, but not limited to, a mobile device, a personal digital assistant, a mobile computing device, a smartphone, a cellular phone, a handheld computer, a server, a server array or server farm, a web server, a network server, a blade server, an internet server, a workstation, a minicomputer, a mainframe computer, a supercomputer, a network appliance, a web appliance, a distributed computing system, a multi-processor system, or a combination thereof. Operating environment 600 can be deployed in a network environment, a distributed environment, a multi-processor environment, or in a standalone computing device with access to remote or local storage devices.

[0065] Computing devices 602 and 604 may include one or more processors 608 and 640, one or more communication interfaces 610 and 642, one or more storage devices 612 and 646, one or more input / output devices 614 and 644, and one or more memory devices 616 and 648. Processors 608 and 640 may be any commercially available or custom processors and may include dual microprocessor and multi-processor architectures. Communication interfaces 610 and 642 facilitate wired or wireless communication between computing devices 602 and 604 and other devices. Storage devices 612 and 646 may be computer-readable media that do not include propagated signals, such as modulated data signals transmitted via a carrier wave. Examples of storage devices 612 and 646 include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CDROM, digital versatile disks (DVDs) or other optical storage, magnetic cassettes, magnetic tape, and magnetic disk storage, all of which do not include propagated signals, such as modulated data signals transmitted via a carrier wave. Computing devices 602 and 604 may include multiple storage devices 612 and 646. Input / output devices 614 and 644 may include a keyboard, a mouse, a pen, a voice input device, a touch input device, a display, a speaker, a printer, and any combination thereof.

[0066] Memory devices 616 and 648 can be any non-transitory computer-readable storage medium that can store executable programs, applications, and data. A computer-readable storage medium does not involve a propagation signal, such as a modulated data signal transmitted via a carrier wave. It can be any type of non-transitory memory device (e.g., random access memory, read-only memory, etc.), magnetic storage device, volatile storage device, non-volatile storage device, optical storage device, DVD, CD, floppy disk drive, etc., which does not involve a propagation signal, such as a modulated data signal transmitted via a carrier wave. Memory devices 616 and 648 can also include one or more external storage devices or remotely located storage devices, which do not involve a propagation signal, such as a modulated data signal transmitted via a carrier wave.

[0067] Memory devices 616 and 648 may include instructions, components, and data. A component is a software program that performs a specific function and is also referred to as a module, program, component, and / or application. Memory device 616 may include an operating system 618, a source code library 620, a static analyzer 622, a pull request engine 624, a neural encoder 626, a neural decoder 628, a clustering engine 630, one or more groups 632, a hint generator 634, and other applications and data 636. Memory device 648 may include an operating system 650, a large language model 652, and other applications and data 654.

[0068] The computing devices 602 and 604 may be communicatively coupled via a network 606. The network 606 may be configured as an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless WAN (WWAN), a metropolitan area network (MAN), the Internet, a portion of a public switched telephone network (PSTN), a plain old telephone service (POTS) network, a wireless network, network, or any other type of network or combination of networks.

[0069] The network 606 may employ various wired and / or wireless communication protocols and / or technologies. The generations of different communication protocols and / or technologies that may be employed by the network may include, but are not limited to, Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Enhanced Data GSM Environment (EDGE), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access 2000 (CDMA-2000), High Speed Downlink Packet Access (HSDPA), Long Term Evolution (LTE), Universal Mobile Telecommunications System (UMTS), Evolution-Data Optimized (EV-DO), Worldwide Interoperability for Microwave Access (WiMax), Time Division Multiple Access (TDMA), Orthogonal Frequency Division Multiplexing (OFDM), Ultra-Wideband (UWB), Wireless Application Protocol (WAP), User Datagram Protocol (UDP), Transmission Control Protocol / Internet Protocol (TCP / IP), any portion of the Open Systems Interconnection (OSI) model protocols, Session Initiation Protocol / Real-time Transport Protocol (SIP / RTP), Short Message Service (SMS), Multimedia Messaging Service (MMS), or any other communication protocols and / or technologies.

[0070] in conclusion

[0071] A system is disclosed, comprising: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. The one or more programs include instructions for performing the following actions: obtaining a source code snippet having a vulnerability of a vulnerability type; obtaining a few-shot example associated with the vulnerability type, wherein the few-shot example represents a code change made to correct the vulnerability of the vulnerability type; generating a hint for a large language model to generate a fix code to correct the vulnerability in the source code snippet, wherein the hint includes the few-shot example and the source code snippet having the vulnerability; providing the hint to the large language model; and receiving the fix code from the large language model.

[0072] In one aspect, the few-shot example includes a code diff of code changes made to correct a vulnerability. In one aspect, the hint includes instructions for the large language model to generate fixed code. In one aspect, the hint includes instructions for the large language model to output the fixed code in a code diff format. In one aspect, the hint includes the context of the source code snippet. In one aspect, the large language model is a neural transformer model that only considers the decoder. In one aspect, the large language model is a neural transformer model pre-trained with attention on natural language text and source code.

[0073] A computer-implemented method is disclosed, comprising: providing a group having a plurality of code changes made to correct a first type of software vulnerability; obtaining a source code snippet having a vulnerability in the first type of software vulnerability; selecting at least one code change from the group associated with the first type of software vulnerability; generating a hint for a large language model to generate a repair code for correcting the vulnerability in the source code snippet, wherein the hint includes the source code snippet having the vulnerability and the at least one selected code change; and obtaining the repair code from the large language model given the hint.

[0074] In one aspect, the computer-implemented method further comprises: extracting, from a source code repository, a first code change from a plurality of code changes for correcting a first type of software vulnerability; and clustering the first code change and a description of the first code change into a group of closely related code changes. In another aspect, the computer-implemented method further comprises: identifying vulnerabilities in the first type of vulnerability by statically analyzing source code fragments.

[0075] In one aspect, the computer-implemented method further includes generating an embedding of the first code change; generating a description of the first code change; and grouping the embedding of the first code change and the description of the code change into similar groups, wherein the similar groups are based on similar code change embeddings and code change descriptions.

[0076] In one aspect, selecting at least one code change from a group associated with a first type of software vulnerability further comprises selecting at least one code change from the group based on a closest distance to an average value of the group. In one aspect, the hint comprises context of a source code snippet. In one aspect, the hint comprises instructions describing an output format of the code change, wherein the output format comprises a code diff format. In one aspect, the large language model is a neural transformer model with attention.

[0077] A computer-implemented method is disclosed, comprising: obtaining a group including a plurality of code differences associated with a first type of software vulnerability, wherein a code difference in the plurality of code differences is associated with a vector including an embedding of the corresponding code difference and a description of the corresponding code difference; selecting a selected one of the code differences in the group based on closest similarity of the vector of the selected one of the code differences to an average intra-cluster distance and an average nearest cluster distance of each vector in the group; obtaining a source code snippet having the first type of vulnerability; creating a hint for a large language model to generate fixed code for the source code snippet, wherein the hint includes the source code snippet and the selected one of the code differences in the group associated with the first type of vulnerability; and obtaining fixed code for the source code snippet from the large language model given the hint.

[0078] In one aspect, the computer-implemented method further comprises: generating an embedding of the corresponding code difference using an encoder neural transformer model with attention given to the corresponding code difference. In one aspect, the computer-implemented method further comprises: generating a description of the corresponding code difference using a decoder neural transformer model with attention given to the corresponding code difference.

[0079] In one aspect, the prompt includes the context of the source code snippet and instructions indicating the output format of the repaired code. In another aspect, the large language model includes a neural transformer model with attention.

[0080] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

[0081] It will be understood that, unless otherwise indicated, the representative methods do not necessarily have to be performed in the order presented or in any particular order. Furthermore, the various activities described with respect to the methods can be performed in a serial or parallel manner, or any combination of serial and parallel operations. In one or more aspects, the methods illustrate the operation of the systems and devices disclosed herein.

Claims

1. A system comprising: one or more processors; as well as A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the following actions: Get a source code snippet of a vulnerability with the vulnerability type; Obtaining a few sample examples associated with the vulnerability type, wherein the few sample examples represent code changes made to correct the vulnerability of the vulnerability type; generating a hint for a large language model to generate a fix for the vulnerability in the source code snippet, wherein the hint includes the few-shot example and the source code snippet having the vulnerability; providing the prompt to the large language model; as well as The repaired code is received from the large language model. 2 . The system of claim 1 , wherein the few-shot examples include code differences in the code changes made to correct the vulnerability.

3. The system of claim 1, wherein the hint comprises instructions instructing the large language model to generate the repaired code. 4 . The system of claim 1 , wherein the hint comprises instructions instructing the large language model to output the repaired code in a code difference format. The system of claim 1 , wherein the hint comprises a context of the source code snippet.

6. The system of claim 1, wherein the large language model is a decoder-only neural transformer model with attention.

7. The system of claim 1, wherein the large language model is a neural transformer model with attention pre-trained on natural language text and source code.

8. A computer-implemented method comprising: providing a group having a plurality of code changes made to correct a first type of software vulnerability; Obtaining a source code fragment having a vulnerability among the first type of software vulnerabilities; selecting at least one code change from the group associated with the first type of software vulnerability; generating a hint for a large language model to generate code for a fix to correct the vulnerability in the source code snippet, wherein the hint includes the source code snippet having the vulnerability and the selected at least one code change; as well as Given the hint, the repaired code is obtained from the large language model.

9. The computer-implemented method of claim 8, further comprising: extracting from a source code repository a first code change of the plurality of code changes for correcting the first type of software vulnerability; as well as The first code change and the description of the first code change in the plurality of code changes are clustered into a group of closely related code changes.

10. The computer-implemented method of claim 8, further comprising: The vulnerability in the first type of vulnerability is identified by static analysis of the source code fragment.

11. The computer-implemented method of claim 9, further comprising: generating an embedding of the first code change; generating a description of the first code change; as well as The embeddings of the first code change and the descriptions of the code changes are clustered into similar groups, wherein the similar groups are based on similar code change embeddings and code change descriptions.

12. The computer-implemented method of claim 9, wherein selecting the at least one code change from the group associated with the first type of software vulnerability further comprises: The at least one code change is selected from the group based on a closest distance to a mean value of the group.

13. The computer-implemented method of claim 8, wherein the hint comprises a context of the source code snippet.

14. The computer-implemented method of claim 8, wherein the hint comprises instructions describing an output format of the code change, wherein the output format comprises a code diff format.

15. The computer-implemented method of claim 8, wherein the large language model is a neural transformer model with attention.