System and method for generating search extension patches for automatic program repair
The search extension patch generation framework enhances automatic program repair by using a hybrid retriever and CodeT5 model to retrieve and synthesize effective code patches, addressing inefficiencies in existing tools and improving accuracy and efficiency.
Patent Information
- Application Number
- JP2024568263
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-26
- Filing Date
- 2023-05-05
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-05-05
AI Technical Summary
Existing automatic program repair tools face limitations in accuracy and efficiency due to reliance on heuristic rules and redundancy assumptions, particularly in complex program repair scenarios, leading to inefficient and inaccurate code patch generation.
A search extension patch generation framework using a hybrid patch retriever that combines lexical and semantic matching, integrated with a pre-trained Transformer-based encoder-decoder model like CodeT5, to retrieve and synthesize repair patches from a codebase, leveraging both sparse and dense searches.
The framework significantly improves the accuracy and efficiency of code repair by providing relevant and diverse patch generation, outperforming existing methods in benchmarks, especially in handling various error types and correction patterns.
Smart Images

Figure 2025521113000001_ABST
Abstract
Description
Technical Field
[0001] [Cross - Reference] This disclosure claims priority to U.S. Non - Provisional Patent Application No. 17 / 896,873, filed Aug. 26, 2022, which claims priority to U.S. Provisional Patent Application No. 63 / 343,264, filed May 18, 2022, under 35 U.S.C. § 119, and these are hereby incorporated by reference in their entirety.
[0002] [Technical Field] Embodiments generally relate to machine learning and automatic code generation, and more specifically to systems and methods for automatic program repair (APR) using retrieval - augmented patch generation (RAP - Gen).
Background Art
[0003] Software developers often spend a great deal of time and energy debugging and repairing source code, making software development expensive and time - consuming. Among existing automatic program repair tools, there are some that reduce the difficulty and cost of program repair in use cases that involve searching for patches during development, build, or execution. For example, in some search - based (also called generate - and - validate) approaches, repairs can be searched for based on manually - mined correction patterns or redundancy - based techniques. Redundancy - based techniques generally make the redundancy assumption that in many cases, the modified patch can be found (or re - constructed) from other places (donor code snippets) within the codebase. Thus, these conventional search - based techniques have limited accuracy and efficiency when repairing programs.
[0004] Therefore, a more efficient method for automatic program repair is needed.
Brief Description of the Drawings
[0005]
Figure 1
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
[0006] In these figures, elements having the same reference numerals have the same or similar functions.
DETAILED DESCRIPTION OF THE INVENTION
[0007] As used herein, the term "network" may comprise any hardware or software-based framework that includes any artificial intelligence network or system, neural network or system, and / or any training or learning model implemented thereon or therewith.
[0008] As used herein, the term "module" may include a hardware or software-based framework that executes one or more functions. In some embodiments, the module may be implemented on one or more neural networks.
[0009] Existing automatic program repair systems can reduce manual debugging efforts and improve software reliability. Conventional search-based techniques typically rely on heuristic rules or redundancy assumptions to mine repair patterns. Some deep learning-based approaches can automate the program repair process by training a learning model to generate code repair patches. However, the performance of such learning models is often limited by a fixed set of parameters for modeling the very complex search space of program repair.
[0010] In view of the need for an efficient and accurate code repair system, the embodiments described herein provide a search extension patch generation framework for searching for code patches using a patch retriever based on related repair patterns. Specifically, a hybrid patch retriever can be configured for repair pattern mining that takes into account both lexical matching and semantic matching through sparse and dense searches based on raw source code. Also, since this retriever does not require language-specific features such as abstract syntax trees, it is a language-independent retriever. One improvement over previous repair pattern mining models is that instead of clustering various repair templates, the retriever utilizes the top one related bug-fix pair as a guiding repair pattern for each buggy patch. This strategy is consistent with the debugging behavior of human developers, who often explore related bug-fix examples to extract several repair leads for bug fixing.
[0011] In one embodiment, a pre-trained Transformer-based encoder-decoder model (e.g., the CodeT5 model) can be employed as the base patch generator. CodeT5 is a generic programming language model pre-trained on a large-scale source code corpus using a code-aware language modeling objective. A two-stage training strategy can be used to train the pre-trained encoder-decoder model to connect the patch retriever and the CodeT5 patch generator. The patch retriever first explores related bug-fix patterns and then passes them to the patch generator to synthesize the repaired patches based on both the source buggy code and the external (searched) bug-fix knowledge. Then, the retrieved repair patterns can be directly added to the source buggy patch. In this way, the retriever can be integrated with any sequence-to-sequence learning-based model for searching in repair pattern mining for program repair.
[0012] FIG. 1 is a schematic diagram of a computing device 100 for implementing an automatic program repair framework shown in FIG. 3 according to some embodiments. As shown in FIG. 1, the computing device 100 includes a processor 110 coupled to a memory 120. The operation of the computing device 100 is controlled by the processor 110. Also, although the computing device 100 is shown with only one processor 110, the processor 110 can represent one or more central processing units, multi-core processors, microprocessors, microcontrollers, digital signal processors, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), graphics processing units (GPUs), etc. within the computing device 100. It is understood that the computing device 100 can be implemented as a stand-alone subsystem, as a board added to a computing device, and / or as a virtual machine.
[0013] The memory 120 can be used to store software executed by the computing device 100 and / or one or more data structures used during the operation of the computing device 100. The memory 120 can include one or more types of machine-readable media. Some common forms of machine-readable media include floppy (registered trademark) disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, any other optical media, punch cards, paper tapes, any other physical media having a pattern of holes, RAM, PROM, EPROM, FLASH-EPROM, any other memory chip or cartridge, and / or any other media adapted to be read by a processor or computer.
[0014] Processor 110 and / or memory 120 can be arranged in any suitable physical configuration. In some embodiments, processor 110 and / or memory 120 can be implemented on the same board, within the same package (e.g., system-in-package), on the same chip (e.g., system-on-chip), etc. In some embodiments, processor 110 and / or memory 120 can include distributed, virtualized, and / or containerized computing resources. Consistent with such embodiments, processor 110 and / or memory 120 can be located in one or more data centers and / or cloud computing facilities.
[0015] In some examples, memory 120 can include a non-transitory tangible machine-readable medium that includes executable code that, when executed by one or more processors (e.g., processor 110), can cause the one or more processors to execute the methods described in more detail herein. For example, as illustrated, memory 120 can include instructions for an automatic program repair module 130 that can be used to implement and / or emulate a system and model and / or to implement any of the methods further described herein. The automatic program repair module 130 can receive an input 140 that includes an input such as a program bug via a data interface 115. The automatic program repair module 130 can generate an output 150 such as a code patch.
[0016] In some embodiments, the automatic program repair module 130 includes a retriever encoder sub-module 131, a patch retriever sub-module 132, and a patch generator sub-module 133. In one embodiment, the automatic program repair module 130 and its sub-modules 131-133 can be implemented by hardware, software, and / or a combination thereof.
[0017] Some examples of computing devices, such as computing device 200, may include a non-transitory tangible machine-readable medium that includes executable code that, when executed by one or more processors (e.g., processor 110), can cause the one or more processors to execute a process of a method. Some common forms of machine-readable media that may include a process of a method include, for example, a floppy (registered trademark) disk, a flexible disk, a hard disk, a magnetic tape, any other magnetic medium, a CD-ROM, any other optical medium, a punch card, a paper tape, any other physical medium having a pattern of holes, a RAM, a PROM, an EPROM, a FLASH-EPROM, any other memory chip or cartridge, and / or any other medium adapted to be read by a processor or a computer.
[0018] Figure 2 is a simplified block diagram of a networked system suitable for implementing the automatic program repair framework described in Figure 3 and other embodiments described herein. In one embodiment, block diagram 200 shows a system that includes a user device 210 that can be operated by a user 240, database vendor servers 245, 270, and 280, a server 230, and other forms of device, server, and / or software components that operate to execute various methodologies according to the described embodiments. Exemplary devices and servers can include devices similar to the computing device 100 described in FIG. 1 that operate an OS such as a MICROSOFT® OS, UNIX® OS, LINUX® OS, or other suitable device and / or server-based OS, stand-alone, and enterprise-class servers. It can be understood that the devices and / or servers shown in FIG. 2 may be deployed in other ways, and the operations performed and / or the services provided by such devices and / or servers may be combined or separated for a given embodiment and may be performed by more or fewer devices and / or servers. One or more devices and / or servers may be operated and / or maintained by the same or different entities.
[0019] The user device 210, database vendor servers 245, 270, and 280, and server 230 can communicate with each other via a network 260. The user device 210 can be utilized by a user 240 (e.g., a driver, system administrator, etc.) to access various functions available on the user device 210, which may include processes and / or applications associated with the server 230 to receive an output data anomaly report.
[0020] The user device 210, the database vendor server 245, and the server 230 may each include one or more processors, memories, and other appropriate components for executing instructions such as program code and / or data stored on one or more computer-readable media to implement the various applications, data, and steps described herein. For example, such instructions may be stored on one or more computer-readable media, such as memories or data storage devices internal and / or external to the various components of the system 200 and / or accessible via the network 260.
[0021] The user device 210 may be implemented as a communication device that utilizes appropriate hardware and software configured for wired and / or wireless communication with the database vendor server 245 and / or the server 230. For example, in one embodiment, the user device 210 may be implemented as an autonomous vehicle, a personal computer (PC), a smartphone, a laptop / tablet computer, a wristwatch having appropriate computer hardware resources, glasses having appropriate computer hardware (e.g., GOOGLE GLASS (registered trademark)), other types of wearable computing devices, an embedded communication device, and / or other types of computing devices capable of transmitting and / or receiving data such as an IPAD (registered trademark) of APPLE (registered trademark). Although only one communication device is shown, multiple communication devices may function similarly.
[0022] The user device 210 of FIG. 2 includes a user interface (UI) application 212 and / or other applications 216, which may correspond to applications having executable processes, procedures, and / or related hardware. For example, the user device 210 may receive a message indicating buggy code and / or fixed code from the server 230 and display the message via the UI application 212. In other embodiments, the user device 210 may include additional or different modules having dedicated hardware and / or software as needed.
[0023] In various embodiments, the user device 210 includes other applications 216 that may be desired in certain embodiments to provide functionality to the user device 210. For example, other applications 216 may include a security application for implementing client-side security features, a program client application for interfacing with an appropriate application programming interface (API) via the network 260, or other types of applications. Other applications 216 may also include communication applications such as email, texting, voice, social networking, and IM applications that enable the user to send and receive email, calls, texts, and other notifications via the network 260. For example, other applications 216 may be an email or instant messaging application that receives prediction result messages from the server 230. Other applications 216 may include a device interface and other display modules that may receive input and / or output information. For example, other applications 216 may include a software program for asset management executable by a processor that includes a graphical user interface (GUI) configured to provide the user 240 with an interface for viewing buggy code and / or fixed code.
[0024] The user device 210 may further include a database 218 stored in the temporary and / or non-temporary memory of the user device 210 that stores various applications and data and can be utilized during the execution of various modules of the user device 210. The database 218 may store a user profile regarding the user 240, predictions previously viewed or saved by the user 240, historical data received from the server 230, and the like. In some embodiments, the database 218 may be local to the user device 210. However, in other embodiments, the database 218 may be external to the user device 210 and accessible by the user device 210 including a cloud storage system and / or a database accessible via the network 260.
[0025] The user device 210 includes at least one network interface component 219 adapted to communicate with the database vendor server 245 and / or the server 230. In various embodiments, the network interface component 219 may include DSL (e.g., Digital Subscriber Line) modems, PSTN (Public Switched Telephone Network) modems, Ethernet® devices, broadband devices, satellite devices, and / or various other types of wired and / or wireless network communication devices including microwave, radio frequency, infrared, Bluetooth®, and near-field communication devices.
[0026] The database vendor server 245 may correspond to a server that hosts one or more of the databases 203a - n (or collectively referred to as 203) to provide the server 230 with a training data set including pairs of buggy code and corrected code. The database 203 may be implemented by one or more relational databases, distributed databases, cloud databases, and the like.
[0027] The database vendor server 245 includes at least one network interface component 226 adapted to communicate with the user device 210 and / or the server 230. In various embodiments, the network interface component 226 can include a DSL (e.g., Digital Subscriber Line) modem, a PSTN (Public Switched Telephone Network) modem, an Ethernet® device, a broadband device, a satellite device, and / or various other types of wired and / or wireless network communication devices including microwave, radio frequency, infrared, Bluetooth®, and near field communication devices. For example, in one implementation, the database vendor server 245 can transmit asset information from the database 203 to the server 230 via the network interface 226.
[0028] The server 230 can house the automatic program repair module 130 and its sub-modules described in FIG. 1. In some implementations, the module 130 can receive data from the database 219 in the database vendor server 245 via the network 260 to generate a modified patch of the code. The generated modified patch of the code can be transmitted to the user device 210 for review by the user 240 via the network 260.
[0029] The database 232 can be stored in the temporary and / or non-temporary memory of the server 230. In one implementation, the database 232 can store data obtained from the database vendor server 245. In one implementation, the database 232 can store parameters of the automatic program repair module 130. In one implementation, the database 232 can store previously generated modified patches of the code and corresponding input feature vectors.
[0030] In some embodiments, database 232 may be local to server 230. However, in other embodiments, database 232 is external to server 230 and may be accessible by server 230 including a cloud storage system and / or database accessible via network 260.
[0031] Server 230 includes at least one network interface component 233 adapted to communicate with user device 210 and / or database vendor servers 245, 270, or 280 via network 260. In various embodiments, network interface component 233 may include various other types of wired and / or wireless network communication devices including, for example, a DSL (e.g., Digital Subscriber Line) modem, a PSTN (Public Switched Telephone Network) modem, an Ethernet® device, a broadband device, a satellite device, and / or microwave, radio frequency (RF), and infrared (IR) communication devices.
[0032] Network 260 may be implemented as a single network or a combination of multiple networks. For example, in various embodiments, network 260 may include the Internet or one or more intranets, landline communication networks, wireless networks, and / or other suitable types of networks. Accordingly, network 260 may correspond to a small-scale communication network such as a private or local area network accessible by various components of system 200, or a large-scale network such as a wide area network or the Internet.
[0033] FIG. 3 is an exemplary block diagram showing an exemplary architecture of an automatic program repair framework 300, also referred to as the RAP-Gen framework 300, according to the embodiments described herein. The RAP-Gen framework 300 is aimed at generating a target program patch based on an input bug-containing patch, along with related bug-fixing patterns through search.
[0034] The task formulation of search-augmented patch generation for automatic program repair is described as follows.
Number
Number
[0035] In some embodiments, the original input sequence X i 308 is extended with the retrieved bug-fixing pairs to form a new input sequence 312, e.g., [Number] Next, a patch generator 306 (for example, one that uses a sequence-to-sequence (seq2seq) generator, also called a sequence generator 306) autoregressively [Number] from Y i 316 can be generated. The framework 300 uses the patch generator 306 parameterized by θ to probabilistically [Number] can learn, where Y i,1 :Y i,k-1 is the sequence before the previous of the k-th token, and n indicates the number of tokens in the target sequence Yi. In some embodiments, the external codebase C302 can be regarded as a non-parametric memory, and the retrieved bug fix pair 310 can be regarded as a guiding correction pattern for the patch generation model 306. Probabilistically, the search Z j (B j ,F j ) can be formulated as a latent variable, which can be approximated by top-1 in some cases. Formally, [Number] is, where [Number] is the top-1 retrieved output from the retriever P φ (Z j X i ). Marginalization for k > 1 makes training and inference complex and inefficient, so a top-1 approximation can be adopted for efficiency improvement. In some embodiments, top-k (e.g., k = 2, 3, 5) using the Freshness Initiation Distance (FiD) method can be used.
[0036] As shown in the example of FIG. 3, the RAP-Gen framework 300 includes a patch retriever 304 and a code recognition pre-trained patch generator 306. The patch retriever 304 is configured to search for relevant correction patterns useful for automatic program repair. This calculates the relevance between a (query) bug-containing patch Xi 308 and a previous (key) bug-containing patch B φ (X i , B j ) in the codebase C302 based on the relevance scoring function f j . In various embodiments, the patch retriever 304 may include a lexical-based retriever (e.g., BM25) and / or a semantic-based retriever (e.g., Dense Passage Retrieval (DPR)). In the example of FIG. 3, the patch retriever 304 includes a neural network model (e.g., retriever encoder 318) and uses a hybrid approach to combine a lexical-based retriever (e.g., BM25) and a semantic-based retriever (e.g., DPR) to take into account both lexical and semantic information.
[0037] Lexical-based Retriever. In some embodiments, a lexical-based retriever (e.g., BM25) may be implemented using a term-based retriever and may use a sparse vector representation for lexical matching. The lexical-based retriever converts each code patch into a bag-of-words representation and can calculate the lexical similarity between the query patch X i and the candidate patch B j . The calculated similarity score is f φ (X i , B j ) = BM25(X i , B j) is represented as. In one example, a sparse term-based retriever may be sensitive to the choice of identifier naming in source code that does not affect the semantics of the code.
[0038] Semantic-based Retriever. In some embodiments, the semantic-based retriever may be implemented using a Dense Passage Retriever (DPR), and may search for relevant patches by measuring their semantic similarity. In some embodiments, an encoder (e.g., a Transformer-based encoder) may be used to map each patch to a fixed-size dense vector in order to encode the code patches. The DPR may be initialized from the encoder of a pre-trained Transformer-based neural network model (e.g., Code Bidirectional Encoder Representations from Transformers (CodeBERT), etc.). The encoder may be pre-trained using a large code repository of one or more programming languages (e.g., a GitHub code repository of six programming languages). In one example, the final layer hidden state of the [CLS] token from the encoder is used as the patch representation. In some embodiments, a shared DPR is used to separately encode the query patch X i 308 and candidate patch B in C j respectively,
Number
Number
[0039] In some embodiments, a shared DPR is used to separately encode the query patch X iCandidate patches F within 308 and C j are, respectively, [Number] encoded separately as. Then, as follows, the similarity is calculated by the inner product between these two patch representations: [Number]
[0040] The description in this specification generally uses the similarity between X i and B j (e.g., f φ (X i , B j )), but the similarity used for searching may include the similarity between X i and B j (e.g., f φ (X i , B j ))), the similarity between X i and F j (e.g., f φ (X i , F j ))), and / or combinations thereof. Note that
[0041] In some embodiments, a meaning-based retriever (e.g., DPR) is further trained using a training dataset that includes pairs of buggy patches and fixed patches. In one example, a codebase 302 including bug fix pairs can be used by treating the buggy code B j as a query and the corresponding fixed code F j as a key. This can be performed based on the assumption that buggy patches and their fixed patches share similar semantics (e.g., identifiers, data flow, and code structure) in many cases. This technique can be used to avoid the huge amount of manual annotation work required to curate a bug-to-bug search dataset.
[0042] In an example where the bug-fix pair is used as a query and a corresponding key, a contrastive learning method 314 using in-batch negatives is used to train a meaning-based retriever, and the in-batch negatives are used to optimize a contrastive loss (e.g., InfoNCE contrastive loss) as follows: [Number] Here, M is the current mini-batch, and N represents the number of positive training examples in the mini-batch. This objective aims to maximize the similarity between positive examples while minimizing the similarity between negative examples. Each positive example can have |M| - 1 negative samples. Various contrastive learning techniques, such as in-batch negative strategies, hard negative mining strategies, etc., can be used. However, it should be noted that in some embodiments, the contrastive learning using in-batch negatives described above provides better performance than the hard negative mining strategy for noisy training data.
[0043] In some embodiments, at the inference stage, when a query bug-containing patch X i 308 is given, a meaning-based retriever (e.g., DPR) calculates the similarity between X i (query) and B j (key) to retrieve the relevant bug-fix pair (B j , F j ). In some embodiments, the meaning-based retriever can retrieve the relevant bug-fix pair based on the similarity between X i and F j , and / or a combination of the similarity between X i (query) and B j (key).
[0044] Hybrid Retriever. As shown in the example of Figure 3, in some embodiments, a hybrid approach that combines a lexical retriever (e.g., BM25) and a semantic retriever (e.g., DPR) is utilized to take into account both lexical information and semantic information. For example, the similarity score can be calculated as
Number
Number
[0045] In the example of Figure 3, the RAP-Gen framework 300 includes a patch generator 306 for generating a modified code patch. In some embodiments, an input bug-containing patch 308, also referred to as a source bug-containing patch 308 or a query bug-containing patch 308 (X i shown as), and a retrieved bug-fixing pattern 310 (B j , F j shown as) are used by the code recognition pre-training patch generator 306 to generate an extended bug-containing code patch 312, for example, using concatenation such as:
Number
[0046] In some embodiments, the patch generator 306 includes a code recognition programming language model pre-trained on a large-scale source code corpus. In one example, the sequence generator uses CodeT5, an integrated pre-trained Transformer-based encoder-decoder model that achieves state-of-the-art (SoTA) results in multiple code intelligence tasks such as defect detection and code refinement. This can be pre-trained on 8.3 million functions in 8 different programming languages (including JavaScript® and Java®) collected from GitHub. CodeT5 can adopt identifier recognition pre-training purposes to incorporate code-specific knowledge into the language model. This can provide a code-specific byte pair encoding (BPE) tokenizer optimized for code and may be able to avoid the Out-of-Vocabulary (OoV) problem. CodeT5 can be used in the patch generator 306 that can provide strong code understanding capabilities.
[0047] As shown in the example of FIG. 3, the search expansion input 312 to the patch generator 306 (e.g., CodeT5) is
Number
Number
[0048] In various embodiments, the RAP-Gen framework 300 leverages general code understanding knowledge encoded through pre-training on a large code corpus (e.g., using CodeT5). For example, the source input sequence 312 may be generated by concatenating the original bug-containing code patch 308 and the top-ranked bug-fix pairs 310 from the patch retriever 304. In some embodiments, the extended source input bug-containing patch 312 may be generated by concatenating the top-k (e.g., k=2, 3, 5) retrieved bug-fix pairs to the input bug-containing patch 308.
[0049] 4A is an exemplary logic flow diagram illustrating a method for training a search extension patch generation framework for automatic program repair as shown in FIG. 3 according to some embodiments described herein. One or more of the processes of the method 400 may be implemented, at least in part, in the form of executable code stored on a non-transitory tangible machine-readable medium that, when executed by one or more processors, may cause the one or more processors to perform one or more of the processes. In some embodiments, the method 400 corresponds to the operation of the automatic program repair module 130 (e.g., FIG. 1) for performing automatic program repair using search extension patch generation.
[0050] In step 402, a patch retriever including a retriever encoder is provided. In the example of FIG. 3, in the search extension patch generation framework 300, a patch retriever 304 including a retriever encoder 318 is provided. In some embodiments, as shown in step 404, the retriever encoder 318 is pre-trained using, for example, a first training dataset including a large-scale programming language corpus (e.g., a GitHub code repository or other suitable code repository in one or more programming languages).
[0051] In step 406, a patch generator including a sequence generator neural network model is provided. In the example of FIG. 3, in the search extension patch generation framework 300, the patch generator 306 includes a sequence generator neural network model, specifically, a Transformer-based encoder-decoder model including a generator encoder 318 and a generator decoder 320. In some embodiments, as shown in step 404, the patch generator 306 is pre-trained using, for example, a second training dataset including a large-scale programming language corpus (e.g., a GitHub code repository or other suitable code repository in one or more programming languages).
[0052] In step 410, a RAP-Gen framework including a patch retriever and a patch generator (e.g., the RAP-Gen framework of FIG. 3) can be trained using, for example, a two-stage training process. The two-stage training process includes step 412 where the first-stage training is performed by training the patch retriever using a third training dataset. In some embodiments, the third training dataset can use bug-fix pairs within the codebase 302. For example, when using the third training dataset to train the semantic retriever of the patch retriever 304, the buggy code Bj of the bug-fix pair within the codebase can be regarded as a query, and the corresponding fixed code Fj can be regarded as a key. In another example, the fixed code Fj of the bug-fix pair within the codebase can be regarded as a query, and the corresponding buggy code Bj can be regarded as a key. This is based on the assumption that buggy patches and their fixed patches share similar semantics (e.g., identifiers, data flow, and code structure) in many cases. By using bug-fix pairs within the codebase 302 for the third training dataset, the huge amount of manual annotation work required to curate the bug-to-bug search dataset as the third training dataset can be avoided. In one example, the first-stage training can use a contrastive learning algorithm by optimizing the contrastive loss.
[0053] The two-stage training process includes step 414 where the second-stage training is performed by training the patch generator using a fourth training dataset using the patch retriever trained by the first-stage training. In one example, when the input to the patch generator is generated using the original input buggy code patch and the top-ranked bug-fix pairs from the trained patch retriever, a teacher-forcing algorithm is used to minimize the language modeling loss.
[0054] In an example where, during the second stage of training, a fourth training set is generated from the bug-fixed pair codebase, the patch retriever (already trained using the first stage of training) is not permitted to access the ground truth bug-fixed pairs. Otherwise, the training loss would easily drop close to zero since the patch generator could directly copy the retrieved fixes as the target output. In that example, each sample of the fourth training set is a bug-containing patch of the corresponding bug-fixed pair from the codebase (also called the ground truth bug-fixed pair), and the corresponding ground truth is the fixed patch of the corresponding bug-fixed pair. For each sample bug-containing patch input, another bug-fixed pair (not the ground truth one) is retrieved from the codebase by the patch retriever. The retrieved bug-fixed pair is added to the bug-containing patch input to generate an extended sequence input for the patch generator. Note that the requirement of no access to the ground truth bug-fixed pairs applies only to the second stage of training when the codebase is used to provide the fourth training set, and not to the first stage of training the patch retriever when the codebase is used to provide the third training set.
[0055] How are the third and fourth datasets generated? Recall that there is a bug-fixed pair for each downstream dataset, which is exactly the third training set.
[0056] Referring to FIG. 4B, an exemplary logic flow diagram showing a method 450 of an inference process using a trained search extension patch generation framework according to some embodiments described herein is shown. One or more of the processes of method 450, when executed at least in part by one or more processors, can be implemented in the form of executable code stored on a non-transitory tangible machine-readable medium that can cause the one or more processors to execute one or more of the processes. In some embodiments, method 450 corresponds to the operation of an automatic program repair module 130 (e.g., FIG. 1) for performing automatic program repair using search extension patch generation.
[0057] In step 452, a first bug-containing patch is received by the trained search extension patch generation framework. In the example of FIG. 3, the trained search extension patch generation framework 300 receives the first bug-containing patch 308 and provides it as an input to its trained patch retriever 304.
[0058] In step 454, one or more bug-fix pairs are provided based on the first bug-containing patch. In the example of FIG. 3, the trained patch retriever 304 receives the first bug-containing patch 308 and searches for one or more bug-fix pairs from the codebase 302, for example, based on the similarity between the first bug-containing patch 308 and the bug-fix pairs. In various embodiments, the similarity is determined based on the similarity between the first bug-containing patch 308 and the bug-containing patch of the bug-fix pair, the similarity between the first bug-containing patch 308 and the fixed patch of the bug-fix pair, or a combination thereof. The similarity can include lexical similarity, semantic similarity, or a combination thereof.
[0059] In step 456, a first extended bug-containing patch is generated based on the first bug-containing patch and the one or more bug fix pairs retrieved. In the example of FIG. 3, the first extended bug-containing patch 312 is generated using the first bug-containing patch 308 and the one or more bug fix pairs 310 provided by the patch retriever 304. The first extended bug-containing patch 312 is provided to the patch generator 306.
[0060] In step 458, a first corrected patch for the first bug-containing patch is generated using the first extended bug-containing patch. In the example of FIG. 3, the patch generator 306 receives the first extended bug-containing patch 312 and generates a first corrected patch 316 based on the first extended bug-containing patch 312. Exemplary data experiments and performance
[0061] Referring to Fig. 5, in several experiments, the RAP-Gen framework was evaluated on two common APR datasets, namely, TFix in JavaScript (Berkay Berabi, Jingxuan He, Veselin Raychev, and Martin T. Vechev, TFix: Learning to Fix Coding Errors with a Text-to-Text Transformer, Proceedings of Machine Learning Research (PMLR), Vol. 139, 780-791) and Code Refinement in Java (Michele Tufano, Cody Watson, Gabriele Bavota, Massimiliano Di Penta, Martin White, and Denys Poshyvanyk, An Empirical Study on Learning Bug-Fixing Patches in the Wild via Neural Machine Translation, ACM Trans. Softw. Eng. Methodol. 28, 4 (2019), 19:1-19:29). Both datasets were originally collected from GitHub commits, but there is a major difference that while the bug-fix pairs in TFix can be verified by a static analyzer, the pairs in Code Refinement are verified by checking whether the commit messages contain keywords such as "fix bug". The data statistics of the TFix and Code Refinement benchmarks are shown in Table 1 of Fig. 5.
[0062] TFix. Specifically, TFix is a large-scale program repair dataset containing JavaScript code patch pairs curated from 5.5 million GitHub commits. This comprehensively covers 52 unique error types detected by the static analyzer ESLint. In addition to error types, it provides rich error annotations such as error messages and localized error lines, eliminating the need for fault localization as in conventional work. In TFix, the APR task is tackled as a text-to-text generation problem using T5-large. In the source input sequence, they combine all error information with the bug-containing code patch into a single text: fix {error type} {error message} {error context} Here, the error context consists of the given localized error line, and its two adjacent code lines are used to form the bug-containing code patch. The target sequence is to replace the error line with the corrected line in the error context. The same data format is adopted in the experiments, and data examples can be found in the source input of Figure 6 (showing one bug-fixing example in the TFix test set where the RAP-Gen framework correctly fixes the bug).
[0063] During data processing, duplication problems within and between data splits are observed. Specifically, there are 114, 2, and 4 duplicates in the training, validation, and test splits respectively. Regarding duplication between splits, there are 28, 34, and 4 duplicates between the training and test, training and test, and validation and test splits respectively. These duplicates (243) are filtered, and the deduplicated version TFix (Dedup) is shown in Table 1 of Figure 5.
[0064] Code Refinement. Tufano et al. released two code refinement datasets containing bug-fix pairs at the function level, collected from the GitHub Archive (https: / / www.gharchive.org / ) published between March 2011 and October 2017. To ensure the quality of the collected bug-fix function pairs, the Google BigQuery API was used to identify all Java commits with messages containing the patterns of (「fix」 or 「solve」) and (「bug」 or 「issue」 or 「problem」 or 「error」). The functions were normalized by obfuscating the identifiers using indexed tokens such as TYPE1, VAR1, METHOD1, etc. One data example can be seen in Figure 7 (showing one bug-fix example in the Refinement Small test set, where the RAP-Gen framework is making a correct prediction). The two data subsets are determined by the number of tokens, i.e., for the small set, the number of code tokens ≤ 50, and for the medium set, 50 < the number of code tokens ≤ 100.
[0065] In some embodiments, the RAP-Gen framework 300 can be fine-tuned for each benchmark using, for example, the AdamW optimizer (Ilya Loshchilov and Frank Hutter Decoupled Weight Decay Regularization, ICLR, 2019) with sequence-to-sequence generation loss (e.g., for 30 epochs). Grid search can be performed for hyperparameter tuning using various batch sizes (e.g., 16, 32, 64) and learning rates (e.g., 1e-4, 5e-5, 2e-5). For example, a learning rate of 1e-4 and a batch size of 64 can be used for TFix, and a learning rate of 5e-5 and a batch size of 32 can be used for Code Refinement. In one example, the training time for RAP-Gen-base in each benchmark using one A100 GPU is within 2 days. During inference, beam search can be employed with a beam size of 5 to generate a ranked list of synthesized and modified patches.
[0066] In some embodiments, bug-fix pairs in the training set are adopted as the search codebase for constructing the patch retriever 304. For the vocabulary-based retriever, an exemplary open-source Python library (e.g., https: / / pypi.org / project / rank-bm25 for BM25) can be used. As a sparse term-based retriever, the choice of tokenizer greatly affects the search performance. In the experiment, the CodeT5 tokenizer, which is a code-specific BPE tokenizer optimized for code tokenization, is adopted. The BM25 search engines for the benchmarks TFix and Code Refinement are applied on a 95CPU machine with 600G of memory. Each experiment finishes within 1 hour with multiprocessing.
[0067] In the experiment, for the meaning-based retriever, CodeBERT initialized with DPR is used to encode each patch into a dense vector for semantic matching. Also, the InfoNCE contrastive loss is used to fine-tune the DPR model for each benchmark with 50 epochs. A batch size of 64 and a learning rate of 2e-5 are used for fine-tuning on one A100 GPU with 40G memory. The training time for TFix and Code Refinement is about 9 and 5 GPU hours respectively.
[0068] For the hybrid retriever, the BM25 and DPR ranking scores are calculated, and these normalized scores are linearly combined with equal weights to construct the hybrid retriever, namely "Hybrid". For all retrievers, the CodeT5 tokenizer is used to encode patches with a maximum sequence length of 256.
[0069] Evaluation Metrics. For evaluation metrics, the smoothed BLEU-4 (Chin-Yew Lin and Franz Josef Och, ORANGE: a Method for Evaluating Automatic Evaluation Metrics for Machine Translation, COLING, 2004) score and exact match (EM) accuracy are used to evaluate program repair performance (e.g., following Yue Wang, Weishi Wang, Shafiq R. Joty, and Steven C. H. Hoi, CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation, EMNLP, Association for Computational Linguistics, 8696-8708). BLEU-4 is a looser metric for evaluating the degree of subword overlap, and EM is a stricter metric that requires the prediction to be identical to the ground truth patch in the actual commit. Since buggy programs can have different repair methods, the Error Removal metric (such as that used in TFix) is used to account for various forms of fixes. A prediction is counted as correct for Error Removal if the existing error is removed and no new error is introduced after the fix. For all metrics, the results are presented on a scale of 0 to 100 (%), with higher scores indicating better performance.
[0070] Baseline Models. The RAP-Gen framework is compared with learning-based models in two program repair benchmarks. CoCoNuT is a context-aware neural machine translation framework based on a convolutional encoder-decoder model. SequenceR is an LSTM-based sequence-to-sequence generation model with a copy mechanism. Additionally, the RAP-Gen framework is compared with a programming language model pre-trained based on the Transformer architecture. One group of these models are encoder-only models such as RoBERTa (code), CodeBERT, and GraphCodeBERT. These encoder-only models require a decoder randomly initialized for the program repair task.
[0071] Furthermore, the RAP-Gen framework is compared with an encoder-decoder Transformer model. PLBART is an integrated pre-trained model with the purpose of noise removal including token masking, token deletion, and token infilling. TFix is initialized with a T5-large checkpoint and continues fine-tuning on the TFix dataset. CoTexT is another T5-based model pre-trained on both text and code. NSEdit is a language model where the encoder and decoder are initialized from CodeBERT and CodeGPT respectively. This is fine-tuned to generate repairs via neural symbolic editing sequences and ranked as the current SoTA model on the Code Refinement benchmark. Results from all baseline models are obtained from their original papers.
[0072] In the experiment, we verify that search-augmented patch generation is an effective approach for program repair. To compare RAP-Gen with the traditional learning-based methods on two benchmarks, we conduct comprehensive experiments. First, the CodeT5 model is evaluated against TFix, and the evaluation is improved by providing a deduplicated version of the dataset and a more reasonable metric, as well as additionally introducing a looser metric of BLEU-4 score that matches exact matches. As a result, CodeT5-base establishes new SoTA performance on this task, improving 49.70 of T5-large to 53.57 in EM and 76.98 to 78.85 in BLEU-4. Furthermore, the RAP-Gen model is evaluated using both datasets of TFix and Code Refinement. It is observed that RAP-Gen using vocabulary and meaning-based retrievers significantly improves performance. Specifically, RAP-Gen-base using "Hybrid" improves exact matches beyond the best-performing baseline in TFix (49.70→54.15), and RAP-Gen-base using "Hybrid" improves exact matches in the small set of the Code Refinement benchmark (24.04→24.80) and also improves exact matches in the medium set (14.18→15.84). All these results verify that search-augmented patch generation (RAP-Gen) is an effective approach for APR.
[0073] This experiment shows that search expansion patch generation using CodeT5 is an effective approach for program repair. First, CodeT5 is compared with traditional APR techniques on the TFix benchmark and improved using a deduplicated version of the data and a more appropriate evaluation metric. Next, the RAP-Gen framework integrated with two sizes of CodeT5 is evaluated on the TFix and Code Refinement benchmarks. Additionally, this experiment shows that the patch retriever finds relevant patches with respect to lexical and semantic similarity. In addition, a case study is provided to show how the retrieved bug fix patterns are useful for program repair. Moreover, as shown by the experiment, the RAP-Gen framework provides improved performance for various error types and correction patterns. A breakdown of the detailed performance for 52 error types is listed and error types that do not benefit from search expansion in RAP-Gen are examined. Furthermore, how the model functions using a trivial but dominant error line removal correction pattern that simply removes the error line from the buggy code is studied.
[0074] This experiment shows that search expansion patch generation using CodeT5 is an effective approach for program repair. First, the TFix evaluation improves. The original TFix benchmark uses the direct average of exact match (EM) accuracy across 52 error types as the main evaluation metric. However, as shown in Table 7 of Figure 14, these error types have a rather unbalanced distribution. For example, there are 16,217 instances of the major error type "no-invalid-this", while there are only 10 instances of the minor error type "no-new-symbol". Therefore, in some embodiments, a weighted average is adopted to take into account the error type distribution. Further, upon examining the released code on how TFix calculates exact matches, there is another limitation that a predicted fix is considered an incorrect exact match if it contains one or more whitespace such as a space or a new line more than the ground truth fix. However, in the JavaScript language, extra whitespace does not affect the correctness of the program. Therefore, a better metric of the weighted average of EM w / o spaces is proposed, where whitespace is normalized before calculating EM to eliminate the effect of mismatches in multiple whitespace. Since the TFix dataset has a duplication problem, the results regarding its deduplicated version are also included. Apart from the exact match accuracy, a looser metric of the BLEU-4 score is used to measure the partial subsequence overlap between the predicted fix and its ground truth fix. Note that the BLEU-4 score is also calculated after whitespace normalization.
[0075] As shown in Table 2 of Figure 9, the CodeT5 model is compared with other learning-based baselines by TFix. One main observation is that in the case of the original average EM w / spaces metric, CodeT5-base (50.88) yields better accuracy than T5-large (49.33), even though the model size of T5-large is much larger (about 3.5 times that of CodeT5-base). Furthermore, looking at the reasonable direct average EM w / o spaces, CodeT5-base improves the absolute accuracy by about 5 compared to T5-large (49.35 → 54.30), significantly enhancing the performance. Based on the weighted average EM w / o spaces, both CodeT5-small (50.31) and CodeT5-base (53.57) outperform all baselines including T5-large (49.70). This indicates that the CodeT5 model with code recognition pre-training on a large-scale code corpus understands the program better. For the TFix evaluation, EM is used to denote the weighted average EM w / o spaces unless otherwise specified. For the BLEU-4 metric, the consistency with the exact match metric is good, and CodeT5-base also exhibits state-of-the-art (SoTA) performance of 78.85 against the original TFix.
[0076] Next, the ablation study observations are explained. In the deduplicated TFix dataset, the performance across various metrics consistently decreases slightly. This is an expected phenomenon because the overlap (34 instances) between the training and test splits in the original data leads to a data leakage problem, which inappropriately increases the performance. When the error information including error types and error messages is removed, a consistent performance degradation is observed in both the CodeT5-small model and the CodeT5-base model, revealing that it is useful to notify what types of errors need to be corrected for the program repair model.
[0077] Referring to Table 3 in Figure 10, the RAP-Gen evaluation for TFix will be described. Table 3 shows the results of the RAP-Gen framework for the deduplicated version of the TFix benchmark. First, a Random baseline is established by randomly searching for bug-fix pairs from the codebase. The performance degradation of both RAP-Gen-small and RAP-Gen-base by random search means that the randomly searched correction patterns cannot provide useful guiding signals for program repair. Next, different retrievers integrated with RAP-Gen are compared, including the vocabulary-based retriever BM25, the semantic-based retriever DPR based on dense vector matching, and two ensemble methods that combine them. As a result, all search expansion approaches significantly improved the performance for both exact match and BLEU-4 for both the small model and the base model. This indicates that search expansion generation is a feasible and effective approach for APR, and that both semantic information and vocabulary information are important for searching relevant correction patterns. In the case of the ensemble method, RAP-Gen-base using "Hybrid" brings the best improvement beyond T5-large (49.58→54.15 EM). This verifies that the ensemble approach considering both vocabulary information and semantic information can combine the best of both worlds. Another observation is that the performance gain using search expansion is larger for RAP-Gen-small than for RAP-Gen-base, which means that the improvement tends to reach a saturation point with the increase in model size. Both RAP-Gen-small and RAP-Gen-base use the RAP-Gen framework with different patch generator backbones, specifically CodeT5-base and CodeT5-small with different model sizes respectively.
[0078] In some embodiments, there may be multiple ways to fix bugs. Therefore, an exact match with one ground truth patch is a metric that is too strict to account for correct fixes in other forms. To address this, a looser evaluation using an error removal metric according to TFix is used. Under this metric, a fixed patch is considered correct as long as it resolves the errors in the source bug-containing patch and does not introduce new errors (detected by the static analyzer ESLint). Attempting to reproduce this metric for 10,465 test instances presents two difficulties: (1) Applying ESLint requires the full file context for each code patch, but it was found that 95 code files could not be retrieved. (2) For some data samples, applying ESLint with the released configuration (https: / / github.com / eth-sri / TFix) results in parser errors. As a result, by excluding those unavailable code files and samples with parser errors, a curated subset of 6,793 instances is obtained, where it was also found that the fixes generated from TFix tend to have more parser errors. Referring to Table 4 in Figure 11, an error removal comparison is shown. It is observed that the RAP-Gen-small model is significantly more performant than the T5-large model in error removal, which means that the RAP-Gen model is more capable of synthesizing different forms of good fixes. Furthermore, in RAP-Gen-small, there is an observed misalignment between the error removal metric and the exact match metric, where the exact match is low but the error removal accuracy is high. Such a misalignment is also observed in TFix.
[0079] Referring to Table 5 in Figure 12, the Code Refinement results and the comparison with the previously described method are shown. All baseline results (including the CodeT5 model) are directly obtained from their original papers. In "Naive Copy", the BLEU-4 score is quite high, but the exact match (EM) is zero, which is observed to indicate that the overlap between the buggy code and its fix is large and that the exact match should be adopted as the primary evaluation metric. Among the baselines, NSEdit is a very competitive one that gives the best result (24.04 EM) in the small subset, and CodeT5-base with multi-task training gives the best result (14.18 EM) in the medium set.
[0080] From the RAP-Gen model comparison, it is observed that RAP-Gen using various retrievers consistently improves performance over their CodeT5 counterparts. The best model establishes new SoTA results in both subsets (24.80 EM in the small case and 15.84 EM in the medium case), and in particular, in the more difficult medium set, it exceeds NSEdit by about 2 absolute points. This reconfirms that the retrieved repair patterns provide useful signals for guiding program repair. Among the various retrievers, DPR gives better results than BM25 for both RAP-Gen-small and RAP-Gen-base, revealing that semantic information can play a more important role than the semantic information for this benchmark. Furthermore, "Hybrid" outperforms BM25 and DPR, meaning that the hybrid ensemble method is a more robust retriever for balancing both semantic and semantic information for this benchmark.
[0081] In summary, comprehensive experiments are conducted to compare RAP-Gen with traditional learning-based methods on two benchmarks. First, the CodeT5 model is evaluated against TFix, providing a deduplicated version of the dataset and a more reasonable metric, and further improving the evaluation by introducing a looser metric of BLEU-4 score that matches exact matches. As a result, CodeT5-base established new SoTA performance on this task, improving from 49.70 of T5-large to 53.57 in EM and from 76.98 to 78.85 in BLEU-4. Next, the RAP-Gen model is evaluated on both the TFix dataset and the code refinement dataset, and it is observed that RAP-Gen using vocabulary and meaning-based retrievers significantly improves performance. Specifically, RAP-Gen-base using "Hybrid" improves exact matches (49.70 → 54.15) beyond the best-performing baseline in TFix, and RAP-Gen-base using "Hybrid" improves exact matches in the small set of the Code Refinement benchmark (24.04 → 24.80) and also improves exact matches in the medium set (14.18 → 15.84). All these results verify that retrieval-augmented patch generation (RAP-Gen) using CodeT5 is an effective approach for APR.
[0082] Next, experiments are conducted to evaluate whether the patch retriever can find relevant correction patterns useful for program repair. First, an automatic evaluation is provided to measure the relevance regarding the lexical and semantic similarity between the query and the retrieved patch. Further, specific cases are provided to understand how the retrieved correction patterns contribute to better APR.
[0083] Referring to Table 6 in Figure 13, the evaluation of the retriever is shown. The retriever is analyzed regarding the lexical and semantic matching between the query and the top retrieved patches. In the case of lexical matching, the BLEU-4 score is used to measure their sub-token overlap, and in the case of semantic matching, the cosine similarity (CosSim) between their dense vectors encoded by the fine-tuned DPR retriever is used. Table 6 in Figure 13 shows the performance of the patch retriever in both the TFix and Code Refinement benchmarks. The first row shows the lower bound performance by randomly retrieving bug-fix pairs from the codebase, and it is observed that this Random baseline achieves much lower scores in both lexical and semantic matching. In the case of lexical matching, BM25 performs better than DPR in TFix but worse in the two Code Refinement subsets, which may be due to the data difference between TFix and Code Refinement, and the latter adopts obfuscated identifiers (e.g., VAR1, VAR2, …) that hinder the performance of the vocabulary-based BM25 retriever. The hybrid retriever achieves the best lexical matching in all datasets, revealing that semantic information can complement lexical matching.
[0084] In the case of semantic matching, DPR achieves the best results for all datasets, which is not surprising since it is optimized towards the same goal. In particular, the hybrid retriever achieves slightly lower results than DPR but much better results than BM25, which means that the hybrid retriever can balance both lexical and semantic information and may be more robust than the vocabulary-based retriever that is sensitive to the choice of identifier naming.
[0085] Referring back to FIGS. 6 and 7, case studies such as those regarding TFix (FIG. 6) and Code Refinement (FIG. 7) are used to show how the correction patterns retrieved in program repair are useful, where the RAP-Gen model with search expansion predicts the correct correction, while CodeT5 without search expansion cannot do so. As shown in FIG. 6, the retrieved bug correction pattern is exactly what is needed to repair the source bug-containing code. Without search expansion, CodeT5 probably wrongly removes ".classify()" from the buggy line through learning from the previous adjacent lines. In the case of Code Refinement in FIG. 7, the retrieved bug correction pair provides sufficient information to guide the RAP-Gen model to correct the source bug-containing code. Without search expansion, CodeT5 performs a wrong repair by simply removing the last line of the code.
[0086] Therefore, to evaluate the performance of the patch retriever and the corresponding automatic program repair system, both quantitative (Table 6 in FIG. 13) and qualitative (FIGS. 6 and 7) results are obtained. As a result, it is shown that the hybrid patch retriever is more robust and can find vocabulary and semantically related patches to assist the program repair system.
[0087] Referring to Table 7 in Figures 8 and 15, the performance of RAP-Gen for various error types and repair patterns will be described. First, regarding the breakdown of its performance for different error types, the detailed program repair performance breakdown for the deduplicated TFix dataset is listed in Table 7 of Figure 15. CodeT5-base outperforms the previous SoTA T5-large in 44 / 52 error types. In particular, for the major error type "no-invalid-this", CodeT5-base improves its exact match from 37.48 of T5-large to 43.57, which corresponds to fixing 98 more instances. T5-large can fix at least 50% of the bugs for 44% of the 52 error types, but CodeT5-base significantly increases this percentage to 60%, and RAP-Gen-small further improves it to 63%. Overall, RAP-Gen-base correctly repairs 478 more buggy programs than T5-large with a much smaller model size.
[0088] Furthermore, analyze the effect of search expansion in RAP-Gen compared to the CodeT5 model for various error types. As shown in Table 7 of Figure 14, the search expansion technique has different effects on different error types, and it is observed that it can even degrade the program repair performance for a specific subset of error types. Specifically, in the experiment, it degrades the APR performance for 10 error types of CodeT5-small and 18 error types of CodeT5-base. Based on the number of exact match corrections, the largest performance degradation of RAP-Gen-small is for the error type "no-extra-semi" (497→490), and the largest performance degradation of RAP-Gen-base is for the error type "no-console" (228→220) is observed.
[0089] To investigate the reasons why search expansion may hinder the exact match performance in the RAP-Gen model, a case study of the "no-console" error type is provided in Figure 8. In this case, the ground truth fix is to directly remove the error line in the buggy patch, and the RAP-Gen model repairs it in a different form based on the retrieved fix patterns that are counted as incorrect predictions regarding the exact match. This reaffirms the limitations of the exact match metric for evaluating program repair systems.
[0090] Next, the TFix benchmark is used to analyze which fix patterns are executed by the model. After manually inspecting the bug fix pairs, it is observed that most of the fixes are composed of deletion operations compared to code insertion and replacement operations. The bug fix actions are composed of code insertion (12.5%), replacement (8.1%), deletion (47.9%), insertion and replacement (6.9%), insertion and deletion (8.2%), replacement and deletion (7.2%), and all three methods (9.2%). Previous studies also reflect that the deletion operation is one of the most common fix patterns. Among the deletion actions, one dominant bug fix pattern is error line removal, which is simply to remove the error line from the buggy code (as in the example shown in Figure 8). This obvious fix pattern accounts for approximately 23% in the deduplicated TFix test set. To further analyze this pattern, different models are compared using the error line removal pattern to show how they function, and the results are shown in Table 8. It is observed that using search expansion, RAP-Gen-base achieves the lowest number of false detections, 56 (equivalent to the highest accuracy of 97.09), compared to 67 for CodeT5-base and 71 for T5-large. This indicates that RAP-Gen can learn more diverse bug fix patterns rather than relying overly on the obvious error line removal pattern. Furthermore, RAP-Gen-small achieves the best recall rate and F1 score, but at the cost of generating more false detection predictions.
[0091] In summary, the difficulty of program repair varies for each error type. The best RAP-Gen-base in the experiment can repair 456 more buggy programs than the baseline T5-large that exhibits the best performance. Error analysis is performed to analyze the reason why search expansion may downgrade performance, and a case study is provided to show that it may be due to the limitation of the exact match metric. Furthermore, one high-frequency correction pattern of error line removal is investigated, showing that when dealing with this pattern, RAP-Gen-base gives the best accuracy score, and RAP-Gen-small achieves the best recall and F1 scores.
[0092] This description and the accompanying drawings showing aspects, embodiments, implementations, or applications of the present invention should not be construed as limiting. Various mechanical, compositional, structural, electrical, and operational changes can be made without departing from the spirit and scope of this specification and the claims. In some instances, well-known circuits, structures, or techniques are not shown or described in detail so as not to obscure the embodiments of the present disclosure. Like numbers in two or more figures represent the same or similar elements.
[0093] In this description, specific details are set forth that describe several embodiments consistent with the present disclosure. Numerous specific details are set forth to provide a complete understanding of the embodiments. However, it will be apparent to those skilled in the art that some embodiments may be practiced without some or all of these specific details. The specific embodiments disclosed herein are illustrative and not limiting. Those skilled in the art can implement other elements not specifically described herein but within the scope and spirit of the present disclosure. Additionally, to avoid unnecessary repetition, one or more features illustrated and described in connection with one embodiment may be incorporated into other embodiments if not specifically described otherwise or if one or more features would render the embodiment non-functional.
[0094] Exemplary embodiments have been shown and described, but a wide range of modifications, changes, and substitutions are contemplated in the foregoing disclosure, and in some instances, some features of the embodiments may be employed without the corresponding use of other features. Those skilled in the art will recognize many variations, alternatives, and modifications. Accordingly, the scope of the present invention should be limited only by the following claims, and the claims should be broadly construed so as to be consistent with the scope of the embodiments disclosed herein.
Claims
1. A method for automatic program repair, comprising: Receiving a first bug-containing code patch; Generating a first representation of the first bug-containing code patch using a retriever encoder of a patch retriever; Searching, using the patch retriever, for a first bug-fixing code pair from a first plurality of bug-fixing code pairs based on the first representation; Generating a first extended bug-containing code patch based on the first bug-containing code patch and the first bug-fixing code pair; Generating a modified code patch based on the first extended bug-containing code patch via a patch generator A method comprising.
2. The method according to claim 1, wherein the patch retriever is configured to perform the search based on at least one of a lexical similarity and a semantic similarity with the first bug-containing code patch.
3. The method according to claim 2, wherein the patch retriever is configured to perform the search based on a combination of a lexical similarity and a semantic similarity with the first bug-containing code patch.
4. The method according to claim 1, wherein the patch generator includes a Transformer-based neural network model for sequence generation.
5. The method according to claim 1, wherein the first bug-fixing pair of the first extended bug-containing is used as a guide correction pattern for the patch generator.
6. Before receiving the first bug-containing code patch, performing a two-stage training process, comprising: Training the patch retriever in a first stage using a first training set; and Training the patch generator in a second stage using the trained patch retriever and a second training set A step comprising The method according to claim 1, further comprising.
7. The method according to claim 6, wherein the first training set includes bug-containing patches and corresponding modified patches.
8. A non-transitory machine-readable medium including a plurality of machine-readable instructions that, when executed by one or more processors, cause the one or more processors to: Receive a first bug-containing code patch; Generating a first representation of the first bug-containing code patch using a retriever encoder of a patch retriever; Searching for a first bug-fixing code pair from a first plurality of bug-fixing code pairs based on the first representation using the patch retriever; Generating a first extended bug-containing code patch based on the first bug-containing code patch and the first bug-fixing code pair; Generating a modified code patch based on the first extended bug-containing code patch via a patch generator A non-transitory machine-readable medium adapted to cause a method including the above steps to be executed.
9. The non-transitory machine-readable medium according to claim 8, wherein the patch retriever is configured to perform a search based on one or more of a lexical similarity and a semantic similarity with the first bug-containing code patch.
10. A system, comprising: A non-transitory memory; One or more hardware processors coupled to the non-transitory memory and configured to read instructions from the non-transitory memory to cause the system to execute a method The method includes: Receiving a first bug-containing code patch; Generating a first representation of the first bug-containing code patch using a retriever encoder of a patch retriever; Searching for a first bug-fixing code pair from a first plurality of bug-fixing code pairs based on the first representation using the patch retriever; Generating a first extended bug-containing code patch based on the first bug-containing code patch and the first bug-fixing code pair; Generating a modified code patch based on the first extended bug-containing code patch via a patch generator A system including the above steps.
11. The system according to claim 10, wherein the patch retriever is configured to perform a search based on at least one of a lexical similarity and a semantic similarity with the first bug-containing code patch.
12. The system according to claim 11, wherein the patch retriever is configured to perform a search based on a combination of the lexical similarity and the semantic similarity with the first bug-containing code patch.
13. The system according to claim 10, wherein the patch generator includes a Transformer-based neural network model for sequence generation.
14. The system according to claim 10, wherein the first bug-fixed pair containing the first extended bug is used as a guide correction pattern for the patch generator.
15. The method further includes the step of performing a two-stage training process before receiving the first bug-containing code patch, training the patch retriever in a first stage using a first training set, and training the patch generator in a second stage using the trained patch retriever and a second training set as steps further included in the system according to claim 10.
Citation Information
Patent Citations
Enhancing software development using bug data
US20180276103A1
Methods and apparatus for automatic detection of software bugs
US20210182031A1
Automated program repair tool
US20210357307A1