Secure ai-assisted code generation system with iterative patching

CN122804215APending Publication Date: 2026-09-22GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202580014233.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-11
Filing Date
2025-12-10
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

本文描述的一些实现方式解决了尤其是在大型且复杂的代码库中高效生成并测试代码变更的问题,以修复漏洞或实现新特征

Benefits of technology

[0003]这些智能体中的至少一些智能体在安全的沙盒环境内操作,该沙盒环境可以采取各种形式,诸如虚拟机(VM)或者可以或可以不在单独VM上运行的微虚拟机(microVM)。这些沙盒环境提供了强大的隔离性,从而防止不受信任的代码访问敏感资源。在各种实现方式中,采用本公开的选定方面配置的沙盒环境可以允许对虚拟机状态进行快照和恢复,从而能够通过智能体轨迹的分叉和合并来有效地探索多种代码变更策略。这有助于迭代开发和测试,其中后续代码变更是基于先前迭代的结果而构建的,所有这些都在安全且可恢复的沙盒环境内进行。在一些实现方式中,从漏洞检测到代码生成和测试的整个过程都可以由编排智能体进行管理,从而允许灵活地组合程序化智能体和动态智能体以在开发过程中执行不同任务。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122804215A_ABST
    Figure CN122804215A_ABST
Patent Text Reader

Abstract

Implementations for generating and testing code changes in a codebase are described herein. In some implementations, one or more agents can identify one or more constituent source files in the codebase that are amenable to patch trigger data. One or more of the agents can generate a plurality of candidate source code changes using one or more generative models. The candidate source code changes can be based on the one or more constituent source code files. One or more of the agents can generate a plurality of unit tests using one or more of the generative models. One or more of the agents can run different permutations of the unit tests and the candidate source code changes. Based on the results of the different permutations of the unit tests and the candidate source code changes, one of the candidate source code changes can be selected and incorporated into a multi-file patch.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Using machine learning for automated code generation promises to alleviate the burden on programmers when writing source code, freeing them from tedious code processing tasks and allowing them to focus on more creative aspects of software engineering design. The powerful output of automated code generation can be used for a variety of purposes, such as code change analysis, automated testing, and integrating new functionality into existing code. Code generation tools can utilize machine learning, including generative models such as Large Language Models (LLMs), to generate useful code from large codebases. Current generative model code generation solutions are typically used to generate synthetic source code in small fragments. For example One snippet at a time is generated. While this saves at least some manpower compared to writing code entirely by hand, creating code snippets one at a time can still be a tedious process. Furthermore, there is no guarantee that any synthesized source code generated by the generative model will function correctly. Phantom, defective code, and synthesized source code that compiles but does not perform its intended function can still occur. This leaves the tasks of testing (including unit testing) and debugging to the programmer. Summary of the Invention

[0002] This paper provides implementations for using artificial intelligence (AI) agents to assist software development. Some of the implementations described address the problem of efficiently generating and testing code changes, particularly in large and complex codebases, to fix bugs or implement new features. The various implementations described provide a framework that orchestrates multiple AI agents, each potentially employing a programmatic, dynamic, or hybrid approach, or a combination thereof, to collaboratively solve code modification tasks. One agent, referred to in this paper as the "editor" agent, can be configured to generate multiple candidate code changes and corresponding unit tests, allowing the selection of the best option based on test results. Other agents can be responsible for identifying relevant files in the codebase and planning the necessary code changes.

[0003] At least some of these agents operate within a secure sandbox environment, which can take various forms, such as a virtual machine (VM) or a microVM that may or may not run on a separate VM. These sandbox environments provide strong isolation, preventing untrusted code from accessing sensitive resources. In various implementations, sandbox environments configured with selected aspects of this disclosure allow for snapshotting and restoring the VM state, enabling efficient exploration of multiple code change strategies through forking and merging of agent trajectories. This facilitates iterative development and testing, where subsequent code changes are built upon the results of previous iterations, all within a secure and recoverable sandbox environment. In some implementations, the entire process from vulnerability detection to code generation and testing can be managed by orchestration agents, allowing for the flexible combination of procedural and dynamic agents to perform different tasks during development.

[0004] The implementation described in this paper allows for iterative improvements to code changes, where the application's state is captured after each successful iteration, providing a starting point for the next iteration. Systems configured with selected aspects of this disclosure can run in a cloud environment, remotely processing code changes and returning them for merging into the main repository. Users can interact with the system at a high level, specifying tasks and receiving improved code changes without needing to understand the underlying agent interactions or sandbox execution. The system can employ multiple agents, each with a specific role (…). For example The process involves identifying relevant files, generating a natural language description of the required changes, generating candidate code changes, and testing. The output of one agent can be used as input to another, creating an action pipeline that iteratively performs multiple rounds of code generation and testing, improving the code changes in each iteration. In some implementations, the workflow begins by processing "patch trigger data" (such as error messages), identifying relevant files, generating a natural language description of the required changes, generating candidate code changes and unit tests, and evaluating these changes to identify the best candidate code changes to include in a multi-file patch. The resulting multi-file patch can then be applied (or "executed") to the codebase. Attached Figure Description

[0005] Figure 1 This diagram illustrates a code knowledge system that interacts with multiple clients and their codebases using machine learning models and programming language corpora.

[0006] Figure 2 The diagram illustrates a patch generation agent that coordinates multiple agents to generate multi-file patches, including file identification, code generation, unit test generation, and iterative testing and improvement.

[0007] Figure 3 A matrix is ​​presented, which shows how multiple candidate source code edits are cross-evaluated for multiple unit tests to select the best candidate.

[0008] Figure 4 The paper describes the use of multiple microvirtual machines to iteratively generate and test code changes, where each microvirtual machine captures and restores the application state to enable efficient exploration of multiple code change strategies.

[0009] Figure 5 A flowchart illustrating selected aspects of this disclosure is shown.

[0010] Figure 6 A flowchart illustrating selected aspects of this disclosure is shown.

[0011] Figure 7 A block diagram of an example computing device capable of performing the system and method is shown.

[0012] Figure 8 An example of how an agent can be implemented according to various implementation methods is shown.

[0013] Figure 9 An example of how an agent can be implemented according to various implementation methods is shown.

[0014] Figure 10 An example of how an agent can be implemented according to various implementation methods is shown.

[0015] Figure 11 An example of how an agent can be implemented according to various implementation methods is shown. Detailed Implementation

[0016] The implementation disclosed in this paper involves generating multi-file patches for a codebase used in an application. One or more agents identify one or more component source files in the codebase suitable for patch triggering data. Patch triggering data may include, for example, error message data generated via application execution, problem statements issued by the agents (…). For example "The application keeps crashing while processing <specific data>" wait One or more agents use one or more generative models to generate multiple candidate source code changes and multiple unit tests. The agents then run different permutations of the unit tests and candidate source code changes. Based on the results, a candidate source code change is selected and incorporated into a multi-file patch.

[0017] The implementation method disclosed in this article can alleviate ( For exampleThis addresses various shortcomings of existing technologies. For example, as mentioned above, generating multiple candidate code changes and their corresponding unit tests solves the problem of generating single code snippets that may contain vulnerabilities or fail to run. As another example, the iterative process of generating and testing code changes, starting from the application state after each successful iteration, overcomes the limitations of single-step code generation and testing. As yet another example, using a secure sandbox environment to execute code changes and unit tests prevents the main codebase from being compromised and allows for the safe exploration of multiple code change strategies.

[0018] As a non-limiting example of some implementations disclosed in this paper, consider a scenario where a user encounters an error message indicating that the web application crashed while processing a specific type of data. This error message serves as patch trigger data. An agent (possibly a dynamic agent utilizing an LLM) analyzes the error message and identifies relevant files within the application's codebase. For example (Files related to data processing). Then, another agent (possibly a hybrid agent combining procedural and dynamic methods) generates multiple candidate code changes for these files, each designed to resolve the crash. A third agent generates corresponding unit tests for each candidate code change. These candidates and tests are then executed in separate snapshottable sandbox environments. Each sandbox environment starts with a snapshot of the application state before the code changes are applied. The results of the unit tests in each sandbox environment are compared, and the candidate code change that passes the most tests and / or successfully resolves the problem data type is selected as the best patch. This patch is then applied to the codebase.

[0019] Figure 1 Example environments in which selected aspects of this disclosure can be implemented according to various implementation methods are illustrated. Figure 1 Any computing device depicted elsewhere in the accompanying drawings may include one or more microprocessors, such as those that execute computer-readable instructions stored in memory. For example The logic of a central processing unit or “CPU”, a graphics processing unit or “GPU”, a tensor processing unit or “TPU”, a neural processing unit or “NPU”, or other types of logic such as an application-specific integrated circuit (“ASIC”), a field-programmable gate array (“FPGA”), etc. Figure 1 Some of the systems depicted, such as the code knowledge system 102, can be implemented using one or more server computing devices, which are sometimes referred to as “cloud infrastructure,” although this is not required.

[0020] A code knowledge system 102 may be provided to assist clients 110-1 to 110-P in managing their respective code repositories 112-1 to 112-P. The code knowledge system 102 and clients 110-1 to 110-P may be communicatively coupled via one or more computer networks generally designated 199. The code knowledge system 102 may include multiple agents 104, etc., configured to perform selected aspects of this disclosure to assist one or more clients 110-1 to 110-P (e.g., via multi-file patching) in managing and / or modifying one or more corresponding code repositories 112-1 to 112-P. Each client 110 may be, for example, an entity or organization such as an enterprise (…). For example Financial institutions and banks wait This includes non-profit organizations, clubs, universities, government agencies, or any other organization that operates one or more software systems. For example, a bank may operate one or more software systems to manage funds under its control, including tracking deposits and withdrawals, tracking loans, tracking investments, and so on. An airline may operate one or more software systems for booking / cancelling / rebooking flight orders, managing flight delays or cancellations, managing flight-related personnel such as passengers, crew, and ground staff, and managing airport gates, etc.

[0021] Multiple agents 104 may each be configured to execute one or more selected aspects of this disclosure to assist clients 110-1 to 110-P in editing, updating, replatformizing, migrating, or otherwise acting on their codebases 112-1 to 112-P. For example (Applying patches). Agents 104 can be programmatic, dynamic, or a combination thereof, and do not necessarily have to be of the same type. In some implementations, one or more agents 104 can be configured to perform some or all of the tasks described below with reference to a single code change candidate, and other agents 104 can be configured to perform similar tasks on other code change candidates. In other implementations, some agents can perform all or part of the tasks described with reference to a single code change candidate, some agents can perform some tasks, and some agents can perform all tasks.

[0022] Among the various implementations, the code knowledge system 102 may include machine learning ( Figure 1The "ML" database 105 in the document includes data indicating one or more trained machine learning models 106-1 to 106-N. These trained machine learning models 106-1 to 106-N can take various forms, including generative models. Generative models themselves can take various forms, such as large language models (LLMs), and / or can take other forms, such as neural networks and / or other models trained to perform certain tasks, such as generating candidate code changes, testing candidate code changes, and / or analyzing source files of code changes to identify candidate code changes. Generative models can be encoder-decoder, encoder-decoder-encoder, autoencoder, and / or other forms. Some generative models can take the form of a base model. The base model can act as an encoder or decoder of various forms. In some implementations, generative models may contain combinations of other models.

[0023] Generative models can have varying numbers of parameters. For example, in some implementations, a generative model may include tens of thousands to hundreds of thousands of parameters and can be used on resource-constrained devices, such as client devices operated by client 110. Other implementations may have other numbers of parameters, such as hundreds of thousands to millions of parameters, while others may have billions, tens of billions, hundreds of billions, or even more parameters.

[0024] In some implementations, the code knowledge system 102 may also access one or more programming language-specific corpora 108-1 to 108-M. In some implementations, these programming language-specific corpora 108-1 to 108-M may be used, for example, to train, fine-tune, or perform in-situ learning using one or more of machine learning models 106-1 to 106-N. In some implementations, the programming language-specific corpora 108-1 to 108-M may include examples of source code ( For example The entire codebase, the library. wait ), inline comments, and text metadata associated with the source code ( For exampleSubmissions include documents such as textbooks and programming manuals, language-specific discussion threads, presentations, academic papers, etc. In some implementations, at least some of the machine learning models 106-1 to 106-N can be trained using at least some of the language-specific corpora 108-1 to 108-M. In some implementations, some of the language-specific corpora 108-1 to 108-M can be used as training data for generating models 106-1 to 106-N, while other language-specific corpora 108-1 to 108-M can be used as test data for generating models 106-1 to 106-N. In some implementations, some of the programming language-specific corpora 108-1 to 108-M can be used exclusively for training, while other programming language-specific corpora 108-1 to 108-M can be used exclusively for testing.

[0025] Figure 2 This illustration depicts an example of how various agents implemented by the code knowledge system 102 can collaborate within the patch generation agent 200 to generate a multi-file patch 230 for an application. In this example, the patch generation agent 200 begins with patch trigger data 220. Patch trigger data 220 can be issued by the user ( For example (as a natural language problem statement), or it can be generated through the execution of the application. For example This can be either application-generated output or an error message. Patch trigger data 220 can be processed by both a procedural agent 204A and a dynamic agent 204B. The procedural agent 204A can be configured to programmatically (…). For example Using pre-existing instructions and logic wait The system identifies one or more component source code files 222A in the codebase 212 of the application being patched, which can be used to patch the trigger data 220. The dynamic agent 204B can also be configured to identify one or more component source code files 222B, but unlike the application, the dynamic agent 204B does not identify them programmatically; instead, it uses one or more machine learning models. For example , Figure 1 The generative models 106-1 to 106-N in the model are used to identify the constituent source code file 222B. In many cases, the constituent source code file 222A and the constituent source code file 222B may overlap at least partially. In some such implementations, the overlapping constituent source code files may be generated by... Figure 2 The other components described in the source code file are selected for downstream processing.

[0026] Each of these component source code files 222A and 222B can be passed to the editing agent 204C. The editing agent 204C can be configured to generate multiple candidate source code edits 222-1-N and multiple unit tests based on the component source code files 222A and 222B. Figure 2 (UT in the text) 223-1-M. Multiple candidate source code edits 222-1-N and multiple unit tests ( Figure 2 Unit test 223-1-M can be provided to test agent 204D. Test agent 204D can be configured to run different arrangements of unit test 223-1-M and candidate source code editors 222-1-N. Based on the results of these different arrangements of unit test 223-1-M and candidate source code editors 222-1-N, a specific source code editor 222E can be selected to be included in the multi-file patch 230. As indicated by the dashed lines, in some implementations, there may be multiple test agents 204D, 204D', each configured to run different arrangements of candidate source code editors and unit tests for different source code files. Figure 2 For example, the additional test agent 204D generates additional candidate source code edits 222F to be included in the multi-file patch 230.

[0027] exist Figure 2 In the example, test agent 204D operates multiple sub-agents to perform iterative improvements on source code changes. Tester agent 224 can be configured to run unit tests 223-1-M, for example, in a sandbox environment configured with selected aspects of this disclosure, using candidate source code edits 222-1-N. Based on the results of such testing, repair agent 204E (which in some implementations may be the same as dynamic agent 204B) can be configured to generate improved candidate source code edits 222-1-X, which can be provided back to tester agent 224 for additional testing. This loop can be repeated until one or more conditions are met ( For example The predetermined number of loops, the confidence metric for candidate source code editing meeting the threshold, and the absence of compiler errors or crashes are all confirmed. wait At this point, the test agent can select a candidate source code change 222E that meets another criterion, such as the candidate source code change 222E that can pass the maximum number of unit tests 223. As described above, this selected candidate source code change 222E can... For example Together with one or more additional candidate source code changes 222F generated by one or more additional test agents, they are incorporated into multi-file patch 230.

[0028] In some implementations, the additional planner agent 204F may be configured to process the patch trigger data 220 and / or the constituent source code files 222A-B using one or more generative models to generate one or more natural language statements 226. The natural language statements 226 may include descriptions of functional changes made to the codebase 212 and / or the constituent source code files 222A-B based on the patch trigger data 220. In some implementations, candidate source code changes 222-1-N may be generated based on one or more of the natural language descriptions 226. Additionally or alternatively, in some implementations, multiple unit tests 223-1-M may be generated based on one or more of the natural language descriptions 226.

[0029] As an example, patch trigger data 220 may include, for instance, an error code generated as a byproduct of an unexpected application termination. This error code can be processed by a procedural agent 204A and a dynamic agent 204B to identify component source code files 222A and 222B that may be applicable to the error code. The component source code files 222A and 222B can then be provided to an editing agent 204C and a planner agent 204F. The planner agent 204F can use one or more generative models to process the error code to generate one or more natural language statements 226 that describe one or more causes of the error. For example (vulnerability) and / or make one or more functional changes to the constituent source code files 222A and / or 222B to fix the root cause of the error. Then, the natural language statement 226 can be processed by the editing agent 204C to... For example For each of the constituent source code files 222A-B, multiple candidate source code editors 222-1-N and multiple unit tests 223-1-M are generated. The candidate source code editors 222-1-N and unit tests 223-1-M can then be tested by the test agent 204D as described above. The resulting candidate source code editors 222E that meet one or more criteria can... For example This, along with any additional edits 222F generated by one or more additional test agents 204', is incorporated into the multi-file patch 230. The multi-file patch 230 can then be applied to the codebase 212, and this application can be executed again. The same or different errors can trigger [further action / action]. Figure 2 Another iteration of the presentation process.

[0030] Figure 3Examples are depicted of how different permutations of unit tests 223-1-M and candidate source code edits 223-1-N can be evaluated to select one of candidate source code edits 222E for inclusion in the multi-file patch 230. In various implementations, the test agent 204D can be configured to cross-evaluate these different permutations of unit tests 223-1-M and candidate source code edits 223-1-N, for example, in a sandbox environment. As shown in the matrix, multiple candidate source code edits A through J are cross-evaluated against multiple unit tests A through J. For example, candidate source code edit A is tested using each unit test A through J, candidate source code edit B is tested using each unit test A through J, and so on. As illustrated, if a candidate source code edit passes a specific unit test, the corresponding cell in the matrix is ​​shaded. If a candidate source code edit fails a specific unit test, the corresponding cell in the matrix remains unshaded.

[0031] In this example, candidate source code editor A passed unit tests A, C, and G, but failed the others. Candidate source code editor B passed unit tests B, C, D, and G, a total of six tests, but failed the others. Candidate source code editor C passed unit tests C and I through J, but failed the others. Candidate source code editor D passed unit tests B, C, and G, but failed the others, and so on. It can be seen that among all candidate source code editors A through J, source code editor H passed the most unit tests (A through F, H), a total of seven. Therefore, in the implementation where the candidate source code editor selection criterion is based on the candidate source code editor who passed the most unit tests, source code editor H will be the selected candidate source code editor to be included in multi-file patch 230.

[0032] Figure 4 The illustration schematically depicts an example of how, according to various aspects of this disclosure, a sandboxed environment in the form of microvirtual machines 452-1 to 452-7 can be used to create, save, and / or restore snapshots of application state to help clients 110-1 to 110-P iteratively generate multi-file patches for codebase 212. Specifically, Figure 4 This demonstrates how iterative improvements to multi-file patch 230 can be performed on various component source code files in codebase 212 using microvirtual machines 452-1 to 452-7. It should be noted that, although... Figure 4 The document describes six microvirtual machines, 452-1 to 452-7, but other implementations are not limited to these. For example... Figure 4 As shown by the arrows, time flows from left to right.

[0033] Starting from the top left corner, the first microvirtual machine 452-1 executes the application until it reaches the first application state 450-1, and experiences an unexpected termination or "crash," as represented by a starburst diagram. At this point, the first microvirtual machine 452-1 takes a snapshot of the first application state 450-1 and provides this snapshot to the orchestration agent 404. The orchestration agent 404 can then initiate a second microvirtual machine 452-2 to implement an instance of the patch generation agent 200. The patch generation agent 200 can use the same techniques described above within the second microvirtual machine 452-2. For example ,refer to Figure 2 The first multi-file patch 430-1 was generated for the application and incorporated into the codebase 212.

[0034] At this point, the third microvirtual machine 452-3 can begin executing the application from the first application state 450-1, which can be resumed until the application execution reaches the second application state 450-2 and another unexpected termination occurs, as shown in the starburst diagram. As before, a snapshot of the second application state 450-2 can be created and provided to the orchestrator agent 404, and the orchestrator agent 404 can initiate a fourth microvirtual machine 452-4 to implement an instance of the patch generation agent 200. At this point, the patch generation agent 200 can use the same techniques described above within the fourth microvirtual machine 452-4. For example ,refer to Figure 2 A second multi-file patch 430-2 was generated and incorporated into codebase 212.

[0035] Similar to the previous description, the fifth microvirtual machine 452-5 can begin executing the application from the second application state 450-2, which can be resumed until the application execution reaches the third application state 450-3 and another unexpected termination occurs, as shown in the starburst diagram. As before, a snapshot of the third application state 450-3 can be created and provided to the orchestrator agent 404, which can then initiate a sixth microvirtual machine 452-6 to implement an instance of the patch generation agent 200. At this point, the patch generation agent 200 can use the same techniques described above within the sixth microvirtual machine 452-6. For example ,refer to Figure 2 A third multi-file patch 430-3 is generated and incorporated into codebase 212. At this point, the seventh microvirtual machine 452-7 can begin executing the application from the third application state 450-3, which can be restored. Eventually, expected termination may occur, at which point the multi-file patch can complete its iterative improvements.

[0036] Figure 5Example method 500 for performing selected aspects of this disclosure is depicted. For convenience, the operations of the flowchart are described with reference to a system performing the operations. This system may include various components of various computer systems, such as code knowledge system 102 and / or client system 110. Furthermore, although the operations of method 500 are shown in a specific order, this is not intended to be limiting. One or more operations may be reordered, omitted, or added.

[0037] At box 502, the system can cause one or more agents ( Figure 1 104) Identification code library ( Figure 2 212) is applicable to patch trigger data ( Figure 2 One or more component source files (220 in the middle) Figure 2 (222A, 222B). This could involve the use of programmatic intelligent agents ( Figure 2 204A in the middle) or dynamic intelligent agent ( Figure 2 204B in the document) or both, thereby potentially utilizing machine learning models ( Figure 1 106-1 to 106-N in the corpus) and / or a corpus specific to the programming language ( Figure 1 (108-1 to 108-M in the original text). Patch trigger data may include data generated by the execution of the application, such as error messages indicating unexpected termination or natural language problem statements.

[0038] At box 504, the system can cause one or more of the agents ( Figure 1 104 in the middle) uses one or more generative models ( Figure 1 Multiple candidate source code changes (from 106-1 to 106-N) were generated. Figure 2 (222-1-N in the original text). These models can be LLMs or other types of generative models. In some implementations, one or more generative models may include those for generating candidate source code changes (…). For example The system generates a first generative model (presented as "candidate patches") and a second generative model for generating unit tests. In other implementations, the same generative model can be used to generate both. The generation of candidate source code changes can be based on identified component source files, and in some implementations, can refer to data generated by other intelligent agents (…). example like , Figure 2 The planner agent 204 F generates a natural language description of the required changes.

[0039] At box 506, the system can cause one or more of the agents ( Figure 1 104) uses one or more of the generative models. Figure 1Multiple unit tests are generated from 106-1 to 106-N. Figure 2 (223-1-M in [reference]). In some implementations, each unit test can be generated based on a corresponding candidate source code change; and in some cases, at least one unit test generated based on one candidate source code change can be used to test another candidate source code change. In some implementations, the generation of unit tests can refer to the actions of other intelligent agents ([reference]). For example , Figure 2 The planner agent 204F generates a natural language description of the required changes. Figure 2 226 in the middle.

[0040] At box 508, the system can cause one or more of the agents ( Figure 1 104 in the middle) For example In one or more sandbox environments such as VMs or microvirtual machines ( Figure 4 Run unit tests in sections 452-1 to 452-7. Figure 2 223-1-M) and candidate source code changes ( Figure 2 Different permutations of 222-1-N in [the original text]. Some such VMs or microvirtual machines can be snapshottable and recoverable, such as... Figure 4 This is illustrated in the diagram. The iterative process can involve multiple rounds of code generation and testing, with each iteration starting from the application state after the previous iteration successfully completed. Figure 4 It begins with 450-1, 450-2, and 450-3. The sandbox environment allows for safe and efficient exploration of various code change strategies.

[0041] At box 510, the system can be based on a sandbox environment ( Figure 4 The results of running different permutations of (452-1 to 452-7) are used to select one of the candidate source code changes. Figure 2 222E) and incorporated it into the multi-file patch ( Figure 2 In section 230), this selection can be based on which candidate source code change satisfies the most unit tests. Figure 3 Or other standards. Selected candidate source code changes can be applied to codebase 212. In some implementations, the system may include an orchestrator ( Figure 4 404 in the middle) to manage intelligent agents ( Figure 1 (104 in the text), the agent can be programmable ( Figure 2 (204A in the figure), dynamic or hybrid. At box 512, the system can apply multi-file patches (230 in the figure) to the codebase ( Figure 2 (212 in the middle).

[0042] Figure 6An example method 600 for performing selected aspects of this disclosure is depicted. For convenience, the operations of the flowchart are described with reference to a system performing the operations. This system may include various components of various computer systems, such as code knowledge system 102 and / or client system 110. Furthermore, although the operations of method 600 are shown in a specific order, this is not intended to be limiting. One or more operations may be reordered, omitted, or added.

[0043] At box 602, the system can run on one or more virtual machines ( For example , Figure 4 The application executes in a microvirtual machine (452-1 to 452-7) until one or more predetermined events occur. These events may include error messages, unexpected termination, application functionality failure, or a specific target state of the application. wait .

[0044] At box 604, the system can determine whether one or more of the predetermined events have been detected. If the answer is no, method 600 can return to box 602. However, after the first event of the predetermined events is detected, at box 606, the system can capture the first application state corresponding to the detection of the first event. Figure 4 A snapshot (450-1) of the application. This snapshot captures the application state at the time of the first event, thus saving the state for later use as the "current application state".

[0045] At box 608, based on the first event, the system can cause a second virtual machine ( Figure 4 452-2) Managed patch generates intelligent agents ( Figure 2 The patch generates an agent from an instance of 200. For example Use unit testing ( Figure 2 (223-1-M) Iteratively generates and tests candidate source code changes ( Figure 2 222-1-N in the codebase (for the application) Figure 2 (212) Generate multi-file patch ( Figure 2 (230 in the text). In some implementations, the patch-generating agent can utilize a generative model ( Figure 1 (106-1 to 106-N) and other intelligent agents ( Figure 2 (204A, 204B, 204C, 204D, 204E), as previously mentioned.

[0046] At box 610, the system can apply a multi-file patch to the codebase to generate an updated application. This updated application incorporates the changes suggested by the agent generated from the patch. Then, method 600 can return to box 602, at which point the system can proceed from the "current" application state (which will be...) after one iteration. Figure 4Starting with 450-1 in the first virtual machine ( Figure 4 452-1 in the middle) or a third virtual machine ( Figure 4 In step 452-3), the updated application is executed until one or more of the predetermined events occur. Method 600 can be performed as described above and can be iterated until one or more stopping conditions are met, such as the application completing execution or reaching a certain desired state without errors after a certain number of iterations. wait .

[0047] Figure 7 A block diagram of an example computing device 700 is depicted. The computer system 710 typically includes a processor 714 that communicates with a plurality of peripheral devices via a bus subsystem 712. These peripheral devices may include a storage subsystem 724 (including, for example, a memory subsystem 725 and a file storage subsystem 726), a user interface output device 720, a user interface input device 722, and a network interface subsystem 716. The input and output devices allow users to interact with the computer system 710. The network interface subsystem 716 provides an interface to an external network and is coupled to corresponding interface devices in other computer systems.

[0048] User interface input device 722 may include a keyboard, pointing device (such as a mouse, trackball, touchpad, or graphics tablet), scanner, touchscreen integrated into a display, audio input device (such as a voice recognition system, microphone), and / or other types of input device. Generally, user interface input device 722 may include any means for inputting information into computer system 710.

[0049] User interface output device 720 may include a display subsystem, a printer, a fax machine, or a non-visual display (such as an audio output device). The display subsystem may include a cathode ray tube (CRT), a flat panel device (such as a liquid crystal display (LCD)), a projection device, or some other mechanism for producing visible images. The display subsystem may also provide non-visual displays, such as via an audio output device. Generally, user interface output device 720 may include any means for outputting information from computer system 710 to a user or to another machine or computer system.

[0050] Storage subsystem 724 provides some or all of the functional programming and data construction of the modules described herein. For example, storage subsystem 724 may include functions for performing... Figure 5 and Figure 6The selection of the method's logic. These software modules are typically executed by the processor 714 alone, or in combination with other processors. The processor 714 can take various forms, such as a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a neural processing unit (NPU), etc.

[0051] The memory 725 used in the storage subsystem 724 may include multiple memories, including a main random access memory (RAM) 730 for storing instructions and data during program execution and a read-only memory (ROM) 732 for storing fixed instructions therein. The file storage subsystem 726 provides persistent storage for program and data files and may include hard disk drives, floppy disk drives, and associated removable media, CD-ROM drives, optical disk drives, or removable media cartridges. Modules implementing certain functionalities of the implementation may be stored by the file storage subsystem 726 within the storage subsystem 724, or in other machines accessible to the processor 714.

[0052] The bus subsystem 712 provides a mechanism for enabling the various components and subsystems of the computer system 710 to communicate with each other as intended. Although the bus subsystem 712 is schematically shown as a single bus, alternative implementations of the bus subsystem may use multiple buses.

[0053] Computer systems 710 can be of different types, including workstations, servers, computing clusters, blade servers, server farms, or any other data processing systems or computing devices. Due to the constantly changing nature of computers and networks, therefore... Figure 7 The description of the computer system 710 depicted herein is intended only as a specific example for illustrating some implementation methods. Many other configurations of the computer system 710 are possible, and these configurations have the same characteristics as... Figure 7 The computer system depicted in the text has more or fewer components compared to the one described.

[0054] Figures 8 to 11 The illustration depicts non-restrictive examples of how agents can be implemented according to various implementation methods. Figure 8 This describes how the planner agent 804 can process data according to various implementations. The list of processes on the left provides a list of different types of agents (or "processes" or "tools") available to the orchestrator agent 804. In this example, these types of agents include data loading, data cleaning, data preparation, data analysis, data visualization, feature engineering, model training, model optimization, model evaluation, data exploration, data segmentation, and data preparation.

[0055] Among the various implementation methods, task data ( For example Patch trigger data 220) and a list of worker processes can be provided to planner agent 804. Planner agent 804 (which can be implemented in one of the sandbox environments mentioned elsewhere in this document) can be configured as follows: For example The task data and list of work processes are processed using program logic and machine learning models (such as trained classifiers and generative models) to generate output. In this example, the output may include a plan. For instance, the plan may include and / or identify work processes in the list that should be initiated at each of several steps.

[0056] Figure 9 The diagram schematically illustrates how orchestrator agent 904 can operate in several implementations. Task data, plans generated by planner agent 804, and a history of generated / executed subtasks can be provided to orchestrator agent 904. Orchestrator agent 904 can then... For example The data is processed using one or more generative models to generate one or more subtasks. As shown by the arrows pointing from the subtasks to the task data, and as indicated at the bottom, in some implementations, these subtasks can be generated incrementally, and the newly generated subtask can be added to the history to generate the next subtask.

[0057] Figure 10 With Figure 9 Different methods are illustrated to illustrate how the orchestrator agent 904 can operate in some implementations. Starting from the left, the orchestrator agent 904 can process states to identify which worker process should be triggered next ("data_load" in the first instance). The triggered data_load worker process ( Figure 10 The output of the "summary" can be added to the state. During the next iteration, the updated state can trigger the data cleanup worker process. For example This process handles any missing data and / or inconsistencies previously added to the state. The result is an updated state that includes an updated summary.

[0058] Figure 11This paper schematically depicts an example of how a worker agent can be implemented. The state can include the high-level task being executed, a specific subtask assigned to the worker agent, and a plan for executing that subtask. The specific worker agent has two actions available: `Code_Block` (to generate a code block) and `Complete_Task` (to complete code generation). In this paper, the result of the first state is that the worker agent selects the `Code_Block` action, at which point data indicating inference, data frames, and code execution is added to the state. This is repeated during the next iteration, and additional data indicating inference, data frames, and code execution is added to the state again. Finally, when the selected action is `Complete_Task`, data indicating inference is added to the state.

[0059] Among various implementations, one method is provided for generating multi-file patches for an application's codebase. One or more agents can identify one or more component source files in the codebase that are suitable for patch trigger data. Patch trigger data may include, for example, error message data generated by application execution, or problem statements issued by the agents. One or more agents can use one or more generative models to generate multiple candidate source code changes and multiple unit tests. One or more agents can run different arrangements of unit tests and candidate source code changes. Based on the results, a candidate source code change can be selected and incorporated into the multi-file patch.

[0060] In various implementations, multi-file patches can be applied to a codebase. One or more of the generative models can include one or more large language models. A first generative model can be used to generate multiple candidate multi-file patches, and a second generative model can be used to generate multiple unit tests. Iterative generation and testing of candidate source code changes can be performed, where each iteration can begin from the application state after the previous iteration has successfully completed.

[0061] Different arrangements of unit tests and candidate source code changes can be executed in one or more sandbox environments. Sandbox environments may include virtual machines. Virtual machines may be snapshottable and recoverable. Patch trigger data may include data generated via application execution. Patch trigger data may include error messages generated based on unexpected application termination. Patch trigger data may include natural language problem statements.

[0062] One or more agents can generate natural language descriptions of changes to the codebase based on patch-triggered data. Multiple candidate source code changes can be generated based on one or more of these natural language descriptions. Multiple unit tests can be generated based on one or more of these natural language descriptions. Each of these unit tests can be generated based on a corresponding candidate source code change. At least one of the unit tests generated based on one of the candidate source code changes can be used to test another candidate source code change. These one or more component files can be recognized by a programmatic agent. These one or more component files can be recognized by a dynamic agent.

[0063] In other implementations, a system and / or a temporary or non-temporary computer-readable storage medium may be provided for performing any of the methods described above. The system may include one or more processors and a memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform any of the methods described above. In various implementations, a temporary or non-temporary computer-readable storage medium may be provided. The computer-readable storage medium may include instructions that, when executed by the one or more processors, cause the one or more processors to perform any of the methods described above.

[0064] On the other hand, a system for generating and testing code changes in a codebase is provided. This system may include multiple agents, at least one of which can identify relevant files in the codebase, generate multiple candidate patches based on those files, and generate multiple unit tests for each candidate patch. An evaluator can evaluate different permutations of the unit tests and candidate patches to select a given patch from the multiple candidate patches to include in a multi-file patch. At least one of the agents may include a large language model. The agents may be selected from a group of procedural agents, dynamic agents, and hybrid agents. A sandbox environment may be provided to execute the candidate patches and unit tests. The sandbox environment may include a snapshottable and recoverable microvirtual machine (microVM).

[0065] While several implementations have been described and illustrated herein, a variety of other means and / or structures may be utilized to perform functions and / or obtain results and / or one or more of the advantages described herein, and each of these variations and modifications is considered to be within the scope of the implementations described herein. More generally, all parameters, dimensions, materials, and configurations described herein are intended to be exemplary, and actual parameters, dimensions, materials, and / or configurations will depend on one or more specific applications for which this teaching is used. Those skilled in the art will recognize, or confirm, many equivalents of the implementations described herein using only conventional experimentation. Therefore, it should be understood that the foregoing implementations are presented by way of example only, and it should be understood that implementations may be practiced in ways other than those specifically described and claimed within the scope of the appended claims and their equivalents. The implementations of this disclosure relate to each individual feature, system, article of manufacture, material, kit, and / or method described herein.

Claims

1. A method for generating multi-file patches for an application's codebase, the method comprising: Enable one or more intelligent agents to identify one or more component source files in the codebase that are suitable for patch triggering data; One or more of the intelligent agents generate multiple candidate source code changes using one or more generative models, wherein the candidate source code changes are based on the one or more constituent source code files; One or more of the intelligent agents generate multiple unit tests using one or more of the generative models; One or more of the agents may run different arrangements of the unit tests and the candidate source code changes; as well as Based on the results of the unit tests and the different arrangements of the candidate source code changes, one of the candidate source code changes is selected and incorporated into the multi-file patch.

2. The method of claim 1, further comprising applying the multi-file patch to the codebase.

3. The method as described in claim 1 or 2, wherein, One or more of the generative models include one or more large language models.

4. The method as described in any of the preceding claims, wherein, The first generation model is used to generate the plurality of candidate multi-file patches, and the second generation model is used to generate the plurality of unit tests.

5. The method as described in any of the preceding claims, further comprising: Candidate source code changes are generated and tested iteratively, with each iteration starting from the state of the application after the previous iteration has been successfully completed.

6. The method as described in any of the preceding claims, wherein, The unit tests and the different arrangements of the candidate source code changes are executed in one or more sandbox environments.

7. The method of claim 6, wherein, The sandbox environment includes virtual machines.

8. The method of claim 7, wherein, The virtual machine is snapshottable and recoverable.

9. The method as described in any of the preceding claims, wherein, The patch trigger data includes data generated via the execution of the application.

10. The method of claim 9, wherein, The patch trigger data includes error messages generated based on the unexpected termination of the application.

11. The method as described in any of the preceding claims, wherein, The patch trigger data includes natural language problem statements.

12. The method as described in any of the preceding claims, further comprising causing one or more of the agents to generate a natural language description of the changes to be made to the codebase based on the patch trigger data.

13. The method of claim 12, wherein, The multiple candidate source code changes are generated based on one or more of the natural language descriptions.

14. The method of claim 12, wherein, The multiple unit tests are generated based on one or more of the natural language descriptions.

15. The method as described in any of the preceding claims, wherein, Each of the multiple unit tests is generated based on the corresponding candidate source code changes.

16. The method of claim 15, wherein, At least one of the unit tests generated based on one of the candidate source code changes is used to test the other of the candidate source code changes.

17. The method as described in any of the preceding claims, wherein, The one or more component files are identified by a programmatic intelligent agent.

18. The method as described in any of the preceding claims, wherein, The one or more component files are identified by the dynamic intelligent agent.

19. A system comprising one or more processors and a memory, the memory storing instructions that, in response to execution by the one or more processors, cause the one or more processors to perform the method as claimed in any one of claims 1 to 18.

20. A computer-readable medium comprising at least one transient or non-transitory instruction, which, in response to execution by one or more processors, causes the one or more processors to perform the method as claimed in any one of claims 1 to 18.

21. A system for generating and testing code changes in a codebase, comprising: Multiple intelligent agents, wherein at least one intelligent agent identifies relevant files in the code repository, at least one intelligent agent generates multiple candidate patches based on the relevant files, and at least one intelligent agent generates multiple unit tests for the candidate patches; as well as An evaluator evaluates different permutations of the unit tests and the candidate patches to select a given patch from the plurality of candidate patches to be included in a multi-file patch.

22. The system of claim 21, wherein, At least one of the intelligent agents includes a large language model.

23. The system as claimed in claim 21 or 22, wherein, The agent is selected from a group consisting of programmed agents, dynamic agents, and hybrid agents.

24. The system according to any one of claims 21 to 23, further comprising: A sandbox environment used to execute the candidate patches and the unit tests.

25. The system of claim 24, wherein, The sandbox environment includes a snapshot-enabled and recoverable microVM.

26. A method for iteratively generating and testing code changes for an application, the method comprising: (a) Establish multiple sandbox virtual machine environments, each of which includes a snapshottable and recoverable state of the application; (b) In at least one of the sandbox virtual machine environments, multiple candidate code changes are generated based on a triggering event; (c) In at least one of the sandbox virtual machine environments, generate multiple unit tests corresponding to the multiple candidate code changes; (d) In at least one of the multiple candidate code changes and multiple unit tests, perform at least one permutation of the multiple candidate code changes and the multiple unit tests; (e) Restore the state of at least one of the sandboxed virtual machine environments from the snapshot; as well as (f) Based on the results of step (d), repeat steps (b) to (e).

27. A method for iteratively generating and testing code changes, comprising: Create multiple sandbox virtual machine environments, each of which includes a snapshottable and recoverable state of the application; In at least one of the sandbox virtual machine environments, multiple candidate code changes are generated; In at least one of the sandbox virtual machine environments, multiple unit tests corresponding to the multiple candidate code changes are generated; In at least one of the sandbox virtual machine environments, at least one permutation of the plurality of candidate code changes and the plurality of unit tests is executed; Restore the state of at least one of the sandboxed virtual machine environments from the snapshot; as well as Based on the results of the execution steps, repeat the generation, generation, and execution steps.

28. A method implemented using one or more processors, comprising: Execute the application in one or more virtual machines until one or more predetermined events occur; Upon detecting the first event in the predetermined events, a snapshot of the first application state of the application corresponding to the detection of the first event is captured; Based on the first event, an instance of a second virtual machine-managed patch generation agent is made, which iteratively generates and tests candidate source code changes to generate multi-file patches for the application's codebase. The multi-file patch is applied to the codebase to generate an updated application; as well as The updated application is executed in the first virtual machine or the third virtual machine, starting from the first application state, until one or more of the predetermined events occur.

29. The method of claim 28, further comprising: Upon detecting another given event among the predetermined events, capture a snapshot of the second application state of the application corresponding to the detection of the other given event; Based on the other given event, the same instance of the patch generating agent hosted by the second virtual machine or a different instance of the patch generating agent hosted by the fourth virtual machine iteratively generates and tests candidate source code changes to generate another multi-file patch for the application's codebase. as well as The other multi-file patch is applied to the codebase to generate a further updated application.

30. The method of claim 28 or 29, wherein, One or more instances of the virtual machine include microvirtual machines.

31. The method of any one of claims 28 to 30, wherein, One or more of the predetermined events include error messages.

32. The method according to any one of claims 28 to 31, wherein, One or more of the predetermined events include the unexpected termination of the application.

33. The method according to any one of claims 28 to 32, wherein, One or more of the predetermined events include a functional failure of the application.

34. A system comprising one or more processors and a memory, the memory storing instructions that, in response to execution by the one or more processors, cause the one or more processors to perform the method as claimed in any one of claims 28 to 33.

35. A computer-readable medium comprising at least one transient or non-transitory instruction, which, in response to execution by one or more processors, causes the one or more processors to perform the method as described in any one of claims 28 to 33.