Method, device and computer program for validating code generated by artificial intelligence using abstract syntax trees (validating code generated by artificial intelligence using abstract syntax trees)
The use of an abstract syntax tree for comparing and regenerating AI-generated code during migration addresses the challenge of verifying and improving the accuracy of code conversion from legacy to modern languages, ensuring functional equivalence and maintainability.
Patent Information
- Application Number
- JP2024147906
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2024-08-29
- Publication Date
- 2025-07-10
AI Technical Summary
Migrating code from a legacy programming language to a modern language using artificial intelligence is laborious and requires thorough verification of the converted code's accuracy and functionality, especially when dealing with large codebases.
Utilizing an abstract syntax tree (AST) to compare the equivalence between the original and converted code, identifying non-equivalence, and iteratively regenerating the code until verification success is achieved, with the AI language model being retrained based on the verification results.
Enhances the reliability and maintainability of code conversion by ensuring functional equivalence, allowing for accurate assessment and iterative improvement of the AI-generated code.
Smart Images

Figure 2025105427000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method, an apparatus, and a product for verifying code generated by artificial intelligence using an abstract syntax tree.
Background Art
[0002] By migrating the functions of legacy source code to a more modern programming language, the maintainability and readability of the source code can be improved, and the system performance can be enhanced.
Summary of the Invention
Problems to be Solved by the Invention
[0003] However, such migration can be a laborious task that may involve writing, testing, verifying, and debugging a huge amount of code.
Means for Solving the Problems
[0004] According to embodiments of the present disclosure, various methods, apparatuses, and products for verifying code generated by artificial intelligence using an abstract syntax tree are described herein. In some aspects, an artificial intelligence (AI) language model is used to remap application source code from an original codebase to a target codebase while maintaining the same functionality. In some aspects, an abstract syntax tree (AST) is used to verify the conversion of original application source code to AI-generated source code. The equivalence mapping between the AST of the input source code and the AST of the output source code indicates the conversion accuracy of the AI-generated code. The verification result can be identified from the degree of equivalence shown in the equivalence mapping. In some aspects, if it is discovered that the AST contains non-equivalence, the AI language model can be prompted to regenerate the code. In some aspects, improvements to the language model can be measured using AST comparison as a verification metric. Thus, AST comparison facilitates code verification when migrating from an original codebase to a new codebase using AI-generated code, for example, from a first programming language to a second programming language, or from a legacy system to a modernized system.
[0005] In certain embodiments, a method of verifying code generated by artificial intelligence using an abstract syntax tree comprises an artificial intelligence (AI) language model generating output source code based on input source code. The method also comprises determining an equivalence mapping between a first abstract syntax tree (AST) constructed for the input source code and a second AST constructed for the output source code. The method further comprises indicating a verification result for the output source code based on the equivalence mapping. In some examples, the input source code is implemented in a first programming language and the output source code is implemented in a second programming language different from the first programming language. In this way, even if the ASTs show different structures, it can be determined whether the ASTs of the input source code and the output source code are functionally equivalent. All permutations of core functions can be relatedly mapped for AST comparison.
[0006] In some variations, the verification result indicates verification failure if the equivalence mapping shows at least one non-equivalent element in at least one of the first AST and the second AST, and the verification result indicates verification success if the equivalence mapping shows equivalence for all elements of the first AST and the second AST. In other variations, the verification result indicates a degree of equivalence. The purpose of code conversion is for the AI-generated output source code to have an AST equivalent to the AST of the input source code, and thus the verification result can be pass / fail; however, the degree of equivalence is useful when evaluating the progress of the AI language model before and after retraining and for determining whether the identified non-equivalence results in a change in functionality.
[0007] In some embodiments, determining the equivalence mapping between a first AST constructed for input source code and a second AST constructed for output source code involves splitting the first AST and the second AST into subtrees, and identifying the equivalence between the subtrees of the first AST and the second AST. In this way, complex ASTs constructed from hundreds of thousands, if not millions, of lines of code are unitized for comparison, such that the independent subtrees of one AST are rearranged to match the structure of the other AST, facilitating comparison and determining equivalence.
[0008] In some embodiments, identifying the equivalence between the subtrees of the first AST and the second AST involves identifying one or more equivalent permutations of a first subtree of one AST, and determining that a second subtree in the other AST matches one of the one or more equivalent permutations. In some examples, one or more equivalent permutations of the first subtree are generated by one or more of a variable name change and a rearrangement of independent statements. In this way, subtrees that have different shapes but the same path can be used to determine whether one tree is functionally equivalent to the other subtree. This enables rearrangement such that elements in one AST match the shape of the other AST.
[0009] In some embodiments, the method also comprises indicating the positions of non - equivalent elements found in at least one of the input source code and the output source code. In this way, a software engineer can analyze the conversion error and determine whether the error (i.e., indicated by non - equivalence in the AST) is a critical error. The engineer can further use this information to identify ways to retrain an AI language model.
[0010] In some variations, the method also comprises the AI language model regenerating new output source code from the input source code in response to the verification result. In this way, the AI language module can iteratively regenerate the output source code until an acceptable verification score is achieved.
[0011] In some variations, the method also comprises, following retraining of the AI language model, determining a second verification result for the regenerated output source code. In these variations, the method further comprises quantifying an improvement to the AI language model based at least on the verification result and the second verification result. In this way, the accuracy and reliability of the AI language model can be evaluated, and the results of the retraining of the AI language model can be measured.
[0012] In some aspects, the apparatus may comprise a processing device; and a memory operably coupled to the processing device, the memory storing computer program instructions that, when executed, configure the processing device to perform the operations described above. In some aspects, a computer program product comprising a computer-readable storage medium may store computer program instructions that, when executed, configure a computer to perform the operations described above.
Brief Description of the Drawings
[0013]
Figure 1
[0014]
Figure 2
[0015]
Figure 3
[0016]
Figure 4
[0017]
Figure 5
[0018]
Figure 6
[0019]
Figure 7
[0020] In the world of software development, there is an increasing need to modernize a codebase from one programming language to another. For example, the source code for an application can be migrated from a legacy programming language (e.g., COBOL) to a modern programming language (e.g., Java (registered trademark)). Motivations for such migrations can include facilitating easier maintenance and readability of the source code, enhancing security and error handling, improving software and / or hardware performance, and other advantages that will be recognized by those skilled in the art.
[0021] According to the present disclosure, artificial intelligence (AI) is used to transplant or migrate the source code of an application into a different programming language. To develop a generative AI that can output source code based on an input or prompt, a large language model (LLM) is trained on a dataset containing a vast amount of source code. That is, an AI language model is used to generate new source code based on the input of the original source code. For example, the AI language model may be given a prompt such as "generate Java code that achieves the same purpose as the following COBOL code", in which case the legacy COBOL source code is provided as the input. In response, the AI language model may ideally output AI-generated Java source code that performs the same function as the legacy and produces the same output.
[0022] However, migrating a codebase to a new language presents significant challenges in ensuring the accuracy and functionality of the converted code, especially when leveraging AI for automatic conversion. The difficulty lies in the verification of AI-generated code conversion and in checking whether the converted code retains the intended logic, functionality, and structure of the original code. The inherent complexity of programming languages, combined with the ways in which developers express their logic, presents challenges in reliably verifying the accuracy and similarity of AI-generated conversions. Furthermore, verifying the output source code converted from the input source code requires analyzing hundreds of thousands if not millions of lines of code.
[0023] The present disclosure particularly focuses on enhancing the reliability and maintainability of automatic code conversion through the comparison of Abstract Syntax Trees (ASTs), and addresses issues associated with verifying the accuracy of AI-generated code conversion. According to the present disclosure, the AST is utilized to verify the AI-converted code by decomposing the input source code and the AI-generated source code into elements such as if statements, loops, function declarations, variable declarations, etc. represented in the AST. Next, both ASTs are traversed, and the types and properties of each node are compared to identify functional equivalence. Program-based fuzzing and complexity obfuscation are applied to the shapes in the AST, whereby all permutations of the core functions in the program can be relationally mapped when performing AST comparison. In some examples, the AST is unitized by decomposing the AST into subtrees for direct comparison. In some examples, the verification results identify overfitting and underfitting in the AI-generated code, and this metric is used for the analysis of conversion accuracy.
[0024] Referring now to FIG. 1, an exemplary computing environment in accordance with aspects of the present disclosure is shown. Computing environment 100 includes an example of an environment for the execution of at least a portion of the computer code involved in performing the various methods described herein, such as code analysis module 107. In addition to code analysis module 107, computing environment 100 includes, for example, computer 101, wide area network (WAN) 102, end user device (EUD) 103, remote server 104, public cloud 105, and private cloud 106. In the present embodiment, computer 101 includes a processor set 110 (including processing circuitry 120 and cache 121), a communication fabric 111, volatile memory 112, persistent storage 113 (including operating system 122 and code analysis module 107 as specified above), a set of peripheral devices 114 (including a set of user interface (UI) devices 123, storage 124, and a set of Internet of Things (IoT) sensors 125), and a network module 115. Remote server 104 includes remote database 130. Public cloud 105 includes gateway 140, cloud orchestration module 141, a set of host physical machines 142, a set of virtual machines 143, and a set of containers 144.
[0025] Computer 101 can take the form of a desktop computer, laptop computer, tablet computer, smartphone, smartwatch, or other wearable computer, mainframe computer, quantum computer, or any other form of computer or mobile device that is currently known or will be developed in the future and that can execute programs, access a network, or query a database such as remote database 130. As is well understood in the field of computer technology and depending on the technology, the execution of computer implementation methods can be distributed among multiple computers and / or multiple locations. On the other hand, in this description of computing environment 100, for the sake of simplicity as much as possible, the detailed discussion focuses on a single computer, specifically computer 101. Although computer 101 is not shown within the cloud in FIG. 1, it can be located within the cloud. On the other hand, computer 101 is not required to exist within the cloud except within any arbitrarily shown range.
[0026] Processor set 110 includes one or more computer processors of any type that are currently known or will be developed in the future. Processing circuit 120 can be distributed across multiple packages, for example, multiple integrated circuit chips that have been adjusted. Processing circuit 120 can implement multiple processor threads and / or multiple processor cores. Cache 121 is a memory located within the processor chip package and is typically used for high-speed access to data or code that should be available to the threads or cores executing on processor set 110. Cache memory is typically organized into multiple levels depending on its relative proximity to the processing circuit. Alternatively, some or all of the cache for the processor set can be located "off-chip". In some computing environments, processor set 110 can be designed to operate using qubits and execute quantum computing.
[0027] Computer-readable program instructions are typically loaded onto computer 101 and executed by a set of processors 110 of computer 101 in a series of operational steps, thereby enabling a computer-implemented method. As a result, the instructions so executed will instantiate the method specified in the flowchart and / or description of the computer-implemented method included in this document. These computer-readable program instructions are stored in various types of computer-readable storage media such as cache 121 and other storage media discussed below. The program instructions and associated data are accessed by the set of processors 110 to control and direct the execution of the computer-implemented method. In computing environment 100, at least some of the instructions for executing the computer-implemented method may be stored in code analysis module 107 within persistent storage 113.
[0028] Communication fabric 111 is a signal conduction path that enables various components of computer 101 to communicate with each other. Typically, this fabric is created by switches and conductive paths such as buses, bridges, physical input / output ports, and switches and conductive paths that make up the equivalent, etc. Other types of signal communication paths such as optical fiber communication paths and / or wireless communication paths may be used.
[0029] Volatile memory 112 is any type of volatile memory, known currently or to be developed in the future. Examples include dynamic random access memory (RAM) or static RAM. Typically, volatile memory 112 is characterized by random access, although this is not required unless affirmatively indicated. In computer 101, volatile memory 112 is located within a single package and exists within computer 101. Alternatively or additionally, volatile memory may be distributed across multiple packages and / or located external to computer 101.
[0030] The persistent storage 113 is any form of non-volatile storage for a computer that is currently known or will be developed in the future. The non-volatility of this storage means that the stored data is maintained regardless of whether power is directly supplied to the computer 101 and / or to the persistent storage 113. The persistent storage 113 can be read-only memory (ROM), but usually at least a part of the persistent storage enables writing of data, deletion of data, and rewriting of data. Some well-known forms of persistent storage include magnetic disks and solid-state storage devices. The operating system 122 can take multiple forms, such as various known proprietary operating systems or open-source Portable Operating System Interface (POSIX)-type operating systems that employ a kernel. The code included in the code analysis module 107 typically includes at least a part of the computer code involved in the execution of the computer-implemented methods described herein.
[0031] The peripheral device set 114 includes a set of peripheral devices of the computer 101. The data communication connections between the peripheral devices of the computer 101 and other components can be implemented in various ways, such as a Bluetooth (registered trademark) connection, a Near-Field Communication (NFC) connection, a connection by a cable (such as a Universal Serial Bus (USB) type cable), an insertion type connection (for example, a Secure Digital (SD) card), a connection through a local area communication network, and even a connection through a wide area network such as the Internet. In various embodiments, the UI device set 123 may include components such as a display screen, a speaker, a microphone, wearable devices (such as goggles and smartwatches), a keyboard, a mouse, a printer, a touchpad, a game controller, and a haptic device. The storage 124 is an external storage such as an external hard drive or an insertable storage such as an SD card. The storage 124 can be persistent and / or volatile. In some embodiments, the storage 124 can take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments where the computer 101 is required to have a large amount of storage (for example, when the computer 101 locally stores and manages a large-scale database), this storage can be provided by a peripheral storage device designed to store a very large amount of data, such as a storage area network (SAN) shared by a plurality of geographically dispersed computers. The IoT sensor set 125 is composed of sensors that can be used in applications of the Internet of Things. For example, one sensor can be a thermometer, and another sensor can be a motion detector.
[0032] The network module 115 is an aggregate of computer software, hardware, and firmware that enables the computer 101 to communicate with other computers through the WAN 102. The network module 115 may include hardware such as a modem or a Wi-Fi (registered trademark) signal transceiver, software for packetizing and / or depacketizing data for communication network transmission, and / or web browser software for communicating data via the Internet. In some embodiments, the network control function and the network transfer function of the network module 115 are executed on the same physical hardware device. In other embodiments (for example, embodiments that utilize Software-Defined Networking (SDN)), the control function and the transfer function of the network module 115 are executed on physically separate devices such that the control function manages multiple different network hardware devices. Computer-readable program instructions for executing a computer implementation method may typically be downloaded to the computer 101 from an external computer or an external storage device through a network adapter card or a network interface included in the network module 115.
[0033] The WAN 102 is any wide area network (such as the Internet) that can communicate computer data over non-local distances by any technology for communicating computer data that is currently known or will be developed in the future. In some embodiments, the WAN 102 may be replaced and / or supplemented by a local area network (LAN) designed to communicate data between devices located in a local area, such as a Wi-Fi network. The WAN and / or the LAN typically includes computer hardware such as copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and edge servers.
[0034] The end-user device (EUD) 103 is any computer system used and controlled by an end user (e.g., a customer of the enterprise operating the computer 101), and can take any of the forms discussed above in relation to the computer 101. The EUD 103 typically receives beneficial and useful data from the operation of the computer 101. For example, in a virtual case where the computer 101 is designed to provide recommendations to an end user, this recommendation will typically be communicated from the network module 115 of the computer 101, via the WAN 102, to the EUD 103. In this way, the EUD 103 can display or otherwise present the recommendation to the end user. In some embodiments, the EUD 103 can be a client device such as a thin client, a thick client, a mainframe computer, a desktop computer, and the like.
[0035] The remote server 104 is any computer system that provides at least some data and / or functions to the computer 101. The remote server 104 can be controlled and used by the same entity that operates the computer 101. The remote server 104 represents a machine that collects and stores beneficial and useful data for use by other computers such as the computer 101. For example, in a virtual case where the computer 101 is designed and programmed to provide recommendations based on past data, this past data can be provided from the remote database 130 of the remote server 104 to the computer 101.
[0036] The public cloud 105 is any computer system available for use by multiple entities that provides on-demand availability of computer system resources and / or other computing capabilities, particularly data storage (cloud storage) and computing power, without direct active management by the user. Cloud computing typically exploits resource sharing to achieve coherence and economies of scale. The direct active management of the computing resources of the public cloud 105 is performed by the computer hardware and / or software of the cloud orchestration module 141. The computing resources provided by the public cloud 105 are typically implemented by virtual computing environments that run on various computers that make up the host physical machine set 142, which is the universe of physical computers within and / or available in the public cloud 105. Virtual computing environments (VCEs) typically take the form of virtual machines from a virtual machine set 143 and / or containers from a container set 144. It is understood that these VCEs can be stored as images and transferred either as images or after instantiation of the VCE, within and among various physical machine hosts. The cloud orchestration module 141 manages the transfer and storage of images, deploys new instantiations of VCEs, and manages the active instantiation of VCE deployments. The gateway 140 is an aggregate of computer software, hardware, and firmware that enables the public cloud 105 to communicate via the WAN 102.
[0037] Here, some further explanations are provided for a virtualized computing environment (VCE). A VCE can be stored as an "image". A new active instance of a VCE can be instantiated from the image. Two well-known types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to a feature of the operating system where the kernel enables the existence of multiple isolated instances of user space, called containers. These isolated instances of user space typically behave as actual computers from the perspective of the programs running within them. A computer program running on a normal operating system can utilize all the resources of that computer, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, a program running inside a container can only use the contents of the container and the devices allocated to the container, and this feature is known as containerization.
[0038] The private cloud 106 is similar to the public cloud 105, except that computing resources are only available for use by a single enterprise. The private cloud 106 is shown as being in communication with the WAN 102, but in other embodiments, the private cloud may be completely disconnected from the Internet and only accessible via a local / private network. A hybrid cloud is a composite of multiple different types of clouds (e.g., private cloud, community cloud, or public cloud types), and is often implemented by different vendors. Each of the multiple clouds remains a separate discrete entity, but the larger hybrid cloud architecture is coupled by standardized or proprietary technologies that enable orchestration, management, and / or data / application portability between the multiple constituent clouds. In this embodiment, both the public cloud 105 and the private cloud 106 are part of a larger hybrid cloud.
[0039] For further explanation, FIG. 2 shows a flowchart of an exemplary method for verifying code generated by artificial intelligence using an abstract syntax tree, according to some embodiments of the present disclosure. The method of FIG. 2 may be performed by a code analysis module 201, such as the code analysis module 107 of FIG. 1, for example. In some examples, the code analysis module 201 may be implemented as a process or service that includes an AI language model that generates output source code from input source code. In other examples, the code analysis module 201 may be implemented as part of a process or service separate from the process or service that includes the AI language model. In a further example, the code analysis module 201 may be implemented as part of a process or service that monitors the quality of the AI language model to evaluate whether retraining of the AI language model is appropriate or successful.
[0040] The method of FIG. 2 comprises 202 an artificial intelligence (AI) language model 211 generating an output source code 205 based on an input source code 203. The AI language model 211 can be trained with a vast dataset of original source code in a first programming language remapped to source code in different programming languages. Therefore, the AI language model 211 is configured to autonomously convert a block of input source code in one programming language into a block of output source code in a different programming language. In some examples, the input source code and the output source code reflect the migration of the source code of an application from a first programming language (e.g., a legacy codebase) to a second programming language (e.g., a modern codebase). For example, the input source code may include legacy source code written in an older programming language (e.g., COBOL), while the output source code may be implemented in a modern programming language (e.g., Java); however, both the input source code and the output source code are intended to achieve the same purpose, provide the same interface, and produce the same output.
[0041] In some examples, the output source code 205 is generated by prompting the AI language model to generate the output source code based on the input source code 203. For example, the AI language model can be prompted to "generate Java source code from block A of the COBOL source code" when block A is provided as the input source code. In response, the AI language model generates Java source code that is intended to provide the same interface, perform the same function, and produce the same output as the original COBOL source code.
[0042] The method of FIG. 2 also includes determining 204 an equivalence mapping 207 between a first AST 213 constructed for the input source code 203 and a second AST 215 constructed for the output source code 205. The code analysis module 201 determines 204 the equivalence mapping by identifying the AST 213 for the input source code and the AST 215 for the output source code. These ASTs can be constructed by the code analysis module 201, or the code analysis module 201 can adopt ASTs constructed by software development utilities or compilers.
[0043] An AST is a hierarchical tree data structure that represents the syntactic structure of source code in a programming language. The AST abstracts the specific syntactic details, including information that is not necessary for understanding the semantics of the program. The AST focuses on the relationships between language elements such as expressions, statements, and declarations, while removing and considering details such as parentheses, brackets, and other punctuation. The AST is implemented as a tree of nodes that represent different syntactic constructs (e.g., IF statements, loops, function and variable declarations, assignments) in the source code. These nodes are organized in a hierarchical structure that reflects the nested and hierarchical nature of the source code. The AST provides a concise and abstract view of the code that facilitates automatic static code analysis. According to the present disclosure, the comparison of the ASTs of the input source code and the AI-generated output source code is used to verify that the output source code is functionally equivalent to the input source code.
[0044] Here, the equivalence between two ASTs means that for all nodes in the first tree, there are equivalent nodes in the second AST, and for all edges between two nodes in the first tree, there is an edge between two equivalent nodes in the second AST. The equivalence of two ASTs does not require the trees to show the same order if the functionality of the code is not modified by the order of independent statements or the order of independent paths. In the present disclosure, the terms "node" and "element" of an AST can be used interchangeably.
[0045] In some examples, the code analysis module 201 determines 204 the equivalence mapping 207 by identifying whether there are equivalent nodes and edges in the second AST for all nodes and edges in the first AST. In some implementations, the code analysis module 201 identifies equivalent nodes by walking each AST 213, 215 and comparing the types and properties of each node. For example, the type can be an abstract type that is non-syntax-dependent, such as an IF statement, a loop (e.g., "for" or "do...while"), a function declaration, a variable declaration, and the like. The properties of each node can include the degree of the node and the set of nodes to which the node is connected. In some cases, the nodes of the AST may not be labeled with information indicating a type or other non-observable property. In such cases, determining the equivalence mapping can be accomplished by matching patterns in one AST with patterns in the other AST.
[0046] In some examples, the code analysis module 201 unitizes each AST 213, 215 by decomposing the AST into subtrees for comparison, as will be described in more detail below. In these examples, the code analysis module 201 determines 204 the equivalence mapping 207 by identifying subtrees in one AST that match subtrees in the other AST. Once equivalent subtrees are identified, the independent subtrees can be rearranged as needed to reconstruct the AST 215 of the output source code 205 to match the AST 213 of the input source code 203.
[0047] In some examples, the code analysis module 201 determines 204 the equivalence mapping 207 by creating a data structure that indicates the determination of equivalence for each element. If an element of the second AST is identified as equivalent to an element of the first AST, the equivalence is recorded. If no equivalent element is identified in the other AST for an element in a given AST, non-equivalence is recorded. That is, a non-equivalent element in the second AST can be an element that has no corresponding relationship in the first AST, or can be an element of the first AST that is missing in the second AST. In some implementations, the code analysis module 201 records only non-equivalence. The equivalence mapping 207 can be used to verify the output source code 205.
[0048] The method of FIG. 2 also comprises 206 showing a verification result 209 for the output source code based on an equivalence mapping 207. In some examples, the code analysis module 201 generates the verification result 209 based on the degree of equivalence (including non-equivalence) indicated by the equivalence mapping. For example, the verification result can be a ranking of the output source code based on the number of non-equivalent elements found in the equivalence mapping 207. Alternatively, the verification result can be a score based on the ratio of non-equivalent elements to all elements of the second AST. In some implementations, the code analysis module 201 determines whether the output source code passes the verification when this rank or score is lower than a threshold. In other implementations, the verification result indicates verification failure when the equivalence mapping shows that at least one non-equivalent element has been found in either the first AST or the second AST, while the verification result indicates verification success when the equivalence mapping shows equivalence for all elements of the first AST and the second AST. In these implementations, the output source code is successfully verified only if the structure of its AST is equivalent to the structure of the input source code. Thus, in some examples, showing the verification result 209 for the output source code based on the equivalence mapping 207 206 can involve presenting the verification result, such as a pass / fail result, ranking, score, or other verification assessment, to the user or within a data structure. For example, the verification result can be presented in a graphical user interface or a command line interface.
[0049] For further explanation, FIG. 3 shows a flowchart of an exemplary method for verifying code generated by artificial intelligence using an abstract syntax tree, according to some embodiments of the present disclosure. The method of FIG. 3 extends the method of FIG. 2 in that determining the equivalence mapping 207 between the first AST 213 constructed for the input source code 203 and the second AST 215 constructed for the output source code 205, 204 has splitting 302 the first AST 213 and the second AST 215 into subtrees. For example, an AST can be split into subtrees of various granularities and nested subtrees. In some examples, the code analysis module 201 splits each AST 213, 215, 302 by selecting a leaf node and walking up the tree until it identifies the position of the root of the subtree. What constitutes the root of the subtree can be predefined. For example, the root of the subtree can be given as any operator or expression, or the root of the subtree can be more explicitly defined as a particular type of expression, such as a decision point or a function declaration, for example. Once the root of the subtree is identified, that node and all the nodes connected below it are grouped as a subtree.
[0050] As part of determining the equivalence mapping 204, the code analysis module 201 also identifies 304 the equivalence between subtrees of the first AST 213 and the second AST 215. In some examples, the code analysis module 201 identifies 304 the equivalence between subtrees by repeatedly selecting a subtree in one AST (e.g., AST 213) and mapping the first subtree to an equivalent subtree in the other AST (e.g., the second AST 215). For example, the code analysis module 201 can walk the selected subtree and evaluate the type and / or properties of each node to determine whether there is an equivalent subtree with equivalent nodes and structure in the other AST (e.g., the second AST 215). In some implementations, the code analysis module 201 further scores the equivalence between subtrees based on the presence of any non-equivalent elements (e.g., extra or missing nodes), or otherwise flags the equivalence. Using these subtrees as building blocks and by changing the positions of the subtrees, the code analysis module 201 can reconstruct the second AST 215 to be structurally identical to the first AST 213 (or vice versa) through iteration and recursion, revealing any missing or additional elements in either AST.
[0051] For further explanation, FIG. 4 shows a flowchart of an exemplary method for verifying code generated by artificial intelligence using an abstract syntax tree, according to some embodiments of the present disclosure. The method of FIG. 4 extends the method of FIG. 3 in that identifying the equivalence 304 between subtrees of a first AST 213 and a second AST 215 has identifying one or more equivalent permutations 402 of the first subtree. If the code analysis module 201 cannot identify an exact match for the subtree, it may analyze functionally equivalent permutations of the subtree. In some examples, the code analysis module 201 identifies 402 one or more equivalent permutations of the first subtree by rearranging independent statements, for example by swapping the positions of two nodes. In such an example, if the swapped nodes are the root of the subtree, the position of the entire subtree is swapped. Consider an example where a subtree S of tree X includes an ordered path [{A, B, E}, {A, B, D, F, G}, {A, C}]. Permutations of this subtree include the ordered path [{A, B, D, F, G}, {A, B, E}, {A, C}] by swapping nodes E and D, or the ordered path [{A, C}, {A, B, E}, {A, B, D, F, E}] by swapping nodes B and C, and the like. Although structurally different, these permutations are functionally identical. In other examples, identifying 402 one or more equivalent permutations of the first subtree may include modifying the variable names of the nodes when it is confirmed that one node associated with a first variable name corresponds to another node associated with a different variable name in that those nodes are functionally equivalent nodes.
[0052] As part of identifying the equivalence between the subtrees of the first AST213 and the second AST215 304, the method of FIG. 4 also includes determining 404 that the second subtree in the other AST matches one of one or more permutations. In some examples, the code analysis module 201 determines 404 that the second subtree matches one of the permutations of the first subtree by walking the subtrees and comparing the node types and properties. Continuing with the above example, the code analysis module 201 may determine that the subtree T in tree Y matches one of the permutations of the subtree S in tree X, including paths ordered as [{A, B, D, F}, {A, B, E}, {A, C}].
[0053] For further explanation, FIG. 5 shows a flowchart of an exemplary method for verifying code generated by artificial intelligence using an abstract syntax tree, according to some embodiments of the present disclosure. The method of FIG. 5 extends the method of FIG. 2 in that the method of FIG. 5 further comprises indicating 502 the location of non-equivalent elements found in at least one of the input source code and the output source code. In some cases, the code analysis module 201 will identify that the second AST 215 contains elements that do not correspond to elements in the first AST 213, or that the second AST 215 is missing elements that exist in the first AST 213. In such cases, the code analysis module 201 indicates 502 the location of the non-equivalent elements by annotating, highlighting, or otherwise calling attention to the portion of the output source code that contains the non-equivalent or missing elements. For example, the code analysis module 201 may indicate a function that contains a missing statement, or may call attention to additional statements not found in the first AST 213. In some examples, the code analysis module 201 provides verification analysis to a user interface. In such examples, the verification analysis may include the verification results, as well as the locations within the output source code that contain non-equivalent or missing elements. In this way, a software engineer can examine blocks of code and evaluate whether additional or missing elements are critical errors or merely conversion defects.
[0054] For further explanation, FIG. 6 shows a flowchart of an exemplary method for verifying code generated by artificial intelligence using an abstract syntax tree, according to some embodiments of the present disclosure. The method of FIG. 6 extends the method of FIG. 2 in that the method of FIG. 6 further comprises 602 the AI language model 211 regenerating the output source code 205 from the input source code 203 based on the verification result 209. In some examples, the code analysis module 201 determines that the verification result 209 for the output source code indicates that the output source code has failed the verification. Accordingly, the code analysis module 201 determines that the output source code, or some portion of the output source code, needs to be regenerated. In some examples, the code analysis module 201 generates a prompt in substantially the same manner as described above; however, in this case, the prompt indicates to the AI language model that the AI language model needs to generate a different implementation. In such a case, the code analysis module 201 may generate a prompt such as "Regenerate the code for block A" or "Regenerate code for block A that is syntactically different from the previously generated code". In response, the AI language model regenerates alternative code for the input code corresponding to block A. In some implementations, the code analysis module 201 repeatedly re-prompts the AI language model 211 to regenerate new output source code from the input source code, for example, until the output source code passes the verification or until a threshold number of attempts is reached. In some examples, the code analysis module 201 may identify blocks of code related to non-equivalence and prompt the AI language model to regenerate only those blocks.
[0055] In some implementations, the code analysis module 201 adjusts one or more parameters of the AI language model before regenerating the output source code. The AI language model may include configurable parameters that affect the creativity of the model's response to a prompt. For example, a temperature parameter adjusts the probability distribution used to select the next token for the output stream. When selecting the next token for the output stream, a lower temperature causes the language model to select tokens with probabilities within a narrower range, tending to produce a more deterministic output, while a higher temperature causes the language model to select tokens with probabilities within a wider range, tending to produce a more random output. Another exemplary parameter is the top-k parameter that controls the randomness of selecting the next token by telling the language model that it must select from the top k most probable tokens. Yet another exemplary parameter is the top-p parameter that controls the randomness of selecting the next token by telling the language model that it must select from the most probable tokens whose total probability exceeds a p-value.
[0056] In some examples, the code analysis module 201 adjusts one or more parameters of the AI language model in response to determining that one or more iterations of generating the output source code have failed verification. For example, as the number of iterations increases, the parameters controlling the creativity of the AI language model may be adjusted to increase the randomness of the output. In this way, the AI language model can be induced to generate solutions that are not similar to the failed solutions presented in previous iterations. In some examples, adjusting one or more parameters is accomplished by including a statement for adjusting the parameter, such as "set the temperature to 0.8", in the prompt.
[0057] For further explanation, FIG. 7 shows a flowchart of an exemplary method for verifying code generated by artificial intelligence using an abstract syntax tree, according to some embodiments of the present disclosure. The method of FIG. 6 extends the method of FIG. 2 in that the method of FIG. 6 further comprises determining 702 a second verification result for the regenerated output source code following retraining of the AI language model 211. In some examples, the AI language model 211 is retrained with an additional training dataset to improve the quality of the AI code conversion of the input source code. To evaluate whether the AI language model has been improved in terms of the quality and accuracy of the code conversion and to quantify such improvement, the AI language model is prompted to regenerate the output source code based on the input source code for which the verification score has previously been determined. In these examples, the code analysis module 201 determines 702 the second verification result in the manner described above when generating the first verification result 209.
[0058] The method of FIG. 6 also comprises quantifying 704 an improvement to the AI language model 211 based at least on the verification result and the second verification result. In some examples, the code analysis module 201 quantifies 704 the improvement to the AI language model 211 by comparing the first verification result 209 with the second verification result to determine whether the AI language model 211 has generated an output source code that is structurally identical to the input source code.
[0059] Embodiments are useful when migrating or porting an application from one programming language to a different programming language and from a legacy programming language to a more modern programming language, although it will be understood that in some examples the original source code and the new source code may be written in the same programming language.
[0060] In view of the above, by using an abstract syntax tree to verify code generated by artificial intelligence according to the present disclosure, a number of advantages are provided. Embodiments of the present disclosure improve the accuracy and quality of automatic code generation, further improve the reliability and maintainability of the source code generated through automatic code generation, thereby providing a mechanism for addressing the problem of verifying AI-generated source code against the original source code. This provides technical advantages in the field of automatic code conversion and throughout the fields of software development and maintenance. Consistent with these advantages, by comparing the ASTs of the input source code and the output source code, it is determined whether the input source code and the output source code are functionally equivalent, even though the ASTs may exhibit different structures. All permutations of the core functions can be relatedly mapped for AST comparison. Complex ASTs constructed from hundreds of thousands if not millions of lines of code can be unitized for comparison, whereby independent subtrees of one AST are rearranged to match the structure of the other AST to facilitate comparison and determine equivalence. Subtrees with different shapes but the same paths can be used to determine whether one tree is functionally equivalent to another subtree. This enables rearrangement such that elements in one AST match the shape of the other AST. Based on the verification result of the initial generation of the output source code, the AI language module can iteratively regenerate the output source code until an acceptable verification result is achieved. The verification result enables a software engineer to analyze the conversion error and determine whether the error (i.e., indicated by non-equivalence in the AST) is a critical error. The engineer can further use this information to identify ways to retrain the AI language model. Additionally, the verification result can be used to quantify the accuracy and reliability of the AI language model, and the effect of retraining the AI language model can be measured.
[0061] Various aspects of the present disclosure are illustrated by descriptions, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic included in embodiments of a computer program product (CPP). With respect to any flowchart, depending on the technology involved, operations may be performed in an order different from that shown in a given flowchart. For example, again depending on the technology involved, two operations shown in consecutive flowchart blocks may be performed in reverse order, as a single integrated step, simultaneously, or at least partially overlapping in time.
[0062] An embodiment of a computer program product (referred to herein as a "CPP embodiment" or "CPP") is, in the context of the present disclosure, a term used to describe any set of one or more storage media (also referred to as "media") collectively included in a set of one or more storage devices, which collectively contain machine-readable code corresponding to instructions and / or data for performing the computer operations specified in a given CPP claim. A "storage device" is any tangible device capable of holding and storing instructions for use by a computer processor. A computer-readable storage medium can be, but is not limited to, an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these media are floppy disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanically encoded devices (such as punch cards or pits / lands formed on the major surfaces of discs), or any suitable combination of the foregoing. A computer-readable storage medium shall not be construed as storage in the form of a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide, optical pulses passing through an optical fiber cable, electrical signals transmitted through a wire, and / or other transmission media, when the term is used in the context of the present disclosure. As will be understood by those skilled in the art, data is typically moved during normal operation of a storage device, such as during access, defragmentation, or garbage collection, at some random points in time, but the data is not transient while it is stored, and thus the storage device is not considered to be transient for this reason.
[0063] The description of various embodiments of the present disclosure is presented for illustrative purposes and is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used herein are selected to best describe the principles of the embodiments, the practical application, or a technical improvement over the technologies found in the market, or to enable other skilled artisans to understand the embodiments disclosed herein.
Claims
**Claim 1** A step of an artificial intelligence (AI) language model generating output source code based on input source code; A step of determining an equivalence mapping between a first abstract syntax tree (AST) constructed for the input source code and a second AST constructed for the output source code; and A step of showing a verification result for the output source code based on the equivalence mapping A method for verifying code generated by artificial intelligence using an abstract syntax tree, comprising: **Claim 2** The method according to claim 1, wherein the input source code is implemented in a first programming language, and the output source code is implemented in a second programming language different from the first programming language. **Claim 3** The verification result indicates verification failure when the equivalence mapping shows at least one non-equivalent element in at least one of the first AST and the second AST; and The verification result indicates verification success when the equivalence mapping shows equivalence for all elements of the first AST and the second AST The method according to claim 1. **Claim 4** The method according to claim 1, wherein the verification result indicates a degree of equivalence. **Claim 5** The step of determining an equivalence mapping between a first AST constructed for the input source code and a second AST constructed for the output source code comprises: A step of dividing the first AST and the second AST into subtrees; and A step of identifying the equivalence between the subtrees of the first AST and the second AST The method according to claim 1, comprising: **Claim 6** The step of identifying the equivalence between the subtrees of the first AST and the second AST comprises: A step of identifying one or more equivalent permutations of a first subtree of one AST; and A step of determining that a second subtree in the other AST matches one of the one or more equivalent permutations The method according to claim 5, comprising: **Claim 7** The method according to claim 6, wherein the one or more equivalent permutations of the first subtree are generated by one or more of a variable name change and a rearrangement of independent statements. **Claim 8** A step of showing the position of a non-equivalent element found in at least one of the input source code and the output source code The method according to claim 1, further comprising: **Claim 9** The step of the AI language model regenerating new output source code from the input source code in response to the verification result The method according to claim 1, further comprising this step
10. After the retraining of the AI language model, determining a second verification result for the regenerated output source code; and Quantifying the improvement of the AI language model based on at least the verification result and the second verification result The method according to claim 1, further comprising this step
11. A memory; and A processing device operably coupled to the memory An apparatus comprising: the processing device is configured to An artificial intelligence (AI) language model generates output source code based on input source code; Determine an equivalence mapping between a first abstract syntax tree (AST) constructed for the input source code and a second AST constructed for the output source code; and Indicate a verification result for the output source code based on the equivalence mapping It is configured as follows Apparatus
12. To determine an equivalence mapping between a first AST constructed for the input source code and a second AST constructed for the output source code, the processing device is configured to Divide the first AST and the second AST into subtrees; and Identify the equivalence between the subtrees of the first AST and the second AST The apparatus according to claim 11, which is configured as follows
13. To identify the equivalence between the subtrees of the first AST and the second AST, the processing device is configured to Identify one or more equivalent permutations of a first subtree of one AST; and Determine that a second subtree in the other AST matches one of the one or more equivalent permutations The apparatus according to claim 12, which is configured as follows
14. The processing device is configured to Indicate the positions of non-equivalent elements found in at least one of the input source code and the output source code The apparatus according to claim 11, which is further configured as follows
15. The processing device is configured to The AI language model regenerates new output source code from the input source code in response to the verification result The apparatus according to claim 11, which is further configured as follows
16. The processing device is configured to Following the retraining of the AI language model, determining a second verification result for the regenerated output source code; and Quantifying the improvement of the AI language model based on at least the verification result and the second verification result The apparatus according to claim 11, further configured as such.
17. In a processing device: A procedure for determining an equivalence mapping between a first Abstract Syntax Tree (AST) constructed for input source code and a second AST constructed for output source code, where the output source code is generated by an artificial intelligence (AI) language model based on the input source code; and A procedure for indicating a verification result for the output source code based on the equivalence mapping A computer program for causing the execution.
18. The computer program according to claim 17, wherein the input source code is implemented in a first programming language, and the output source code is implemented in a second programming language different from the first programming language.
19. In order to determine an equivalence mapping between a first AST constructed for the input source code and a second AST constructed for the output source code, in the processing device: A procedure for dividing the first AST and the second AST into subtrees; and A procedure for identifying the equivalence between the subtrees of the first AST and the second AST The computer program according to claim 17, causing the execution.
20. In the processing device: A procedure for determining a second verification result for the regenerated output source code following the retraining of the AI language model; and A procedure for quantifying the improvement of the AI language model based on at least the verification result and the second verification result The computer program according to claim 17, further causing the execution.