Method, device and computer program (using complexity metrics to assess code generated using artificial intelligence)

By using an AI language model to translate code with complexity metric verification, the challenges of codebase migration are addressed, ensuring accurate and maintainable translations with iterative regeneration for improved reliability.

JP2025105468APending Publication Date: 2025-07-10INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024195435
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-28
Filing Date
2024-11-07
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Migrating codebases from one programming language to another using artificial intelligence is laborious and challenging due to the need for extensive writing, testing, and debugging, with difficulties in validating the accuracy and functionality of translated code, especially in ensuring the translated code retains the intended logic and structure.

Method used

Utilizing an AI language model to remap application source code while maintaining functionality, employing complexity metrics to verify the translation accuracy by comparing complexity scores of input and output source code, and iteratively regenerating code until acceptable verification scores are achieved.

Benefits of technology

Enhances the reliability and maintainability of automated code translations by ensuring the translated code reproduces the logical flow and structural complexity of the original code, reducing human verification effort and improving the accuracy of AI-generated code.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025105468000001_ABST
    Figure 2025105468000001_ABST
Patent Text Reader

Abstract

To resolve the problem that a migration is an arduous task that may include writing, testing, validating, and debugging massive amounts of code.SOLUTION: A method of using complexity metrics to assess code generated using artificial intelligence is provided, comprising: generating output source code based on input source code using an artificial intelligence (AI) language model; identifying respective complexity scores for the input source code and the output source code using one or more complexity metrics; and generating a validation score based on an evaluation of the respective complexity scores for the output source code.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a method, apparatus, and product for assessing code generated using artificial intelligence using complexity metrics. By migrating the functionality of legacy source code to a more modern programming language, the maintainability and readability of the source code can be enhanced, and system performance can be improved. However, such migrations can be laborious tasks that may involve writing, testing, validating, and debugging large amounts of code.

Summary of the Invention

Problems to be Solved by the Invention

[0002] Migration can be a laborious task that may involve writing, testing, validating, and debugging large amounts of code.

Means for Solving the Problems

[0003] According to embodiments of the present disclosure, various methods, apparatuses, and products for evaluating code generated using artificial intelligence using complexity metrics are described herein. In some aspects, an artificial intelligence (AI) language model is used to remap application source code from an original codebase to a target codebase while maintaining the same functionality. In some aspects, complexity metrics are used to verify the translation from the original application source code to the AI-generated source code. Using the assumption that the complexity of the input source code and the output source code should be somewhat similar, the translation accuracy of the AI-generated code is indicated by the respective complexity metric scores of the input source code and the output source code. In some aspects, when the complexity metric scores are different, the AI language model is prompted to regenerate the code. In this way, by comparing the complexity scores, it becomes easier to verify the code when migrating from the original codebase to a new codebase using the AI-generated code, for example, from a first programming language to a second programming language, or from a legacy system to a modernized system.

[0004] In certain embodiments, a method of auditing code generated using artificial intelligence using a complexity metric includes a stage where an artificial intelligence (AI) language model generates output source code based on input source code. The method also includes identifying, using one or more complexity metrics, a respective complexity score for the input source code and the output source code. The method further includes generating a verification score for the output source code based on an evaluation of the respective complexity scores. In this way, the accuracy of automated code generation can be audited using a comparison of the complexity of the input source code and the output source code, thereby determining whether the control flow and structure of the original source code are maintained. For example, the input source code may be implemented in a first programming language, and the output source code may be implemented in a second programming language different from the first programming language. The one or more complexity metrics may include one or more of a cyclomatic complexity metric, one or more Halstead metrics, a live variable metric, a knot metric, and a complexity index based on multiple complexity metrics.

[0005] In some variations, the stage of identifying a respective score for the input source code and the output source code using one or more complexity metrics includes calculating a first complexity score for the input source code using multiple complexity metrics, so the first complexity score represents a combination of the multiple complexity metrics. This variation also includes calculating a second complexity score for the output source code using the multiple complexity metrics, so the second complexity score represents a combination of the multiple complexity metrics. In this way, multiple complexity metrics can be represented by a single score for comparison.

[0006] In some variants, the step of generating a verification score for the output source code based on the evaluation of each complexity score includes adjusting the weight of at least one of the input source code and the output source code based on the programming language. In this way, the inherent differences in the complexity of different programming languages are compensated for.

[0007] In some variants, the method also includes the step of regenerating the output source code from the input source code by the AI language model based on the verification score. In this way, the AI language module can iteratively regenerate the output source code until an acceptable verification score is obtained.

[0008] In some variants, the method also includes the step of indicating that the verification score deviates from an acceptable tolerance. In this way, a software engineer can be warned when the automated code generation fails to accurately reproduce the input source code.

[0009] In some variants, the method also includes the step of generating a second verification score for the regenerated output source code after retraining the AI language model. This variant further includes the step of quantifying the improvement of the AI language model based on at least the verification score and the second verification score. In this way, the accuracy and reliability of the AI language model can be evaluated, and the results of retraining the AI language model can be measured.

[0010] In some aspects, the apparatus may include a processing device; and a memory operably coupled to the processing device, the memory storing computer program instructions that configure the processing device to perform the above operations when executed. In some aspects, a computer program product comprising a computer-readable storage medium may store computer program instructions that configure a computer to perform the above operations when executed.

Brief Description of the Drawings

[0011]

Figure 1

[0012]

Figure 2

[0013]

Figure 3

[0014]

Figure 4

[0015]

Figure 5

[0016]

Figure 6

[0017] In the world of software development, the need to modernize codebases from one programming language to another is increasing. For example, the source code for an application can be migrated from a legacy programming language (e.g., COBOL) to a modern programming language (e.g., Java (registered trademark)). Motivations for such migrations can include facilitating easier maintenance and readability of the source code, improving security and error handling, enhancing software and / or hardware performance, and other advantages that will be recognized by those skilled in the art.

[0018] According to the present disclosure, artificial intelligence (AI) is used to transplant or migrate the source code of an application to a different programming language. A large language model (LLM) is trained on a dataset containing a large amount of source code in order to develop a generative AI that can output source code based on an input or prompt. That is, an AI language model is used to generate new source code based on the input of the original source code. For example, a prompt such as "Generate Java code that achieves the same purpose as the following COBOL code" can be given to the AI language model, where the legacy COBOL source code is provided as the input. In response, the AI language model can output AI-generated Java source code that, at least ideally, implements the same functionality as the legacy and produces the same output.

[0019] However, migrating a codebase to a new language poses significant challenges in ensuring the accuracy and functionality of the translated code, especially when leveraging AI for automated translation. The difficulty lies in validating AI-generated code translations and confirming whether the translated code retains the intended logic, functionality, and structure of the original code. The inherent complexity of programming languages, combined with the way developers express their logic, presents challenges in reliably validating the correctness and similarity of AI-generated translations. Additionally, verifying the output source code translated from the input source code requires analyzing hundreds of thousands, if not millions, of lines of code.

[0020] This disclosure is particularly focused on addressing the challenges associated with validating the accuracy of AI-generated code translations by enhancing the reliability and maintainability of automated code translations through the comparison of complexity scores. For example, cyclomatic complexity is a quantitative measure of program complexity and serves as a useful metric for assessing the complexity of control flow structures. By comparing the complexity scores of the input source code and the output source code, where it is predicted that the complexity scores will be substantially similar, it is ensured that the translated code not only reproduces the logical flow of the original code but also maintains a similar level of structural complexity. A threshold may be set to verify that they are actually similar (e.g., the difference in scores must be within 5), and if the threshold is not met, the AI language model may regenerate the code until the threshold is met.

[0021] Referring now to FIG. 1, an exemplary computing environment in accordance with aspects of the present disclosure is shown. Computing environment 100 includes an example of an environment for executing at least some of the computer code associated with the implementation of the various methods described herein, such as code analysis module 107. In addition to code analysis module 107, computing environment 100 includes, for example, computer 101, wide area network (WAN) 102, end user device (EUD) 103, remote server 104, public cloud 105, and private cloud 106. In this embodiment, computer 101 includes a processor set 110 (including processing circuit 120 and cache 121), communication fabric 111, volatile memory 112, persistent storage 113 (including operating system 122 and the code analysis module 107 identified above), a set of peripheral devices 114 (including user interface (UI) device set 123, storage 124, and Internet of Things (IoT) sensor set 125), and network module 115. Remote server 104 includes remote database 130. Public cloud 105 includes gateway 140, cloud orchestration module 141, host physical machine set 142, virtual machine set 143, and container set 144.

[0022] Computer 101 may take the form of a desktop computer, laptop computer, tablet computer, smartphone, smartwatch or other wearable computer, mainframe computer, quantum computer, or any other form of computer or mobile device now known or later to be developed that can execute a program, access a network, or query a database such as remote database 130. As is well understood in the field of computer technology and depending on the technology, the implementation of the computer implementation method may be distributed among multiple computers and / or among multiple locations. On the other hand, in this description of computing environment 100, for the sake of simplicity as much as possible, the detailed discussion focuses on a single computer, specifically computer 101. Although computer 101 is not shown within the cloud in FIG. 1, it may be located within the cloud. On the other hand, computer 101 does not need to exist within the cloud except within any arbitrarily shown range.

[0023] Processor set 110 includes one or more computer processors of any type now known or later to be developed. Processing circuitry 120 may be distributed among multiple packages, e.g., multiple integrated circuit chips. Processing circuitry 120 may implement multiple processor threads and / or multiple processor cores. Cache 121 is memory located within a processor chip package and is typically used for data or code that should be available for fast access by a thread or core executing on processor set 110. Cache memory is typically organized into multiple levels depending on its relative proximity to the processing circuitry. Alternatively, some or all of the cache for the processor set may be located "off-chip". In some computing environments, processor set 110 may be designed to operate using qubits to perform quantum computing.

[0024] Computer-readable program instructions are typically loaded onto computer 101 and cause a series of operational steps to be performed by the processor set 110 of computer 101, thereby executing a computer-implemented method, and thus the instructions so executed will instantiate the method specified in the flowchart and / or description of the computer-implemented method included in this document. These computer-readable program instructions are stored in various types of computer-readable storage media such as cache 121 and other storage media discussed below. The program instructions and associated data are accessed by the processor set 110 to control and direct the implementation of the computer-implemented method. In computing environment 100, at least some of the instructions for implementing the computer-implemented method may be stored in the code analysis module 107 of the persistent storage 113.

[0025] Communication fabric 111 is a signal conduction path that enables various components of computer 101 to communicate with each other. Typically, this fabric is created with switches and conductive paths such as buses, bridges, physical input / output ports, and switches and conductive paths that make up the like. Other types of signal communication paths such as optical fiber communication paths and / or wireless communication paths may be used.

[0026] Volatile memory 112 is any type of volatile memory known now or to be developed in the future. Examples include dynamic random access memory (RAM) or static RAM. Typically, volatile memory 112 is characterized by random access, although this is not required unless expressly stated. In computer 101, volatile memory 112 is located within a single package and exists inside computer 101, but alternatively or additionally, the volatile memory may be distributed across multiple packages and / or located external to computer 101.

[0027] The persistent storage 113 is any form of non-volatile storage for a computer, whether currently known or to be developed in the future. The non-volatility of this storage means that the stored data is maintained regardless of whether power is being supplied to the computer 101 and / or directly to the persistent storage 113. The persistent storage 113 may be read-only memory (ROM), but typically at least a portion of the persistent storage enables writing, deleting, and rewriting of data. Some well-known forms of persistent storage include magnetic disks and solid-state storage devices. The operating system 122 may take several forms, such as various known proprietary operating systems that utilize a kernel or open-source portable operating system interface type operating systems. The code included in the code analysis module 107 typically includes at least some of the computer code related to the implementation of the computer-implemented methods described herein.

[0028] The peripheral device set 114 includes a set of peripheral devices of the computer 101. Data communication connections between the peripheral devices of the computer 101 and other components may be implemented in various ways, such as a Bluetooth (registered trademark) connection, a Near Field Communication (NFC) connection, a connection by a cable (such as a Universal Serial Bus (USB) type cable), an insertion type connection (such as a Secure Digital (SD) card), a connection through a local area communication network, and even a connection through a wide area network such as the Internet. In various embodiments, the UI device set 123 may include components such as a display screen, a speaker, a microphone, wearable devices (such as goggles and smartwatches), a keyboard, a mouse, a printer, a touchpad, a game controller, and a haptic device. The storage 124 is an external storage such as an external hard drive or an insertable storage such as an SD card. The storage 124 may be persistent and / or volatile. In some embodiments, the storage 124 may take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments where the computer 101 is required to have a large amount of storage (for example, the computer 101 locally stores and manages a large-scale database), this storage may be provided by a peripheral storage device designed to store a very large amount of data, such as a Storage Area Network (SAN) shared by a plurality of geographically dispersed computers. The IoT sensor set 125 is composed of sensors that can be used in Internet of Things applications. For example, one sensor may be a thermometer, and another sensor may be a motion detector.

[0029] The network module 115 is a collection of computer software, hardware, and firmware that enables the computer 101 to communicate with other computers through the WAN 102. The network module 115 may include hardware such as a modem or a Wi-Fi (registered trademark) signal transceiver, software for packetizing and / or depacketizing data for communication network transmission, and / or web browser software for transmitting data over the Internet. In some embodiments, the network control function and the network transfer function of the network module 115 are implemented on the same physical hardware device. In other embodiments (e.g., embodiments utilizing software-defined networking (SDN)), the control function and the transfer function of the network module 115 are physically implemented on separate devices, and thus the control function manages several different network hardware devices. Computer-readable program instructions for implementing the computer-implemented method are typically downloaded to the computer 101 from an external computer or an external storage device through a network adapter card or a network interface included in the network module 115.

[0030] The WAN 102 is any wide area network (e.g., the Internet) capable of transmitting computer data over a non-local distance by any technique for transmitting computer data known now or to be developed in the future. In some embodiments, the WAN 102 may be replaced and / or supplemented by a local area network (LAN) designed to transmit data between devices located in a local area, such as a Wi-Fi network. The WAN and / or the LAN typically includes computer hardware such as copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and edge servers.

[0031] The end-user device (EUD) 103 is any computer system used and controlled by an end user (e.g., a customer of an enterprise operating the computer 101), and can take any of the forms discussed above in relation to the computer 101. The EUD 103 typically receives beneficial and useful data from the operation of the computer 101. For example, in a hypothetical case where the computer 101 is designed to provide recommendations to an end user, this recommendation would typically be transmitted from the network module 115 of the computer 101 to the EUD 103 via the WAN 102. Thus, the EUD 103 can display or otherwise present the recommendation to the end user. In some embodiments, the EUD 103 may be a client device such as a thin client, a thick client, a mainframe computer, a desktop computer, etc.

[0032] The remote server 104 is any computer system that provides at least some data and / or functionality to the computer 101. The remote server 104 may be controlled and used by the same entity that operates the computer 101. The remote server 104 represents a machine that collects and stores beneficial and useful data for use by other computers such as the computer 101. For example, in a hypothetical case where the computer 101 is designed and programmed to provide recommendations based on historical data, in this case, this historical data may be provided from the remote database 130 of the remote server 104 to the computer 101.

[0033] The public cloud 105 is any computer system available for use by multiple entities that provides on-demand availability of computer system resources and / or other computing capabilities, particularly data storage (cloud storage) and computing power, without direct active management by the user. Cloud computing typically exploits resource sharing to achieve consistency and economies of scale. The direct and active management of the computing resources of the public cloud 105 is performed by the computer hardware and / or software of the cloud orchestration module 141. The computing resources provided by the public cloud 105 are typically implemented by a virtual computing environment that runs on various computers that make up the host physical machine set 142, which is an aggregate of physical computers that are within and / or available to the public cloud 105. The virtual computing environment (VCE) typically takes the form of virtual machines from the virtual machine set 143 and / or containers from the container set 144. These VCEs may be stored as images and it is understood that they can be transferred as images or after instantiation of the VCE among and within various physical machine hosts. The cloud orchestration module 141 manages the transfer and storage of the images, deploys new instantiations of the VCE, and manages the active instantiation of the VCE deployment. The gateway 140 is a collection of computer software, hardware, and firmware that enables the public cloud 105 to communicate via the WAN 102.

[0034] Next, some further explanation is provided regarding a virtualized computing environment (VCE). A VCE can be stored as an "image". A new active instance of a VCE can be instantiated from the image. Two well-known types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to a feature of the operating system that enables the kernel to have multiple isolated user-space instances called containers. These isolated user-space instances typically behave as actual computers from the perspective of the programs running within them. A computer program running on a general operating system can utilize all the resources of that computer, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, a program running inside a container can only use the contents of the container and the devices assigned to that container, and this feature is known as containerization.

[0035] The private cloud 106 is similar to the public cloud 105, except that computing resources are only available for use by a single enterprise. The private cloud 106 is shown as being in communication with the WAN 102, but in other embodiments, the private cloud may be completely disconnected from the Internet and only accessible via a local / private network. A hybrid cloud is a composite of multiple different types of clouds (e.g., private, community, or public cloud types), often implemented by different vendors. Each of the multiple clouds remains a separate discrete entity, but the larger hybrid cloud architecture is coupled together by standardized or proprietary technologies that enable orchestration, management, and / or data / application portability between the constituent clouds. In this embodiment, both the public cloud 105 and the private cloud 106 are part of a larger hybrid cloud.

[0036] For further explanation, FIG. 2 depicts a flowchart of an exemplary method for assessing code generated using artificial intelligence using a complexity metric, according to some embodiments of the present disclosure. The method of FIG. 2 may be implemented by a code analysis module 201, such as the code analysis module 107 of FIG. 1, for example. In some examples, the code analysis module 201 may be implemented as a process or service that includes an AI language model that generates output source code from input source code. In other examples, the code analysis module 201 may be implemented as part of a process or service separate from the process or service that includes the AI language model. In another example, the code analysis module 201 may be implemented as part of a process or service that monitors the quality of the AI language model and assesses whether or not retraining of the AI language model is appropriate or successful.

[0037] The method of FIG. 2 includes a stage 202 where an artificial intelligence (AI) language model 211 generates an output source code 205 based on an input source code 203. The AI language model 211 can be trained with a vast dataset of original source code in a first programming language remapped to source code in different programming languages. Thus, the AI language model 211 is configured to autonomously translate a block of input source code in one programming language into a block of output source code in a different programming language. In some examples, the output source code 205 is generated by prompting the AI language model to generate the output source code based on the input source code 203. For example, the AI language model may be prompted with "Generate Java source code from block A of COBOL source code", where block A is provided as the input source code. In response, the AI language model generates Java source code intended to provide the same interface, perform the same functions, and produce the same output as the original COBOL source code. In some examples, the input source code and the output source code reflect the migration of the source code of an application from a first programming language (e.g., a legacy codebase) to a second programming language (e.g., a modern codebase). For example, the input source code may include legacy source code written in an older programming language (e.g., COBOL), while the output source code may be implemented in a modern programming language (e.g., Java); however, both the input source code and the output source code are intended to achieve the same purpose, provide the same interface, and produce the same output.

[0038] The method of FIG. 2 includes a step 204 of identifying each complexity score for the input source code 203 and the output source code 205 using one or more complexity metrics. In some implementations, as described in more detail below, the code analysis module 201 identifies each complexity score 204 by calculating each complexity score for the input source code 203 and the output source code 205. In other implementations, the code analysis module 201 does not calculate the complexity scores for the input and output source codes, but rather identifies each complexity score 204 by receiving the complexity scores for the input source code 203 and the output source code 205 calculated by a separate complexity analysis utility.

[0039] In some examples, the code analysis module 201 uses cyclomatic complexity as the complexity metric to identify each complexity score for the input and output source codes. Cyclomatic complexity is a software metric used to measure the complexity of a program's control flow. It was developed by Thomas J. McCabe and is sometimes referred to as the McCabe number or McCabe complexity. The cyclomatic complexity of a program is calculated based on the number of linearly independent paths through its source code. This metric is particularly useful when assessing the maintainability and testability of a software system.

[0040] The cyclomatic complexity may be determined by constructing a control flow graph of a module of code (e.g., a function or method), in which case each statement is a node and an edge connects the first node to the second node if control can proceed from the first statement to the second statement. In some examples, the cyclomatic complexity formula can be defined as V = E - N + 2P, where V is the cyclomatic complexity, E is the number of edges in the control flow graph of the program, N is the number of nodes in the control flow graph, and P is the number of connected components (P = 1 for a single linear program). When a function or method is single, the cyclomatic complexity can be defined as V = E - N + 2.

[0041] Put simply, the cyclomatic complexity can be understood as the number of decision points or branches within a program. Thus, in some examples, the cyclomatic complexity can be defined as V = D, where V is the cyclomatic complexity and D is the number of decision points in the code (e.g., the number of conditional statements or branch points). This is an indicator of the structural complexity of the program and is often associated with the number of test cases required to achieve comprehensive test coverage.

[0042] A higher cyclomatic complexity suggests that the program structure is more complex, which can increase the difficulty of understanding, testing, and maintaining the code. As a rule of thumb, a lower cyclomatic complexity is desirable as it tends to indicate simpler and more manageable code. In a set of modules (e.g., methods, classes, subroutines), the complexity of the individual functions they contain may be used to determine the total, average, or maximum cyclomatic complexity. The cyclomatic complexity per line of source code can be expressed as the decision density.

[0043] In some examples, the code analysis module 201 uses one or more Halstead metrics as complexity metrics to identify the complexity scores of the input source code and the output source code. The Halstead complexity metrics developed by Maurice H. Halstead are a set of metrics designed to quantify various aspects of a software program by focusing on the volume and difficulty of the code. These metrics were intended to provide a quantitative assessment of software complexity and assist in predicting the effort of software development.

[0044] To calculate the Halstead metrics, define n1 as the number of distinct operators, n2 as the number of distinct operands, N1 as the total number of operators, and N2 as the total number of operands. In this case, the program vocabulary n is represented as n = n1 + n2. The program length N is represented as N = N1 + N2. The calculated program length N' can be represented as N' = n1 log2 n1 + n2 log2 n2. The program volume V is a metric representing the volume or size of the program and can be calculated as V = N log2 n.

[0045] Since any program must have at least two operators, one for function calls and one for the end of a statement, the ratio (n1) / 2 can be considered the relative difficulty level due to the large number of operators in the program. The ratio (N2) / n2 represents the average number of times an operand is used. This ratio can be larger in programs where variables change more frequently. Since such programs are more difficult to understand, the difficulty D of reading or writing the program can be calculated as D = (n1 * n2) / (2 * n2).

[0046] The effort E is a metric that estimates the amount of time required for a human to write code, which can be calculated as E = D x V, where D is the difficulty metric and V is the program volume discussed above. The time T to write the code can be calculated as T = E / 18 seconds. The number of bugs B introduced can be estimated as B = V / 3000.

[0047] Halstead metrics provide insights into program size, the diversity of operators and operands, and the difficulty of understanding the code. A large program volume may indicate that the program is large and potentially complex, while a high program difficulty suggests that the code may be difficult to understand.

[0048] In some examples, the code analysis module 201 uses a raw metric as a complexity metric to identify the complexity scores of the input source code and the output source code. Some raw metrics, including the lines of code (LOC) in the program, the logical lines of code (LLOC), the source lines of code (SLOC), the percentage of comment lines, and the percentage of blank lines, can be used as indicators of complexity.

[0049] In some examples, the code analysis module 201 uses a live - variable metric as the complexity metric to identify the complexity scores of the input source code and the output source code. The live - variable metric is a measure of program complexity based on the number of live variables associated with statements in the program. This provides a quantitative assessment of the cognitive load and difficulty associated with code understanding and maintenance. In the context of this metric, a live variable refers to a variable whose value is relevant or still needed at a particular point in the execution of the program. The more live variables a program has, the more difficult it can be to understand and maintain the program. Thus, the live - variable metric functions as an indicator of program complexity.

[0050] Specifically, a live variable is a variable whose value is still being used or is needed at a particular point in the program. A variable is considered to "live" from its first reference within a module to its last reference, encompassing all statements in between. A particular statement is considered to be associated with a variable if that statement is located between the first and last occurrences of the live variable within the program. Using static code analysis, the live - variable metric can be calculated by counting the number of live variables associated with each statement. This metric provides insights into the complexity of statements based on the number of live variables each statement contains.

[0051] This metric may be extended across the module by calculating the average of the live variables. The average live variable metric is determined by summing the number of live variables for all executable statement in the module and then dividing this sum by the total number of executable statements. A large average live variable metric indicates that the module is more complex because, on average, there are more variables whose values need to be tracked and understood throughout the execution of the program. This metric provides a quantitative measure of the cognitive load imposed on a programmer attempting to understand or maintain the code.

[0052] In some examples, code analysis module 201 uses a knot metric as a complexity metric to identify the complexity score of the input source code and the output source code. The knot metric represents the complexity and unstructuredness of the control flow of the module. The knot metric can be calculated by counting the number of intersections between control flow paths through the module of code. For illustration, arrows can be drawn from a control transfer point to its destination. The more these arrows become intertwined, the more complex the program becomes.

[0053] In some examples, code analysis module 201 uses a naturalness metric as a complexity metric to identify the complexity score of the input source code and the output source code. The naturalness of a particular statement in the source code is represented by the number of occurrences of that statement in the corpus of training data provided to the AI language model. The fact that a portion of the code has statements that occur infrequently in the training data may indicate that that portion of the code is complex. The naturalness metric can be represented by the percentage of statements within a block of code whose occurrence value is less than a particular threshold.

[0054] In some examples, the code analysis module 201 uses an ultrametric topology metric as a complexity metric to identify the complexity scores of the input source code and the output source code. The ultrametric topology is related to the analysis of hierarchical functional relationships and can be used to model the complexity of a landscape. Land units on a map are connected to a function that indicates the direction of information movement or exchange between pairs of land units. The land units and functions are part of an encompassing landscape unit. Here, the ultrametric topology is adapted to the code by defining code modules (e.g., functions, methods, classes, subroutines) as "land units" or nodes connected to each other via an ultrametric function that indicates the exchange of information during the passage of control flow. The connections between nodes are edges, and thus the ultrametric distance between two nodes is the number of edges that must be traversed to reach one node from another. The sum of the ultrametric distances between all nodes may be used as a score for code complexity. Further, the sum of the degrees (the number of edges connected to a node) of each node may be used as a score for code complexity. Additionally, the cyclomatic complexity of the code may be determined as the number of edges minus the number of nodes, plus one. When constructing a matrix of ultrametric distances, the eigenvectors of this matrix can be calculated and used to determine the "direction" or "influence" of a module, indicating how a change in one module can affect other modules.

[0055] In some examples, the code analysis module 201 identifies 204 the complexity scores of the input source code 203 and the output source code 205 using one or more complexity metrics by calculating a first complexity score for the input source code 203 using a first complexity metric and calculating a second complexity score for the output source code using the first complexity metric. For example, the code analysis module 201 calculates 206 the first complexity score by applying one of the complexity analysis techniques discussed above to the input source code 203 and calculates the second complexity score by applying the same complexity analysis technique to the output source code 205. In some implementations, the calculation of the complexity score is performed by calculating an aggregate complexity score or an average complexity score based on the individual complexity scores of each block of code (e.g., functions, methods, classes, subroutines, etc.) within the source code.

[0056] It will be appreciated that this technique may be repeated using multiple complexity metrics such that multiple complexity scores are calculated for the input source code and multiple complexity scores are generated for the output source code. For example, the code analysis module may calculate a third complexity score for the input source code and a fourth complexity score for the output source code using a second complexity metric. Thus, each complexity score for the input source code and the output source code may include one or more complexity scores based on one or more of a cyclomatic complexity metric, one or more Halstead metrics, one or more raw metrics such as lines of source code, knot metrics, live variable metrics, ultra metric topological metrics, and naturalness metrics.

[0057] In some examples, the first complexity score and the second complexity score represent a set of various complexity scores. Thus, in some examples, the code analysis module 201 calculates 206 a first complexity score for the input source code using a plurality of complexity metrics, where the first complexity score represents a combination of the plurality of complexity metrics, and calculates 208 a second complexity score for the output source code using the plurality of complexity metrics, thereby identifying 204 the complexity score of each of the input source code 203 and the output source code 205 using one or more complexity metrics. For example, the complexity score may be a complexity index calculated from a weighted average calculated using a plurality of complexity metrics. In a particular implementation, eigenvectors are constructed from a plurality of complexity metrics. A base value is calculated from the square root of the sum of the squares of each of these values. Each base value calculated from the input source code and the output source code may be used as each complexity score for comparing the input source code and the AI-generated output source code.

[0058] In different programming languages (e.g., COBOL and Java), different operator vocabularies and different syntaxes are used, which can distort the complexity score if not considered. Thus, in some examples, different complexity metric definitions are used for the input source code and the output source code. For example, when assessing the cyclomatic complexity based on the number of decision points, a conditional statement, a branch statement, or an operator that increments the number of decision points in one programming language must be corresponded to a statement having the same effect in the other programming language. Similarly, it may be revealed by statistical analysis that the number of lines of source code in one programming language is predicted to be a certain percentage larger than the number of lines of source code in the other programming language. Thus, the calculation of complexity may be adjusted based on the differences between the syntaxes of programming languages.

[0059] To identify each complexity score for the input source code and the output source code, it will be understood that any single complexity metric or combination of complexity metrics described above can be used by the code analysis module 201. Further, it will be understood that the code analysis module may use other complexity metrics and mathematical constructs not discussed above in a manner consistent with the present disclosure to quantify the complexity of the input source code and the output source code.

[0060] The method of FIG. 2 also includes a stage 210 of generating a verification score 209 for the output source code 205 based on the evaluation of each complexity score. In some examples, the code analysis module 201 generates the verification score 209 by comparing one or more complexity scores of the input source code with one or more complexity scores of the output source code and determining a verification score 209 that represents their similarity or difference. For example, the verification score 209 may be an absolute or relative deviation of the complexity score of the output source code from the complexity score of the input source code. In some examples, the verification score 209 may be based on the evaluation of multiple complexity scores using multiple complexity metrics for the input source code and the output source code, such as an average or weighted average of various scores. In some examples, the code analysis module 201 may set a tolerance such as a threshold or range to determine whether the output source code passes or fails the verification. For example, the code analysis module may determine that the output source code fails the verification if the difference between the complexity scores exceeds a specific threshold or if the complexity score of the output source code is greater than the complexity score of the input source code. Thus, in some implementations, the verification score 209 may be a binary result such as pass / fail.

[0061] For further explanation, FIG. 3 depicts a flowchart of an exemplary method for assessing code generated using artificial intelligence using a complexity metric, according to some embodiments of the present disclosure. The method of FIG. 3 extends the method of FIG. 2 in that it further includes a step 302 of adjusting the weight of the complexity score of at least one of the input source code 203 and the output source code 205 based on the programming language at stage 210 of generating a verification score for the output source code 205 based on the evaluation of each complexity score. Some programming languages are, by their nature, more complex than other programming languages. For example, a program written in assembly language is expected to be generally more complex than the same program written in Java. To account for this imbalance, in some instances, the code analysis module 201 weights at least one of the input source code and the output source code based on the expected complexity of the programming language in which the source code is written. For example, if the input source code is part of a legacy codebase and the output source code is written in a more modern programming language, the code analysis module 201 may lower the weight of the complexity score of the input source code to account for the expected decrease in complexity when translated to the more modern programming language.

[0062] For further explanation, FIG. 4 depicts a flowchart of an exemplary method for evaluating code generated using artificial intelligence using a complexity metric, according to some embodiments of the present disclosure. The method of FIG. 4 extends the method of FIG. 2 in that it further includes step 402 where the AI language model 211 regenerates the output source code 205 from the input source code 203 based on the verification score 209. In some examples, the code analysis module 201 determines that the verification score 209 for the output code is outside an acceptable tolerance, or otherwise indicates that the output source code has failed verification. Accordingly, the code analysis module 201 determines that the output source code should be regenerated. In some examples, the code analysis module 201 generates the second prompt in substantially the same manner as generating the first prompt; however, in this case, the prompt indicates to the AI language model that the AI language model should generate a different implementation. In such cases, the code analysis module 201 may generate a prompt such as "Regenerate the code for block A" or "Regenerate the code for block A that is syntactically different from the previously generated code". In response, the AI language model regenerates alternative code for the input code corresponding to block A. In some implementations, the code analysis module 201 repeatedly re-prompts the AI language model to regenerate the source code until the output source code passes verification or until a threshold number of attempts is reached.

[0063] In some implementations, the code analysis module 201 adjusts one or more parameters of the AI language model in response to determining that the translation score deviates from an acceptable tolerance within which it is accepted. The AI language model can include configurable parameters that affect the creativity of the model's response to a prompt. For example, a temperature parameter adjusts the probability distribution that can be used to select the next token in the output stream. When selecting the next token in the output stream, if the temperature is lower, the language model tends to select tokens with probabilities within a narrower range, resulting in a more deterministic output, while if the temperature is higher, the language model tends to select tokens with probabilities within a wider range, resulting in a more random output. Another exemplary parameter is the top-k parameter, which controls the randomness of the next token selection by telling the language model that it must select from the top k most probable tokens. Yet another exemplary parameter is the top-p parameter, which controls the randomness of the next token selection by telling the language model that it must select from the most probable tokens whose probabilities sum to a p value or exceed the p value.

[0064] In some examples, the code analysis module 201 adjusts one or more parameters of the AI language model in response to determining that one or more iterations of generating the output source code have failed to meet the tolerance threshold. For example, as the number of iterations increases, the parameter that controls the creativity of the AI language model can be adjusted to increase the randomness of the output. In this way, the AI language model can be induced to generate solutions that are not similar to the failed solutions presented in previous iterations. In some examples, the adjustment of one or more parameters is done by including a statement to adjust the parameter, such as "set the temperature to 0.8", in the prompt. It will be understood that the parameters of the language model may be adjusted at any stage of the process. For example, in some implementations, a preprocessing stage analyzes the original source code before the AI language model generates new source code from the original source code and sets the language model parameters based on the analysis. For example, statistical analysis of the original code can be used to predict how creative or deterministic the language model should be in its output.

[0065] For further illustration, FIG. 5 depicts a flowchart of an exemplary method for assessing code generated using artificial intelligence using a complexity metric, according to some embodiments of the present disclosure. The method of FIG. 5 extends the method of FIG. 2 in that it further includes a stage 502 indicating that the output source code has failed verification, depending on the verification score. In some examples, the code analysis module 201 indicates at 502 that the output source code has failed verification in response to determining that the verification score is outside an acceptable tolerance or indicates a verification failure. Stage 502 indicating that the output source code has failed verification may include flagging the output source code or issuing a warning to the responsible person indicating that the output source code has failed verification.

[0066] For further explanation, FIG. 6 depicts a flowchart of an exemplary method for auditing code generated using artificial intelligence using a complexity metric, according to some embodiments of the present disclosure. The method of FIG. 6 extends the method of FIG. 2 in that it further includes step 602 of generating a second verification score for the regenerated output source code after retraining the AI language model 211. In some examples, the AI language model 211 is retrained with an additional training dataset to improve the quality of the AI code translation of the input source code. To assess whether the AI language model has improved in terms of code translation quality and accuracy, and to quantify the improvement, the AI model is prompted to regenerate the output source code based on the input source code for which the verification score has been previously determined. In these examples, the code analysis module 201 generates the second verification score 602 in the manner described above, using the same complexity metric that was used to generate the initial verification score.

[0067] The method of FIG. 6 also includes step 604 of quantifying the improvement of the AI language model 211 based at least on the verification score and the second verification score. In some examples, the code analysis module 201 quantifies the improvement of the AI language model 211 by comparing the initial verification score and the second verification score to determine whether the AI language model 211 is generating an output source code that is more similar to the input source code in terms of complexity 604.

[0068] Embodiments are useful when migrating or porting applications from one programming language to a different programming language and from a legacy programming language to a more modernized programming language. However, it will be further understood that in some examples, the original source code and the new source code may be written in the same programming language.

[0069] In view of the foregoing, the assessment of AI-generated code using complexity metrics according to the present disclosure offers a number of advantages. Embodiments of the present disclosure improve the accuracy and quality of automated code generation and further improve the reliability and maintainability of the source code generated through automated code generation. The evaluation of the complexity score is advantageous in quantifying the verification of the output source code against the input source code, and further shows not only whether the translated code reproduces the logical flow of the original code, but also whether it maintains a similar level of structural complexity. The evaluation of the complexity score is advantageous in determining whether the AI-generated code needs to be regenerated, thus reducing the human effort required to verify the AI-generated code. Furthermore, the evaluation of the complexity score is useful in quantifying the improvement in the accuracy and ability of the AI language model to translate source code.

[0070] Various aspects of the present disclosure are illustrated by descriptions, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic included in embodiments of a computer program product (CPP). For any flowchart, operations may be performed in an order different from that shown in a given flowchart, depending on the relevant technology. For example, again depending on the relevant technology, two operations shown in consecutive flowchart blocks may be performed in reverse order, as a single integrated step, simultaneously, or at least partially in a temporally overlapping manner.

[0071] An embodiment of a computer program product (referred to herein as a "CPP embodiment" or "CPP") is any set of one or more storage media (also referred to as "mediums") collectively included in a set of one or more storage devices that collectively contain machine-readable code corresponding to instructions and / or data for performing the computer operations specified in a given CPP claim. A "storage device" is any tangible device capable of holding and storing instructions for use by a computer processor. A computer-readable storage medium may be, but is not limited to, an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these media are floppy disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices (such as punch cards or pits / lands formed on the major surfaces of disks), or any suitable combination of the foregoing. A computer-readable storage medium is not to be construed as storage in the form of a transient signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide, an optical pulse passing through an optical fiber cable, an electrical signal transmitted through a wire, and / or other transmission media. As will be understood by those skilled in the art, data is typically moved during some irregular points during the normal operation of a storage device, such as during access, defragmentation, or garbage collection, but the data is not transient while it is stored, so the above notwithstanding, a storage device is not considered to be transient.

[0072] The descriptions of various embodiments of the present disclosure are presented for illustrative purposes and are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will become apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles of the embodiments, the practical application, or a technical improvement to technologies found in the marketplace, or to enable other skilled artisans to understand the embodiments disclosed herein.

Claims

**Claim 1** A method for evaluating code generated using artificial intelligence using a complexity metric, comprising: generating, by an artificial intelligence (AI) language model, output source code based on input source code; identifying, using one or more complexity metrics, a respective complexity score for the input source code and the output source code; and generating, based on an evaluation of the respective complexity scores, a verification score for the output source code A method comprising the steps of: **Claim 2** The method according to claim 1, wherein the input source code is implemented in a first programming language and the output source code is implemented in a second programming language different from the first programming language. **Claim 3** The method according to claim 1 or 2, wherein the one or more complexity metrics include one or more of a cyclomatic complexity metric, one or more Halstead metrics, a live variable metric, a knot metric, an ultrametric topological metric, and a complexity index based on a plurality of complexity metrics. **Claim 4** The step of identifying, using one or more complexity metrics, a respective score for the input source code and the output source code comprises: calculating, using a plurality of complexity metrics, a first complexity score for the input source code, wherein the first complexity score represents a combination of the plurality of complexity metrics; and calculating, using the plurality of complexity metrics, a second complexity score for the output source code, wherein the second complexity score represents a combination of the plurality of complexity metrics The method according to claim 1 or 2, comprising the steps of: **Claim 5** The step of generating, based on an evaluation of the respective complexity scores, a verification score for the output source code comprises: adjusting, based on the programming language, a weight of at least one of the complexity scores of the input source code and the output source code. The method according to claim 1 or 2, comprising the steps of: **Claim 6** The method according to claim 1 or 2, further comprising: regenerating, by the AI language model, the output source code from the input source code based on the verification score. **Claim 7** The method according to claim 1 or 2, further comprising a stage indicating that the verification score deviates from an acceptable tolerance.

8. After retraining the AI language model, generating a second verification score for the regenerated output source code; and Quantifying the improvement of the AI language model based on at least the verification score and the second verification score The method according to claim 1 or 2, further comprising.

9. A memory; and A processing device operably coupled to the memory, the processing device Generating output source code based on input source code by an artificial intelligence (AI) language model; Identifying each complexity score for the input source code and the output source code using one or more complexity metrics; and Generating a verification score for the output source code based on an evaluation of each complexity score Configured to perform An apparatus comprising.

10. The apparatus according to claim 9, wherein the input source code is implemented in a first programming language and the output source code is implemented in a second programming language different from the first programming language.

11. The apparatus according to claim 9 or 10, wherein the one or more complexity metrics include one or more of a cyclomatic complexity metric, one or more Halstead metrics, a survival variable metric, a knot metric, an ultrametric topological metric, and a complexity index based on a plurality of complexity metrics.

12. To identify each score for the input source code and the output source code using one or more complexity metrics, the processing device Calculating a first complexity score for the input source code using a plurality of complexity metrics, wherein the first complexity score represents a combination of the plurality of complexity metrics; and Calculating a second complexity score for the output source code using the plurality of complexity metrics, wherein the second complexity score represents a combination of the plurality of complexity metrics The apparatus according to claim 9 or 10, further configured to perform.

13. To generate a verification score for the output source code based on the evaluation of each complexity score, the processing device adjusts the weight of the complexity score of at least one of the input source code and the output source code based on its programming language The apparatus according to claim 9 or 10, further configured to perform **Claim 14** The processing device is further configured to regenerate the output source code from the input source code by the AI language model based on the verification score The apparatus according to claim 9 or 10, further configured to perform **Claim 15** The processing device generates a second verification score for the regenerated output source code after retraining the AI language model; and quantifies the improvement of the AI language model based on at least the verification score and the second verification score The apparatus according to claim 9 or 10, further configured to perform **Claim 16** A computer program for causing a processing device to specify each complexity score for the input source code and the output source code using one or more complexity metrics, where the output source code is generated by an artificial intelligence (AI) language model based on the input source code; and generate a verification score for the output source code based on the evaluation of each complexity score to execute **Claim 17** The computer program according to claim 16, wherein the output source code is generated by the AI language model in response to the AI language model being prompted to generate the output source code using the input source code as part of a prompt **Claim 18** The computer program according to claim 16 or 17, wherein the input source code is implemented in a first programming language and the output source code is implemented in a second programming language different from the first programming language **Claim 19** A computer program for further causing the processing device to prompt the AI language model to regenerate the output source code from the input source code based on the verification score to execute, the computer program according to claim 16 or 17 **Claim 20** The processing device After retraining the AI language model, a procedure for generating a second verification score for the regenerated output source code; and A procedure for quantifying the improvement of the AI language model based on at least the verification score and the second verification score The computer program according to claim 16 or 17, for further causing the above to be executed.