Tagging deterministic code in artificial intelligence generated code
By identifying and marking deterministic and non-deterministic parts of AI-generated code and generating multi-path code, the problem of introductory errors and performance degradation of non-deterministic code is solved, improving the reliability and performance of the code.
Patent Information
- Application Number
- CN202411677788.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-27
- Filing Date
- 2024-11-22
- Publication Date
- 2025-06-27
AI Technical Summary
There are deterministic and nondeterministic codes in code generated by AI, which may introduce errors or performance degradation and are difficult to identify and process.
By identifying deterministic and nondeterministic parts in the code, marking these parts, and generating multipath codes when necessary to deal with nondeterministic code.
It realizes the identification and distinction between deterministic and non-deterministic parts in AI generated code, reduces the generation and use of non-deterministic code, and improves the reliability and performance of the code.
Smart Images

Figure CN120215907A_ABST
Abstract
Description
Background Art
[0001] The present disclosure relates to methods, apparatuses, and products for tagging deterministic code in artificial intelligence-generated code. Generative artificial intelligence (AI) can be used to generate code such as application source code. Such code can be based on a prompt provided to the generative AI model that describes the function the code should perform. Such code can also include a conversion from one programming language to another programming language performed by the generative AI model. AI-generated code can include deterministic code that always produces the same set of outputs given the same input. However, AI-generated code can also include non-deterministic code that, although using the same set of inputs, can produce unreliable or inconsistent outputs. Such non-deterministic code may introduce undesirable entropy into the system, resulting in errors or performance degradation. Summary of the Invention
[0002] According to embodiments of the present disclosure, various methods, apparatuses, and products for tagging deterministic code in artificial intelligence-generated code are described herein. In some aspects, tagging deterministic code in artificial intelligence-generated code includes: receiving code generated by a generative artificial intelligence (AI) model; identifying at least one portion of the code by identifying at least one of: one or more portions of deterministic code or one or more portions of non-deterministic code; and tagging the at least one identified portion of the code. This provides the advantage of identifying the deterministic or non-deterministic portions of the code in AI-generated code, the importance of which is first recognized by the present disclosure. In some aspects, an apparatus can include a processing device; and a memory operatively coupled to the processing device, where the memory stores computer program instructions that, when executed, cause the processing device to perform the method. In some aspects, a computer program product including a computer-readable storage medium can store computer program instructions that, when executed, perform the method.
[0003] In some aspects, tagging deterministic code in artificial intelligence-generated code can further include generating a portion of multi-path code corresponding to the portion of the deterministic code in response to tagging a portion of the deterministic code. This enables the implementation of multi-path code for portions of known deterministic code, which may be more suitable for multi-path compared to non-deterministic code. In some aspects, tagging deterministic code in artificial intelligence-generated code can further include: providing data describing at least one portion of the identified code to the generative AI model. This provides the advantage of enhancing the generation of deterministic code and reducing the generation of non-deterministic code by the generative AI model. Brief Description of the Drawings
[0004] Figure 1Describes an example computing environment for tagging deterministic code in AI-generated code according to some embodiments of the present disclosure.
[0005] Figure 2 A flowchart of an example method for tagging deterministic code in AI-generated code according to some embodiments of the present disclosure.
[0006] Figure 3 A flowchart of another example method for tagging deterministic code in AI-generated code according to some embodiments of the present disclosure.
[0007] Figure 4 A flowchart of another example method for tagging deterministic code in AI-generated code according to some embodiments of the present disclosure.
[0008] Figure 5 A flowchart of another example method for tagging deterministic code in AI-generated code according to some embodiments of the present disclosure. Detailed Description
[0009] Deterministic code is code that will always produce the same output if given the same set of inputs. In other words, deterministic code is operationally consistent code. Sequences of deterministic code can be used as examples of known facts when training generative artificial intelligence (AI) models to generate code. Generative artificial intelligence (AI) uses models such as neural networks to generate content including code such as source code in response to a prompt. For example, a generative AI model can be used to generate code based on a description of the functionality to be implemented by the generated code. As another example, a generative AI model can be used to convert code from one programming language to another by providing the base code as input to the generative AI model.
[0010] Generative AI models can produce non-deterministic code. For example, although the same set of inputs is used, there may be potential structural or architectural issues with the program that can produce unreliable or inconsistent outputs. When used and accepted as known facts in a system, the system will inherit an entropy level that can lead to errors or defects. Therefore, it may be beneficial to identify the deterministic and non-deterministic portions of AI-generated code to provide feedback to the generative AI model to strengthen learning of its deterministic outputs while mitigating future generation of non-deterministic code.
[0011] When introducing multi-path code into a program or application, the identification of deterministic and non-deterministic code is also useful. Multi-path code describes the use of multiple individual code paths that are configured to perform similar functions. For example, different code paths can be configured or designed to produce the same output when applied to the same input, or can be configured or designed to perform similar functions in other ways. These different code paths that are designed to perform the same function can be implemented by different engineers or teams, such that the resulting code paths are not identical, but are designed to achieve the same result. These different code paths can also be written in different languages, access different libraries, or be different in other ways when designed to perform similar functions. Each path of the multi-path code can be accessed using a shared interface, such as an application programming interface (API) or other interface as may be understood.
[0012] When a portion of multi-path code is encountered during the execution of an application or other software, state information describing the execution state at the point when the portion of multi-path code was encountered can be saved. The state information can describe the values of various registers, memory locations, counters, attributes, etc. If an error occurs in a code path of the multi-path code, the saved state information can be used as a checkpoint for restoring or rolling back the execution state of the application to the point before the code path of the multi-path code where the error occurred. Then, different code paths of the multi-path code can be executed. This process can be repeated until the code paths of the multi-path code are executed without error.
[0013] In cases where a portion of the multi-path code is non-deterministic (e.g., where the path being executed is non-deterministic), it may be computationally difficult to roll back the execution state in the event of an error. For example, in the case where a code path modifies a shared memory location or resource, another process or service may access the shared memory location before the error occurs. To roll back the execution state, the shared memory location or resource should have its value restored to its state before it was modified. However, other processes have accessed potentially erroneous data from that memory location or resource. Therefore, other processes may also need to have their execution state rolled back or restored. As the interdependencies between processes increase, restoring the execution state of non-deterministic code becomes increasingly complex. In contrast, deterministic code lacking these interdependencies will not require these complex steps to roll back its execution state, making it a better candidate for introducing multi-path code. Therefore, identifying deterministic and / or non-deterministic code to identify candidates for multi-path code or to exclude candidates from multi-path code may be beneficial.
[0014] Now refer to Figure 1, which shows an example computing environment in accordance with aspects of the present disclosure. Computing environment 100 includes an example of an environment for executing at least some of the computer code involved in performing the various methods described herein, such as code analysis module 107. In addition to block 107, computing environment 100 includes, for example, computer 101, wide area network (WAN) 102, end user device (EUD) 103, remote server 104, public cloud 105, and private cloud 106. In this embodiment, computer 101 includes a set of processors 110 (including processing circuitry 120 and cache 121), communication fabric 111, volatile memory 112, permanent storage device 113 (including operating system 122 and block 107, as described above), a set of peripheral devices 114 (including a set of user interface (UI) devices 123, storage device 124, and a set of Internet of Things (IoT) sensors 125), and network module 115. Remote server 104 includes remote database 130. Public cloud 105 includes gateway 140, cloud coordination module 141, a set of host physical machines 142, a set of virtual machines 143, and a set of containers 144.
[0015] Computer 101 may take the form of a desktop computer, laptop computer, tablet computer, smart phone, smart watch or other wearable computer, mainframe computer, quantum computer, or any other form of computer or mobile device now known or later developed that is capable of running programs, accessing a network, or querying a database such as remote database 130. As is well known in the computer art and depending on the technology, the performance of computer-implemented methods may be distributed among multiple computers and / or multiple locations. On the other hand, in this presentation of computing environment 100, the discussion focuses in detail on a single computer, specifically computer 101, to keep the presentation as simple as possible. Computer 101 may be located in the cloud, even if Figure 1 it is not shown in the cloud. On the other hand, computer 101 does not need to be in the cloud, unless to any degree that can be affirmatively indicated.
[0016] The set of processors 110 includes one or more computer processors of any type now known or later developed. The processing circuitry 120 may be distributed across multiple packages, such as multiple cooperative integrated circuit chips. The processing circuitry 120 may implement multiple processor threads and / or multiple processor cores. The cache 121 is a memory located within the processor chip package and is generally used for data or code that should be made available for rapid access by threads or cores running on the set of processors 110. Caches are generally organized into multiple levels based on their relative proximity to the processing circuitry. Alternatively, some or all of the caches of the set of processors may be located "off-chip". In some computing environments, the set of processors 110 may be designed to work with qubits and perform quantum computing.
[0017] Computer-readable program instructions are typically loaded onto the computer 101 so that the set of processors 110 of the computer 101 executes a series of operational steps to implement a computer-implemented method such that the instructions so executed will instantiate the method specified in the flowchart and / or narrative description of the computer-implemented method included in this document. These computer-readable program instructions are stored in various types of computer-readable storage media, such as the cache 121 and other storage media discussed below. The program instructions and associated data are accessed by the set of processors 110 to control and direct the execution of the computer-implemented method. In the computing environment 100, at least some of the instructions for performing the computer-implemented method may be stored in block 107 of the persistent storage 113.
[0018] The communication fabric 111 is a signal conduction path that permits the various components of the computer 101 to communicate with one another. Generally, this fabric consists of switches and conductive paths, such as switches and conductive paths that make up a bus, a bridge, a physical input / output port, etc. Other types of signal communication paths may be used, such as fiber optic communication paths and / or wireless communication paths.
[0019] The volatile memory 112 is any type of volatile memory now known or later developed. Examples include dynamic random access memory (RAM) or static RAM. Generally, volatile memory 112 is characterized by random access, but this is not required unless specifically stated. In the computer 101, the volatile memory 112 is located within a single package and inside the computer 101, but, alternatively or additionally, the volatile memory may be distributed across multiple packages and / or located external to the computer 101.
[0020] The permanent memory 113 is any form of non-volatile memory for a computer that is now known or will be developed in the future. The non-volatility of this memory means that the stored data is retained regardless of whether power is supplied to the computer 101 and / or directly to the permanent memory 113. The permanent memory 113 can be a read-only memory (ROM), but typically at least a portion of the permanent memory allows for the writing, deletion, and re-writing of data. Some common forms of persistent storage include magnetic disks and solid-state storage devices. The operating system 122 can take several forms, such as various known proprietary operating systems or operating systems of the open-source portable operating system interface type that employ a kernel. The code included in block 107 generally includes at least some of the computer code involved in performing the computer-implemented methods described herein.
[0021] The set of peripheral devices 114 includes the set of peripheral devices of the computer 101. Data communication connections between the peripheral devices and other components of the computer 101 can be implemented in various ways, such as Bluetooth connections, near-field communication (NFC) connections, connections made by a cable (such as a universal serial bus (USB)-type cable), plug-in connections (e.g., Secure Digital (SD) card), connections made through a local communication network, and even connections made through a wide area network such as the Internet. In various embodiments, the set of UI devices 123 can include components such as a display screen, speakers, microphones, wearable devices (such as goggles and smartwatches), keyboards, mice, printers, touchpads, game controllers, and haptic devices. The storage device 124 is an external storage device, such as an external hard drive, or a pluggable storage device, such as an SD card. The storage device 124 can be permanent and / or volatile. In some embodiments, the storage device 124 can take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments where the computer 101 needs to have a large amount of storage (e.g., in the case where the computer 101 locally stores and manages a large database), this storage can be provided by a peripheral storage device (such as a storage area network (SAN) shared by multiple geographically distributed computers) designed to store a very large amount of data. The set of IoT sensors 125 consists of sensors that can be used in Internet of Things applications. For example, one sensor can be a thermometer, and another sensor can be a motion detector.
[0022] The network module 115 is a collection of computer software, hardware, and firmware that allows the computer 101 to communicate with other computers via the WAN 102. The network module 115 can include hardware such as a modem or a Wi-Fi signal transceiver, software for packetizing and / or depacketizing data transmitted over a communication network, and / or web browser software for transmitting data over the Internet. In some embodiments, the network control function and the network forwarding function of the network module 115 are executed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing software-defined networking (SDN)), the control function and the forwarding function of the network module 115 are executed on physically separate devices such that the control function manages several different network hardware devices. Computer-readable program instructions for performing computer-implemented methods can generally be downloaded to the computer 101 from an external computer or an external storage device via a network adapter or network interface included in the network module 115.
[0023] The WAN 102 is any wide area network (e.g., the Internet) capable of transmitting computer data over non-local distances via any technology now known or hereafter developed for transmitting computer data. In some embodiments, the WAN 102 can be replaced and / or supplemented by a local area network (LAN) designed to transmit data between devices located in a local area (e.g., a Wi-Fi network). The WAN and / or LAN typically includes computer hardware such as copper transmission cables, fiber optic transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and edge servers.
[0024] The end-user device (EUD) 103 is any computer system used and controlled by an end user (e.g., a customer of an enterprise operating the computer 101) and can take any form discussed above in connection with the computer 101. The EUD 103 typically receives useful and actionable data from the operation of the computer 101. For example, in the hypothetical case where the computer 101 is designed to provide recommendations to an end user, the recommendations will typically be transmitted from the network module 115 of the computer 101 via the WAN 102 to the EUD 103. In this way, the EUD 103 can display or otherwise present the recommendations to the end user. In some embodiments, the EUD 103 can be a client device such as a thin client, a thick client, a mainframe computer, a desktop computer, etc.
[0025] The remote server 104 is any computer system that provides at least some data and / or functionality to the computer 101. The remote server 104 can be controlled and used by the same entity that operates the computer 101. The remote server 104 represents a machine that collects and stores useful and valuable data used by other computers such as the computer 101. For example, in the hypothetical case where the computer 101 is designed and programmed to provide recommendations based on historical data, then that historical data can be provided to the computer 101 from the remote database 130 of the remote server 104.
[0026] The public cloud 105 is any computer system that can be used by multiple entities, which provides on-demand availability of computer system resources and / or other computing capabilities (especially data storage (cloud storage) and computing power) without direct active management by the user. Cloud computing typically leverages the sharing of resources to achieve scale consistency and economy. The direct and active management of the computing resources of the public cloud 105 is performed by the computer hardware and / or software of the cloud coordination module 141. The computing resources provided by the public cloud 105 are typically implemented by virtual computing environments running on various computers that make up the set of host physical machines 142, which is the universe of physical computers within and / or available for the public cloud 105. The virtual computing environment (VCE) typically takes the form of virtual machines from the set of virtual machines 143 and / or containers from the set of containers 144. It should be understood that these VCEs can be stored as images and can be transferred between various physical machine hosts either as images or after the instantiation of the VCE. The cloud coordination module 141 manages the transfer and storage of the images, deploys new instantiations of the VCE, and manages the active instantiations of the VCE deployment. The gateway 140 is a collection of computer software, hardware, and firmware that allows the public cloud 105 to communicate via the WAN 102.
[0027] Some further explanations of the virtualized computing environment (VCE) will now be provided. The VCE can be stored as an "image". New active instances of the VCE can be instantiated from this image. Two common types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to an operating system feature where the kernel allows for the existence of multiple isolated user space instances, called containers. From the perspective of the programs running within them, these isolated user space instances typically behave as actual computers. A computer program running on a normal operating system can utilize all the resources of that computer, such as connected devices, files and folders, network shares, CPU capabilities, and quantifiable hardware capabilities. However, a program running within a container can only use the contents of the container and the devices allocated to the container, which is a feature known as containerization.
[0028] The private cloud 106 is similar to the public cloud 105, except that the computing resources are only available for use by a single enterprise. Although the private cloud 106 is depicted as communicating with the WAN 102, in other embodiments, the private cloud can be completely disconnected from the Internet and only accessible through a local / private network. A hybrid cloud is a combination of multiple clouds of different types (e.g., private, community, or public cloud types) that are typically implemented by different vendors. Each of the multiple clouds remains an independent and discrete entity, but the larger hybrid cloud architecture is bound together by standardized or proprietary technologies that enable coordination, management, and / or data / application portability between the multiple constituent clouds. In this embodiment, both the public cloud 105 and the private cloud 106 are part of the larger hybrid cloud.
[0029] For further explanation, Figure 2 FIG. shows a flowchart of an example method of marking deterministic code in AI-generated code according to some embodiments of the present disclosure. Figure 2 The method can be performed, for example, by Figure 1 the code analysis module 107. As an example, in some embodiments, Figure 2 the method can be performed in response to a request to compile a certain AI-generated code or during the compilation of such code. As another example, in some embodiments, the method can be performed in response to receiving a certain code from a generative AI model based on a certain request or prompt. Figure 2 The method.
[0030] Figure 2 The method includes receiving 202 code generated by a generative AI model. In some embodiments, receiving 202 code generated by a generative AI model includes receiving the code as output by the generative AI model. For example, the generative AI model can be configured to output code and provide the code to a process or service that performs Figure 2 the method. In some embodiments, receiving 202 code generated by a generative AI model includes loading the code from a storage device, which is stored in the storage device after being generated by the generative AI model.
[0031] In some embodiments, the code generated by the generative AI model includes code generated based on a prompt that includes a natural language description of some functions that the code should implement or perform. For example, a prompt to the generative AI model can request code to perform a specific task (e.g., opening a network connection to a specific destination, instantiating a database table with specific characteristics, or any other task that can be understood).
[0032] In some embodiments, the code generated by the generative AI model includes code in a certain programming language that is converted from other code in different programming languages. For example, the generative AI model can be provided with some Cobol code and be prompted to convert the Cobol code into Java. In the examples related to code conversion using the generative AI model described below, the code input into the generative AI model is hereinafter referred to as "base code", and the code output by the generative AI model is hereinafter referred to as "converted code".
[0033] Figure 2 The method further includes identifying at least one part of the 204 code by identifying at least one of the following: one or more parts of the deterministic code or one or more parts of the non-deterministic code. As described above, the deterministic code includes code that consistently provides the same output if given the same set of inputs. In contrast, the non-deterministic code includes code that will not provide the same output or has the possibility of not providing the same output if given the same set of inputs. This part of the code can include various granularities or scopes. For example, this part of the code can include defined code blocks, specific functions or methods, multiple code blocks, or combined functions, etc.
[0034] In some embodiments, identifying at least one part of the 204 code can include identifying one or more parts of the 206 deterministic code in the code. In other words, specifically identifying the parts of the 206 deterministic code from the code. In some embodiments, identifying at least one part of the 204 code can include identifying one or more parts of the 208 non-deterministic code in the code. In other words, specifically identifying the parts of the 208 non-deterministic code from the code. In some embodiments, identifying at least a part of the 204 code can include identifying one or more parts of the 206 deterministic code and identifying one or more parts of the 208 non-deterministic code. Those skilled in the art will understand that in some embodiments, some parts of the code may not be identified as deterministic or non-deterministic. For example, the analysis of a given part of the code may not be able to deterministically determine whether the code is deterministic or non-deterministic.
[0035] In some embodiments, identifying at least one part of the 204 code can be based on one or more user inputs for identifying parts of the deterministic and / or non-deterministic code in the code. For example, in some embodiments, the user can highlight or otherwise select a part of the code and use an input to a menu, hotkey, or other input to identify that part of the code as deterministic or non-deterministic. In other words, in some embodiments, identifying at least a part of the 204 code can be performed using the user's manual selection of at least a part of the code.
[0036] In some embodiments, at least a portion of the 204 code can be executed dynamically. For example, in some embodiments, identifying at least a portion of the 204 code can include analyzing the code and applying one or more rules to identify at least a portion of the 204 code. For example, the code can be analyzed to identify a particular portion of the code (e.g., a particular block of code or other sub-portion). One or more rules can be applied to a given portion of the code to determine whether those portions of the code are deterministic or non-deterministic. For example, in some embodiments, the presence of a particular function known to be non-deterministic in a portion of the code (e.g., a function that loads data into or stores data from a shared memory location or resource, loads data into or stores data to such a location or resource without locking or otherwise restricting access, depending on a random number generation) can cause that portion of the code to be identified as non-deterministic. As another example, the absence of a function known to be non-deterministic in a portion of the code can cause that portion of the code to be identified as deterministic. For example, a function that performs a mathematical calculation and provides some output for only some inputs can be identified as deterministic. In some embodiments, the code can be provided to a trained model (e.g., a generative AI model or another model as may be understood), which is configured to identify portions of deterministic code and / or non-deterministic code.
[0037] In some embodiments, identifying at least one portion of the 204 code can include monitoring the execution of the code or portions thereof to identify deterministic or non-deterministic behavior. For example, multiple sets of the same input can be used to repeatedly execute the code or portions thereof to determine whether the output of a particular portion of the code is consistent across all instances of the same input, indicating deterministic code. As another example, the execution of an application or code can be monitored in the context of a larger system or a complete application to detect unexpected or incorrect outputs of functions in the code, which can indicate non-deterministic code or potentially non-deterministic code (e.g., excluding the code being identified as deterministic while not explicitly identifying it as non-deterministic).
[0038] Figure 2The method further includes marking at least a portion of the 210 code. In some embodiments, marking at least a portion of the 210 code may include marking one or more portions of the identified 206 of the deterministic code as deterministic code. In some embodiments, marking at least a portion of the 210 code may include marking one or more portions of the identified 208 of the deterministic code as non-deterministic code. In some embodiments, marking at least a portion of the 210 code may include adding to the code some non-compilable tags or comments indicating portions of the deterministic or non-deterministic code. In some embodiments, marking at least a portion of the 210 code may include generating some data that indicates, in the code, portions identified as deterministic and / or portions identified as non-deterministic. Thus, in some embodiments, at least a portion of marking the 210 code may include generating 212 metadata identifying at least a portion of the code.
[0039] In some embodiments, the metadata may identify, for a given portion of the code, the location of that portion of the code within the code (e.g., using line numbers, line offsets relative to a function name or other component of the code, etc.). In some embodiments, the metadata may identify, for a given portion of the code, whether the portion of the code is identified as deterministic or non-deterministic. In some embodiments, the metadata may indicate why a given portion of the code is identified as deterministic or non-deterministic. For example, the metadata may indicate a specific function within the code portion that may cause it to be marked as non-deterministic, or may indicate specific rules that cause the code portion to be marked as deterministic and / or non-deterministic.
[0040] Identifying the 204 code and marking the 210 as deterministic and / or non-deterministic provides several advantages, which will be described in further detail below. Identifying and marking one or more portions of the 206 deterministic code may allow for the generation of multi-path code for the deterministic code. Additionally, the identified 206 and marked 210 deterministic code may be provided as feedback and / or training data to a generative AI model to enhance the generation of deterministic code. Identifying and marking one or more portions of the 208 non-deterministic code may be used to restrict or prevent the use of multi-path code for non-deterministic code. Additionally, the identified 208 and marked 210 non-deterministic code may be provided as feedback and / or training data to a generative AI model to mitigate or reduce the future generation of non-deterministic code by the generative AI model.
[0041] For further explanation, Figure 3 FIG. [X] illustrates a flowchart of an example method of marking deterministic code in AI-generated code according to some embodiments of the present disclosure. Figure 3 The method of [X] is similar to Figure 2 that of [X] in that Figure 3The method also includes receiving 202 code generated by a generative artificial intelligence (AI) model; identifying 204 at least one part of the code by identifying at least one of the following: one or more parts of deterministic code or one or more parts of non-deterministic code, including: identifying 206 one or more parts of deterministic code in the code; and identifying 208 one or more parts of non-deterministic code in the code; and marking 210 at least one part of the code.
[0042] Figure 3 The method of Figure 2 differs from Figure 3 in that the method of
[0043] also includes, in response to marking a part of the deterministic code, generating 302 a part of the multi-path code corresponding to the part of the deterministic code. As described above, the multi-path code describes the use of multiple separate code paths configured to perform similar functions. When a part of the multi-path code is encountered during the execution of an application or other software, state information describing the execution state at the point where the part of the multi-path code is encountered can be saved. The state information can describe the values of various registers, memory locations, counters, attributes, etc. If an error occurs in a code path of the multi-path code, the saved state information can be used as a checkpoint for restoring or rolling back the execution state of the application to the point before the error occurred in the code path of the multi-path code where the error occurred. Then a different code path of the multi-path code can be executed.
[0044] Although Figure 3 the method of
[0045] describes the use of generating multi-path code in response to a part of the deterministic code being marked, the marked part of the non-deterministic code can also affect the use of the multi-path code. For example, in cases where the multi-path code can be automatically or dynamically added to the code, a part of the code marked as non-deterministic can be prevented from being included in the multi-path code. As another example, when some other action such as compilation or execution is performed on the code including the multi-path code, in response to detecting a multi-path code including some non-deterministic code, a warning can be generated or the compilation can be blocked.Figure 4 A flowchart illustrating an example method of marking deterministic code in AI-generated code in accordance with some embodiments of the present disclosure. Figure 4 The method of Figure 3 is similar to Figure 4 in that the method of
[0046] Figure 4 also includes receiving 202 code generated by a generative artificial intelligence (AI) model; identifying 204 at least one portion of the code by identifying at least one of: one or more portions of deterministic code or one or more portions of non-deterministic code, including: identifying 206 one or more portions of deterministic code in the code; and identifying 208 one or more portions of non-deterministic code in the code; marking 210 at least one portion of the code; and generating 302 a portion of multi-path code corresponding to the portion of the deterministic code in response to marking a portion of the deterministic code. Figure 3 The method of Figure 3 differs from
[0047] in that generating 302 a portion of multi-path code corresponding to the portion of the deterministic code in response to marking a portion of the deterministic code further includes selecting 402 other code from the base code for the portion of the multi-path code. In
[0048] the example method of , it is assumed that the code received from the generative AI model includes transformed code in a certain programming language generated from base code in a different programming language. For example, the base code may include code written in Cobol, which is transformed by the generative AI model into transformed code written in Java.
[0047] Here, the other code for the portion of the multi-path code (e.g., the code path that will be executed if an error occurs when executing the portion of the deterministic code) is selected from the base code from which the deterministic transformed code is generated. Specifically, the other code selected from the base code may include a portion of the base code from which the generative AI model generates the deterministic transformed code. Thus, if the execution of the deterministic transformed code encounters an error, the portion of the base code from which the deterministic transformed code is generated will be executed.
[0048] This method provides additional resilience in AI-generated code. For example, assume that the base code is written in a traditional programming language but is known to be functional and resilient. The base code can be transformed into a different programming language to run more optimally on modern computing systems. If an error occurs in the transformed code, the known reliable and functional base code can alternatively be executed by using multi-path code. Additionally, the method enables the implementation of multi-path code without the need to manually write different, functionally equivalent code by using known functionally equivalent code in a different programming language.
[0049] For further explanation, Figure 5 FIG. is a flowchart illustrating an example method of marking deterministic code in AI-generated code in accordance with some embodiments of the present disclosure. Figure 3 The method of Figure 2 is similar to Figure 3 in that the method of
[0050] Figure 5 also includes receiving 202 code generated by a generative artificial intelligence (AI) model; identifying 204 at least one portion of the code by identifying at least one of: one or more portions of deterministic code or one or more portions of non-deterministic code, including: identifying 206 one or more portions of deterministic code in the code; and identifying 208 one or more portions of non-deterministic code in the code; and marking 210 at least one portion of the code. Figure 2 The method of Figure 5 differs from
[0051] in that the method of
[0052] also includes providing 502 data to the generative AI model that describes at least one portion of the identified code. In some embodiments, the data provided 502 to the generative AI model may include the metadata generated 212 described above. In some embodiments, the data provided 502 to the generative AI model may be based on or include various data points from the metadata generated 212. For example, the data provided 502 to the generative AI model may identify specific portions of deterministic and / or non-deterministic code, provide an indication of why a specific portion of the code was identified as deterministic or non-deterministic, and so on.
[0052] Aspects of the present disclosure are described by narrative text, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic included in computer program product (CPP) embodiments. With respect to any flowchart, depending on the technology involved, operations may be performed in an order different from the order shown in a given flowchart. For example, again depending on the technology involved, two operations shown in consecutive flowchart blocks may be performed in reverse order, as a single integrated step, simultaneously, or in a manner that at least partially overlaps in time.
[0053] A computer program product embodiment (“CPP embodiment” or “CPP”) is a term used in the present disclosure to describe any collection of one or more storage media (also referred to as “media”) collectively included in a set of one or more storage devices, the set of one or more storage devices collectively including machine-readable code corresponding to instructions and / or data for performing computer operations specified in a given CPP claim. A “storage device” is any tangible device that can hold and store instructions used by a computer processor. By way of non-limitation, a computer-readable storage medium can be an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these media include: magnetic disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices such as punched cards or pits / lands formed in the major surface of a disk, or any suitable combination of the foregoing. A computer-readable storage medium, as the term is used in the present disclosure, should not be construed as storing in the form of a transitory signal per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide, optical pulses propagating through an optical fiber cable, or electrical signals transmitted through wires and / or other transmission media. As will be understood by those skilled in the art, data is typically moved at certain incidental points in time during the normal operation of a storage device, such as during access, defragmentation, or garbage collection, but this does not render the storage device transitory because the data is not transitory when it is stored.
[0054] The description of the various embodiments of the present disclosure has been presented for purposes of illustration, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terms used herein were chosen to best explain the principles of the embodiments, the practical application, or the technical improvement present in the marketplace, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A method comprising: Receive code generated by a generative artificial intelligence (AI) model; identifying at least one portion of the code by identifying at least one of: one or more portions of the deterministic code or one or more portions of the non-deterministic code; as well as At least one portion of the identified code is marked.
2. The method according to claim 1, wherein: Identifying at least one portion of a code includes identifying one or more portions of a deterministic code in the code.
3. The method according to claim 1, wherein: Identifying at least one portion of a code includes identifying one or more portions of a non-deterministic code in the code.
4. The method according to claim 1, further comprising: In response to marking a portion of the deterministic code, a portion of the multi-path code is generated corresponding to the portion of the deterministic code.
5. The method according to claim 4, wherein: The portion of the multipath code includes other code in a first programming language, and the portion of the deterministic code is encoded in a second programming language different from the first programming language.
6. The method according to claim 5, wherein: The code generated by the generative AI model includes: converted code in a second programming language converted by the generative AI model from a base code in a first programming language, and wherein the portion generating the multi-path code includes: selecting the other code from the base code.
7. The method according to claim 1, further comprising: The generative AI model is provided with data describing at least one portion of the identified code.
8. The method according to claim 1, wherein: The data is provided as training data for retraining the generative AI model.
9. The method according to claim 1, wherein: Marking at least one portion of the code includes generating metadata identifying at least one portion of the code.
10. An apparatus comprising: Processing equipment; as well as A memory operatively coupled to the processing device, wherein the memory stores computer program instructions which, when executed, cause the processing device to perform a method according to any one of claims 1-9.
11. A computer program product comprising computer program instructions executable by a processor to cause the processor to perform the method according to any one of claims 1 to 9.
Citation Information
Cited By
Tagging deterministic code in artificial intelligence-generated code
US12699552B2