Method, computer program, and device (combinatoric code generation for training artificial intelligence systems)
By reducing and synthesizing code portions to create synthetic programs, the method addresses the complexity and acquisition challenges of training datasets, enhancing the training efficiency and quality of AI systems.
Patent Information
- Application Number
- JP2024195424
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-27
- Filing Date
- 2024-11-07
- Publication Date
- 2025-07-09
AI Technical Summary
Training datasets for artificial intelligence systems are often large and complex, difficult to generate or maintain, and face challenges such as data acquisition limitations and licensing restrictions, especially when data is not freely available.
Generating combined code through combination reduction and synthesis of code portions, followed by training an AI system using synthetic programs derived from reduced subsets of these combinations, which can include constraints like successful compilation and execution.
This approach efficiently generates training data for AI systems, overcoming data generation challenges and ensuring high-quality synthetic programs are used for training, thereby improving the training process.
Smart Images

Figure 2025104256000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to methods, apparatuses, and products for generating combined code for training an artificial intelligence system. Modern computer environments can involve the use of artificial intelligence systems that require extensive configurations for use. For example, some artificial intelligence systems include artificial intelligence or machine learning models composed of various algorithms. These algorithms need to be trained with specific datasets for purposes such as prediction, analysis, or the like. These training datasets are often large, complex, and can be difficult to generate or maintain. There can be limitations regarding how data can be generated or obtained, and how much data can be used by an artificial intelligence or machine learning model, for different types of data such as data structured in some other way rather than raw text. Further, acquiring training data can present additional challenges, especially when the data is not freely available to any user and requires permission or a license to obtain.
Summary of the Invention
Problems to be Solved by the Invention
[0002] Training datasets are often large, complex, and can be difficult to generate or maintain.
Means for Solving the Problems
[0003] According to embodiments of the present disclosure, various methods, apparatuses, and products for generating combined code for training an artificial intelligence system are described herein. In some aspects, generating combined code for training an artificial intelligence system includes reducing, using combination reduction, a combination of a plurality of code portions to a subset of combinations of code portions that satisfy one or more constraints, generating one or more synthetic programs using the subset of combinations of the code portions, and training an artificial intelligence system using the synthetic programs.
Brief Description of the Drawings
[0004]
Figure 1
[0005]
Figure 2
[0006]
Figure 3
[0007]
Figure 4
[0008]
Figure 5
Best Mode for Carrying Out the Invention
[0009] The disclosure of this specification includes a system and method for generating combined code for training an artificial intelligence system. The generation of a synthetic program may include generating a combination of code portions of computer code. In many cases, the set of combinations may quickly constitute a very large number of possible combinations. Thus, the set of combinations may be reduced to another set, and the reduced set may have certain characteristics. The reduced set of combinations may then be combined or recombined in various ways to generate a synthetic program or code file. In particular, the program may then be used as training data for an artificial intelligence system, such as for training a large language model (LLM). The LLM may be, for example, an LLM trained to output other synthetic code generated as a conversion of code from one programming language to code in another programming language.
[0010] FIG. 1 illustrates an exemplary computing environment in accordance with an aspect of the present disclosure. Computing environment 100 includes an example of an environment for the execution of at least some of the computer code involved in performing the various methods described herein, such as synthetic program module 107. In addition to synthetic program module 107, computing environment 100 includes, for example, computer 101, wide area network (WAN) 102, end user device (EUD) 103, remote server 104, public cloud 105, and private cloud 106. In the present embodiment, computer 101 includes a processor set 110 (including processing circuit 120 and cache 121), communication fabric 111, volatile memory 112, persistent storage 113 (including operating system 122 and synthetic program module 107 as specified above), a set of peripheral devices 114 (including a set of user interface (UI) devices 123, storage 124, and a set of Internet of Things (IoT) sensors 125), and network module 115. Remote server 104 includes remote database 130. Public cloud 105 includes gateway 140, cloud orchestration module 141, a set of host physical machines 142, a set of virtual machines 143, and a set of containers 144.
[0011] Computer 101 can be in the form of a desktop computer, laptop computer, tablet computer, smartphone, smartwatch, or any other wearable computer, mainframe computer, quantum computer, or any other form of computer or mobile device that is currently known or will be developed in the future and that can execute a program, access a network, or query a database such as remote database 130. As is well understood in the field of computer technology and depending on the technology, the execution of computer implementation methods can be distributed among multiple computers and / or multiple locations. On the other hand, in this description of computing environment 100, for the sake of simplicity as much as possible, the detailed discussion focuses on a single computer, specifically computer 101. Although computer 101 is not shown within the cloud in FIG. 1, it can be located within the cloud. On the other hand, computer 101 is not required to exist within the cloud except within any arbitrarily shown scope.
[0012] Processor set 110 includes one or more computer processors of any type that are currently known or will be developed in the future. Processing circuit 120 can be distributed across multiple packages, such as multiple packaged integrated circuit chips. Processing circuit 120 can implement multiple processor threads and / or multiple processor cores. Cache 121 is a memory located within the processor chip package and is typically used for high-speed access to data or code that should be available to the threads or cores executing on processor set 110. Cache memory is typically organized into multiple levels depending on its relative proximity to the processing circuit. Alternatively, some or all of the cache for the processor set can be located "off-chip". In some computing environments, processor set 110 can be designed to operate using qubits and execute quantum computing.
[0013] Computer-readable program instructions are typically loaded onto computer 101 and executed by a set of processors 110 of computer 101 in a series of operational steps, thereby enabling a computer-implemented method, and as a result, the instructions so executed will instantiate the method specified in the flowchart and / or description of the computer-implemented method included in this document. These computer-readable program instructions are stored in various types of computer-readable storage media such as cache 121 and other storage media discussed below. The program instructions and associated data are accessed by the set of processors 110 to control and direct the execution of the computer-implemented method. In computing environment 100, at least some of the instructions for executing the computer-implemented method may be stored in synthetic program module 107 within persistent storage 113.
[0014] Communication fabric 111 is a signal conduction path that enables various components of computer 101 to communicate with each other. Typically, this fabric is created with switches and conductive paths such as buses, bridges, physical input / output ports, and switches and conductive paths that make up the equivalent, etc. Other types of signal communication paths such as optical fiber communication paths and / or wireless communication paths may be used.
[0015] Volatile memory 112 is any type of volatile memory known currently or to be developed in the future. Examples include dynamic random access memory (RAM) or static RAM. Typically, volatile memory 112 is characterized by random access, although this is not essential unless affirmatively indicated. In computer 101, volatile memory 112 is located within a single package and exists inside computer 101, but alternatively or additionally, volatile memory may be distributed across multiple packages and / or located external to computer 101.
[0016] The persistent storage 113 is any form of non-volatile storage for a computer that is currently known or will be developed in the future. The non-volatility of this storage means that the stored data is maintained regardless of whether power is directly supplied to the computer 101 and / or to the persistent storage 113. The persistent storage 113 can be read-only memory (ROM), but usually at least a part of the persistent storage enables writing of data, deletion of data, and rewriting of data. Some well-known forms of persistent storage include magnetic disks and solid-state storage devices. The operating system 122 can take several forms, such as various known proprietary operating systems or open-source Portable Operating System Interface (POSIX)-type operating systems that employ a kernel. The code included in the synthetic program module 107 typically includes at least some of the computer code involved in performing the computer-implemented methods described herein.
[0017] The peripheral device set 114 includes a set of peripheral devices of the computer 101. Data communication connections between the peripheral devices of the computer 101 and other components can be implemented in various ways, such as a Bluetooth (registered trademark) connection, a Near-Field Communication (NFC) connection, a connection via a cable (such as a Universal Serial Bus (USB) type cable), an insertion type connection (for example, a Secure Digital (SD) card), a connection through a local area communication network, and even a connection through a wide area network such as the Internet. In various embodiments, the UI device set 123 can include components such as a display screen, a speaker, a microphone, wearable devices (such as goggles and smartwatches), a keyboard, a mouse, a printer, a touchpad, a game controller, and a haptic device. The storage 124 is an external storage such as an external hard drive or an insertable storage such as an SD card. The storage 124 can be persistent and / or volatile. In some embodiments, the storage 124 can take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments where the computer 101 is required to have a large amount of storage (for example, the computer 101 locally stores and manages a large-scale database), this storage can be provided by a peripheral storage device designed to store a very large amount of data, such as a Storage Area Network (SAN) shared by a plurality of geographically dispersed computers. The IoT sensor set 125 is composed of sensors that can be used in applications of the Internet of Things. For example, one sensor can be a thermometer, and another sensor can be a motion detector.
[0018] The network module 115 is a collection of computer software, hardware, and firmware that enables the computer 101 to communicate with other computers through the WAN 102. The network module 115 may include hardware such as a modem or a Wi-Fi (registered trademark) signal transceiver, software for packetizing and / or depacketizing data for communication over a communication network, and / or web browser software for communicating data over the Internet. In some embodiments, the network control function and the network transfer function of the network module 115 are executed on the same physical hardware device. In other embodiments (e.g., embodiments that utilize Software-Defined Networking (SDN)), the control function and the transfer function of the network module 115 are executed on physically separate devices such that the control function manages several different network hardware devices. The computer-readable program instructions for executing the computer implementation method can typically be downloaded to the computer 101 from an external computer or an external storage device through a network adapter card or a network interface included in the network module 115.
[0019] The WAN 102 is any wide area network (e.g., the Internet) that can communicate computer data over a non-local distance by any technology for communicating computer data that is currently known or will be developed in the future. In some embodiments, the WAN 102 can be replaced and / or supplemented by a local area network (LAN) designed to communicate data between devices located in a local area, such as a Wi-Fi network. The WAN and / or LAN typically includes computer hardware such as copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and edge servers.
[0020] The end-user device (EUD) 103 is any computer system used and controlled by an end user (e.g., a customer of the enterprise operating the computer 101) and can take any of the forms discussed above in relation to the computer 101. The EUD 103 typically receives beneficial and useful data from the operation of the computer 101. For example, in a virtual scenario where the computer 101 is designed to provide recommendations to the end user, this recommendation will typically be communicated from the network module 115 of the computer 101 to the EUD 103 through the WAN 102. In this way, the EUD 103 can display the recommendation to the end user or present it in another way. In some embodiments, the EUD 103 can be a client device such as a thin client, a thick client, a mainframe computer, a desktop computer, and the like.
[0021] The remote server 104 is any computer system that provides at least some data and / or functions to the computer 101. The remote server 104 can be controlled and used by the same entity that operates the computer 101. The remote server 104 represents a machine that collects and stores beneficial and useful data for use by other computers such as the computer 101. For example, in a virtual scenario where the computer 101 is designed and programmed to provide recommendations based on past data, this past data can be provided from the remote database 130 of the remote server 104 to the computer 101.
[0022] The public cloud 105 is any computer system that provides on-demand availability of computer system resources and / or other computing capabilities, particularly data storage (cloud storage) and computing power, for use by multiple entities without direct active management by the user. Cloud computing typically exploits resource sharing to achieve coherence and economies of scale. The direct active management of the computing resources of the public cloud 105 is performed by the computer hardware and / or software of the cloud orchestration module 141. The computing resources provided by the public cloud 105 are typically implemented by virtual computing environments that run on various computers that make up the host physical machine set 142, which is the universe of physical computers within and / or available to the public cloud 105. The virtual computing environment (VCE) typically takes the form of virtual machines from the virtual machine set 143 and / or containers from the container set 144. It is understood that these VCEs can be stored as images and transferred as images or after instantiation of the VCEs among and within various physical machine hosts. The cloud orchestration module 141 manages the transfer and storage of the images, deploys new instantiations of the VCEs, and manages the active instantiation of the VCE deployments. The gateway 140 is an aggregate of computer software, hardware, and firmware that enables the public cloud 105 to communicate through the WAN 102.
[0023] Here, some further explanations are provided for a virtualized computing environment (VCE). A VCE can be stored as an "image". A new active instance of a VCE can be instantiated from the image. Two well-known types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to a feature of the operating system where the kernel enables the existence of multiple isolated instances of user space, called containers. These isolated instances of user space typically behave as actual computers from the perspective of the programs running within them. A computer program running on a normal operating system can utilize all the resources of that computer, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, a program running inside a container can only use the contents of the container and the devices allocated to the container, and this feature is known as containerization.
[0024] The private cloud 106 is similar to the public cloud 105, except that computing resources are available only for use by a single enterprise. The private cloud 106 is shown as being in communication with the WAN 102, but in other embodiments, the private cloud may be completely disconnected from the Internet and accessible only through a local / private network. A hybrid cloud is a composite of multiple different types of clouds (e.g., private cloud, community cloud, or public cloud types), and is often implemented by different vendors. Each of the multiple clouds remains a separate discrete entity, but the larger hybrid cloud architecture is coupled together by standardized or proprietary technologies that enable orchestration, management, and / or data / application portability between the constituent clouds. In this embodiment, both the public cloud 105 and the private cloud 106 are part of a larger hybrid cloud.
[0025] FIG. 2 shows a flowchart of an exemplary method for generating combination code for training an artificial intelligence system, according to some embodiments of the present disclosure. The method of FIG. 2 can be executed, for example, by the synthesis program module 107 of FIG. 1. In some embodiments, the synthesis program module 107 can be implemented as a separate process or service from the application or software that implements the generation of combination code for training an artificial intelligence system. For example, the synthesis program module 107 can be implemented by an operating system or other software that monitors the behavior and execution of an application that implements the generation of combination code for training an artificial intelligence system. As another example, in some embodiments, the synthesis program module 107 can be implemented as a process or service that applies updates or patches to an application or code capable of implementing the generation of combination code for training an artificial intelligence system, or as a process or service that monitors or detects updates or patches to an application or code capable of implementing the generation of combination code for training an artificial intelligence system.
[0026] The method of FIG. 2 includes a stage 202 that uses combination reduction to reduce a combination of multiple code portions to a subset of combinations of code portions that satisfy one or more constraints. In some embodiments, the synthesis program module 107 obtains code files to use for generating combinations of code. The code files can be, for example, part of a software suite or platform (e.g., an operating system). The code files can represent complete programs, such as a program for memory management used as part of an operating system.
[0027] In some embodiments, the code file is split up into a plurality of code snippets or code portions. The splitting can be performed in various ways. For example, the code file can be broken down by individual functions, individual code blocks, classes, methods, or the like. As another example, the code file can be split into code portions using a specific character or bit of code as a delimiter (e.g., the curly brace "{" or "}" characters). The synthesis program module 107 can be configured to receive the code file as input and decompose it into various code portions. The synthesis program module 107 can also assign metadata to each code portion. For example, the metadata for a code portion can include a code portion identifier, a parent code file identifier, a description of what the code portion does, the code author, and the like.
[0028] As an example, a memory management code file can be split into several code portions according to individual functions executed by different sections of the memory management code file. For example, different code portions can be generated for sections of the memory management code file such as a memory allocation function, a memory deallocation function for memory release, a garbage collection function, and the like. As a more specific example, the memory management code file can be a COBOL code file.
[0029] In some embodiments, the resulting code portions are recombined into different combinations of code portions that can be designed as a new synthesis program. Each code portion or group of code portions can form a component of a combination of code portions that can later become an executable program. The code portions can be combinable in different ways. Generally, it can be understood that a set of x elements can be combined in x! ways, or x factorial. Further, if p elements are selected from a set of q elements, the number of combinations is the Cartesian product, or p qIt can be (which can also be expressed as p^q, or p to the power of q). The reader will understand that the number that can be brought about by the factorial of a set of elements, or even the Cartesian product, will grow very rapidly and become extremely large as the number of elements increases. As an example, if there are 8 elements or code parts, the number of possible combinations can be 8! or 40,320. For example, if 2 elements are to be selected from a set of 8 elements, the Cartesian product is 256, and if 2 elements are to be selected from a set of only 20 elements, the Cartesian product results in more than 1 million possible combinations. Even if it is possible to generate combinations of a large number of code parts, it may be unrealistic to use this number of combinations as a synthetic program for training an artificial intelligence system. Furthermore, if the combinations are regenerated each time a new code part is added, this can result in an inefficient and burdensome process of constantly regenerating a large number of combinations.
[0030] In some embodiments, the Cartesian product of combinations of code portions can be reduced in one or more ways. The method of FIG. 2 includes a stage 203 that performs an n-wise reduction of the number of combinations of the Cartesian product. In this embodiment, n corresponds to a subset of elements to be selected from the complete set of elements for inclusion in the combination of code portions. By way of example, n can be equal to 2, resulting in a 2-wise or pairwise reduction of the number of combinations of the Cartesian product. Purely by way of example, there can be a set of 8 elements, and each set can be considered to have pairs of elements. That is, there can be 8 sets of code portions, and each set includes two different code portions. As a more specific example, each code portion within a set can correspond to a different variation of a particular function or operation. For example, a set of code portions can correspond to two different functions, such as memalloc1 and memalloc2, each of which performs memory allocation. The two memory allocation functions can be derived from the same source code file, or from different source code files (e.g., from the same or different COBOL source code files). In other words, the variations of the memory allocation functions can each form a component part of a composite program that can later be used as a memory management program.
[0031] Continuing with the above example, recall that for a set of 8 elements, each having 2 elements, the Cartesian product can be 256. When pairwise reduction is performed, the combinatorial formula can be applied as shown below:
Equation
[0032] The method of FIG. 2 also includes step 204 of generating one or more synthesis programs using a subset of the combinations of code parts. The step of generating one or more synthesis programs using a subset of the combinations of code parts can include generating a complete synthesis program or an executable code file using a combination of code parts from the subset of the combinations of code parts generated in step 202. The complete synthesis program can be, for example, a recombination of code parts that results in another memory management program containing different code parts compared to the originally used memory management program that was previously decomposed into code parts. For example, different versions of memory allocation functions such as memalloc1 and memalloc2 can be used as part of different combinations of code parts, resulting in different recombined synthesis programs.
[0033] The method of FIG. 2 also includes a step 206 of training an artificial intelligence system using a synthetic program. The step of training an artificial intelligence system using a synthetic program may include a step of training a generative AI model such as an artificial intelligence model, a machine learning model, a deep learning model, or a large language model (LLM) using the synthetic program as a training data set. The synthetic program module 107 may be configured to train the AI system in different ways. For example, the step of training the AI system may include a step of providing a set of training data sets to the AI system, and a step of causing the AI system to make a decision or generate an output based on the training data set. In an initial stage of training, the AI system may be provided with data tagged or labeled with metadata that helps the AI system understand the type of data and generate an output. The reader will understand that an exemplary AI system can be used to extract code written in one programming language and generate similar code in a second programming language. In such an example, first, the synthetic program module 107 may provide a training data set of a combination of code portions tagged in a specific way, such as using the programming language used to generate the code within the combination of code portions. The synthetic program module 107 may then prompt the AI system to generate another program based on the training data set, but provide an identifier of a programming language different from that used for the training data set within the prompt or via other metadata. In a later training stage, the tagging (e.g., identification of the programming language) may be removed.
[0034] Figure 3 shows a flowchart of another exemplary method for generating combination code for training an artificial intelligence system, according to some embodiments of the present disclosure. The method of Figure 3 is similar to the method of Figure 2 in that it includes a step 202 of using combination reduction to reduce a combination of a plurality of code portions to a subset of combinations of code portions that satisfy one or more constraints, a step 204 of using the subset of combinations of code portions to generate one or more synthetic programs, and a step 206 of using the synthetic programs to train an artificial intelligence system.
[0035] The method of Figure 3 differs from the method of Figure 2 in that it also includes a step 302 of identifying a combination of one or more code portions that compiles successfully. The step 302 of identifying a combination of one or more code portions that compiles successfully may include a step of processing a synthetic program formed from the combination of code portions using a compiler, and a step of determining whether the program compiles successfully. For example, as previously mentioned, the code portions may be COBOL code portions, and thus the synthetic program generated from the code portions may also be a COBOL program. Since COBOL is a compiled language, a COBOL compiler may be used to compile the program, and the synthetic program module 107 may be configured to determine whether the program compiles successfully. In some embodiments, a successfully compiled program may be used for training an LLM model, while other programs created using other combinations of code portions may not be used for training the LLM. Further, only a small sample of the subset of combinations of code portions may be compiled. Based on the results from the compilation of the small sample, the synthetic program module 107 may determine whether to compile the remaining combinations within the subset.
[0036] The method of FIG. 3 also includes a step 304 of identifying a combination of one or more code portions that result in one or more executable programs. The reader will understand that a particular code may not generate an error during compilation, but may generate an error during execution. Thus, the synthetic program module 107 may be configured to execute a synthetic program generated from one or more of the combinations of code portions and check that it has been successfully executed before using the synthetic program as part of a training dataset for the LLM. For example, the combination may be executed using an integrated development environment (IDE) or other execution environment.
[0037] The method of FIG. 3 also includes a step 306 of identifying a combination of one or more code portions that generate one or more expected results when executed. The reader will understand that a particular code may be successfully executed but generate unexpected results. Thus, the synthetic program module 107 may be configured to execute one or more programs generated using a combination of code portions and check whether the generated execution results match the expected results for the program.
[0038] FIG. 4 shows a flowchart of another exemplary method for generating combination code for training an artificial intelligence system, according to some embodiments of the present disclosure. The method of FIG. 4 is similar to the method of FIG. 2 in that it includes a step 202 of using combination reduction to reduce a combination of a plurality of code portions to a subset of combinations of code portions that satisfy one or more constraints, a step 204 of generating one or more synthetic programs using the subset of combinations of code portions, and a step 206 of training an artificial intelligence system using the synthetic programs.
[0039] The method of FIG. 4 differs from the method of FIG. 2 in that it also includes step 402 of selecting a combination of code portions for the subsets such that each code portion is included in a combination of at least one of the subsets. In some embodiments, the synthetic program module 107 may be configured to include a combination of code portions within the subset such that each code portion is part of a combination of at least one code portion, and as a result, all code portions are used. The reader will understand that a code portion can be part of a combination of more than one code portion. To ensure coverage, the synthetic program module 107 may be configured to check that each code portion is part of a combination of at least one code portion. For example, a code portion may have an associated identifier. The synthetic program module 107 may be configured to store (e.g., in some data structure such as a table) a record of the identifier for each code portion included in the combination of code portions as part of the combination operation that generates the combination of code portions. As an example, the synthetic program module 107 may ensure that each n-wise configuration (e.g., a 2-wise configuration for pair-wise reduction, or a 3-wise configuration for 3-wise reduction) is included in a combination of at least one code portion.
[0040] The synthetic program module 107 may further be configured to compare the set of included code portions with the set of code portions created in the initial decomposition of the code portions of the code file. By the comparison, it may become clear whether any code portion is missing and not included in any combination of code portions. If so, the synthetic program module 107 may generate another combination of code portions that includes the missing code portion.
[0041] The method of FIG. 4 also includes a step 404 of including in a subset a combination of code portions having a set of specific code portions. The reader will understand that a specific code portion may form an important component of a synthetic program generated for some purpose. For example, a specific memory allocation function may be an important component of the generated memory management program. Thus, the synthetic program module 107 may be configured to generate a combination of code portions that results in a synthetic program including a specific memory allocation function. Further, in some cases, the synthetic program module 107 may be configured to use a combination engine to generate a combination of code portions. In such a case, the synthetic program module 107 may provide the combination engine with specific constraints such as a constraint that a specific code portion is included in the set of code portions and in any combination of code portions generated by the combination engine.
[0042] FIG. 5 shows a flowchart of another exemplary method for generating combination code for training an artificial intelligence system, according to some embodiments of the present disclosure. The method of FIG. 5 is similar to the method of FIG. 2 in that it includes a step 202 of using combination reduction to reduce a combination of a plurality of code portions to a subset of combinations of code portions that satisfy one or more constraints, a step 204 of using the subset of combinations of code portions to generate one or more synthetic programs, and a step 206 of using the synthetic programs to train an artificial intelligence system.
[0043] The method of FIG. 5 differs from the method of FIG. 2 in that it also includes a step 502 of training a large language model used by an artificial intelligence system. In particular, the step of training the LLM may include different phases, which may be executed individually, as a group, sequentially, or in any suitable configuration. In the pre-training phase, the code portions may be provided as token inputs to the LLM, and the LLM may be configured to predict the next token, such as the next code portion. In another phase, also referred to as the supervised fine-tuning or instruction tuning phase, the LLM may be provided with a full synthetic program (such as a combination of code portions that are complete programs or code files), and sample outputs are provided as target outputs. The LLM may then be trained to generate a full program as a response that minimizes the difference between its output and the provided sample output. In yet another phase, the LLM may be trained using reinforcement learning, which is another fine-tuning step to align the model output with specific preferences or configuration parameters. For example, the LLM may be trained to output a full program that conforms to specific preferences, such as outputting code files in a specific programming language or using other constraints.
[0044] The method of FIG. 5 also includes a step 504 of training the artificial intelligence system using a subset of combinations of code portions while excluding combinations of one or more other code portions from the training. As described above, a large number of combinations of code portions can be created from a set of code portions (for example, the Cartesian product of a set of 8 pairs of code portions can result in 256 combinations of code portions). However, in some embodiments, only a subset of the combinations of code portions resulting from the combination reduction of the entire set of combinations of code portions is used for AI training. For example, only a subset of the combinations of code portions resulting from an n-wise (e.g., pairwise) combination reduction of the entire set of combinations of code portions is used for AI training.
[0045] Various aspects of the present disclosure are illustrated by block diagrams of machine logic included in descriptions, flowcharts, block diagrams of computer systems, and / or embodiments of a computer program product (CPP). With respect to any flowchart, depending on the technology involved, operations may be performed in an order different from that shown in a given flowchart. For example, again depending on the technology involved, two operations shown in consecutive flowchart blocks may be performed in reverse order, as a single integrated step, simultaneously, or at least partially overlapping in time.
[0046] An embodiment of a computer program product (referred to herein as a "CPP embodiment" or "CPP") is, in the context of this disclosure, a term used to describe any set of one or more storage media (also referred to as "media") collectively included in a set of one or more storage devices, which collectively contain machine-readable code corresponding to instructions and / or data for performing the computer operations specified in a given CPP claim. A "storage device" is any tangible device that can hold and store instructions for use by a computer processor. A computer-readable storage medium can be, but is not limited to, an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these media are floppy disks, hard disks, random access memory (RAM), read only memory (ROM), erasable programmable read only memory (EPROM or flash memory), static random access memory (SRAM), compact disk read only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices (such as punch cards or pits / lands formed on the major surfaces of disks), or any suitable combination of the foregoing. A computer-readable storage medium shall not be construed as storage in the form of a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide, optical pulses passing through an optical fiber cable, electrical signals transmitted through a wire, and / or other transmission media, as the term is used in this disclosure. As will be understood by those skilled in the art, data is typically moved during normal operation of a storage device at some irregular points in time, such as during access, defragmentation, or garbage collection, but the fact that the data is not transient while it is stored does not make the storage device a transient one.
[0047] The description of various embodiments of the present disclosure is presented for illustrative purposes and is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used herein are chosen to best explain the principles of the embodiments, the practical application, or a technical improvement found in the marketplace, or to enable other skilled artisans to understand the embodiments disclosed herein.
Claims
Claim 1 A method for generating combination codes for training an artificial intelligence system, comprising: Reducing a combination of a plurality of code parts to a subset of combinations of code parts that satisfy one or more constraints using combination reduction; Generating one or more synthetic programs using the subset of combinations of the code parts; and Training an artificial intelligence system using the one or more synthetic programs A method comprising. Claim 2 The step of reducing the combination of the plurality of code parts to a subset of the combinations of the code parts using the combination reduction: Performing an n-wise reduction of the combination of the plurality of code parts The method according to claim 1, further comprising. Claim 3 The step of reducing the combination of the plurality of code parts to a subset of the combinations of the code parts using the combination reduction: Identifying a combination of one or more successfully compiled code parts The method according to claim 1 or 2, further comprising. Claim 4 The step of reducing the combination of the plurality of code parts to a subset of the combinations of the code parts: Identifying a combination of one or more code parts that results in one or more executable programs The method according to claim 1 or 2, further comprising. Claim 5 The step of reducing the combination of the plurality of code parts to a subset of the combinations of the code parts: Identifying a combination of one or more code parts that generates one or more expected results when executed The method according to claim 1 or 2, further comprising. Claim 6 The combination of the plurality of code parts is generated using a plurality of code parts, and the step of reducing the combination of the plurality of code parts to the subset: Selecting a combination of code parts for the subset such that each code part is included in a combination of at least one code part of the subset The method according to claim 1 or 2, further comprising. Claim 7 The step of reducing the combination of the plurality of code parts to a subset of the combinations of the code parts: Including in the subset a combination of code parts having a specific set of code parts The method according to claim 1 or 2, further comprising. Claim 8 The step of training the artificial intelligence system includes the step of training a large language model used by the artificial intelligence system, according to the method of claim 1 or 2.
9. The step of training the artificial intelligence system includes the step of excluding a combination of one or more other code portions from the training, while using a subset of the combination of the code portions of the combination of the code portions to train the artificial intelligence system, according to the method of claim 1 or 2.
10. On a computer: A procedure for reducing a combination of a plurality of code portions to a subset of a combination of code portions that satisfy one or more constraints using combination reduction; A procedure for generating one or more synthetic programs using the subset of the combination of the code portions; and A procedure for training an artificial intelligence system using the one or more synthetic programs A computer program for causing the above to be executed.
11. On the computer A procedure for performing n-wise reduction of the combination of the plurality of code portions The computer program according to claim 10, for further causing the above to be executed.
12. The procedure for reducing the combination of the plurality of code portions to a subset of the combination of the code portions using combination reduction is: A procedure for identifying a combination of one or more code portions that are successfully compiled The computer program according to claim 10 or 11, further including the above.
13. The procedure for reducing the combination of the plurality of code portions to a subset of the combination of the code portions is: A procedure for identifying a combination of one or more code portions that result in one or more executable programs The computer program according to claim 10 or 11, further including the above.
14. The procedure for reducing the combination of the plurality of code portions to a subset of the combination of the code portions is: A procedure for identifying a combination of one or more code portions that generate one or more expected results when executed The computer program according to claim 10 or 11, further including the above.
15. The combination of the plurality of code portions is generated using a plurality of code portions, and the procedure for reducing the combination of the plurality of code portions to the subset is: A procedure for selecting a combination of code parts for the subset such that each code part is included in a combination of at least one code part of the subset The computer program according to claim 10 or 11, further comprising the procedure
16. A processing device; and A memory operably coupled to the processing device An apparatus comprising: when executed, the memory causes the processing device to Use combination reduction to reduce a combination of a plurality of code parts to a subset of combinations of code parts that satisfy one or more constraints; Generate one or more synthetic programs using the subset of combinations of the code parts; and Train an artificial intelligence system using the one or more synthetic programs Store computer program instructions The apparatus
17. When executed The apparatus according to claim 16, further having computer program instructions for performing an n-wise reduction of the combination of the plurality of code parts
18. The instructions for reducing the combination of the plurality of code parts to the subset of combinations of the code parts using the combination reduction are: The apparatus according to claim 16 or 17, further comprising instructions for identifying a combination of one or more successfully compiled code parts
19. The instructions for reducing the combination of the plurality of code parts to the subset of combinations of the code parts are: The apparatus according to claim 16 or 17, further comprising instructions for identifying a combination of one or more code parts that results in one or more executable programs
20. The instructions for reducing the combination of the plurality of code parts to the subset of combinations of the code parts are: The apparatus according to claim 16 or 17, further comprising instructions for identifying a combination of one or more code parts that generates one or more expected results when executed