Context aware transcoding and optimization based on neural network
Through neural network model and context data processing technology, the source code is parsed and converted, and the accuracy and inefficiency of code conversion between programming languages in the prior art are solved, and the code conversion effect of high fidelity and semantic integrity is achieved.
Patent Information
- Application Number
- CN202411731524.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-14
- Filing Date
- 2024-11-29
- Publication Date
- 2025-06-17
AI Technical Summary
The prior art has problems of accuracy and inefficiency in code conversion between programming languages, especially in the preservation of semantic and logical integrity of source code.
By using neural network models and context data processing techniques, the source code is parsed into intermediate representations (IRs), and the structural and semantic models of the source code are established through static analysis. Then, the program architecture of the object code is constructed, including generating context-aware placeholders, converting IR to a single static allocation (SSA) form, building a system dependency graph (SDG), and using a large language model (LLM) to convert the sorted tasks into the target language.
Improves the accuracy and efficiency of code conversion, ensures the syntax and semantic integrity of the target code, and can effectively convert the entire application and adapt to third-party libraries and runtime behavior.
Smart Images

Figure CN120162048A_ABST
Abstract
Description
Background Art
[0001] The present invention generally relates to neural-network-based automatic code conversion between programming languages, and more particularly to improving the accuracy and efficiency of converted code by leveraging a neural network model and context data processing to impart control in programming language conversion using a large language model (LLM).
[0002] Traditionally, fine-tuning pre-trained models has been the mainstay in code conversion tasks, typically involving converting source code to a standard intermediate representation (IR), such as single static assignment (SSA), and training a model to convert from that IR to the target language. This conventional approach utilizes a corpus of a large number of source code snippets and target code snippets to repeatedly update the model via gradient descent, typically requiring a large number of labeled examples for model convergence. While efficient in some respects, this approach has significant drawbacks, including the need for large task-specific datasets and the risks of poor generalization and reliance on potentially misleading features of the training data. Additionally, the standard SSA-based IR, while simplifying the code for efficient data flow analysis and facilitating compiler optimizations, often strips away nuanced language constructs that are crucial for understanding the intent and logic of the original code. This loss of language-specific properties, such as object-oriented programming elements, and the introduction of synthetic variables create a significant semantic gap that may impede the readability and convertibility of the target code.
[0003] Furthermore, the rise of large language models (LLMs) has introduced new methodologies for in-context learning. Although they are successful in natural language processing tasks, they face considerable challenges when applied to code conversion. The unique semantics and characteristics of programming languages make it difficult to specify the correct context for conversion, with the risk of token overflow or incorrect conversion results. Even though LLMs have shown the potential to perform chain-of-thought reasoning conditioned on a few examples, converting an entire application or adapting to third-party libraries and runtime behavior remains a complex task beyond the capabilities of traditional conversion or fine-tuning methods. Summary of the Invention
[0004] According to an embodiment of the present invention, there is provided a computer-implemented method for efficiently converting program code from a source language to a target language. An input source code is parsed into an intermediate representation (IR) using a processor device. By applying static analysis to the IR, a structural and semantic model of the source code is established, and a program architecture of the target code is constructed based on the IR, including generating context-aware placeholders. The IR is transformed into a single static assignment (SSA) form, and a system dependence graph (SDG) is constructed from the SSA form. The SDG is traversed to sort the conversion tasks, and a large language model (LLM) is used to convert the sorted tasks into the target language. A converted program is generated by integrating the converted code segments into a coherent program structure of the target language.
[0005] According to an embodiment of the present invention, there is provided a system for efficiently converting program code from a source language to a target language. An input source code is parsed into an intermediate representation (IR) using a processor device. By applying static analysis to the IR, a structural and semantic model of the source code is established, and a program architecture of the target code is constructed based on the IR, including generating context-aware placeholders. The IR is transformed into a single static assignment (SSA) form, and a system dependence graph (SDG) is constructed from the SSA form. The SDG is traversed to sort the conversion tasks, and a large language model (LLM) is used to convert the sorted tasks into the target language. A converted program is generated by integrating the converted code segments into a coherent program structure of the target language.
[0006] According to an embodiment of the present invention, there is provided a non-transitory computer-readable storage medium including a computer-readable program operably coupled to a processor device for efficiently converting program code from a source language to a target language. An input source code is parsed into an intermediate representation (IR) using a processor device. By applying static analysis to the IR, a structural and semantic model of the source code is established, and a program architecture of the target code is constructed based on the IR, including generating context-aware placeholders. The IR is transformed into a single static assignment (SSA) form, and a system dependence graph (SDG) is constructed from the SSA form. The SDG is traversed to sort the conversion tasks, and a large language model (LLM) is used to convert the sorted tasks into the target language. A converted program is generated by integrating the converted code segments into a coherent program structure of the target language.
[0007] These and other features and advantages will become apparent from the following detailed description of its illustrative embodiments, which is to be read in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The following description will provide details of the preferred embodiments with reference to the following drawings, in which:
[0009] Figure 1FIG. is an exemplary processing system for code conversion to which this principle can be applied, showing an embodiment according to the present invention;
[0010] Figure 2 FIG. is a diagram showing an advanced system and method for code conversion optimized via a transformer-based neural network using a large language model (LLM) according to an embodiment of the present invention;
[0011] Figure 3 FIG. is a diagram showing a system and method for model adjustment and prompt adjustment in language model training for neural network-based code conversion according to an embodiment of the present invention;
[0012] Figure 4 FIG. is a diagram showing a system and method for efficient code conversion using an exemplary frozen language model with frozen weights and learnable weights according to an embodiment of the present invention;
[0013] Figure 5 FIG. is a diagram showing a system and method for neural network-based conversion of programming code using IR and standardizing source code through a language model for precise conversion according to an embodiment of the present invention;
[0014] Figure 6 FIG. is a block diagram / flowchart showing a method for neural network-based code conversion using a large language model (LLM) according to an embodiment of the present invention;
[0015] Figure 7 FIG. is a block diagram / flowchart showing a method for neural network-based code conversion including target language architecture generation according to an embodiment of the present invention;
[0016] Figure 8 FIG. is a generalized diagram of an exemplary neural network system for neural network-based context-aware code conversion and optimization according to an embodiment of the present invention;
[0017] Figure 9 FIG. is a hardware diagram of an exemplary artificial neural network (ANN) system for neural network-based context-aware code conversion and optimization according to an embodiment of the present invention;
[0018] Figure 10 FIG. is a block diagram of an exemplary neuron in a neural network system for neural network-based context-aware code conversion and optimization according to an embodiment of the present invention;
[0019] Figure 11 FIG. is a diagram of an exemplary hierarchical neural network system for neural network-based context-aware code conversion and optimization according to an embodiment of the present invention;
[0020] Figure 12 is a diagram showing a system for neural network-based context-aware code transformation and optimization according to an embodiment of the present invention; and
[0021] Figure 13 is a diagram showing an exemplary computing environment for executing at least some of a computer code for context-aware code transformation and optimization based on an adaptive neural network according to an embodiment of the present invention. Detailed Description
[0022] According to aspects of the present invention, systems and methods for context-aware code transformation and optimization based on an adaptive neural network are provided.
[0023] In various embodiments, according to aspects of the present invention, the present invention may include neural network-based systems and methods for transforming programming code, which significantly improve the accuracy and efficiency of code transformation across various programming languages by leveraging neural network models of the prior art and complex context data processing techniques.
[0024] The present invention may include an adaptive neural network system that can seamlessly integrate and process different programming languages and contexts. The system incorporates an innovative mechanism for neural network adaptation, employing a dual-structured model that divides network weights into "locked" and "trainable" copies. This unique configuration enables the system to fine-tune its transformation capabilities for specific tasks without compromising the underlying strength and generality of the pre-trained model.
[0025] In addition to neural network adaptation, the present invention may employ advanced techniques in program analysis, such as the use of intermediate representation (IR) and dedicated processing modules. These techniques ensure that the source code is not only accurately transformed into the target language but also retains the structural and functional integrity of the original program. The present invention addresses and overcomes challenges commonly encountered in conventional code transformation methods, such as the semantic gap, readability issues, and the complexity of transforming third-party libraries and custom APIs.
[0026] In conventional practice, transforming programming code from one language to another typically involves an intermediate representation (IR), such as single static assignment (SSA). While this approach helps optimize the code for runtime benefits and simplifies certain compiler optimizations, it often loses high-level constructs, resulting in a semantic gap that hinders the readability and fidelity of the transformed code to the original logic. Additionally, transforming an entire application, especially one involving third-party libraries, remains beyond the capabilities of standard converters and may result in code that is unreadable or unmaintainable by humans.
[0027] Traditional methods of training neural network models for code generation tasks rely on converting source code into a standardized IR and fine-tuning the model to convert from the IR to the target language. Despite leveraging extensive labeled datasets and iterative gradient updates, this method faces challenges such as the need for large, task-specific datasets, poor generality, and the utilization of irrelevant training data features.
[0028] Pre-trained large language models (LLMs) have brought progress through their ability to perform in-context learning, adapting to new tasks during inference by adjusting a few examples. This method has been successful in natural language processing tasks. However, the application of this method to code conversion has become complex due to the difficulty in specifying the context for conversion between programming languages, the limited token size handling ability of LLMs, and the tendency of LLMs to generate hallucinations or unfaithful outputs.
[0029] In view of these challenges, the present invention improves conventional systems and methods by integrating a neural network model with advanced context data processing strategies. In some embodiments, the present invention employs an adaptive system that improves the accuracy and efficiency of code conversion by using context-aware neural networks and optimization techniques. According to aspects of the present invention, the system and method can maintain the integrity of the logic and semantics of the original code while optimizing the converted code for specific operating conditions and performance requirements of the target environment.
[0030] In various embodiments, the present invention can be used to overcome the limitations of traditional methods by providing a more robust and context-intelligent conversion process. It leverages the advantages of neural networks to understand code structures and patterns and introduces sophisticated context manipulation strategies to achieve high-fidelity conversion across various programming languages. The system of the present invention encapsulates the principles of "frozen and opaque" and "frozen and semi-transparent" LLMs, utilizing innovative modules such as ControlNet for in-context learning, prompt embedding, and prompt enhancement to enrich conversion prompts. According to aspects of the present invention, this comprehensive framework significantly advances the rules of code conversion, providing a solution that is sensitive to the nuances of different programming contexts and capable of delivering optimized code for the target environment.
[0031] As will be understood by those skilled in the art, aspects of the present invention may be implemented as a system, method or computer program product that can be executed on local and / or remote computing devices. Accordingly, aspects of the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.) or an embodiment combining software and hardware aspects. In addition, embodiments of the present invention may take the form of a computer program product embodied in one or more computer-readable media on one or more computing devices having computer-readable program code embodied thereon. The embodiments described herein may be entirely hardware, entirely software or include both hardware and software elements. In some embodiments, in accordance with aspects of the present invention, the present invention is implemented in software, which includes but is not limited to firmware, resident software, microcode, etc.
[0032] Any combination of one or more computer-readable media may be utilized. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. Other examples of a computer-readable storage medium may include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any combination thereof. In this document, a computer-readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with a computing system, apparatus, or device.
[0033] The program code implemented on a computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, fiber optic cable, etc. or any combination thereof. The computer program code for performing the operations of the aspects of the present invention can be written in any combination of one or more programming languages, including but not limited to any general-purpose programming language (e.g., PHP, Java, C++, etc.) and / or domain-specific programming language (e.g., HTML, SQL, etc.), blockchain-specific programming language (e.g., solidity, rust, Java, python, etc.). The program code can be executed entirely on the user's computer / mobile device, partially on the user's computer / mobile device, executed as standalone software, partially on the user's computer / mobile device and partially on a remote computer / mobile device, executed entirely on a remote computer or server, and / or executed using a blockchain. The remote computer can be connected to the user's computer through any type of network (e.g., local area network (LAN), wide area network (WAN), connection to an external computer (e.g., using an Internet service provider via the Internet), etc.).
[0034] Aspects of the present invention will now be described with reference to the flowchart illustrations and / or block diagrams of methods, systems, and computer program products according to embodiments of the present invention. Note that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions.
[0035] These computer program instructions can be sent to a processor of any type of computing system (e.g., a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine), such that the instructions executed by the processor of the computing system create a means for implementing the functions / instructions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer program instructions can also be stored in a computer-readable medium, which can direct any computing device to act in a particular manner, such that the instructions stored in the computer-readable medium produce an article of manufacture including instructions for implementing the functions / instructions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0036] The computer program instructions can also be loaded onto a computer, mobile device, other programmable data processing device, or other device to cause a series of operational steps to be performed on any computing system to produce a computer-implemented process, such that the instructions executed on the computer or other programmable device provide a process for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0037] A computer-readable signal medium can include a propagated data signal (e.g., a baseband, a portion of a carrier wave, etc.) that contains computer-readable program code. Such a propagated signal can take any of a variety of forms, including but not limited to electromagnetic, optical, or any combination thereof. A computer-readable signal medium can be any computer-readable medium that is not a computer-readable storage medium and can communicate, propagate, or transmit a program for use by or in connection with a computing system, apparatus, or device.
[0038] A data processing system suitable for storing and / or executing program code can include at least one processor directly or indirectly coupled to memory elements through a system bus. The memory elements can include local memory, mass storage, and cache memory employed during actual execution of the program code, and the cache memory provides temporary storage of at least some program code to reduce the number of times code is retrieved from mass storage during execution. Input / output or I / O devices (including but not limited to keyboards, displays, pointing devices, etc.) can be coupled to the system directly or through an intermediate I / O controller.
[0039] A network adapter can also be coupled to the system to enable the data processing system to be coupled to other data processing systems, remote printers, storage devices, blockchains, etc. through an intermediate private or public network. Modems, cable modems, and Ethernet cards are just a few of the currently available types of network adapters.
[0040] The flowcharts and block diagrams in the figures illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in the flowchart or block diagram can represent a module, segment, or portion of code that includes one or more executable instructions for implementing the specified logical function, and in some alternative implementations of the present invention, the functions noted in the blocks may not occur in the order noted in the figures. For example, depending on the functionality of a particular embodiment, two consecutive blocks shown may actually be executed substantially simultaneously, sometimes in reverse order, or in any other order.
[0041] It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by a special-purpose hardware system that performs a specific function / action, or by a combination of special-purpose hardware and computer instructions in accordance with the principles herein.
[0042] Now referring to the drawings, where like numbers represent like or similar elements, and first referring to Figure 1, an exemplary processing system 100 for code conversion to which the principles of the present invention can be applied is illustratively depicted according to an embodiment of the present invention. The processing system 100 may include at least one processor (CPU) 104 operably coupled to other components via a system bus 102. A cache 106, a read-only memory (ROM) 108, a random access memory (RAM) 110, an input / output (I / O) adapter 120, a sound adapter 130, a network adapter 140, a user interface adapter 150, and a display adapter 160 may be operably coupled to the system bus 102.
[0043] A first storage device 122 and a second storage device 124 may be operably coupled to the system bus 102 via the I / O adapter 120. The storage devices 122 and 124 may be any one of a disk storage device (e.g., a magnetic disk or an optical disk storage device), a solid-state magnetic device, etc. The storage devices 122 and 124 may be the same type of storage device or different types of storage devices.
[0044] A speaker 132 may be operably coupled to the system bus 102 via the sound adapter 130. According to the present invention, the speaker 132 may be used to provide a sound alert or some other indication related to flexible battery charging. A transceiver 142 may be operably coupled to the system bus 102 via the network adapter 140. A display device 162 may be operably coupled to the system bus 102 via the display adapter 160.
[0045] A first user input device 152 and a second user input device 154 may be operably coupled to the system bus 102 via the user interface adapter 150. The user input devices 152, 154 may be any one of a keyboard, a mouse, a keypad, an image capture device, a motion sensing device, a microphone, a device having the functions of at least two of the foregoing devices, etc. Of course, other types of input devices may also be used while maintaining the spirit of the present invention. The user input devices 152, 154 may be the same type of user input device or different types of user input devices. The user input devices 152, 154 may be used to input information to and output information from the system 100. According to an aspect of the present invention, the system 100 may include a system dependency graph builder / traverser / stack popper in block 156, and a code transformer / converter / generator in block 164, which will be described in further detail hereinafter.
[0046] Of course, as would be readily appreciated by those skilled in the art, the processing system 100 may also include other elements (not shown), and certain elements may be omitted. For example, as would be readily understood by those of ordinary skill in the art, various other input devices and / or output devices may be included in the processing system 100, depending on the particular implementation of the processing system 100. For example, various types of wireless and / or wired input and / or output devices may be used. Additionally, as would be readily understood by those of ordinary skill in the art, additional processors, controllers, memories, etc. may be utilized in various configurations. Given the teachings of the present invention provided herein, these and other variations of the processing system 100 would be readily appreciated by those of ordinary skill in the art.
[0047] In addition, it should be understood that the systems 200, 300, 400, 500, 800, 900, 1000, 1100, 1200, and 1300 described below with respect to Figure 2 , 3 , 4, 5, 8, 9, 10, 11, 12, and 13 are systems for implementing various embodiments of the present invention. Portions or all of the processing system 100 may be implemented in one or more elements of the systems 200, 300, 400, 500, 800, 900, 1000, 1100, and 1200 of Figure 2 , 3 , 4, 5, 8, 9, 10, 11, 12, and 13 respectively.
[0048] Furthermore, it should be understood that the processing system 100 may perform at least a portion of the methods described herein, including, for example, at least a portion of the methods 200, 300, 400, 500, 600, 700, and 1300 for Figure 2 , 3 , 4, 5, 6, 7, and 13 respectively. Similarly, Figure 2 , 3 , 4, 5, 8, 9, 10, 11, 12, and 13 respectively, portions or all of the systems 200, 300, 400, 500, 800, 900, 1000, 1100, 1200, and 1300 may be used to perform Figure 2 , 3 , 4, 5, 6, 7, and 13 respectively, at least a portion of the methods 200, 300, 400, 500, 600, 700, and 1300.
[0049] Now referring to Figure 2 , a diagram illustrative of a neural network (NN) system and method 200 for code transformation optimized using a large language model (LLM) is depicted in accordance with an embodiment of the present invention.
[0050] In various embodiments, at block 202, the input may represent the initial code (or other type of data) to be transformed. According to aspects of the present invention, such input 202 may be any of a variety of programming languages. Here, the source code may be tokenized into a format that can be efficiently processed by a neural network, transforming the raw code into a structured sequence of tokens encapsulating both syntax and semantics. At this stage, the code may not just be raw text, but an array of carefully tokenized data points, each representing a discrete, recognizable element of the programming language from which the system can glean semantic and syntactic meaning.
[0051] After the input at block 202, an encoder 204 can be utilized to process the input 202. The encoder 204 may not be a single entity, but rather a complex of multiple transformer layers 203, …, 205 that can serve as the backbone of a transformer-based architecture. Their role can be to break down the input into its constituent parts and apply a self-attention mechanism to determine the context and relationships within the input sequence. Each layer 203, …, 205 can refine this understanding, passing a richer, more context-aware representation to the next layer, ultimately arriving at a comprehensive encoded state.
[0052] In some embodiments, at block 206, the state may capture the essence of the source code as understood by the encoder. It can be a refined representation, a vector embodying the collective understanding of the structure and intent of the code informed by the attention and analysis layers it has passed through. This state can be a bridge between the original complexity of the input 202 and the desired clarity of the output 210. State 706 can be an encoded version of the input, enriched with context information flowing from the encoder to the decoder. This state can serve as a comprehensive representation of the source code, carrying all the necessary information for accurate transformation.
[0053] In various embodiments, the decoder at block 208 may structurally mirror the encoder, but not functionally. The decoder 208 may include a new stack of transformer layers 205, …, 207, where each layer can perform the dual task of decoding the state into the target language and refining the output in real time. According to aspects of the present invention, the decoder layers may contain cross-attention modules, which can enable the system to juxtapose the context of the source code with the emergent transformation, ensuring that each new token produced is consistent with what has come before. The decoder at block 208 may mirror the structure of the encoder, but with additional mechanisms to handle the output of the encoder. It can progressively construct the target sequence (the transformed code). The decoder layers, including the attention mechanism 212 and the feed-forward network 216, can generate predictions for each token of the output sequence based on the encoder's representation and what has been generated so far.
[0054] In some embodiments, the end result of the process can be the output in block 210. This is a tangible result of the system's process, including, for example, the output now forming a token sequence of the converted code. The output can be the final word of the decoder 208, a piece of code translated into a new language, etc., which can then be compiled, executed, or further refined. The output 210 can include a token sequence representing the code in the target language generated by the decoder 208. According to aspects of the present invention, the output 210 can result in an accurate transformation of the input code both syntactically and semantically.
[0055] According to various embodiments, the expanded view 209 of the transformer layers 201, 203, 205, 207 illustratively depicts the internal workings of both the encoder 204 and decoder 208 layers. This detailed breakdown reveals further details of the transformer layers 201, 203, 205, 207, which can include an attention mechanism 212, add and normalize functions 214, 218, and a feed-forward network 216. Each function within this microcosm can contribute to transforming the encoded state into a translation sequence.
[0056] The attention mechanism 212 within these layers can be the core of the transformer model, which can enable the network to focus on different parts of the input sequence and weight them. This mechanism is important for code translation, where dependencies can span the entire sequence, and understanding the context is key to maintaining the integrity of the translation. When predicting each token of the output sequence, this component can also enable the model to focus on different parts of the input sequence. For code translation, given the dependencies that can span the entire codebase, this can mean that the attention mechanism helps the model align code segments with their corresponding translations.
[0057] The add and normalize functions 214, 218 can be layer mechanisms for stabilizing the output, ensuring that values do not escalate to extreme values and that gradients do not vanish or explode, which are common problems in deep neural networks. The add and normalize functions 214, 218 can apply residual connections, followed by layer normalization. This process helps to stabilize learning and allows for deeper networks by alleviating the vanishing gradient problem.
[0058] In some embodiments, the feed-forward network 216 can represent a series of linear transformations with activation functions, which can be used to sequentially process data, transforming the complex relationships and patterns identified by the attention mechanism into a new representation suitable for output generation. According to aspects of the present invention, the feed-forward network 216 can further transform the output of the attention mechanism before passing it to the next layer or, in the case of the decoder, generating the output token.
[0059] This transformer-based neural network architecture enhanced by the vast knowledge and context awareness provided by the LLM represents a significant leap in the field of code translation. The architecture is designed not only to accurately translate code but also to understand its broader context, thanks to the extensive pre-training on a large corpus of code in various programming languages via the LLM. The system is adaptive, capable of direct translation, and can be further refined through prompt tuning and fine-tuning, leveraging the LLM's ability to learn from examples and improve its translations. Thus, according to aspects of the present invention, the transformer-based neural network serves as a robust and sophisticated solution for complex code translation tasks.
[0060] In various embodiments, when integrated with the LLM, the architecture of the transformer utilizes a large number of parameters and pre-training on a wide code corpus. The LLM, which has learned the patterns, structures, and semantics of multiple programming languages, guides the training process of the transformer network for the specific task of code translation. The pre-trained knowledge base of the LLM significantly improves the transformer's ability to understand and translate complex code constructs by providing a broad understanding of programming language syntax and semantics. The combination of the transformer's architecture with the extensive pre-training of the LLM enables more accurate and contextually relevant code translation compared to what might be achieved by using either component alone.
[0061] In the complete translation process, the source code is first tokenized and passed through the encoder. Then, the output of the encoder is used as guidance for the decoder, which generates the translated code. Throughout this process, the pre-trained knowledge of the LLM helps to accurately predict tokens that are semantically correct and syntactically valid in the target programming language.
[0062] The ability of the LLM to provide context learning (where the LLM can understand and generate code based on a few examples) can be particularly beneficial, as the LLM allows for few-shot learning, where the model can efficiently translate even with a limited number of examples from the source and target languages. This is at least partially due to the transformer's ability to utilize the broad context understanding built into the LLM. In summary, according to aspects of the present invention, the combination of the transformer architecture and the LLM creates a powerful model for code translation, capable of capturing and leveraging the complexity of programming languages to produce accurate and efficient translations.
[0063] Now referring to Figure 3 , FIG. 300 illustratively depicts a system and method for model adaptation and prompt tuning in language model training for neural network-based code translation, according to an embodiment of the present invention. Figure 3 Two different methodologies for adapting a pre-trained neural network model with 11 billion parameters (11B parameters) to perform a specific task are shown, according to aspects of the present invention.
[0064] In various embodiments, block 301 depicts a model adaptation method where separate models are trained for different tasks and different models are calibrated separately for a specific task, with each model utilizing a complete set of the parameters of a pre-trained model. In block 303, a prompt adaptation method according to aspects of the present invention is depicted, which simplifies the adaptation process by adjusting a small subset of the model parameters for various tasks, as will be described in further detail below.
[0065] In various embodiments, the model adaptation in block 301 can include task-specific batches for processing. Blocks 302, 304, and 306 respectively represent task-specific data batches for tasks A, B, and C. According to aspects of the present invention, for illustrative purposes, each batch contains different examples, including a1 and a2 for task A batch 302, b1 for task B batch 304, and c1 and c2 for task C batch 306, which can be used to fine-tune a dedicated instance of a neural network model.
[0066] Block 302 shows the dataset for task A, which contains unique examples a1, a2 that represent the specific nuances and requirements of task A. The data within this batch can be curated to encapsulate the diversity of the programming scenarios that task A is expected to encounter. Similarly, block 304 can be filled with data example b1 for task B, which reflects the different characteristics and challenges associated with that specific transformation task. Block 306 reserved for task C includes examples c1, c2, ensuring that the dataset comprehensively represents the breadth of the domain of task C, which can provide a robust training basis for a model dedicated to that task.
[0067] It should be understood that although three (3) batches are shown for simplicity of illustration, any number of batches can be employed according to various embodiments of the present invention.
[0068] In various embodiments, a pre-trained machine learning model characterized by 11 billion parameters is illustratively depicted in block 305. Such extensive parameterization enables the model to have a broad knowledge base applicable to a wide array of tasks. The pre-trained model can be used as a starting point for further task-specific tuning processes. Its architecture can be highly adaptive, allowing for subsequent refinement through both model adaptation and prompt tuning techniques. The model is capable of understanding and processing complex data patterns and can aid in enabling the system's ability to adapt after performing a specific task. Blocks 308, 310, and 312 illustrate separate task-specific models for post-tuning of tasks A, B, and C. The task-specific models 308, 310, and 312, labeled as the task A model, task B model, and task C model respectively, can represent the result of the model adaptation process. Each model that is now fine-tuned can embody the complexity of its corresponding task and can be optimized to perform with high fidelity within its designated domain. Notably, in accordance with aspects of the present invention, these models can retain the full set of parameters of the original pre-trained model but can be optimized for their respective tasks.
[0069] In various embodiments, in accordance with aspects of the present invention, prompt tuning in block 303 can simplify the adaptation process by adjusting a small subset of the model parameters for various tasks. Block 314 depicts a mixed task batch, which can combine data samples from all tasks and be used as a comprehensive input to the pre-trained model during the prompt tuning process. Block 316 presents an innovative prompt embedding strategy, where prompt tokens can be concatenated with a set of tunable context tokens to create a rich input that can be processed by the pre-trained model. This technique can enable the model to retain its large parameter set while being fine-tuned for task-specific outputs and can provide a more efficient utilization of the pre-trained model's knowledge by selectively adjusting a subset of the parameters associated with each task.
[0070] In some embodiments, at the heart of the prompting adjustment strategy, box 314 presents a mixed task batch, an aggregation of examples from all tasks A, B, and C. The method can leverage the diversity of the collective dataset to inform the prompting adjustment process. To advance this technique, box 316 can include performing prompt embedding, where the prompt tokens for each task can be merged with a set of tunable context tokens. This enriched input can interact dynamically with the parameters 318 of the pre-trained model, steering the output of the model towards the desired translation task objective. In box 318, the 'pre-trained model (11B parameters)' can be used as an advanced processing unit capable of fine-tuning its response based on a diverse array of input signals. This box can interact with the "mixed task batch 316", receiving a compilation of various task prompts designed to guide the model through a specialized adjustment process. During this prompting adjustment phase, the model can dynamically adjust its internal parameters, which can be represented by a parameter count of 11 billion, to enhance its ability to interpret and process the complex nuances of mixed task inputs. The interaction between the pre-trained model and the task prompts can provide a refined transformation from source code to target language, leveraging extensive pre-training to accommodate a wide range of programming tasks and ensure a high degree of accuracy and context relevance of the output. According to aspects of the present invention, the model 318 can maintain its original extensive parameter set, but can demonstrate improved performance with fewer resource requirements in task-specific performance through updated context tokens.
[0071] Now referring Figure 4 , FIG. 400 illustratively depicts a system and method for efficient code translation using an exemplary frozen language model having frozen weights and learnable weights, in accordance with an embodiment of the present invention.
[0072] In various embodiments, a frozen language model 402 can be used for code translation tasks. In this embodiment, the language model (LM) 402 can be pre-trained with a fixed or "frozen" set of parameters to retain its initial training on a large data corpus. The fixed parameters are indicated by numerals 404 and 408, which represent the frozen weights of the LM 402.
[0073] The LM 402 can interface with a target 410, which can represent the expected output of the language model. This output can be generated after processing input data through the model, where the target is a sequence of tokens in a programming language. In some embodiments, learnable weights 405 can be included within the model and can be uniquely identified and adjusted. These weights represent parameters that can be fine-tuned to adapt the output of the LM 402 to the specific requirements of the translation task, even while the remaining weights of the model remain unchanged. According to aspects of the present invention, such fine-tuning can achieve a degree of customization and adaptability without the need to retrain the entire model, thereby saving computational resources and time.
[0074] In various embodiments, each input parameter 401, 403, 405, 407, 409, 411, 413, 415 can represent an embedding of an input token, which can be processed by the frozen weights of the LM 402. The model can utilize these parameters to interpret the input data and transform it into a form that can be processed to generate the desired output. The weights W represented by boxes 401, 403, 407, 409, 411, 413, and 415 respectively p1 , W p2 , ……W pm-1 , W pm , W x1 , W x2 , ……w n-1 can be pre-established parameters of the LM that encode language information. Since they are "frozen", these weights do not undergo further modification and maintain the learned representations. Each of these parameters corresponds to a specific token from the input prompt fed into the LM. They can be embedded into a vector space by the model, which allows the LM to process and understand the input.
[0075] In practice, the LM can obtain the input tokens, process them by freezing the weights 404, 408 and with the help of the learnable weight Wu 405, and can produce an output that is a transformed version of the input data. This weight 405 can be adjusted during the prompt tuning process to better align the output of the model with a specific transformation task without affecting the integrity of the rest of the model. This output can be represented by a sequence of weights 417, 419, 421, 423 that direct the target 410. These weights can correspond to the final layer of the model, which shapes the transformed sequence into its final form. These weights associated with the output layer of the LM can be used to generate the final transformed output by transforming the processed information from the internal layers of the LM into a structured sequence corresponding to the target language.
[0076] As Figure 4 shown, the configuration of the LM 402 allows for efficient code transformation by leveraging the general knowledge of the pre-trained model, while also providing the flexibility to incorporate task-specific nuances through the learnable weight 405. This represents a significant innovation in the fields of machine learning and language processing, enabling more accurate and contextually relevant transformations without the overhead of retraining the model from scratch.
[0077] In various embodiments, the interaction of these components can include input tokens represented by their corresponding weights (e.g., 401, 403, 407, etc.), which can be fed into the LM 402. The LM 402 can use its frozen weights to process these inputs to generate an internal representation of the input data. The learnable weights 406 can be used to fine-tune the model's response to a specific input prompt, improving the accuracy of the output without changing the underlying knowledge encoded by the frozen weights. In some embodiments, next, the LM 402 can generate an output sequence, which can be a translation of the input prompt into the target language. According to aspects of the present invention, this sequence can be represented by output weights represented by blocks 417, 419, 421, and 423, which can also be part of the frozen parameters of the model, ensuring that the translation follows the syntactic and semantic structure of the target language.
[0078] Now refer Figure 5 , a diagram illustrating a system and method 500 for performing neural network-based transformation of programming code using IR and normalizing source code for accurate transformation by a language model is illustratively depicted according to an embodiment of the present invention.
[0079] In various embodiments, at block 502, code can be input for processing and can represent the original source code to be transformed. The code can be processed through static analysis to discern its structure and semantics before it is fed into the transformation framework. This block is a repository of the original code that will undergo transformation and can include the necessary instructions, declarations, and constructs that define the operational logic and functionality of the program in its original language. At block 504, the original source code from block 502 can be processed using a static analysis framework. This framework can apply static analysis techniques to parse the code and construct an intermediate representation (IR), which abstracts high-level language details and extracts the code into a form that can be systematically analyzed and manipulated. Block 506 introduces a static analysis tool framework (e.g., Abstract Syntax Tree (AST) framework, parse tree, code tree, WALA projection (Cast) entity, etc.), which can represent an abstraction layer for transforming the input code into an intermediate representation (IR) using the static analysis tool framework. Note that the above frameworks are presented for illustrative purposes and other similar frameworks can be employed according to aspects of the present invention. This entity 506 can be used to understand the syntax and semantics of the source language, and it can facilitate the extraction of structural and behavioral patterns within the code.
[0080] In various embodiments, in block 508, nodes can be derived from the static analysis entity 506. These nodes 508 can represent the basic elements of the program's control flow and data structures, which are broken down into granular components that can be analyzed and transformed individually. In block 510, metadata can be used in conjunction with the nodes to provide additional information about each node, including, for example, data types, scopes, and variable dependencies. This metadata can be used to ensure that the transformed code retains the functionality and logic of the original code.
[0081] In block 512, the nodes 508 and metadata 510 can be used to construct a program architecture to build an architectural version of the target program. This architecture can form a blueprint for the transformed code and can include placeholders for logic and data that can be filled in by the LLM. The target language architecture in block 514 can represent the structured format of the transformed code as it begins to take shape. Here, the basic elements of the target program can be laid out, ready to be filled with the actual code generated by the LLM. In block 516, transformation context placeholders can be strategically positioned within the target language architecture 514. In accordance with aspects of the present invention, these placeholders can be filled with contextually relevant code snippets to ensure that the transformation is not only syntactically correct but also functionally equivalent to the original code.
[0082] In various embodiments, in accordance with aspects of the present invention, in block 518, an intermediate representation (IR) can be utilized, and in some embodiments, it can be enhanced by converting the IR to single static assignment (SSA) IR form, which can simplify the transformation by providing a clear and explicit representation of variable assignments and dependencies. In block 520, a system dependence graph (SDG) can be generated, and the dependencies within the code can be mapped out, providing a visual representation of the execution flow and data relationships. This graph can be utilized to ensure the logical coherence of the transformed code.
[0083] In block 522, a graph traversal can be performed by systematically navigating the SDG, which can identify the sequence in which code segments (e.g., basic blocks) should be transformed to ensure that the transformed code reflects the expected behavior of the original program. Block 524 shows the "SDG stack pop and traversal" action, where the basic blocks identified during the graph traversal can be sequentially processed for transformation. This step can be used to maintain the order and dependencies of the components of the program. Blocks 526 and 528 represent the SDG stack, which can be a dynamic structure that holds the basic code blocks as they are processed. In accordance with aspects of the present invention, the stack can be utilized to ensure that each block is transformed in the correct order, and once transformed, the block can be pushed back onto the program stack to reconstruct the target program.
[0084] In various embodiments, system 500 ensures that the converted code blocks maintain syntactic and semantic integrity during code conversion. Stack management can be dynamic, allowing the stack to be refilled with the converted segments, which can then be integrated back into the target program stack. This management can be used to preserve the original program structure and ensure functional correctness.
[0085] In some embodiments, the described framework can be source and language agnostic, meaning that the entity and context information provided to the LLM can be uniform and replicable across different programming languages. The use of static analysis produces verifiable transformation guarantees for large codebases, providing determinism and functional correctness in the converted code. By constructing a system dependency graph and traversing it to produce concrete and deterministic transformation units, the framework can provide a simulation of pragmatically derived chain-of-thought reasoning, leveraging additional transformation context to enrich the IR. According to aspects of the present invention, as Figure 5 shown, the framework represents the code transformation method of the present invention, leveraging the capabilities of the LLM to manage the complexity of programming languages while ensuring that the converted code adheres to the functional equivalence of the original code.
[0086] Now refer to Figure 6 , a block diagram / flowchart illustrative of a method 600 for neural network-based code transformation using a large language model (LLM) is depicted according to an embodiment of the present invention.
[0087] In various embodiments, in block 602, source code can be received into a processing module for analysis and processing. This can include initiating a conversion process, feeding the source code into the system, and utilizing a static analysis tool. The tool parses the source code into an intermediate representation (IR), which is an important step in dissecting the source code to identify its syntactic and semantic features. This detailed analysis of the structure and meaning of the source code can be crucial for ensuring accurate conversion in later stages. In block 604, the processing system proceeds to manage the parsed source code and can transform it into nodes (e.g., projection nodes, function call nodes, control flow nodes, assignment nodes, etc.) and associated metadata. This task can be performed using a static analysis tool framework (e.g., Abstract Syntax Tree (AST) framework, parse tree, code tree, WALA projection entities, etc.), which can represent an abstraction layer that uses the static analysis tool framework to convert the input code into an intermediate representation (IR). Note that the above frameworks are presented for illustrative purposes and other similar frameworks can be employed in accordance with aspects of the present invention. The framework can analyze the IR and can use static analysis entities to derive nodes, where each node represents a different code element, such as a variable, function, or control structure. Additionally, metadata can be associated with each node, detailing the type, scope, and interrelationships of the code elements within the source code. This detailed metadata can be crucial for maintaining the logical and functional properties of the source code during the conversion process.
[0088] In various embodiments, in block 606, a program architecture construction module can be used to construct an architecture framework for the target program based on these nodes and metadata. This can include: assembling a skeletal structure of the target program that maps out the basic architecture and flow of the source code, incorporating placeholders within the architecture that mark locations for context-driven code insertion, and ensuring that the framework serves as a guide for subsequent conversion tasks to populate the target language code. In various embodiments, in block 608, the intermediate representation (IR) can be utilized, and in some embodiments, it can be enhanced by converting the IR into a single static assignment (SSA) IR form using an SSA module. In accordance with aspects of the present invention, this can include transforming the detailed IR into SSA form to simplify the conversion process, where the SSA form can simplify the IR by ensuring that each variable is assigned once, thus resolving ambiguities and facilitating more straightforward conversion by the LLM, and setting the SSA form to a clean state for the LLM to perform its conversion function.
[0089] In some embodiments, in block 610, the execution flow and dependencies of the code are mapped visually and functionally to a System Dependency Graph (SDG) module. In accordance with aspects of the present invention, this can include constructing an SDG in IR or SSAIR form that visually represents the execution flow of the source program, capturing all functional dependencies and control structures to guide the transformation process, and ensuring that the SDG accurately reflects the execution logic of the program to inform the correct ordering of transformation tasks. In block 612, the SDG can be navigated via graph traversal, and the SDG stack can be populated with basic blocks determined for transformation. In accordance with aspects of the present invention, this can include algorithmically traversing the SDG to identify the sequence in which the basic blocks of the program will be transformed, populating the SDG stack with these blocks that can represent individual functional units within the source code, and ordering the blocks for transformation in a manner that preserves the integrity of the control flow of the original program.
[0090] In block 614, the order and context of basic blocks and method segments can be managed and controlled during transformation using the SDG stack module. This can include dynamically adjusting the SDG stack as blocks are transformed and integrated back into the target program structure, supervising the context integrity of each transformed block to ensure that the target program remains functionally coherent, and applying an organized approach to re-integrate the transformed segments to maintain the original execution logic of the program.
[0091] In block 616, a Large Language Model (LLM) can be used to transform the basic blocks into the target language. In accordance with aspects of the present invention, this can include deploying the LLM to interpret and transform each basic block from the source language into the target language, using context-appropriate placeholders within the program architecture to guide the LLM to generate functionally equivalent code, and integrating the transformed segments into a coherent target program structure, where the LLM can utilize additional context and runtime dependencies from the source code environment to improve the accuracy of the transformation and the functional correctness.
[0092] Now referring to Figure 7 , a block diagram / flowchart illustrative of a method 700 for neural network-based code transformation including target language architecture generation is depicted in accordance with an embodiment of the present invention.
[0093] In accordance with aspects of the present invention, the method 700 can leverage the capabilities of advanced neural network techniques coupled with strategic context data processing. It can integrate different methodologies of "frozen and opaque" and "frozen and semi-transparent" Large Language Models (LLMs), as well as sophisticated context manipulation strategies. The result is a robust system and method 700 that is capable of providing high-fidelity transformation across various programming languages, addressing the nuanced requirements of software development and computational linguistics.
[0094] In various embodiments, in block 702, a neural network block representing a pre-configured neural network architecture can be used to initiate the transformation process. The network can be in a "frozen" state, meaning its weights are set and immutable, ensuring the stability and predictability of the initial transformation mechanism. An input signal (e.g., x) can be received and processed methodically through the network, resulting in an output (e.g., y). This output can mirror the condition of the original model before any adjustments and serve as a baseline for subsequent transformations.
[0095] In block 704, the integrity of the neural network model can be protected and maintained. In its "locked" state, the neural network can receive the same input "x" and can now be subjected to newly introduced context signals (e.g., c), which can be extracted from a repository of statically derived context data. The integration of this context signal with the input can initiate complex chain-of-thought reasoning within the LLM. This process can be used to accurately and efficiently adapt the capabilities of the model to the specific nuances of the particular transformation task being performed. In block 706, according to aspects of the present invention, a neural network (NN) training process can be initiated, which in various embodiments can include using a transformer-based NN for LLM processing, as described in further detail above with reference to Figure 2 Further described in detail.
[0096] In various embodiments, in block 708, prompt embedding and context adjustment can be initiated. This can include the merger of the original input prompt with a curated set of context tokens. According to aspects of the present invention, these tokens can traverse the processing path of the LLM as standard inputs but can be further uniquely designed to undergo selective training. Such training can be configured such that it only adjusts the parameters associated with the context tokens, thus refining the prompt response of the LLM without modifying the core neural network weights. In block 710, prompt enhancement for transformation can be performed. This can include weaving an expandable vocabulary of out-of-dictionary or non-natural language tokens into the standard prompt structure. This enhancement can significantly broaden the descriptive capabilities of the prompt, providing the LLM with a deeper and more nuanced understanding of the code context, resulting in a significant improvement in the fidelity and precision of the code transformation process.
[0097] In block 712, the output from the advanced prompt enhancement in block 710 can be captured and integrated by merging the output of the enhanced prompt (e.g., a code segment transformed with a rich understanding of both the source and target languages). The output can be transformed code that is not only syntactically transformed from the source language to the target language but also semantically rich, thus taking into account the complexity and context of the programming paradigm. According to aspects of the present invention, this transformed code can be integrated into the target program environment, ensuring that the transformed application behaves as expected in its new ecosystem, reflecting the logic, performance expectations, and operational dependencies of the original program.
[0098] Now referring to Figure 8 , an illustrative general diagram showing an exemplary neural network system 800 for neural network-based context-aware code transformation and optimization is depicted according to an embodiment of the present invention.
[0099] An artificial neural network (ANN) is an information processing system inspired by biological nervous systems such as the brain. An element of an ANN is the structure of the information processing system, which includes a large number of highly interconnected processing elements (called "neurons") that work in parallel to solve a specific problem. Additionally, a training data set is used to train the ANN, where learning involves adjusting the weights existing between the neurons. Through such a learning process, the ANN is configured for a specific application, such as pattern recognition or data classification.
[0100] Although a specific structure of the ANN is shown, which has three layers and a set number of fully connected neurons, it should be understood that this is only for illustrative purposes. In practice, the present embodiment can take any suitable form, including any number of layers and any one or more connection patterns therebetween.
[0101] ANNs demonstrate the ability to derive meaning from complex or imprecise data and can be used to extract patterns and detect trends that are too complex to be detected by humans or other computer-based systems. It is generally known that the structure of a neural network has input neurons 802 that provide information to one or more "hidden" neurons 804. The connections 808 between the input neurons 802 and the hidden neurons 804 are weighted, and these weighted inputs are then processed by the hidden neurons 804 according to a particular one in the hidden neurons 804. There can be any number of layers of hidden neurons 804, as well as neurons that perform different functions. There are also different neural network structures, such as convolutional neural networks, max-out networks, etc., which can vary according to the structure and function of the hidden layers and the weight patterns between the layers. Each layer can perform a specific function and can include convolutional layers, pooling layers, fully connected layers, softmax layers, or any other suitable type of neural network layer. Finally, a set of output neurons 806 receives and processes the weighted inputs from the last set of hidden neurons 804.
[0102] This represents a "forward pass" computation where information propagates from the input neurons 802 to the output neurons 806. When the forward pass computation is complete, the output is compared to the desired output obtainable from the training data. The error with respect to the training data is then processed in a "backpropagation" computation where the hidden neurons 804 and the input neurons 802 receive information about the error backpropagating from the output neurons 806. Once the backpropagation of error has been completed, a weight update is performed where the weighted connections 808 are updated to account for the error received. It should be noted that the three modes of operation (forward pass, backpropagation, and weight update) do not overlap with each other. This represents only one form of ANN computation and any suitable form of computation may alternatively be used.
[0103] To train the ANN, the training data can be partitioned into a training set and a test set. The training data consists of pairs of inputs and known outputs. During training, the inputs of the training set are fed into the ANN using forward pass propagation. After each input, the output of the ANN is compared to the corresponding known output. The difference between the output of the ANN and the known output associated with that particular input is used to generate an error value which can be backpropagated through the ANN, after which the weight values of the ANN can be updated. This process continues until the pairs in the training set are exhausted.
[0104] After training is complete, the ANN can be evaluated against the test set to ensure that the training does not result in overfitting. If the ANN can generalize to new inputs beyond those it has been trained on, then it is ready for use. If the ANN does not accurately reproduce the known outputs of the test set, then additional training data may be required or the hyperparameters of the ANN may need to be adjusted.
[0105] The ANN can be implemented in software, hardware, or a combination of both. For example, each weight 808 can be characterized as a weight value stored in a computer memory, and the activation function of each neuron can be implemented by a computer processor. The weight values can store any suitable data values, such as real numbers, binary values, or values selected from a fixed number of possibilities, which are multiplied by the associated neuron output. Alternatively, the weights 808 (e.g., priority list weights, attribute weights for generating writing style / personality, etc.) can be implemented as resistive processing units that generate a predictable current output when an input voltage is applied according to a settable resistance.
[0106] Now referring Figure 9 , a hardware diagram illustrative of an exemplary artificial neural network (ANN) system 900 for neural network-based context-aware code transformation and optimization in accordance with an embodiment of the present invention is depicted.
[0107] It should be understood that this architecture is purely exemplary and other architectures or types of neural networks may alternatively be used. The hardware embodiments described herein are included to illustrate in a high-level generality the general principles of neural network computations and should not be construed as limiting in any way.
[0108] Furthermore, the layers of neurons and the weights connecting them described below are described in a general manner and may be replaced by any type of neural network layer with any appropriate degree or type of interconnectivity. For example, the layers may include convolutional layers, pooling layers, fully connected layers, softmax layers, or any other appropriate type of neural network layer. Additionally, layers may be added or removed as needed, and the weights described herein may be replaced with more complex forms of interconnectivity.
[0109] During the forward feedback operation, the input neurons 902 provide input voltages in parallel with the respective rows of the weights 904. In the hardware embodiments described herein, the weights 904 each have a settable resistance value such that current outputs flow from the weights 904 to the respective hidden neurons 906. Thus, the current output of the weights 904 represents a weighted input to the hidden neurons 906.
[0110] After the hardware embodiment, the current output of a given weight 904 is determined as where V is the input voltage from the input neuron 902 and r is the set resistance of the weight 904. The currents from each of the weights 904 (e.g., priority list weights, attribute weights for generating writing style / personality, etc.) are summed column-wise and flow to the hidden neuron 906.
[0111] A set of reference weights 907 have fixed resistances and combine their outputs into a reference current provided to each of the hidden neurons 906. Since conductance values can only be positive, some reference conductance is needed to encode both positive and negative values in the matrix. The current generated by the weights 904 is continuously valued and positive, and thus the reference weights 907 are used to provide a reference current, currents above which are considered to have positive values and currents below which are considered to have negative values. In software embodiments, the use of reference weights 907 is not required, where the output and weight values can be obtained precisely and directly. As an alternative to using the weights 907, another embodiment may use an array of separate weights 904 to capture negative values.
[0112] The hidden neurons 906 use the currents from the array of weights 904 and the reference weights 907 to perform some computations. The computation may be, for example, any appropriate activation function and may be implemented in hardware or in software using appropriate circuitry.
[0113] The hidden neuron 906 then outputs its own voltage to another array of weights 904 based on an activation function. This array performs its weighted calculation in the same way, where the columns of weights 904 receive voltages from their respective hidden neurons 906 to produce a weighted current output, which is summed row by row and provided to the output neuron 908.
[0114] It should be understood that any number of these stages can be achieved by inserting additional layers of arrays and hidden neurons 906. It should also be noted that some neurons can be constant neurons 909, which provide a constant output to the array. The constant neurons 909 can exist between the input neurons 902 and / or the hidden neurons 906 and are only used during the forward feedback operation.
[0115] During backpropagation, the output neuron 908 provides a voltage backwards across the array of weights 904. The output layer compares the generated network response with the training data and calculates an error. The error is applied to the array as a voltage pulse, where the height and / or duration of the pulse is modulated proportionally to the error value. In this example, the rows of weights 904 receive the voltage from the respective output neurons 908 in parallel and convert this voltage into a current, which is summed column by column to provide an input to the hidden neurons 906. The hidden neurons 906 combine the weighted feedback signal with the derivative of their forward feedback calculation and store the error value before outputting the feedback signal voltage to the corresponding columns of their weights 904. This backpropagation travels through the entire network 900 until all hidden neurons 906 and input neurons 902 have stored the error values.
[0116] The weight update process will depend on how the weights 904 are implemented. For programmable resistors including phase change materials, the input neurons 902 and the hidden neurons 906 can apply a first weight update voltage forward through the network 900, and the output neuron 908 and the hidden neurons 906 can apply a second weight update voltage backwards through the network 900. The combination of these voltages can produce a state change within each weight 904, for example by raising the temperature of the weight 904 above a threshold and thus changing its resistance, causing the weight 904 to assume a new resistance value. In this way, the weights 904 can be trained to adapt the neural network 900 to the errors in its processing.
[0117] As described above, the weight 904 can be implemented in software or hardware, such as using a relatively complex weighting circuit or using a resistive cross-point device. Such a resistive device can have switching characteristics that have non-linearity that can be used to process data. The weight 904 can belong to a class of devices known as a resistive processing unit (RPU). The RPU device can be implemented using resistive random access memory (RRAM), phase change memory (PCM), programmable metallization cell (PMC) memory, or any other device having non-linear resistive switching characteristics. Such an RPU device can also be considered a memristive system.
[0118] Now referring to Figure 10 , and continuing to refer to Figure 9 , a block diagram illustrative of an exemplary neuron 1000 in a neural network system for context-aware code conversion and optimization based on a neural network, in accordance with an embodiment of the present invention, is depicted.
[0119] In various embodiments, the neuron can represent any one of an input neuron 902, a hidden neuron 906, or an output neuron 908, as Figure 9 shown. It should be noted that Figure 10 shows components that address all three operational phases: forward feedback, backpropagation, and weight update. However, since the different phases do not overlap, there must necessarily be some form of control mechanism within the neuron 1000 to control which components are active. Accordingly, it should be understood that, in accordance with aspects of the present invention, there can be switches and other structures not shown in the neuron 1000 to handle the switching between modes.
[0120] In the forward feedback mode, the difference block 1002 determines the value of the input from the array by comparing it to a reference input. This sets the magnitude and sign (e.g., + or -) of the input from the array to the neuron 1000. Block 1004 performs a calculation based on the input, and its output is stored in the storage device 1005. It is specifically contemplated that block 1004 computes a non-linear function and can be implemented as an analog or digital circuit, or can be executed in software. The value determined by the functional block 1004 is converted to a voltage at the forward feedback generator 1006, which applies the voltage to the next array. Signals propagate in this manner through multiple layers of the array and neurons until they reach the final output layer of the neuron. In block 1008, the input is also applied to the derivative of the non-linear function, and its output is stored in the memory 1009.
[0121] During the backpropagation mode, an error signal is generated. The error signal can be generated at the output neuron 908, or can be computed by a separate unit that receives the input from the output neuron 908 and compares the output with the correct output based on the training data. Otherwise, if neuron 1000 is a hidden neuron 906, it receives backpropagation information from the array of weights 904 and compares the received information with a reference signal at the difference box 1010 to provide a continuously valued signed error signal. This error signal is multiplied by the derivative of the non-linear function from the previous forward feedback step stored in the memory 1009 using the multiplier 1012, and the result is stored in the memory 1013. The value determined by the multiplier 1012 is converted into a backpropagation voltage pulse proportional to the computed error at the backpropagation generator 1014, which applies the voltage to the previous array. The error signal propagates through multiple layers of the array and neurons in this way until it reaches the input layer of neuron 902.
[0122] During the weight update mode, after both the forward and backward passes are completed, each weight 904 is updated in proportion to the product of the signals passing through the weights during the forward and backward passes. The update signal generator 1016 provides voltage pulses in both directions (but note that for the input and output neurons, only one direction will be available). According to aspects of the present invention, the shape and amplitude of the pulses from the update generator 1016 are configured to change the state of the weights 904 (e.g., priority list weights, attribute weights for generating writing style / personality, etc.) such that the resistance of the weights 904 is updated.
[0123] Now refer to Figure 11 , a diagram illustratively depicting an exemplary hierarchical neural network system 1100 in a neural network for context-aware code transformation and optimization based on a neural network according to an embodiment of the present invention is shown.
[0124] In a hierarchical neural network, the nodes are arranged in layers. An exemplary simple neural network has an input layer 1120 of source nodes 1122, and a single computational layer 1130 with one or more computational nodes 1132 that also act as output nodes, where there is a single computational node 1132 for each possible class into which the input example can be classified. The input layer 1120 can have a number of source nodes 1122 equal to the number of data values 1112 in the input data 1110. The data values 1112 in the input data 1110 can be represented as a column vector. Each computational node 1132 in the computational layer 1130 generates a linear combination of weighted values based on the input data 1110 fed into the input nodes 1120, and applies a non-linear activation function that is differentiable with respect to the sum. The exemplary simple neural network can perform classification on linearly separable examples (e.g., patterns).
[0125] A deep neural network, such as a multi-layer perceptron, may have an input layer 1120 of source nodes 1122, one or more computational layers 1130 with one or more computational nodes 1132, and an output layer 1140, where for each possible class into which an input example can be classified, there is a single output node 1142. The input layer 1120 may have a number of source nodes 1122 equal to the number of data values 1112 in the input data 1110. The computational nodes 1132 in the computational layer 1130 may also be referred to as hidden layers because they are between the source nodes 1122 and the output nodes 1142 and are not directly observable. Each node 1132, 1142 in the computational layer generates a linear combination of weighted values based on the values output from the nodes in the previous layer, and applies a non-linear activation function that is differentiable over the range of the linear combination. The weights applied to the values from each previous node may be represented, for example, by w1, w2, … w n-1 , w n . The output layer provides the overall response of the network to the input data. The deep neural network may be fully connected, where each node in the computational layer is connected to all other nodes in the previous layer, or may have other connection configurations between the layers. If a link between nodes is lost, the network is referred to as partially connected.
[0126] Training a deep neural network may involve two phases, a forward phase where the weights of each node are fixed and the input is propagated through the network, and a backward phase where the error values are propagated backward through the network and the weight values are updated.
[0127] The computational nodes 1132 in one or more computational (hidden) layers 1130 perform a non-linear transformation on the input data 1112 to generate a feature space. Classes or categories may be more easily separable in the feature space than in the original data space.
[0128] In various embodiments, the present invention may customize the architecture and training process of a neural network to optimize the context-aware transformation of programming code. This involves using a hierarchical network to complexly map and interpret syntax and semantics from one programming language to another, thereby improving the accuracy and efficiency of code transformation. The ability of a deep neural network to discern subtle pattern differences in data through multiple computational layers allows for fine-tuning during the training phase, ensuring that the output of the model aligns with a complex coding framework and facilitating the development of neural network-based code transformation techniques.
[0129] Now referring Figure 12 , a block diagram illustrating a system 1200 for context-aware code transformation and optimization based on a neural network is illustratively depicted in accordance with an embodiment of the present invention.
[0130] In various embodiments, in block 1202, a source code input device is depicted, which serves as an initial interface for receiving source code to be transformed. It can process the raw code input to prepare for subsequent analysis and transformation phases.
[0131] In block 1204, a context data processor can serve as a mediator for enriching the source code with context information. This processor can enhance the code by embedding relevant data (such as comments, documentation, and metadata) that may influence the transformation process, ensuring a more accurate and context-aware transformation output. The transformation control unit, represented by block 1206, can control and orchestrate the entire transformation workflow. It can direct the source code being processed to the appropriate components within the system, manage the interactions between the neural network trainer, prompt embedder, and other key elements to facilitate seamless transformation operations.
[0132] In block 1208, the neural network trainer is the component responsible for the adaptive learning aspect of the system. It can use the provided context information to fine-tune the neural network parameters, adapting the model to the specific nuances and requirements of the source code being transformed. In block 1210, the prompt embedder can integrate policy prompts into the transformation process. These prompts can guide the focus of the neural network during transformation, efficiently steering the output of the model towards the desired target language constructs. The program architecture builder in block 1212 can construct a base template for the target program structure. It can create an architectural framework that outlines the main components and flow of the target program, setting placeholders where the transformed code can ultimately be integrated.
[0133] In block 1214, the prompt enhancement processor can use additional context tokens to enhance the input prompts. These tokens can be utilized to improve the language model's understanding of the code context, resulting in a transformation with higher fidelity and relevance to the target programming language. In various embodiments, an NN (e.g., a transformer-based NN architecture) 1216 can be integrated with an LLM such that the architecture of the transformer can leverage a large number of parameters and pre-training on a wide code corpus. According to aspects of the present invention, an LLM that has learned the patterns, structures, and semantics of multiple programming languages can guide the training process of the transformer network for the specific task of code transformation.
[0134] The pre-trained knowledge base of the LLM can significantly improve the transformer's ability to understand and transform complex code constructs by providing the transformer with a broad understanding of programming language syntax and semantics. According to aspects of the present invention, the combination of the structure of the transformer with the extensive pre-training of the LLM can provide a more accurate and contextually relevant code transformation compared to what may be achievable by using either component alone.
[0135] In block 1218, an IR generator can generate an intermediate representation (IR) of the source code. This IR can be a standardized format for abstract code, making it easier and more efficient for the system to analyze and transform across all different programming languages. The transformation optimizer in block 1220 can apply various algorithms to refine the transformed code. It ensures that the output not only maintains functional equivalence but also adheres to the idiomatic nuances and performance considerations of the target language.
[0136] In block 1222, an SDG builder / stack popper and traverser can be used to construct a system dependence graph (SDG) that maps out the dependencies and execution flow of the source code. It can also manage the traversal and transformation of code segments, maintaining the logic and execution integrity of the original program.
[0137] In block 1224, an output integration and verification unit can be used to integrate the code into the target program architecture. It can verify the syntactic and semantic accuracy for the transformed segments, ensuring that the final program is not only correct but also optimized for the target execution environment.
[0138] As an illustrative example, according to aspects of the present invention, in some embodiments, a high-level representation of one-shot learning is depicted as Algorithm 1. A simple task description can be provided to the model, and a representative example of successfully completing the task is depicted below:
[0139] Algorithm 1: One-Shot Learning
[0140] “Convert Java to Python” / / Prompt
[0141] public class Foo inherits Bar => class Foo(Bar) / / Example 1
[0142] private static void spam(int x) => __________??? _________ / / Task
[0143] In some embodiments, the one-shot method can be extended to few-shot learning by providing a simple task description and a few representative samples of successfully completing the task, as shown in Algorithm 2 below:
[0144] Algorithm 2: Few-Shot Learning
[0145] “Convert Java to Python” / / Prompt
[0146] public class Foo inherits Bar => class Foo(Bar) / / Example 1
[0147] public String baz(Foo foo) => def baz(foo:Foo) -> String / / Example 2 ...
[0149] private static void spam(int x) => __________??? _________ / / Task
[0150] In some embodiments, context learning guided by partial labels (PL) can be performed according to Algorithm 3 as follows:
[0151] Algorithm 3: Context Learning Guided by Partial Labels (PL)
[0152] “Convert Java to Python” / / Hint
[0153] public class Foo inherits Bar => class Foo(Bar) / / Example 1
[0154] public String baz(Foo foo) => def baz(foo:Foo) -> String / / Example 2 ...
[0156] private static void spam(int x) => __________??? _________ / / Task
[0157] public class Foo inherits Bar / / Source 2
[0158] “create a {MODIFIER:public} {TEMPLATE:class} called {NAME:Foo}
[0159] in {TARGET:python} inheriting {INHERITS:Bar}” / / Hint 1
[0160] class Foo(Bar) / / Target 1 ...
[0162] private static void spam(int x) / / Source 2
[0163] “add a {MODIFIER:private} {TEMPLATE:static method} {NAME:
[0164] spam} with {ARGS:[{TYPE:int} x]} that {RETURNS:void}” / / Hint 2
[0165] @staticmethod def _spam(x: int) -> None / / Target 2
[0166] According to various embodiments, “one-shot” and “few-shot” learning examples demonstrate the system's advanced learning algorithms, which are integrated into the code conversion process of the present invention. The present invention enhances code conversion by adopting a system with advanced learning techniques. “One-shot” and “few-shot” learning examples are shown herein, demonstrating the system's proficiency in adapting to programming language constructs with minimal input.
[0167] In some embodiments, at block 1202, a single prompt and task pair can be presented to the system, exemplifying the “one-shot” learning ability of the model. The system can utilize this single example to understand the structure and semantics and convert it from Java to Python. The prompt “Convert Java to Python” coupled with the class inheritance example guides the model to formulate the correct conversion, encapsulating the syntax and object-oriented principles of the source code. Here, the system can utilize a single example to learn and convert the code syntax. It can infer the structural pattern from the Java class inheritance example and can apply this learning to generate an equivalent Python class, demonstrating the system's ability to capture and convert object-oriented concepts after observing only one instance.
[0168] Extending to block 1204, “few-shot” learning can be implemented, where the system assimilates several task descriptions and examples. This approach can enhance the model's understanding of the coding language and refine its prediction accuracy for code conversion. By processing a broader set of examples, the system can discern language patterns and nuances, resulting in a more robust conversion. This approach extends the system's learning by analyzing multiple examples, allowing it to identify a wider array of programming constructs. The system can extract patterns from multiple instances, leading to a more nuanced understanding and a more refined conversion output. According to aspects of the present invention, this approach can enable the system to conform to the variables found in different programming tasks and accordingly adapt its conversion mechanism.
[0169] In various embodiments, the described one-shot learning scenario and few-shot learning scenario can be integrated into a larger framework of the system, connected to block 1212, where the program architecture builder can utilize these learned patterns to construct the infrastructure of the transformed code. Similarly, in block 1214, the prompt enhancement processor can utilize the insights from these examples to enhance the prompt structure, facilitating the overall transformation efficacy of the system. These two learning examples both support the system's ability to process and transform code with high accuracy, reducing the need for extensive datasets. In accordance with aspects of the present invention, they can be used for transformation optimization of the system, enabling it to quickly adapt to new languages and coding paradigms and ensuring that the transformed code is syntactically and semantically accurate.
[0170] Now referring Figure 13 , an exemplary computing environment for performing at least some of the computer code for adaptive neural network-based context-aware code transformation and optimization is illustratively depicted in accordance with an embodiment of the present invention.
[0171] Aspects of the present disclosure are described by narrative text, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic included in computer program product (CPP) embodiments. With respect to any flowchart, depending on the technology involved, operations may be performed in an order different from that shown in a given flowchart. For example, again depending on the technology involved, two operations shown in consecutive flowchart blocks may be performed in reverse order, as a single integrated step, simultaneously, or in a manner that at least partially overlaps in time.
[0172] A computer program product embodiment (“CPP embodiment” or “CPP”) is a term in the present disclosure for any collection of one or more storage media (also referred to as “media”) collectively included in a set of one or more storage devices, which collectively include machine-readable code corresponding to instructions and / or data for performing the computer operations specified in a given CPP claim. A “storage device” is any tangible device that can retain and store instructions used by a computer processor. By way of non-limitation, a computer-readable storage medium can be an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these media include: magnetic disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanically encoded devices such as punched cards or pits / ridges formed in the main surface of a disk, or any suitable combination of the foregoing. As used in the present disclosure, the term computer-readable storage medium should not be construed as storing in the form of a transient signal per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide, optical pulses passing through an optical fiber cable, or electrical signals transmitted through wires and / or other transmission media. As will be understood by those skilled in the art, during the normal operation of a storage device, such as during access, defragmentation, or garbage collection, data typically moves at some occasional points in time, but this does not make the storage device transient because the data is not transient when it is stored.
[0173] Computing environment 1300 contains an example of an environment for executing at least some of the computer code involved in performing the methods of the present invention, such as partially label-guided context learning 1350. This partially label-guided context learning 1350 can include using partial labels and context learning for code transformation. It can include providing detailed hints with specific programming elements (e.g., class definitions, method modifiers) and their corresponding transformations. Given the nuances of programming constructs and their context within the source and target languages, this code can be used to more precisely guide the transformation process. In various embodiments, in accordance with aspects of the present invention, this can be integrated into the advanced learning capabilities of the system to achieve accurate and more context-aware transformation of programming code across different languages.
[0174] In addition to the box 1350, the computing environment 1300 also includes, for example, a computer 1301, a wide area network (WAN) 1302, an end user device (EUD) 1303, a remote server 1304, a public cloud 1305, and a private cloud 1306. In this embodiment, the computer 1301 includes a processor set 1310 (including processing circuitry 1320 and a cache 1321), a communication fabric 1311, volatile memory 1312, a permanent storage device 1313 (including an operating system 1322 and block 200, as identified above), a peripheral set 1314 (including a user interface (UI) device set 1323, a storage device 1324, and an Internet of Things (IoT) sensor set 1325), and a network module 1315. The remote server 1304 includes a remote database 1330. The public cloud 1305 includes a gateway 1340, a cloud orchestration module 1341, a set of host physical machines 1342, a set of virtual machines 1343, and a set of containers 1344.
[0175] The computer 1301 can take the form of a desktop computer, a laptop computer, a tablet computer, a smart phone, a smart watch, or other wearable computer, a mainframe computer, a quantum computer, or any other form of computer or mobile device now known or hereafter developed that is capable of running a program, accessing a network, or querying a database, such as the remote database 1330. As is well understood in the computer art and depending on the technology, the performance of a computer-implemented method can be distributed among multiple computers and / or multiple locations. On the other hand, in this presentation of the computing environment 1300, the detailed discussion focuses on a single computer, specifically the computer 1301, to keep the presentation as simple as possible. The computer 1301 can be located in the cloud, even if it is not shown in the Figure 1 cloud shown. On the other hand, unless it can be positively indicated to any extent, it is not required that the computer 1301 not be in the cloud.
[0176] The processor set 1310 includes one or more computer processors of any type now known or hereafter developed. The processing circuitry 1320 can be distributed across multiple packages, e.g., multiple coordinated integrated circuit chips. The processing circuitry 1320 can implement multiple processor threads and / or multiple processor cores. The cache 1321 is a memory located within the processor chip package and is typically used for data or code that should be made available for rapid access by threads or cores running on the processor set 1310. Cache memory is typically organized into multiple levels based on its relative proximity to the processing circuitry. Alternatively, some or all of the cache for the processor set can be located "off-chip". In some computing environments, the processor set 1310 can be designed to work with qubits and perform quantum computing.
[0177] Computer-readable program instructions are typically loaded onto computer 1301 to cause the processor set 1310 of computer 1301 to execute a series of operational steps to implement a computer-implemented method such that the instructions so executed will instantiate the method specified in the flowchart and / or the narrative description of the computer-implemented method included in this document (collectively referred to as "the method of the present invention"). These computer-readable programs are stored in various types of computer-readable storage media such as cache 1321 and other storage media discussed below. The program instructions and associated data are accessed by the processor set 1310 to control and direct the execution of the method of the present invention. In computing environment 1300, at least some of the instructions for executing the method of the present invention may be stored in block 200 of the permanent storage device 1313.
[0178] The communication structure 1311 is a signal conduction path that allows the various components of computer 1301 to communicate with each other. Generally, this structure is made up of switches and conductive paths such as those that make up a bus, a bridge, a physical input / output port, etc. Other types of signal communication paths may be used such as fiber optic communication paths and / or wireless communication paths.
[0179] The volatile memory 1312 is any type of volatile memory known now or developed in the future. Examples include dynamic type random access memory (RAM) or static type RAM. Generally, the volatile memory 1312 is characterized by random access, but this is not required unless affirmatively indicated. In computer 1301, the volatile memory 1312 is located in a single package and inside computer 1301, but alternatively or additionally, the volatile memory may be distributed across multiple packages and / or located external to computer 1301.
[0180] The permanent storage device 1313 is any form of non-volatile storage device for a computer known now or developed in the future. The non-volatility of this storage device means that the stored data is retained whether or not power is supplied to computer 1301 and / or directly to the permanent storage device 1313. The permanent storage device 1313 may be a read-only memory (ROM), but generally at least a portion of the permanent storage device allows data to be written, deleted, and rewritten. Some familiar forms of the permanent storage device include magnetic disks and solid state storage devices. The operating system 1322 may take several forms such as various known proprietary operating systems or open-source portable operating system interface type operating systems that employ a kernel. The code included in block 200 generally includes at least some of the computer code involved in executing the method of the present invention.
[0181] The peripheral device set 1314 includes the peripherals of a set of computers 1301. The data communication connection between the peripheral devices and other components of the computer 1301 can be implemented in various ways, such as a Bluetooth connection, a Near Field Communication (NFC) connection, a connection by a cable (such as a Universal Serial Bus (USB) type cable), a plug-in connection (e.g., a Secure Digital (SD) card), a connection through a local communication network, and even a connection through a wide area network (such as the Internet). In various embodiments, the UI device set 1323 may include components such as a display screen, a speaker, a microphone, wearable devices (such as goggles and smartwatches), a keyboard, a mouse, a printer, a touchpad, a game controller, and a haptic device. The storage device 1324 is an external storage device, such as an external hard disk drive, or a pluggable storage device, such as an SD card. The storage device 1324 can be persistent and / or volatile. In some embodiments, the storage device 1324 may take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments where the computer 1301 needs to have a large amount of memory (e.g., in the case of locally storing and managing a large database on the computer 1301), the memory can be provided by a peripheral storage device designed to store a very large amount of data, such as a Storage Area Network (SAN) shared by multiple geographically distributed computers. The IoT sensor group 1325 consists of sensors that can be used in Internet of Things applications. For example, one sensor can be a thermometer, and another sensor can be a motion detector.
[0182] The network module 1315 is a collection of computer software, hardware, and firmware that allows the computer 1301 to communicate with other computers via the WAN 1302. The network module 1315 may include hardware (such as a modem or a Wi-Fi signal transceiver), software for packetizing and / or depacketizing data for transmission over a communication network, and / or web browser software for communicating data over the Internet. In some embodiments, the network control function and the network forwarding function of the network module 1315 are executed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing Software-Defined Networking (SDN)), the control function and the forwarding function of the network module 1315 are executed on physically separate devices, such that the control function manages several different network hardware devices. The computer-readable program instructions for performing the methods of the present invention can generally be downloaded to the computer 1301 from an external computer or an external storage device through a network adapter card or a network interface included in the network module 1315.
[0183] The WAN 1302 is any wide area network (e.g., the Internet) that can transfer computer data over non-local distances through any technology, whether currently known or developed in the future, for transferring computer data. In some embodiments, the WAN 1302 may be replaced and / or supplemented by a local area network (LAN), such as a Wi-Fi network, that is designed to transfer data between devices located in a local area. The WAN and / or LAN typically includes computer hardware, such as copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and edge servers.
[0184] The end user device (EUD) 1303 is any computer system used and controlled by an end user (e.g., a customer of an enterprise operating the computer 1301) and may take any form discussed above in connection with the computer 1301. The EUD 1303 typically receives helpful and useful data from the operation of the computer 1301. For example, in the hypothetical case where the computer 1301 is designed to provide recommendations to an end user, the recommendation will typically be transferable from the network module 1315 of the computer 1301 to the EUD 1303 via the WAN 1302. In this way, the EUD 1303 can display or otherwise present the recommendation to the end user. In some embodiments, the EUD 1303 may be a client device, such as a thin client, a thick client, a mainframe computer, a desktop computer, etc.
[0185] The remote server 1304 is any computer system that provides at least some data and / or functionality to the computer 1301. The remote server 1304 may be controlled and used by the same entity that operates the computer 1301. The remote server 1304 represents a machine that collects and stores helpful and useful data for use by other computers, such as the computer 1301. For example, in the hypothetical case where the computer 1301 is designed and programmed to provide recommendations based on historical data, the historical data may be provided to the computer 1301 from the remote database 1330 of the remote server 1304.
[0186] A public cloud 1305 is any computer system that can be used by multiple entities and provides on-demand availability of computer system resources and / or other computing capabilities, especially data storage (cloud storage) and computing power, without direct active management by the user. Cloud computing typically utilizes resource sharing to achieve coherence and economies of scale. The direct and active management of the computing resources of the public cloud 1305 is performed by the computer hardware and / or software of the cloud orchestration module 1341. The computing resources provided by the public cloud 1305 are typically implemented by a virtual computing environment that runs on various computers of a set of host physical machines 1342, which is a collection of physical computers in and / or available for the public cloud 1305. The virtual computing environment typically takes the form of virtual machines from a set of virtual machines 1343 and / or containers from a set of containers 1344. It should be understood that these VCEs can be stored as images and can be transferred among and between various physical machine hosts as images or after instantiation of the VCE. The cloud orchestration module 1341 manages the transfer and storage of the images, deploys new instantiations of the VCE, and manages the active instantiations of the VCE deployment. The gateway 1340 is a collection of computer software, hardware, and firmware that allows the public cloud 1305 to communicate via the WAN 1302.
[0187] Some further explanations of the virtualized computing environment (VCE) will now be provided. A VCE can be stored as an "image". New active instances of the VCE can be instantiated based on the image. Two common types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to an operating system feature where the kernel allows for multiple isolated user space instances (called containers). From the perspective of the programs running within them, these isolated user space instances typically appear as real computers. A computer program running on a normal operating system can utilize all the resources of that computer, such as connected devices, files and folders, network shares, CPU capabilities, and quantifiable hardware capabilities. However, a program running within a container can only use the contents of the container and the devices allocated to the container, which is a characteristic known as containerization.
[0188] A private cloud 1306 is similar to a public cloud 1305, except that the computing resources are only available to a single enterprise. Although the private cloud 1306 is depicted as communicating with the WAN 1302, in other embodiments, the private cloud can be completely disconnected from the Internet and only accessible through a local / private network. A hybrid cloud is a combination of multiple clouds of different types (e.g., private cloud, community cloud, or public cloud type) that are typically implemented by different vendors. Each of the multiple clouds remains a separate and discrete entity, but the larger hybrid cloud architecture is tied together through standardized or proprietary means to enable orchestration, management, and / or data / application portability between the multiple constituent clouds. In this embodiment, both the public cloud 1305 and the private cloud 1306 are part of a larger hybrid cloud.
[0189] As used herein, the term "hardware processor subsystem" or "hardware processor" can refer to a processor, memory, software, or a combination thereof that cooperate to perform one or more specific tasks. In useful embodiments, the hardware processor subsystem can include one or more data processing elements (e.g., logic circuits, processing circuits, instruction execution devices, etc.). The one or more data processing elements can be included in a central processing unit, a graphics processing unit, and / or a separate controller based on a processor or computing element (e.g., logic gates, etc.). The hardware processor subsystem can include one or more on-board memories (e.g., caches, dedicated memory arrays, read-only memories, etc.). In some embodiments, the hardware processor subsystem can include one or more memories that can be on-board or off-board, or can be dedicated to use by the hardware processor subsystem (e.g., ROM, RAM, basic input / output system (BIOS), etc.).
[0190] In some embodiments, the hardware processor subsystem can include and execute one or more software elements. The one or more software elements can include an operating system and / or one or more applications and / or specific code to achieve a specified result.
[0191] In other embodiments, the hardware processor subsystem can include dedicated specialized circuitry that performs one or more electronic processing functions to achieve a specified result. Such circuitry can include one or more application-specific integrated circuits, field-programmable gate arrays, and / or programmable logic arrays.
[0192] These and other variations of the hardware processor subsystem are also contemplated in embodiments of the present invention.
[0193] The present invention can be a system, method, and / or computer program product at any possible level of integration of technical details. The computer program product can include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement aspects of the present invention.
[0194] A computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium can be, by way of example and not limitation, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed as a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through an optical fiber cable) or an electrical signal transmitted through a wire.
[0195] The computer-readable program instructions described herein can be downloaded to a respective computing / processing device from a computer-readable storage medium or can be downloaded to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the respective computing / processing device.
[0196] The computer-readable program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network connection, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit in order to perform aspects of the present invention.
[0197] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0198] These computer-readable program instructions may be provided to a processor of a computer or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium storing the instructions comprises an article of manufacture including instructions for implementing aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0199] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, such that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0200] References in the specification to "one embodiment" or "an embodiment" of the present invention and other variations thereof mean that the specific features, structures, characteristics, etc. described in connection with that embodiment are included in at least one embodiment of the present invention. Thus, the phrase "in one embodiment" or "in an embodiment" and any other variations that occur throughout the specification do not necessarily refer to the same embodiment.
[0201] It should be understood that, for example, in the case of "A / B", "A and / or B", and "at least one of A and B", the use of any one of the following, namely " / ", "and / or", and "at least one of", is intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As another example, in the case of "A, B, and / or C" and "at least one of A, B, and C", such phrases are intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or the selection of only the first and second-listed options (A and B), or the selection of only the first and third-listed options (A and C), or the selection of only the second and third-listed options (B and C), or the selection of all three options (A and B and C). As will be apparent to those of ordinary skill in the art and related fields, this can be extended to any number of items listed.
[0202] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of an instruction that includes one or more executable instructions for implementing the specified logical function. In some alternative embodiments, the functions labeled in the blocks may not occur in the order labeled in the figures. For example, two consecutive blocks shown may actually be completed as one step, executed simultaneously, executed substantially simultaneously, executed in a partially or fully time-overlapped manner, or these blocks may sometimes be executed in the reverse order, depending on the functions involved. It will also be noted that each block of the block diagrams and / or flowcharts and combinations of blocks in the block diagrams and / or flowcharts can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or implements a combination of dedicated hardware and computer instructions.
[0203] Preferred embodiments of neural network-based systems and methods for efficiently converting and transforming program code from a source language to a target language have been described (which are intended to be illustrative and not restrictive), noting that those skilled in the art can make modifications and variations based on the above teachings. Accordingly, it should be understood that changes can be made in the specific embodiments disclosed, which are within the scope of the invention as outlined by the appended claims. Aspects of the invention have been described in the detail and specificity required by patent law, and the claims and desired content protected by the patent certificate are set forth in the appended claims.
Claims
1. A method for efficiently converting program code from a source language to a target language, the method comprising: Parse the input source code into an intermediate representation IR using a processor device; building a structural and semantic model of the source code by applying static analysis to the IR; constructing a program architecture of the target code according to the IR, including generating context-aware placeholders; Transforming the IR into a single static allocation SSA form, and constructing a system dependency graph SDG from the SSA form; traversing the SDG to sort the conversion tasks, and converting the sorted tasks into the target language using a large language model (LLM); and A converted program is generated by integrating the converted code segments into a coherent program structure in the target language. 2 . The method according to claim 1 , further comprising receiving a mixed task batch code segment as the input for processing. 3 . The method of claim 1 , wherein applying the static analysis further comprises generating a node and corresponding metadata for each input code segment. 4 . The method of claim 1 , wherein the transforming comprises populating placeholders with context-sensitive transformations derived from the static analysis.
5. The method of claim 1 further comprising enriching the IR with runtime dependency data from a source computing environment.
6. The method of claim 1, wherein said converting the sequenced tasks into the target language comprises adapting the converted code segments to conform to runtime constraints of a target computing environment to enable execution of the converted program within the target environment.
7. The method of claim 1 , wherein converting the sequenced tasks into the target language comprises adapting the converted code segments to meet specific performance metrics and resource constraints of a target computing environment, the specific performance metrics and resource constraints comprising one or more of memory usage, processing speed, and integration with existing software infrastructure.
8. The method of claim 1 , further comprising displaying the converted code segment on a user interface, receiving user input for code editing, compiling the edited code, and executing the compiled code to implement a transformation of a state within a machine, wherein executing the compiled code causes the machine to perform a series of operations resulting in physical changes indicating the functionality of the code in the target language.
9. A system for efficiently converting program code from a source language to a target language, comprising: A processor device operatively coupled to a computer-readable storage medium, the processor being configured to perform the method according to any one of claims 1-8.
10. A non-transitory computer-readable storage medium comprising a computer-readable program operably coupled to a processor device for efficiently converting program code from a source language to a target language, wherein the computer-readable program, when executed on a computer, causes the computer to perform the method according to any one of claims 1-8.
11. A computer program product comprising instructions executable by a processor to cause the processor to perform the method according to any one of claims 1 to 8.
Citation Information
Cited By
Code language conversion method and device, electronic equipment, storage medium and computer program product
CN120803467A
Code language conversion method and device, electronic equipment, storage medium and computer program product
CN120803467B