Method, system, and computer-readable program (neural network-based context-aware code translation and optimization)

By parsing source code into an intermediate representation and using a large language model for translation, the method addresses the challenges of conventional code translation, achieving improved accuracy and efficiency while maintaining code integrity.

JP2025096185APending Publication Date: 2025-06-26INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024209834
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-14
Filing Date
2024-12-02
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Conventional code translation methods face challenges such as the need for large task-specific datasets, poor generalization, and the loss of nuanced language constructs, leading to semantic gaps and readability issues in translated code.

Method used

The method involves parsing source code into an intermediate representation (IR), establishing a structure and semantic model through static analysis, and using a large language model (LLM) to translate the code while maintaining context-aware placeholders and system dependency graphs to ensure accurate and efficient translation.

Benefits of technology

This approach improves the accuracy and efficiency of code translation by leveraging advanced contextual data processing and neural network models, ensuring that the translated code retains the integrity and logic of the original code.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025096185000001_ABST
    Figure 2025096185000001_ABST
Patent Text Reader

Abstract

To efficiently translate program code from a source language to a target language.SOLUTION: Input source code is parsed, using a processor device, into an Intermediate Representation (IR). A structural and semantic model of the source code are established by applying static analysis to the IR, and a program skeleton of the target code is constructed from the IR, including generating context-aware placeholders. The IR is transformed into a Single Static Assignment (SSA) form, and a System Dependency Graph (SDG) is built from the SSA form. The SDG is traversed to order translation tasks, and the ordered tasks are translated into the target language using a Large Language Model (LLM). A translated program is generated by integrating translated code segments into a coherent program structure in the target language.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to neural network-based automated code translation between programming languages, and more specifically, to using neural network models and contextual data processing to distribute control in programming language translation to a large language model (LLM) to improve the accuracy and efficiency of code translation.

Summary of the Invention

Problems to be Solved by the Invention

[0002] Conventionally, fine-tuning a pre-trained model is an essential element in the code translation task, and typically involves converting source code into a standard intermediate representation (IR) such as single static assignment (SSA), and training the model to translate from this IR to the target language. This conventional method utilizes a vast corpus of source and target code snippets to repeatedly update the model by gradient descent, and often requires extensive labeled examples for model convergence. While effective in some respects, this approach has significant drawbacks including the need for large task-specific datasets, the risk of poor generalization, and reliance on potentially error-prone features of the training data. Additionally, standard SSA-based IR simplifies the code for efficient data flow analysis and promotes compiler optimizations, but often removes nuanced language constructs essential for understanding the intent and logic of the original code. This loss of language-specific characteristics, such as object-oriented programming elements, creates a significant semantic gap that can impede the readability and translatability of the target code, along with the introduction of synthetic variables.

[0003] Moreover, the rise of large language models (LLMs) has brought about new methods for in-context learning, which, while successful in natural language processing tasks, face significant challenges when applied to code translation. Due to the unique semantics and idiosyncrasies of programming languages, it is difficult to identify the correct context for translation, and there is a risk of token overflow or incorrect translation results. Although LLMs have shown the potential to perform judgment by a chain of thought by conditioning on a few examples, translating an entire application or adapting the behavior of third-party libraries and runtimes remains a complex task beyond the capabilities of conventional transpilation or fine-tuning methods.

Means for Solving the Problems

[0004] According to one embodiment of the present invention, a computer-implemented method for efficiently translating program code from a source language to a target language is provided. The input source code is parsed into an intermediate representation (IR) using a processor device. The structure and semantic model of the source code are established by applying static analysis to the IR, and a program skeleton of the target code is constructed from the IR, which includes generating context-aware placeholders. The IR is transformed into a single static assignment (SSA) form, and a system dependency graph (SDG) is assembled from the SSA form. The SDG is traversed to order translation tasks, and the ordered tasks are translated into the target language using a large language model (LLM). The translated program is generated by integrating the translated code segments into a coherent program structure in the target language.

[0005] According to one embodiment of the present invention, a system for efficiently translating program code from a source language to a target language is provided. The input source code is parsed into an intermediate representation (IR) using a processor device. The structure and semantic model of the source code are established by applying static analysis to the IR, and a program skeleton of the target code is constructed from the IR, which involves generating context-aware placeholders. The IR is transformed into a single static assignment (SSA) form, and a system dependence graph (SDG) is assembled from the SSA form. The SDG is traversed to order translation tasks, and the ordered tasks are translated into the target language using a large language model (LLM). The translated program is generated by integrating translated code segments into a coherent program structure in the target language.

[0006] According to one embodiment of the present invention, a non-transitory computer-readable storage medium comprising a computer-readable program operably coupled to a processor device is provided for efficiently translating program code from a source language to a target language. The input source code is parsed into an intermediate representation (IR) using a processor device. The structure and semantic model of the source code are established by applying static analysis to the IR, and a program skeleton of the target code is constructed from the IR, which involves generating context-aware placeholders. The IR is transformed into a single static assignment (SSA) form, and a system dependence graph (SDG) is assembled from the SSA form. The SDG is traversed to order translation tasks, and the ordered tasks are translated into the target language using a large language model (LLM). The translated program is generated by integrating translated code segments into a coherent program structure in the target language.

[0007] These features and advantages, and other features and advantages, will become apparent from the following detailed description of their exemplary embodiments, which should be read in conjunction with the accompanying drawings.

Brief Description of the Drawings

[0008] The following description provides details of the preferred embodiments with reference to the following drawings.

[0009]

Figure 1

[0010]

Figure 2

[0011]

Figure 3

[0012]

Figure 4

[0013]

Figure 5

[0014]

Figure 6

[0015]

Figure 7

[0016]

Figure 8

[0017]

Figure 9

[0018]

Figure 10

[0019]

Figure 11

[0020]

Figure 12

[0021]

Figure 13

[0022] According to an aspect of the present invention, there is provided a system and method for code translation and optimization that recognizes context based on an adaptive neural network.

[0023] According to an aspect of the present invention, in various embodiments, the present invention comprises a neural network-based system and method for translating programming code, and by leveraging state-of-the-art neural network models and sophisticated contextual data processing techniques, the accuracy and efficiency of code translation across various programming languages can be significantly improved.

[0024] The present invention may comprise an adaptive neural network system that can seamlessly integrate and process various programming languages and contexts. The system incorporates an innovative mechanism for neural network adaptation and employs a dual-structure model that separates the weights of the network into "locked" copies and "trainable" copies. This unique configuration enables the system to fine-tune its translation ability for specific tasks without compromising the underlying strength and generalization ability of the pre-trained model.

[0025] In addition to neural network adaptation, the present invention may employ advanced techniques in program analysis, such as the use of intermediate representations (IRs) and dedicated processing modules. These techniques ensure that the source code is not only accurately translated into the target language but also retains the structural and functional integrity of the original program. The present invention addresses and overcomes the challenges commonly encountered in conventional code translation methods, such as the semantic gap, readability issues, and the complexity of translating third-party libraries and custom APIs.

[0026] In conventional implementations, the translation of programming code from one language to another typically involves an intermediate representation (IR) such as single static assignment (SSA). While this approach helps optimize the code for runtime benefits and simplifies certain compiler optimizations, it often results in a loss of high-level constructs and creates a semantic gap that can hinder the readability of the translated code and its faithfulness to the original logic. Additionally, translating an entire application, especially one with third-party libraries, still exceeds the capabilities of standard transpilers and can result in code that is unreadable and maintainable by humans.

[0027] Conventional approaches to training neural network models for code generation tasks have relied on converting source code into a standard IR and a fine-tuning model and translating from the IR to the target language. Despite the use of extensive labeled datasets and iterative gradient updates, this method faces challenges such as the need for large task-specific datasets, poor generalization, and the utilization of irrelevant training data features.

[0028] Pre-trained large language models (LLMs) have evolved in their ability to perform in-context learning and adapt to new tasks during conditional inference based on a few examples. This approach has been successful in natural language processing tasks. However, its application to code translation is difficult due to the challenge of specifying context for translation between programming languages, the limited token size processing ability of LLMs, and the tendency of LLMs to produce hallucinatory or dishonest outputs.

[0029] Taking these challenges into account, the present invention improves conventional systems and methods by integrating neural network models with advanced contextual data processing strategies. In some embodiments, the present invention employs an adaptive system that uses context-aware neural networks and optimization techniques to improve accuracy and efficiency in code translation. According to aspects of the present invention, this system and method can optimize the translated code for specific operating conditions and performance needs of the target environment while maintaining the integrity of the logic and semantics of the original code.

[0030] In various embodiments, the present invention can be utilized to overcome the limitations of conventional methods by providing a more robust and contextually intelligent translation process. It leverages the strengths of neural networks to understand coding structures and patterns and introduces sophisticated context manipulation strategies to achieve high-fidelity translations across various programming languages. The system of the present invention encapsulates the principles of both "Frozen and Opaque" LLMs and "Frozen and Translucent" LLMs and utilizes innovative modules such as ControlNet for context learning, prompt embedding, and prompt expansion to enhance translation prompts. According to aspects of the present invention, this comprehensive framework significantly advances the field of code translation and provides solutions capable of supplying code that is sensitive to the nuances of various programming contexts and optimized for the target environment.

[0031] As will be understood by those skilled in the art, aspects of the present invention may be embodied as a system, method, or computer program product that can be executed on local and / or remote computing devices. Accordingly, aspects of the present invention may take the form of all hardware embodiments, all software embodiments (including firmware, resident software, microcode, etc.), or embodiments combining software and hardware aspects. Further, embodiments of the present invention may take the form of a computer program product embodied in one or more computer-readable media within one or more computing devices in which computer-readable program code is embodied. The embodiments described herein may all be hardware, all software, or may include both hardware and software elements. According to aspects of the present invention, in some embodiments, the present invention is implemented in software including, but not limited to, firmware, resident software, microcode, etc.

[0032] Any combination of one or more computer-readable media may be employed. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. Other examples of the computer-readable storage medium may include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any combination thereof. In this document, the computer-readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with a computing system, apparatus, or device.

[0033] Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, fiber optic cable, or any combination thereof. The computer program code for performing the operations regarding aspects of the present invention may be written in any combination of one or more programming languages, including but not limited to any general-purpose programming language (e.g., PHP, Java®, C++, etc.) and / or domain-specific programming languages (e.g., HTML, SQL, etc.), blockchain-specific programming languages (e.g., solidity, rust, Java®, python®, etc.). The program code may be executed entirely on the user's computer / mobile device, partially on the user's computer / mobile device, as stand-alone software, partially on the user's computer / mobile device and partially on a remote computer / mobile device, entirely on a remote computer or server, and / or using a blockchain. The remote computer may be connected to the user's computer by any type of network (e.g., local area network (LAN), wide area network (WAN), connection to an external computer (e.g., via the Internet using an Internet service provider), etc.).

[0034] Aspects of the present invention will be described below with reference to the flowchart illustrations and / or block diagrams of methods, systems, and computer program products according to embodiments of the present invention. It should be noted that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions.

[0035] These computer program instructions can be sent to a processor of any type of computing system (e.g., a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus for creating a machine), and as a result, means for implementing the functions / instructions / operations specified within one or more blocks of a flowchart and / or block diagram are produced by the instructions executed by the processor of the computing system. These computer program instructions, which can be instructed to function in a specific manner on any computing device, may also be stored in a computer-readable medium so as to produce a product that includes instructions stored in the computer-readable medium for implementing the functions / instructions / operations specified within one or more blocks of a flowchart and / or block diagram.

[0036] Also, the computer program instructions may be loaded onto a computer, a mobile device, other programmable data processing apparatus, or other devices to cause a series of operational steps to be executed on any computing system, thereby creating a computer-implemented process, and as a result, the instructions executed on the computer or other programmable device result in a process for implementing the functions / operations specified in one or more blocks of a flowchart and / or block diagram.

[0037] A computer-readable signal medium may include a propagated data signal (e.g., a baseband, a part of a carrier wave, etc.) having computer-readable program code embodied therein. Such a propagated signal may take any of a variety of forms, including but not limited to electromagnetic, optical, or any combination thereof. A computer-readable signal medium may be any computer-readable medium that is not a computer-readable storage medium and that can communicate, propagate, or transmit a program for use by or in connection with a computing system, apparatus, or device.

[0038] A suitable data processing system for storing and / or executing program code may include at least one processor directly or indirectly coupled to a memory element by a system bus. The memory element may include local memory used during actual execution of the program code, mass storage, and cache memory that provides at least some temporary storage of program code to reduce the number of times code is retrieved from mass storage during execution. Input / output or I / O devices (including, but not limited to, keyboards, displays, pointing devices, etc.) may be coupled to the system either directly or via intervening I / O controllers.

[0039] A network adapter may be coupled to the system such that the data processing system can be coupled to other data processing systems, remote printers, storage devices, blockchains, etc. via intervening private or public networks. Modems, cable modems, and Ethernet® cards are but a few examples of currently available types of network adapters.

[0040] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in the flowchart or block diagram may represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical function, and in some alternative implementations of the present invention, the functions described within the block may be performed out of the order described in the figures. For example, two blocks shown in succession may, in fact, be executed substantially simultaneously depending on the functionality of a particular embodiment, may be executed in the reverse order, or in any other order, as the case may be.

[0041] Note also that each block of the block diagram and / or flowchart diagram, and combinations of blocks in the block diagram and / or flowchart diagram, can be implemented by a specific application hardware system that performs a specific function / operation according to this principle, or by a combination of dedicated hardware and computer instructions.

[0042] Here, referring to a plurality of drawings in which like reference numerals represent the same or similar elements, and first referring to FIG. 1, an exemplary processing system 100 for code translation to which the principles of the present invention can be applied is illustratively shown in accordance with an embodiment of the present invention. The processing system 100 may include at least one processor (CPU) 104 operably coupled to other components via a system bus 102. A cache 106, read-only memory (ROM) 108, random access memory (RAM) 110, input / output (I / O) adapter 120, audio adapter 130, network adapter 140, user interface adapter 150, and display adapter 160 may be operably coupled to the system bus 102.

[0043] The first storage device 122 and the second storage device 124 may be operably coupled to the system bus 102 by the I / O adapter 120. The storage devices 122 and 124 may be any of a disk storage device (e.g., a magnetic or optical disk storage device), a solid state magnetic device, etc. The storage devices 122 and 124 may be the same type of storage device, or different types of storage devices.

[0044] Speaker 132 may be operably coupled to system bus 102 by voice adapter 130. According to the present invention, speaker 132 may be used to provide an audible alarm or some other indication with respect to resilient battery charging. Transceiver 142 may be operably coupled to system bus 102 by network adapter 140. Display device 162 may be operably coupled to system bus 102 by display adapter 160.

[0045] First user input device 152 and second user input device 154 may be operably coupled to system bus 102 by user interface adapter 150. User input devices 152, 154 may be any of a keyboard, a mouse, a keypad, an imaging device, a motion sensing device, a microphone, a device incorporating at least two functions of the foregoing devices, etc. Of course, other types of input devices may also be used while maintaining the spirit of the present invention. User input devices 152, 154 may be the same type of user input device, or different types of user input devices. User input devices 152, 154 may be used to input information into system 100 and then output it. According to an aspect of the present invention, system 100 may include a system-dependent graph builder / traverser / stack popper at block 156, and a code transformer / translator / generator at block 164, which will be described in more detail below in this specification.

[0046] Of course, the processing system 100 may also include other elements (not shown), and certain elements may be omitted, as readily contemplated by those skilled in the art. For example, as will be readily understood by those skilled in the art, various other input devices and / or output devices may be included in the processing system 100 depending on its particular implementation. For example, various types of wireless and / or wired input and / or output devices may be used. Also, additional processors, controllers, memories, etc. in various configurations may be utilized as will be readily understood by those skilled in the art. These and other variations of the processing system 100 will be readily contemplated by those skilled in the art given the teachings of the present invention provided herein.

[0047] Also, it will be understood that the systems 200, 300, 400, 500, 800, 900, 1000, 1100, 1200, and 1300, described below with respect to FIGS. 2, 3, 4, 5, 8, 9, 10, 11, 12, and 13 respectively, are systems for implementing respective embodiments of the present invention. Some or all of the processing system 100 may be implemented in one or more of the elements of the systems 200, 300, 400, 500, 800, 900, 1000, 1100, 1200, and 1300 of FIGS. 2, 3, 4, 5, 8, 9, 10, 11, 12, and 13 respectively.

[0048] Furthermore, it will be understood that the processing system 100 may execute at least a portion of the methods described herein, including, for example, at least a portion of methods 200, 300, 400, 500, 600, 700, and 1300 of FIGS. 2, 3, 4, 5, 6, 7, and 13. Similarly, some or all of systems 200, 300, 400, 500, 800, 900, 1000, 1100, 1200, and 1300 of FIGS. 2, 3, 4, 5, 8, 9, 10, 11, 12, and 13 may be used to execute at least a portion of methods 200, 300, 400, 500, 600, 700, and 1300 of FIGS. 2, 3, 4, 5, 6, 7, and 13.

[0049] Referring now to FIG. 2, a diagram of a high-level transformer-based neural network (NN) system and method 200 for optimized code translation using a large language model (LLM) is illustratively shown in accordance with an embodiment of the present invention.

[0050] In various embodiments, at block 202, the input may represent the initial code (or other type of data) to be translated. According to aspects of the present invention, such an input 202 may be in any of a plurality of types of programming languages. Here, the source code may be tokenized into a format that can be effectively processed by the neural network, transforming the raw code into a structured sequence of tokens that encapsulate both syntax and semantics. At this stage, the code may be not just raw text, but a carefully tokenized array of data points, each of which represents a distinct and identifiable element of the programming language from which the system can gather semantic and syntactic meaning.

[0051] Following the input at block 202, an encoder 204 may be utilized to process the input 202. The encoder 204 need not be a single entity; rather, a composite of multiple transformer layers 201, ..., 203 may function as the backbone of a transformer-based architecture. Their role may be to break the input into its components, apply a self-attention mechanism, and identify the context and relationships within the input sequence. Each layer 201, ..., 203 may refine this understanding, pass on a richer, more context-aware representation to the next, and ultimately result in a comprehensively encoded state.

[0052] In some embodiments, at block 206, the state may capture the essence of the source code as understood by the encoder. It may be a distilled representation, i.e., a vector, that embodies the collective understanding of the code's structure and intent informed by the layers of attention and parsing the code has undergone. This state may act as an intermediary between the raw complexity of the input 202 and the expected clarity of the output 210. The state 206 may be an encoded version of the input, enhanced with context information flowing from the encoder to the decoder. This state may function as a comprehensive representation of the source code that conveys all the information necessary for an accurate translation.

[0053] In various embodiments, the decoder at block 208 may mirror the encoder in structure but not necessarily in function. The decoder 208 may include a fresh stack of transformer layers 205, ..., 207 where each layer can perform the dual task of decoding the state into the target language and refining the output in real time. According to aspects of the present invention, the decoder layer may incorporate a cross-attention module, which may enable the system to juxtapose the context of the source code against the generated translation and ensure that each newly generated token is consistent with what came before it. The decoder at block 208 mirrors the structure of the encoder but may have additional mechanisms to attend to the output of the encoder. It may incrementally build the target sequence (translated code). The decoder layer, including the attention mechanism 212 and the feed-forward network 216, may generate predictions for each token of the output sequence based on both the encoder's representation and what has been generated so far.

[0054] In some embodiments, the end point of this process may be the output at block 210. This is a tangible result of the system's process, including, for example, outputting a sequence of tokens that form the translated code here. This output may be, for example, the final word of the decoder 208, i.e., a piece of code transformed into a new language, which may then be compiled, executed, or further refined. The output 210 may include a sequence of tokens representing the code in the target language generated by the decoder 208. According to aspects of the present invention, this output 210 may result in a syntactically and semantically accurate translation of the input code.

[0055] According to various embodiments, the enlarged view 209 of the transformer layers 201, 203, 205, 207 illustratively shows the workings inside both the encoder 204 and decoder 208 layers. This detailed breakdown reveals further details of the transformer layers 201, 203, 205, 207, which may include the attention mechanism 212, add and normalization (Add&Norm) functions 214, 218, and the feed-forward network 216. Each function within this microcosm may contribute to transforming the encoded state into the translated sequence.

[0056] The attention mechanism 212 within these layers may be at the core of the transformer model, which may enable the network to focus on and weight different parts of the input sequence. This mechanism is important for code translation where dependencies can spread widely across the sequence and understanding the context is key to maintaining the integrity of the translation. Further, this component may enable the model to focus on different parts of the input sequence when predicting each token of the output sequence. In the context of code translation, this may mean that the attention mechanism helps the model align code segments with their corresponding translations taking into account dependencies that can spread across the entire codebase.

[0057] The add and normalization functions 214, 218 may be layer mechanisms to ensure the general problems in deep neural networks, such as stabilizing the output, preventing values from rising extremely, and ensuring that the gradients neither vanish nor explode rapidly. The add and normalization functions 214, 218 may apply residual connections followed by layer normalization. This process helps to stabilize learning and enables deeper networks by alleviating the vanishing gradient problem.

[0058] In some embodiments, the feedforward network 216 may represent a series of linear transformations using activation functions, which may function to sequentially process data and translate it into a new representation suitable for output generation of complex relationships and patterns recognized by the attention mechanism. According to an aspect of the present invention, the feedforward network 216 may further transform the output of the attention mechanism before passing it to the next layer or generating output tokens in the case of a decoder.

[0059] This transformer-based neural network architecture is enhanced by the vast knowledge and context recognition provided by the LLM, representing a significant leap in the field of code translation. The architecture is carefully designed to translate code not only accurately but also with a broader understanding of its context facilitated by the extensive pre-training of the LLM on a vast corpus of code in various programming languages. This system is adaptable and capable of both direct translation and further refinement through prompt tuning and fine-tuning by leveraging the LLM's ability to learn from examples and improve its translation. Thus, according to an aspect of the present invention, the transformer-based neural network exists as a robust and sophisticated solution for the complex task of code translation.

[0060] In various embodiments, when integrated with an LLM, the Transformer architecture leverages pre-training on a vast amount of parameters and a broad code corpus. An LLM that has learned the patterns, structures, and semantics of multiple programming languages guides the training process of the Transformer network for a specific task of code translation. The pre-trained knowledge base of the LLM significantly improves the Transformer's ability to understand and translate complex code structures by providing it with a broad understanding of programming language syntax and semantics. The combination of the Transformer's structure and the LLM's extensive pre-training enables more accurate and contextually relevant code translation than would be possible with either component alone.

[0061] In the complete translation process, the source code is first tokenized and passed through an encoder. The output of the encoder then serves as a guide for the decoder, which generates the translated code. Throughout this process, the pre-trained knowledge of the LLM helps to accurately predict semantically correct and syntactically valid tokens in the target programming language.

[0062] The ability of the LLM to provide in-context learning in which it can understand and generate code based on a few examples is particularly beneficial in that it enables few-shot learning in which the model can effectively translate even using a limited number of examples from the source and target languages. This is at least partially due to the ability of the Transformer to leverage the broad contextual understanding incorporated into the LLM. Overall, according to aspects of the present invention, the combination of the Transformer architecture and the LLM produces a powerful model for translating code that can capture and utilize the complexity of programming languages to produce accurate and efficient translations.

[0063] Referring now to FIG. 3, a diagram illustrating a system and method 300 for model tuning and prompt tuning in language model training for neural network-based code translation is illustratively shown in accordance with an embodiment of the present invention. FIG. 3 illustrates two separate methods for adapting a neural network model pre-trained with 11 billion parameters (11B params) to perform a specific task, according to aspects of the present invention.

[0064] In various embodiments, block 301 represents a model tuning approach where separate models are trained for different tasks, the separate models are individually calibrated for a specific task, and each model utilizes a complete set of the pre-trained model's parameters. According to aspects of the present invention, in block 303, a prompt tuning approach is shown, which simplifies the adaptation process by tuning a small subset of the model's parameters for various tasks, as will be described in more detail below in this specification.

[0065] In various embodiments, the model tuning in block 301 may include task-specific batches for processing. Blocks 302, 304, and 306 each represent a batch of task-specific data for tasks A, B, and C, respectively. According to aspects of the present invention, for illustrative purposes, each batch includes separate examples including a1 and a2 for task A batch 302, b1 for task B batch 304, and c1 and c2 for task C batch 306, which can be used to fine-tune a dedicated instance of the neural network model.

[0066] Block 302 illustrates a dataset for task A that includes unique examples a1, a2 that represent specific nuances and requirements of task A. The data within this batch may be curated to encapsulate the diverse programming scenarios that task A is expected to encounter. Similarly, block 304 may be populated with data example b1 for task B, which reflects the distinct characteristics and challenges associated with this particular translation task. Block 306, which is reserved for task C, includes examples c1, c2 and ensures that the dataset comprehensively represents the breadth of the domain of task C, which may provide a robust training ground for a model dedicated to this task.

[0067] According to various embodiments of the present invention, although three batches are shown for the sake of simplicity of illustration, it will be understood that any number of batches may be employed.

[0068] In various embodiments, a pre-trained machine learning model characterized by 11 billion parameters is illustratively shown at block 305. This extensive parameterization enables the model to have a broad knowledge base suitable for a variety of tasks. The pre-trained model may function as a starting point for a further task-specific tuning process. Its architecture is highly adaptable and may allow for subsequent refinement by both model tuning and prompt tuning techniques. The model is capable of understanding and processing complex data patterns and may assist in enabling the system's ability to perform dedicated tasks after tuning. Blocks 308, 310, and 312 illustrate independent task-specific models for tasks A, B, and C after tuning. The task-specific models 308, 310, and 312 are labeled as task A model, task B model, and task C model, respectively, and may represent the results of the model tuning process. Each model is now fine-tuned and may embody the complexity of its respective task and be optimized to execute with high fidelity within its specified domain. In particular, according to aspects of the present invention, these models may retain the complete parameter set of the original pre-trained model but may be optimized for their respective tasks.

[0069] In various embodiments, according to aspects of the present invention, prompt tuning in block 303 can simplify the adaptation process by tuning a small subset of the model's parameters for various tasks. Block 314 shows a mixed task batch that can combine data samples from all tasks and serve as an inclusive input for the pre-trained model during the prompt tuning process. Block 316 presents an innovative prompt embedding strategy where prompt tokens can be concatenated with a set of tunable context tokens to produce an enhanced input that can be processed by the pre-trained model. This technique can enable the model to be fine-tuned for task-specific outputs while retaining its vast parameter set, and by selectively tuning a subset of the parameters associated with each task, more efficient utilization of the knowledge of the pre-trained model can be achieved.

[0070] In some embodiments, central to the prompt tuning strategy, block 314 presents a mixed task batch, i.e., a set of examples from all of tasks A, B, and C. This approach can utilize diverse aggregate datasets to inform the prompt tuning process. As the technique progresses, block 316 may include performing prompt embedding, where the prompt tokens for each task can be merged with a tunable set of context tokens. This enhanced input can interact dynamically with the parameters 318 of the pre-trained model and guide the output of the model towards the desired translation task goal. In block 318, the "pre-trained model (11 billion parameters)" can function as an advanced processing unit capable of fine-tuning its response based on a wide variety of input signals. This block may interact with the "mixed task batch 316" and receive a collection of various task prompts designed to induce the model by a dedicated tuning procedure. During this prompt tuning phase, the model may dynamically adjust its internal parameters, which can be represented by an 11 billion parameter count, to improve its ability to interpret and process the complex nuances of the mixed task input. The interaction between the pre-trained model and the task prompts results in a refined translation of the source code into the target language, leveraging extensive pre-training to handle a wide range of programming tasks and ensuring a high degree of accuracy and contextual relevance in the output. According to an aspect of the present invention, model 318 may maintain its original extensive parameter set but exhibit improved performance with fewer resource requirements in task-specific performance due to the updated context tokens.

[0071] Referring now to FIG. 4, a diagram illustrating a system and method 400 using an exemplary frozen language model having frozen weights and learnable weights for effective code translation is illustratively shown in accordance with an embodiment of the present invention.

[0072] In various embodiments, the frozen language model 402 can be utilized for code translation tasks. In this embodiment, the language model (LM) 402 may be pre-trained using a set of fixed or "frozen" parameters in order to retain its initial training based on a large corpus of data. The fixed parameters are indicated by reference numerals 404 and 408, which represent the frozen weights of the LM 402.

[0073] LM 402 may be interfaced with a target 410 that can represent the expected output of the language model. This output can be generated after the model processes the input data, and the target is a sequence of tokens in a programming language. In some embodiments, learnable weights 405 may be included within the model and can be uniquely identified and adjustable. These weights represent parameters that can be fine-tuned to adapt the output of LM 402 to the specific requirements of the translation task even when the rest of the model's weights remain unchanged. According to aspects of the present invention, this fine-tuning enables a certain degree of customization and adaptability without the need to retrain the entire model, thereby saving computational resources and time.

[0074] In various embodiments, each input parameter 401, 403, 405, 407, 409, 411, 413, 415 can represent an embedding of input tokens that can be processed by the frozen weights of the LM 402. These parameters can be utilized by the model to interpret the input data and transform it into a processable form to generate the desired output. Weights W p1 , W p2 ,...W pm-1 , W pm , W x1 , W x2 ,...W xn-1They can be pre - established parameters of the LM that encodes linguistic information, represented by blocks 401, 403, 407, 409, 411, 413, and 415 respectively. Since they are "frozen", these weights do not undergo further modification and maintain the learned representation. Each of these parameters corresponds to a specific token from the input prompt fed into the LM. They can be embedded into the vector space by the model, enabling the LM to process and understand the input.

[0075] In practice, the LM can obtain input tokens, process them with the frozen weights 404, 408, and with the help of the learnable weight W u 405, produce an output that is the translated version of the input data. This weight 405 can be adjusted during the prompt - tuning process to better match the model's output to a specific translation task without affecting the rest of the model's integrity. This output can be represented by a sequence of weights 417, 419, 421, 423 that yield the target 410. These weights may correspond to the final layer of the model that shapes the translated sequence into its final form. These weights associated with the output layer of the LM can be utilized to generate the final translated output by transforming the processed information from the internal layers of the LM into a structured sequence corresponding to the target language.

[0076] The configuration of the LM402 shown in Figure 4 enables effective translation of code by leveraging the general knowledge of the pre - trained model while also providing flexibility to incorporate task - specific nuances through the learnable weight 405. This represents a significant innovation in the fields of machine learning and language processing, enabling more accurate and context - relevant translations without the overhead of retraining the model from scratch.

[0077] In various embodiments, the interaction of these components can include that the input tokens are represented by their corresponding weights (e.g., 401, 403, 407, etc.) that can be fed to LM402. LM402 can process these inputs using its frozen weights to generate an internal representation of the input data. The learnable weights 406 can be utilized to fine-tune the model's response to a particular input prompt and can improve the accuracy of the output without changing the basic knowledge encoded by the frozen weights. In some embodiments, LM402 can then create an output sequence that can be a translation of the input prompt into the target language. According to an aspect of the present invention, this sequence is represented by output weights represented by blocks 417, 419, 421, and 423, which can also be part of the frozen parameters of the model, and can ensure that the translation is faithful to the syntactic and semantic structure of the target language.

[0078] Referring now to FIG. 5, a diagram illustrating a system and method 500 for neural network-based translation of programming code that utilizes IR and standardizes source code for precise translation by a language model is exemplarily shown in accordance with an embodiment of the present invention.

[0079] In various embodiments, at block 502, code can be input for processing and can represent the raw source code to be translated. This code can be processed by static analysis so that its structure and semantics can be determined before it is fed into the translation framework. This block is a repository of raw code that will undergo translation and can include the necessary instructions, declarations, and constructs that define the program's operational logic and functionality in its original language. At block 504, a static analysis framework can be utilized to process the raw source code from block 502. This framework can apply static analysis techniques to parse the code and construct an intermediate representation (IR), which abstracts away high-level language details and distills the code into a form that can be analyzed and manipulated by the system. Block 506 introduces a static analysis tool framework (e.g., an Abstract Syntax Tree (AST) framework, parse tree, symbol tree, WALA cast entity, etc.), which can represent an abstraction layer that uses the static analysis tool framework to convert the input code into an intermediate representation (IR). Note that the frameworks mentioned above are presented for illustrative purposes and that other similar frameworks can be employed in accordance with aspects of the present invention. This entity 506 can be utilized to understand the syntax and semantics of the source language, which can facilitate the extraction of structural and behavioral patterns within the code.

[0080] In various embodiments, at block 508, nodes can be obtained from the static analysis entity 506. These nodes 508 can represent essential elements of the program's control flow and data structures and are decomposed into fine-grained components that can be individually analyzed and translated. At block 510, metadata is utilized in conjunction with the nodes to provide additional information about each node, including, for example, data types, scopes, and variable dependencies. This metadata can be utilized to ensure that the translated code preserves the functionality and logic of the original code.

[0081] In block 512, a program skeleton is constructed using the node 508 and metadata 510, and a skeletal version of the target program can be assembled. This skeleton can form a blueprint of the translated code and can include placeholders for logic and data that can be filled by the LLM. The target language skeleton in block 514 can represent a structured format of the translated code as it begins to take shape. Here, the basic elements of the target program can be laid out, and it is ready to be populated with the actual code generated by the LLM. In block 516, the translation context placeholders can be strategically positioned within the target language skeleton 514. According to aspects of the present invention, these placeholders can be filled with context-related code snippets to ensure that the translation is not only syntactically correct but also functionally equivalent to the original code.

[0082] According to aspects of the present invention, in various embodiments, in block 518, an intermediate representation (IR) can be utilized, and in some embodiments, it can be improved by converting the IR into a single static assignment (SSA) IR form, which can simplify the translation by providing a clear and unambiguous representation of variable assignments and dependencies. In block 520, a system dependence graph (SDG) can be generated, mapping and showing the dependencies within the code and providing a visual representation of the execution flow and data relationships. This graph can be utilized to ensure the logical coherence of the translated code.

[0083] In block 522, a graph traversal can be performed by systematically navigating the SDG, which can identify the sequence in which code segments (e.g., basic blocks) are to be translated, ensuring that the translated code reflects the intended behavior of the original program. Block 524 represents a "SDG stack pop and traverse" action, where the basic blocks identified during the graph traversal can be sequentially processed for translation. This stage can be utilized to maintain the order and dependencies of the program components. Blocks 526 and 528 represent the SDG stack, which can be a dynamic structure that holds the basic blocks of the code as they are processed. According to aspects of the present invention, the stack can be utilized to ensure that each block is translated in the correct order and, when translated, the blocks can be pushed onto the program stack so that they can return to reassemble the target program again.

[0084] In various embodiments, system 500 ensures that translated code blocks maintain syntactic and semantic integrity during code translation. Stack management can be dynamic, which allows for repopulating the stack with translated segments, which can then be integrated to return to the target program stack. This management can be utilized to preserve the original program structure and ensure functional correctness.

[0085] In some embodiments, the described framework can be source and language agnostic, meaning that the entities and context information provided to the LLM can be uniform and replicable across various programming languages. The use of static analysis creates a verifiable translation guarantee for large codebases and provides deterministic and functional correctness in the translated code. By assembling a system dependency graph and traversing it to create well-defined deterministic translation units, the framework can provide an analogue of judgment by the actual chain of thought obtained and can enhance the IR with additional translation context. According to aspects of the present invention, this framework represents an inventive approach to code translation, as shown in FIG. 5, leveraging the capabilities of an LLM to manage the complexity of the programming language while ensuring that the translated code is faithful to the functional equivalence of the original code.

[0086] Referring now to FIG. 6, a block / flow diagram illustrating a method 600 for neural network-based code translation using a large language model (LLM) is exemplarily shown in accordance with an embodiment of the present invention.

[0087] In various embodiments, at block 602, source code may be received by a processing module for analysis and processing. This may include starting the translation process, feeding the source code into the system, and utilizing static analysis tools. The tools may parse the source code into an intermediate representation (IR), which is an important step of breaking the source code into pieces to identify its syntactic and semantic features. This detailed analysis of the structure and meaning of the source code can be extremely important to ensure accurate translation at later stages. At block 604, the processing system proceeds to manage the parsed source code and may convert it into nodes (e.g., cast nodes, function call nodes, control flow nodes, assignment nodes, etc.) and associated metadata. This task may be performed using a static analysis tool framework (e.g., an Abstract Syntax Tree (AST) framework, parse tree, symbol tree, WALA cast entity, etc.), which may represent an abstraction layer that uses the static analysis tool framework to convert the input code into an intermediate representation (IR). Note that the frameworks mentioned above are presented for illustrative purposes, and other similar frameworks may be adopted in accordance with aspects of the present invention. This framework may analyze the IR and derive nodes using static analysis entities, where each node represents a separate code element such as a variable, function, or control structure. Additionally, metadata may be associated with each node to indicate details of the type, scope, and interrelationships of the code elements within the source code. This detailed metadata can be extremely important to maintain the logical and functional characteristics of the source code during the translation process.

[0088] In various embodiments, at block 606, the skeletal framework of the target program can be constructed from these nodes and metadata using a program skeleton construction module. This involves assembling the minimal structure of the target program that maps and shows the essential architecture and flow of the source code, incorporating placeholders within this skeleton that mark positions for context-driven code insertion, and ensuring that this framework functions as a guide for subsequent translation tasks to populate the target language code. In various embodiments, at block 608, an intermediate representation (IR) can be utilized, and in some embodiments, it can be enhanced by converting the IR to a single static assignment (SSA) IR form using an SSA module. According to aspects of the present invention, this can involve transforming the detailed IR into an SSA form to simplify the translation process, where the SSA form ensures that each variable is assigned once, thereby eliminating ambiguity and facilitating clearer translation by the LLM, and simplifying the IR by setting up the SSA form as a clean slate for the LLM to perform its translation function.

[0089] In some embodiments, at block 610, a system dependence graph (SDG) module may be used to generate a visual and functional map of the code execution flow and dependencies. According to aspects of the present invention, this may include assembling an SDG that visually represents the execution flow of a source program using IR or SSA IR form, capturing all functional dependencies and control structures to guide the translation process, and ensuring that the SDG accurately reflects the program's execution logic to inform the correct sequencing of the translation tasks. At block 612, the SDG may be navigated by graph traversal, and the SDG stack may be populated with basic blocks determined for translation. According to aspects of the present invention, this may include traversing the SDG using an algorithm to identify the sequence in which the program's basic blocks are translated, populating these blocks, which may represent individual units of functionality within the source code, onto the SDG stack, and queuing the blocks for translation in a manner that preserves the integrity of the control flow of the original program.

[0090] At block 614, the order and context of the basic blocks and method segments may be managed and controlled during translation using the SDG stack module. This may include dynamically adjusting the SDG stack as blocks are translated and integrated back into the target program structure, overseeing the contextual integrity of each translated block to ensure that the target program remains functionally coherent, and applying systematic techniques to the reintegration of translated segments to maintain the original execution logic of the program.

[0091] In block 616, the basic blocks can be translated into the target language using a large language model (LLM). According to aspects of the present invention, this involves deploying the LLM to interpret each basic block and translate them from the source language to the target language, using contextually appropriate placeholders within the program skeleton to induce the LLM when generating functionally equivalent code, and integrating the translated segments into a coherent target program structure, where the LLM can utilize additional context and runtime dependencies from the source code environment to improve the accuracy and functional correctness of the translation.

[0092] Referring now to FIG. 7, a block / flow diagram illustrating a method 700 for neural network-based code translation including the generation of a target language skeleton is exemplarily shown in accordance with an embodiment of the present invention.

[0093] According to aspects of the present invention, this method 700 can utilize the power of advanced neural network techniques combined with strategic context data processing. It can integrate separate approaches of "Frozen and Opaque" large language models (LLMs) and "Frozen and Translucent" large language models (LLMs) using sophisticated context manipulation strategies. The result is a robust system and method 700 that provides high-fidelity translations across a wide variety of programming languages and is capable of addressing the nuanced requirements of software development and mathematical linguistics.

[0094] In various embodiments, at block 702, a neural network block representing a pre-configured neural network architecture can be engaged to initiate the translation process. The network can be in a "frozen" state, which means its weights are set and cannot be changed, ensuring the stability and predictability of the initial translation mechanism. An input signal (e.g., x) can be received and systematically processed by the network to produce an output (e.g., y). This output mirrors the conditions of the original model prior to any adaptation and can function as a baseline for later-occurring deformations.

[0095] At block 704, the integrity of the neural network model can be protected and maintained. In its "locked" state, the neural network can receive the same input "x" and here receive a newly introduced context signal (e.g., c) that can be extracted from a repository of statically derived context data. The integration of this context signal with the input can initiate a judgment based on the sophisticated chain of thought within the LLM. This process can be utilized to accurately and efficiently adapt the model's capabilities to the specific nuances of the particular translation task being performed. According to an aspect of the present invention, at block 706, a neural network (NN) training process can be initiated, which can include utilizing a transformer-based NN for LLM processing in various embodiments, as described in further detail above with reference to FIG. 2.

[0096] In various embodiments, at block 708, prompt embedding and context tuning can be initiated. This can include fusing the original input prompt with a carefully curated set of context tokens. According to aspects of the present invention, these tokens can traverse the processing path of the LLM as standard inputs, but can be uniquely further designed to undergo optional training. Such training can be configured to adjust only the parameters associated with the context tokens, thereby refining the LLM's prompt response without modifying the weights of the underlying neural network. At block 710, prompt expansion can be performed for improved translation. This can include weaving an extensible vocabulary of out-of-dictionary or non-natural language tokens into the standard prompt structure. This expansion can significantly broaden the descriptive power of the prompt, provide the LLM with a deeper and more nuanced understanding of the coding context, and result in a significant improvement in the fidelity and accuracy of the code translation process.

[0097] At block 712, the output from the advanced prompt expansion at block 710 can be captured and integrated by merging the output of the expanded prompt (e.g., a translated code segment with enhanced understanding of both the source and target languages). The output can be not only syntactically transformed from the source to the target language but also semantically rich and thus a translated code that takes into account the complexity and context of the programming paradigm. According to aspects of the present invention, this translated code can be integrated into the target program environment to ensure that the translated application behaves as intended in its new ecosystem and reflects the logic, performance expectations, and operational dependencies of the original program.

[0098] Referring now to FIG. 8, a generalized diagram illustrating an exemplary neural network system 800 for neural network-based context-aware code translation and optimization is shown in accordance with one embodiment of the present invention.

[0099] An artificial neural network (ANN) is an information processing system inspired by biological nervous systems such as the brain. One element of an ANN is the structure of the information processing system, which includes a large number of highly interconnected processing elements (referred to as "neurons") that work in parallel to solve a specific problem. An ANN is further trained using a set of training data, with learning involving the adjustment of weights that exist between the neurons. An ANN is configured for specific applications such as pattern recognition or data classification through such a learning process.

[0100] A specific structure of an ANN with three layers and a set number of fully connected neurons is shown, but it should be understood that this is intended for illustrative purposes only. In reality, the present embodiment can take any suitable form, including any number of layers and their connections between any one or more patterns.

[0101] ANNs can show the ability to derive meaning from complex or inaccurate data and can be used to extract patterns and detect such trends that are too complex to be detected by humans or other computer-based systems. The structure of a neural network generally is known to have input neurons 802 that provide information to one or more "hidden" neurons 804. The connections 808 between the input neurons 802 and the hidden neurons 804 are weighted, and these weighted inputs are then processed by the hidden neurons 804 according to some function within the hidden neurons 804. Any number of layers of hidden neurons 804, as well as neurons performing different functions, may exist. Different neural network structures such as convolutional neural networks, maxout networks, etc. also exist, and these can vary according to the structure and function of the hidden layers, as well as the pattern of weights between the layers. Individual layers can perform a specific function and can include convolutional layers, pooling layers, fully connected layers, softmax layers, or any other appropriate type of neural network layer. Finally, a set of output neurons 806 receives and processes the weighted inputs from the last set of hidden neurons 804.

[0102] This represents a "feedforward" calculation where information propagates from the input neurons 802 to the output neurons 806. Once the feedforward calculation is complete, the output is compared to the desired output available from the training data. Next, the error regarding the training data is processed in a "backpropagation" calculation where the hidden neurons 804 and the input neurons 802 receive information regarding the error that propagates in the reverse direction from the output neurons 806. Once the error backpropagation is complete, weight updates are performed and the weighted connections 808 are updated to take into account the received error. Note that the three modes of operation, namely feedforward, backpropagation, and weight update, do not overlap with each other. This represents just one variant of ANN calculation and any suitable form of calculation can be used instead.

[0103] To train the ANN, the training data can be split into a training set and a test set. The training data includes pairs of inputs and known outputs. During training, the inputs of the training set are fed into the ANN using forward propagation. After each input, the output of the ANN is compared with the respective known output. The discrepancy between the output of the ANN associated with that particular input and the known output is used to generate an error value, which can be backpropagated through the ANN, after which the weight values of the ANN can be updated. This process continues until the pairs within the training set are exhausted.

[0104] After training is complete, the ANN can be evaluated against the test set to ensure that the training has not led to overfitting. If the ANN can generalize to new inputs other than those it has already been trained on, it is ready for use. If the ANN does not accurately reproduce the known outputs of the test set, additional training data may be required or the hyperparameters of the ANN may need to be adjusted.

[0105] The ANN can be implemented in software, hardware, or a combination of the two. For example, each weight 808 may be characterized as a weight value stored in computer memory, and the activation function of each neuron may be implemented by a computer processor. The weight values may store any suitable data values such as real numbers, binary values, or values selected from a fixed number of possibilities, which are multiplied against the associated neuron output. Alternatively, the weights 808 (e.g., priority list weights, attribute weights for generating descriptive styles / personalities, etc.) may be implemented as a resistive processing unit (RPU) that generates a predictable current output when an input voltage is applied according to a settable resistance value.

[0106] Referring now to FIG. 9, a hardware diagram illustrating an exemplary artificial neural network (ANN) system 900 for code translation and optimization that recognizes neural network-based contexts is illustratively shown in accordance with an embodiment of the present invention.

[0107] It should be understood that this architecture is merely exemplary and that other architectures or types of neural networks may be used instead. The hardware embodiments described herein are included to illustrate the general principles of neural network computations at a high level of generality and should in no way be construed as limiting.

[0108] Furthermore, the layers of neurons and the weights connecting them described below are presented in a general fashion and may be replaced with any type of neural network layer having any suitable degree or type of interconnectivity. For example, the layers can include convolutional layers, pooling layers, fully connected layers, softmax layers, or any other suitable type of neural network layer. Further, layers may be added or removed as needed, and the weights described herein may be replaced with more complex forms of interconnectivity.

[0109] During the feed-forward operation, each of the input neurons 902 provides an input voltage in parallel to each row of weights 904. In the hardware embodiments described herein, each of the weights 904 has a configurable resistance value such that a current output flows from the weights 904 to the respective hidden neurons 906. Thus, the current output by the weights 904 represents the weighted input to the hidden neurons 906.

[0110] Following the hardware embodiment, the current output by a given weight 904 is

Equation

[0111] The set of reference weights 907 has fixed resistance values and combines their outputs to form a reference current, which is provided to each of the hidden neurons 906. Since the conductivity values can only be positive numbers, some reference conductivity is required to encode both the positive and negative values within the matrix. The current created by the weights 904 is continuous and positive, so the reference weights 907 are used to provide a reference current. If the current exceeds this reference current, it is considered to have a positive value, and if it is below this reference current, it is considered to have a negative value. The use of the reference weights 907 is not necessary in software embodiments where the output and weight values can be obtained accurately and directly. As an alternative to using the reference weights 907, another embodiment may use a separate array of weights 904 to capture negative values.

[0112] The hidden neurons 906 use the current from the array of weights 904 and the reference weights 907 to perform some calculation. This calculation can be, for example, any suitable activation function and can be implemented in hardware using a suitable circuit or in software.

[0113] Next, the hidden neurons 906 output their own voltages to another array of weights 904 based on the activation function. This array performs the weighted calculation in the same way, where the columns of weights 904 receive the voltages from their respective hidden neurons 906 and create a weighted current output, which is added row by row and provided to the output neuron 908.

[0114] It should be understood that any number of these stages can be implemented by interposing additional layers of arrays and hidden neurons 906. It should also be noted that some neurons can be constant neurons 909 that provide a constant output to the array. The constant neurons 909 can be present within the input neurons 902 and / or the hidden neurons 906 and are only used during the feed-forward operation.

[0115] During backpropagation, output neuron 908 provides a voltage that returns across the array of weights 904. The output layer compares the generated network response with the training data and calculates an error. The error is applied to the array as a voltage pulse, where in this case the height and / or duration of the pulse is modulated in proportion to the error value. In this example, the rows of weights 904 receive voltages in parallel from their respective output neurons 908 and convert that voltage to a current that is summed column by column to provide an input to hidden neurons 906. Hidden neurons 906 combine the weighted feedback signal with the derivative of their feedforward calculation and store the error value prior to outputting the feedback signal voltage to its respective column of weights 904. This backpropagation propagates through the entire network 900 until all hidden neurons 906 and input neurons 902 have stored the error value.

[0116] The weight update process will depend on how the weights 904 are implemented. For configurable resistors that include phase change materials, throughout network 900, input neurons 902 and hidden neurons 906 may apply a first weight update voltage in the forward direction and output neurons 908 and hidden neurons 906 may apply a second weight update voltage in the reverse direction. The combination of these voltages creates a state change within each weight 904, for example, by raising the temperature of weight 904 above a threshold and thereby changing its resistance, causing weight 904 to assume a new resistance value. In this way, weights 904 can be trained to adapt the neural network 900 to the error in its processing.

[0117] As described above, the weight 904 can be implemented in software or hardware, for example, using a relatively complex weighting circuit or using a resistive cross-point device. Such a resistive device can have switching characteristics with non-linearity that can be used to process data. The weight 904 can belong to a class of devices referred to as a resistive processing unit (RPU). The RPU device can be implemented with a resistive change type memory (RRAM (registered trademark)), a phase change memory (PCM), a programmable metallization cell (PMC) memory, or any other device having non-linear resistive switching characteristics. Such an RPU device can also be regarded as a memory system.

[0118] Continuing to refer to FIG. 9 and referring to FIG. 10, a block diagram showing an exemplary neuron 1000 within a neural network system for code translation and optimization that recognizes a neural network-based context is exemplarily shown in accordance with an embodiment of the present invention.

[0119] In various embodiments, this neuron can represent any of the input neuron 902, the hidden neuron 906, or the output neuron 908 as shown in FIG. 9. Note that FIG. 10 shows components that handle all three phases of operation: feedforward, backpropagation, and weight update. However, since the different phases do not overlap, there must necessarily be some form of control mechanism within the neuron 1000 to control which components are active. Thus, according to aspects of the present invention, it should be understood that switches and other structures not shown may exist within the neuron 1000 to handle switching between modes.

[0120] In the feedforward mode, the difference block 1002 determines the value of the input from the array by comparing it with a reference input. Thereby, both the magnitude and sign (e.g., + or -) of the input from the array to the neuron 1000 are set. Block 1004 performs calculations based on the input, and this output is stored in the storage 1005. It is particularly contemplated that block 1004 calculates a non-linear function and can be implemented as an analog or digital circuit, or can be executed by software. The value determined by the function block 1004 is converted to a voltage in the feedforward generator 1006, and the feedforward generator 1006 applies that voltage to the next array. The signal propagates in this way by passing through multiple layers of the array and neurons until it reaches the final output layer of the neuron. The input is also applied to the derivative of the non-linear function of block 1008, and this output is stored in the memory 1009.

[0121] During the backpropagation mode, an error signal is generated. The error signal can be generated at the output neuron 908, or can be calculated by a separate unit that receives the input from the output neuron 908 and compares the output with the correct output based on the training data. Otherwise, if the neuron 1000 is the hidden neuron 906, it receives the information backpropagating from the array of weights 904, and compares the received information with a reference signal in the difference block 1010 to provide a signed error signal of continuous values. This error signal is multiplied using the multiplier 1012 by the derivative of the non-linear function from the immediately preceding feedforward stage stored in the memory 1009, and the result is stored in the storage 1013. The value determined by the multiplier 1012 is converted in the backpropagation generator 1014 to a voltage pulse propagating in the reverse direction, proportional to the calculated error, and the backpropagation generator 1014 applies that voltage to the immediately preceding array. The error signal propagates in this way by passing through multiple layers of the array and neurons until it reaches the input layer of the neuron 902.

[0122] During the weight update mode, after both forward and backward passes are completed, each weight 904 is updated in proportion to the product of the signals passed through the weight during the forward and backward passes. The update signal generator 1016 provides voltage pulses in both directions (note that only one direction is available for the input and output neurons). According to an aspect of the present invention, the shape and amplitude of the pulses from the update generator 1016 change the state of the weight 904 (e.g., priority list weight, attribute weight for generating description style / personality), such that the resistance of the weight 904 is configured to be updated.

[0123] Referring now to FIG. 11, a diagram illustrating an exemplary layered neural network system 1100 in a neural network for code translation and optimization that recognizes a neural network-based context is illustratively shown in accordance with one embodiment of the present invention.

[0124] In a layered neural network, nodes are arranged in layers. An exemplary simple neural network has an input layer 1120 of source nodes 1122 and a single computational layer 1130 having one or more computational nodes 1132 that also function as output nodes, where there is a single computational node 1132 for each possible category into which an input example can be classified. The input layer 1120 may have a number of source nodes 1122 equal to the number of data values 1112 in the input data 1110. The data values 1112 in the input data 1110 may be represented as column vectors. Each computational node 1132 in the computational layer 1130 generates a linear combination of weighted values from the input data 1110 fed to the input nodes 1120 and applies a non-linear activation function, which is differentiable, to the sum. An exemplary simple neural network may perform classification for linearly separable examples (e.g., patterns).

[0125] A deep neural network, such as a multi-layer perceptron, may have an input layer 1120 of source nodes 1122, one or more computing layers 1130 having one or more computing nodes 1132, and an output layer 1140, where there is a single output node 1142 for each possible category into which input examples can be classified. The input layer 1120 may have a number of source nodes 1122 equal to the number of data values 1112 in the input data 1110. The computing nodes 1132 in the computing layer 1130 may also be referred to as hidden layers since they are between the source nodes 1122 and the output nodes 1142 and are not directly observable. Each node 1132, 1142 in the computing layer generates a linear combination of weighted values from the values output by the nodes in the previous layer and applies a non-linear activation function that is differentiable over the range of the linear combination. The weights applied to the values from each previous node may be denoted, for example, as w1, w2,... w n-1、 w n as may be indicated by. The output layer provides the response of the network as a whole to the input data. The deep neural network may be fully connected, where each node in the computing layer is connected to all other nodes in the previous layer, or may have other configurations of connections between the layers. If there are missing links between nodes, the network is said to be partially connected.

[0126] Training a deep neural network may involve two phases, namely, a forward phase in which the weights of each node are fixed and the input propagates through the network, and a backward phase in which the error values propagate backward through the network and the weight values are updated.

[0127] The computing nodes 1132 in one or more computing (hidden) layers 1130 perform a non-linear transformation on the input data 1112, thereby generating a feature space. Classes or categories may be more easily separable in the feature space than in the original data space.

[0128] In various embodiments, the present invention can customize the architecture and training process of a neural network to optimize for the recognition of the context of programming code for translation. This involves leveraging a layered network to complexly map and interpret syntax and semantics from one programming language to another, thereby improving the accuracy and efficiency of code translation. The ability of a deep neural network to discern nuanced patterns within data through multiple computational layers allows for refined adjustment during the training phase, ensuring that the output of the model matches a sophisticated coding framework and contributing to the development of neural network-based code translation techniques.

[0129] Referring now to FIG. 12, a block diagram illustrating a system 1200 for neural network-based context-aware code translation and optimization is exemplarily shown according to an embodiment of the present invention.

[0130] In various embodiments, at block 1202, a source code input device is shown that functions as an initial interface for receiving source code to be translated. It can process raw code input and prepare it for subsequent analysis and translation phases.

[0131] At block 1204, a contextual data processor can function as an intermediate stage for enhancing the source code using contextual information. This processor can improve the code by embedding relevant data that can affect the translation process, such as comments, documentation, and metadata, ensuring a more accurate and context-aware translation output. A translation control unit represented by block 1206 can control and orchestrate the entire translation workflow. It can direct the processed source code to appropriate components within the system, manage the interactions between the neural network trainer, prompt embedder, and other critically important elements, and facilitate a seamless translation operation.

[0132] In block 1208, the neural network trainer is a component responsible for the adaptive learning mode of the system. It can use the provided context information to fine-tune the neural network parameters and adapt the model to the specific nuances and requirements of the source code being translated. In block 1210, the prompt embedder can integrate strategic prompts into the translation process. These prompts can guide the attention of the neural network during translation and effectively advance the output of the model towards the desired target language construction. The program skeleton builder in block 1212 can construct a basic template for the target program structure. It can generate a skeletal framework showing the main components and flow of the target program and set up placeholders where the translated code can be integrated as a result.

[0133] In block 1214, the prompt extension processor can improve the input prompt using additional contextual tokens. These tokens are utilized to enhance the language model's understanding of the code context and can lead to translations with higher fidelity and relevance to the target programming language. In various embodiments, the NN (e.g., a transformer-based NN architecture) 1216 can be integrated with the LLM such that the architecture of the transformer can leverage pre-training on a vast amount of parameters and an extensive code corpus. According to aspects of the present invention, an LLM that has learned the patterns, structures, and semantics of multiple programming languages can guide the training process of the transformer network for a specific task of code translation.

[0134] The pre-trained knowledge base of the LLM can significantly improve the ability of the transformer to understand and translate complex code structures by providing it with a broad understanding of programming language syntax and semantics. According to aspects of the present invention, the combination of the structure of the transformer and the extensive pre-training of the LLM can provide more accurate and contextually relevant code translations than either component alone.

[0135] In block 1218, the IR generator can generate an intermediate representation (IR) of the source code. This IR can be a standardized format that abstracts the code and makes it easier and more efficient for the system to analyze and translate across all different programming languages. The translation optimizer in block 1220 can apply various algorithms to refine the translated code. It ensures that the output not only maintains functional equivalence but is also faithful to the idiomatic nuances and performance considerations of the target language.

[0136] In block 1222, the SDG builder / pop and traverser can be utilized to construct a system dependence graph (SDG) that maps and shows the dependencies and execution flow of the source code. It also manages the traversal and translation of code segments and can maintain the logical and execution integrity of the original program.

[0137] In block 1224, the output integration and verification unit can be utilized to integrate the code into the target program skeleton. It verifies the translated segments for syntactic and semantic correctness and ensures that the final program is not only correct but also optimized for the target execution environment.

[0138] According to an aspect of the present invention, as an exemplary example, in some embodiments, a high-level representation of one-shot learning is shown as Algorithm 1. The model may be provided with a simple task description, and representative examples of what successful task completion may look like are shown below in this specification: Algorithm 1: One-shot Learning "Convert Java to Python (registered trademark)" / / Prompt public class Foo inherits Bar => class Foo(Bar) / / Example 1 private static void spam(int x) => __________???_________ / / Task

[0139] In some embodiments, the one-shot approach can be extended to few-shot learning by providing a simple task description, and multiple representative samples of what successful task completion may look like are shown below in Algorithm 2 in this specification: Algorithm 2: Few-shot Learning "Convert Java to Python (registered trademark)" / / Prompt public class Foo inherits Bar => class Foo(Bar) / / Example 1 public String baz(Foo foo) => def baz(foo: Foo) -> String / / Example 2 . . . private static void spam(int x) => __________???_________ / / Task

[0140] In some embodiments, partial label (PL) induced in-context learning can be performed according to Algorithm 3 shown below: Algorithm 3: Partial Label (PL) Induced In-context Learning "Convert Java to Python (registered trademark)" / / Prompt public class Foo inherits Bar => class Foo(Bar) / / Example 1 public String baz(Foo foo) => def baz(foo: Foo) -> String / / Example 2 . . . private static void spam(int x) => __________???_________ / / Task public class Foo inherits Bar / / Source 2 "create a {MODIFIER: public}{TEMPLATE: class} called {NAME: Foo} in {TARGET: python(registered trademark)} inheriting {INHERITS: Bar}" / / Prompt 1 class Foo(Bar) / / Target 1 . . . private static void spam(int x) / / Source 2 "add a {MODIFIER: private} {TEMPLATE: static method} {NAME: spam} with {ARGS: [{TYPE: int} x]} that {RETURNS: void}" / / Prompt 2 @staticmethod def _spam(x: int) -> None / / Target 2

[0141] According to various embodiments, examples of "one-shot" and "few-shot" learning illustrate the advanced learning algorithms of the system, which are essential for the code translation process of the present invention. The present invention improves code translation by systems employing advanced learning techniques. Illustrated herein are examples of "one-shot" and "few-shot" learning, demonstrating the system's ability to conform to programming language constructs using minimal input.

[0142] In some embodiments, at block 1202, the system may be presented with a single prompt and task pair that illustrates the model's ability for "one-shot" learning. The system may utilize this single example to understand the structure and semantics and translate them from Java® to Python®. The prompt "Convert Java to Python®" is combined with a class inheritance example to induce the model to formulate the correct translation, encapsulating the syntax of the source code and object-oriented principles. Here, the system can utilize a single example to learn and translate the code syntax. It can infer structural patterns from the Java® class inheritance example, apply this learning to generate an equivalent Python® class, and demonstrate the system's ability to capture and translate object-oriented concepts after observing only one instance.

[0143] Expanding to block 1204, "few-shot" learning may be implemented, where the system absorbs multiple task descriptions and examples. This approach can amplify the model's understanding of the coding language and refine the accuracy of its predictions regarding code translation. By processing a broader set of examples, the system can discriminate language patterns and nuances, thereby producing more robust translations. This approach expands the system's learning by analyzing multiple examples, enabling it to recognize more diverse programming structures. The system can extract patterns from multiple instances, resulting in a more nuanced understanding and refined translation output. This technique, according to aspects of the present invention, may enable the system to respond to discrepancies seen in different programming tasks and adapt its translation mechanism accordingly.

[0144] In various embodiments, the one-shot and few-shot learning scenarios described can be essential to the larger framework of the system, connecting to block 1212 where the program skeleton builder can utilize these learned patterns to construct an infrastructure for the code to be translated. Similarly, in block 1214, the prompt expansion processor can utilize insights from these examples to improve the prompt structure and contribute to the overall translation effectiveness of the system. Both examples of learning support the system's ability to process and translate code with high accuracy, reducing the need for extensive datasets. According to aspects of the present invention, they can be utilized in the system's translation optimization, enabling it to quickly adapt to new languages and coding paradigms and ensuring that the translated code is syntactically and semantically accurate.

[0145] Referring now to FIG. 13, an exemplary computing environment for the execution of at least a portion of computer code for code translation and optimization that recognizes an adaptive neural network-based context is illustratively shown in accordance with an embodiment of the present invention.

[0146] Various aspects of the present disclosure are described by the block diagrams of machine logic included in the description, flowchart, block diagrams of computer systems, and / or computer program product (CPP) embodiments. With respect to any flowchart, depending on the relevant technology, operations can be performed in an order different from that shown in a given flowchart. For example, again depending on the relevant technology, two operations shown in consecutive flowchart blocks can be performed in reverse order, as a single integrated step, simultaneously, or in a manner that is at least partially temporally overlapping.

[0147] An embodiment of a computer program product (referred to herein as a "CPP embodiment" or "CPP") is used in the present disclosure and refers to any set of one or more storage media (also referred to as "media") collectively included in a set of one or more storage devices, which collectively contain machine-readable code corresponding to instructions and / or data for performing the computer operations specified in a given CPP claim. A "storage device" is any tangible device that can hold and store instructions for use by a computer processor. By way of non-limiting example, a computer-readable storage medium may be an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these media are floppy disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanically encoded devices (such as punch cards or pits / lands formed on the major surfaces of a disk), or any suitable combination of the foregoing. The term computer-readable storage medium as used in the present disclosure should not be construed to include storage in the form of a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide, optical pulses passing through an optical fiber cable, electrical signals communicated through a wire, and / or other transmission media. As would be understood by one of ordinary skill in the art, data typically moves intermittently during the normal operation of a storage device, such as during access, defragmentation, or garbage collection, but the data is not transient while it is stored, and thus this does not make the storage device transient.

[0148] Computing environment 1300 includes an example of an environment for executing at least a portion of computer code involved in performing the methods of the present invention, such as partial label induced in-context learning 1350. This partial label induced in-context learning 1350 may include utilizing partial labels and in-context learning for code translation. It may include providing detailed prompts having specific programming elements (e.g., class definitions, method modifiers) and their corresponding translations. This code may be utilized to more accurately guide the translation process by taking into account the nuances of programming constructs and their context within the source and target languages. According to aspects of the present invention, in various embodiments, this may be essential for the advanced learning capabilities of the system and may enable accurate and context-aware translation of programming code across different languages.

[0149] In addition to block 1350, computing environment 1300 includes, for example, computer 1301, wide area network (WAN) 1302, end user device (EUD) 1303, remote server 1304, public cloud 1305, and private cloud 1306. In this embodiment, computer 1301 includes a processor set 1310 (including processing circuitry 1320 and cache 1321), communication fabric 1311, volatile memory 1312, persistent storage 1313 (including the operating system 1322 and block 200 identified above), a set of peripheral devices 1314 (including a set of user interface (UI) devices 1323, storage 1324, and a set of Internet of Things (IoT) sensors 1325), and network module 1315. Remote server 1304 includes remote database 1330. Public cloud 1305 includes gateway 1340, cloud orchestration module 1341, a set of host physical machines 1342, a set of virtual machines 1343, and a set of containers 1344.

[0150] Computer 1301 may take the form of a desktop computer, laptop computer, tablet computer, smartphone, smartwatch or other wearable computer, mainframe computer, quantum computer, or any other form of computer or mobile device that is currently known or will be developed in the future and is capable of executing programs, accessing a network, or querying a database such as remote database 1330. As is well understood in the technical field of computer technology and in accordance with the technology, the execution of computer-implemented methods may be distributed among multiple computers and / or among multiple locations. On the other hand, in this presentation of computing environment 1300, for the sake of keeping the presentation as simple as possible, the discussion focuses on a single computer, specifically computer 1301. Although not illustrated within the cloud in FIG. 1, computer 1301 may be located within the cloud. On the other hand, computer 1301 is not required for the cloud, except for any scope that may be affirmatively shown.

[0151] Processor set 1310 includes one or more computer processors of any type currently known or to be developed in the future. Processing circuitry 1320 may be distributed across multiple packages, such as multiple integrated circuit chips that have been adjusted. Processing circuitry 1320 may implement multiple processor threads and / or multiple processor cores. Cache 1321 is memory located within a processor chip package and is typically used for data or code that should be available for rapid access by threads or cores operating on processor set 1310. Cache memory is typically organized into multiple levels according to its relative proximity to the processing circuitry. Alternatively, some or all of the cache for the processor set may be located "off-chip". In some computing environments, processor set 1310 may be designed to work with qubits to perform quantum computing.

[0152] Typically, computer-readable program instructions are loaded onto computer 1301 and cause a series of operational steps to be performed by the processor set 1310 of computer 1301, thereby resulting in a computer-implemented method, and as a result, the instructions executed therefor will instantiate the method specified in the flowchart and / or description of the computer-implemented method included in this document (collectively referred to as "the method of the present invention"). These computer-readable program instructions are stored in various types of computer-readable storage media such as cache 1321 and other storage media discussed below. The program instructions and associated data are accessed by the processor set 1310 to control and direct the execution of the method of the present invention. In computing environment 1300, at least a portion of the instructions for executing the method of the invention may be stored in block 200 in persistent storage 1313.

[0153] Communication fabric 1311 is a signal conduction path that enables various components of computer 1301 to communicate with each other. Typically, this fabric is made up of switches and conductive paths such as buses, bridges, physical input / output ports, and switches and conductive paths that make up the like. Other types of signal communication paths such as optical fiber communication paths and / or wireless communication paths may be used.

[0154] Volatile memory 1312 is any type of volatile memory known currently or developed in the future. Examples include dynamic random access memory (RAM) or static RAM. Typically, volatile memory 1312 is characterized by random access, but this is not essential unless affirmatively shown. In computer 1301, volatile memory 1312 is located within a single package and inside computer 1301, but alternatively or additionally, the volatile memory may be distributed across multiple packages and / or located external to computer 1301.

[0155] The persistent storage 1313 is any form of non-volatile storage for a computer, known currently or developed in the future. The non-volatility of this storage means that the stored data is maintained regardless of whether power is supplied to the computer 1301 and / or directly to the persistent storage 1313. The persistent storage 1313 can be read-only memory (ROM), but is typically at least a portion of the persistent storage that allows writing of data, deletion of data, and rewriting of data. Some well-known forms of persistent storage include magnetic disks and solid-state storage devices. The operating system 1322 can take multiple forms, such as various known proprietary operating systems that employ a kernel or open-source portable operating system interface type operating systems. The code included in block 200 typically includes at least a portion of the computer code involved in the execution of the method of the invention.

[0156] The peripheral device set 1314 includes a set of peripheral devices of the computer 1301. The data communication connections between the peripheral devices of the computer 1301 and other components may be implemented in various ways such as a Bluetooth (registered trademark) connection, a Near-Field Communication (NFC) connection, a connection by a cable (such as a Universal Serial Bus (USB) type cable), an insertion type connection (for example, a Secure Digital (SD) card), a connection via a local area communication network, and even a connection via a wide area network such as the Internet. In various embodiments, the UI device set 1323 may include components such as a display screen, a speaker, a microphone, wearable devices (such as goggles and smartwatches), a keyboard, a mouse, a printer, a touchpad, a game controller, and a haptic device. The storage 1324 is an external storage such as an external hard drive or an insertable storage such as an SD card. The storage 1324 may be persistent and / or volatile. In some embodiments, the storage 1324 may take the form of a quantum computing storage device for storing data in qubits. In embodiments where the computer 1301 is required to have a large amount of storage (for example, the computer 1301 locally stores and manages a large database), in that case, this storage may be provided by a peripheral storage device designed to store a very large amount of data, such as a Storage Area Network (SAN) shared by a plurality of geographically dispersed computers. The IoT sensor set 1325 is composed of sensors that can be used in Internet of Things applications. For example, one sensor may be a thermometer and another sensor may be a motion detector.

[0157] The network module 1315 is an assembly of computer software, hardware, and firmware that enables the computer 1301 to communicate with other computers via the WAN 1302. The network module 1315 may include hardware such as a modem or a Wi-Fi (registered trademark) transceiver, software for packetizing and / or depacketizing data for communication network transmission, and / or web browser software for communicating data over the Internet. In some embodiments, the network control function and the network transfer function of the network module 1315 are executed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing software-defined networking (SDN)), the control function and the transfer function of the network module 1315 are executed on physically separate devices, such that the control function manages multiple different network hardware devices. The computer-readable program instructions for executing the inventive method may typically be downloaded to the computer 1301 from an external computer or an external storage device through a network adapter card or network interface included in the network module 1315.

[0158] The WAN 1302 is any wide area network (e.g., the Internet) capable of communicating computer data over non-local distances by any technology for communicating computer data, now known or hereafter developed. In some embodiments, the WAN 1302 may be replaced and / or supplemented by a local area network (LAN) designed to communicate data between devices located in a local area, such as a Wi-Fi network. The WAN and / or LAN typically includes computer hardware such as copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and edge servers.

[0159] An end - user device (EUD) 1303 is any computer system that is used and controlled by an end - user (e.g., a customer of an enterprise that operates computer 1301) and can take any of the forms discussed above in relation to computer 1301. The EUD 1303 typically receives useful data from the operation of computer 1301. For example, in a virtual case where computer 1301 is designed to provide recommendations to an end - user, this recommendation is typically communicated from the network module 1315 of computer 1301 to the EUD 1303 via the WAN 1302. In this way, the EUD 1303 can display or otherwise present the recommendation to the end - user. In some embodiments, the EUD 1303 can be a client device such as a thin - client, a thick - client, a mainframe computer, a desktop computer, etc.

[0160] A remote server 1304 is any computer system that provides at least some data and / or functions to computer 1301. The remote server 1304 can be controlled and used by the same entity that operates computer 1301. The remote server 1304 represents a machine that collects and stores useful data for use by other computers such as computer 1301. For example, in a virtual case where computer 1301 is designed and programmed to provide recommendations based on historical data, in that case, this historical data can be provided from the remote database 1330 of the remote server 1304 to computer 1301.

[0161] The public cloud 1305 is any computer system available for use by multiple entities that provides on-demand availability of computer system resources and / or other computer functions, particularly data storage (cloud storage) and computing power, without direct active management by the user. Cloud computing typically leverages resource sharing to achieve coherence and economies of scale. The direct and active management of the computing resources of the public cloud 1305 is performed by the computer hardware and / or software of the cloud orchestration module 1341. The computing resources provided by the public cloud 1305 are typically implemented by virtual computing environments that run on various computers that make up the host physical machine set 1342, which is the population of physical computers within and / or available to the public cloud 1305. The virtual computing environment (VCE) typically takes the form of virtual machines from a virtual machine set 1343 and / or containers from a container set 1344. It is understood that these VCEs can be stored as images and transferred among and between various physical machine hosts either as images or after instantiation of the VCE. The cloud orchestration module 1341 manages the transfer and storage of images, deploys new instantiations of the VCE, and manages the active instantiation of VCE deployments. The gateway 1340 is an aggregate of computer software, hardware, and firmware that enables the public cloud 1305 to communicate via the WAN 1302.

[0162] Here, some further explanation of virtual computing environments (VCEs) is provided. A VCE can be stored as an "image". A new active instance of a VCE can be instantiated from the image. Two well-known types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to a feature of the operating system where the kernel enables the existence of multiple isolated instances of user space, called containers. These isolated instances of user space typically act as actual computers from the perspective of the programs running within them. A computer program running on a normal operating system can utilize all the resources of that computer, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, a program running within a container can only use the contents of the container and the devices assigned to the container. This is a feature known as containerization.

[0163] The private cloud 1306 is similar to the public cloud 1305, except that computing resources are only available for use by a single enterprise. Although the private cloud 1306 is shown as communicating with the WAN 1302, in other embodiments, the private cloud may be completely disconnected from the Internet and only accessible via a local / private network. A hybrid cloud is a composition of multiple clouds of different types (e.g., private, community, or public cloud types), often implemented by different vendors. Each of the multiple clouds remains a separate and distinct entity, but the larger hybrid cloud architecture is tied together by standardized or proprietary technologies that enable orchestration, management, and / or data / application portability between the multiple constituent clouds. In this embodiment, both the public cloud 1305 and the private cloud 1306 are part of a larger hybrid cloud.

[0164] As used herein, the terms "hardware processor subsystem" or "hardware processor" can refer to a processor, memory, software, or a combination thereof that cooperate to perform one or more specific tasks. In useful embodiments, the hardware processor subsystem may include one or more data processing elements (e.g., logic circuits, processing circuits, instruction execution devices, etc.). The one or more data processing elements may be included in a central processing unit, a graphics processing unit, and / or a controller based on a separate processor or computing element (e.g., logic gates, etc.). The hardware processor subsystem may include one or more on-board memories (e.g., cache, dedicated memory arrays, read-only memories, etc.). In some embodiments, the hardware processor subsystem may include one or more memories that may be on-board or on-board, or may be dedicated for use by the hardware processor subsystem (e.g., ROM, RAM, basic input / output system (BIOS), etc.).

[0165] In some embodiments, the hardware processor subsystem may include and execute one or more software elements. The one or more software elements may include an operating system, and / or one or more applications, and / or specific code to achieve a specified result.

[0166] In other embodiments, the hardware processor subsystem may include dedicated, special-purpose circuitry that performs one or more electronic processing functions to achieve a specified result. Such circuitry may include one or more application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), and / or programmable logic arrays (PLAs).

[0167] These and other variations of the hardware processor subsystem according to embodiments of the present invention are also contemplated.

[0168] The present invention may be a system, method, and / or computer program product integrated at any possible technical detail level. The computer program product may include a computer-readable storage medium (or multiple computer-readable storage media) having computer-readable program instructions for causing a processor to execute aspects of the present invention.

[0169] A computer-readable storage medium can be a tangible device that holds and stores instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof, but is not limited thereto. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices such as punch cards or raised structures in grooves having instructions recorded thereon, and any suitable combination of the foregoing. A computer-readable storage medium, as used herein, should not be construed to be a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.

[0170] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on a computer-readable storage medium in each respective computing / processing device.

[0171] Computer-readable program instructions for performing the operations of the present invention may be source code or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or object-oriented programming languages such as Smalltalk®, C++, or the like, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, may be executed partially on the user's computer as a stand-alone software package, may be executed partially on the user's computer and partially on a remote computer, or may be executed entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit in order to perform aspects of the present invention.

[0172] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.

[0173] These computer readable program instructions may be provided to a processor of a computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement the aspect(s) of the function / act specified in one or more blocks of the flowchart and / or block diagram.

[0174] Alternatively, the computer readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0175] References to "one embodiment" or "an embodiment" of the invention within this specification and other variations thereof mean that the particular features, structures, characteristics, etc. described in connection with the embodiment are included in at least one embodiment of the invention. Thus, the phrases "in one embodiment" or "in an embodiment" and the appearance of any other variations that occur in various places throughout this specification are not necessarily all referring to the same embodiment.

[0176] It should be understood that the use of any of the following, namely " / ", "and / or", and "at least one of", for example in the cases of "A / B", "A and / or B", and "at least one of A and B", is intended to encompass the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such expressions are intended to encompass the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or the selection of only the first and second-listed options (A and B), or the selection of only the first and third-listed options (A and C), or the selection of only the second and third-listed options (B and C), or the selection of all three options (A and B and C). This may be extended by the same number of items as are listed, as will be readily apparent to those skilled in the art of this and related arts.

[0177] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions that include one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions described in the blocks may be performed in a different order than that shown in the drawings. For example, two blocks shown in succession may in fact be implemented as one step, executed at the same time, substantially simultaneously, partially or wholly in an overlapping manner in time, or the blocks may, in some cases, be executed in the reverse order depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0178] Preferred embodiments of systems and methods for efficient neural network-based translation and transformation of program code from a source language to a target language have been described (intended to be illustrative and not limiting), but it should be noted that modifications and variations can be made by those skilled in the art in light of the above teachings. It should be understood from this that changes can be made within the scope of the present invention and within the specific embodiments disclosed and outlined by the appended claims. Thus, aspects of the present invention have been described with the detail and specificity required by patent law, but what is desired to be claimed and protected by patent is set forth in the appended claims.

Claims

1. 1. A method for efficiently translating program code from a source language to a target language, comprising: parsing the input source code into an intermediate representation (IR) using a processor device; establishing a structural and semantic model of the input source code by applying static analysis to the IR; constructing a program skeleton of the target code from the IR, which includes generating context-aware placeholders; reducing the IR into a Single Static Assignment (SSA) form and constructing a System Dependence Graph (SDG) from the SSA form; traversing the SDG to order translation tasks and translating the ordered tasks into the target language using a large language model (LLM); and generating a translated program by integrating the translated code segments into a coherent program structure in the target language; A method comprising:

2. The method of claim 1 , further comprising receiving as an input a mixed task batch of code segments for processing.

3. The method of claim 1 , wherein the step of applying the static analysis further comprises generating nodes and corresponding metadata for each input code segment.

4. The method of claim 1 , wherein the translating step comprises populating placeholders with contextually relevant translations obtained from the static analysis.

5. The method of claim 1 , further comprising augmenting the IR with runtime dependency data from a source computing environment.

6. 2. The method of claim 1, wherein translating the ordered tasks into the target language includes adapting the translated code segments to comply with runtime constraints of a target computing environment to enable execution of the translated program within the target computing environment.

7. 2. The method of claim 1, wherein translating the ordered tasks into the target language includes adapting the translated code segments to meet specific performance metrics and resource constraints of a target computing environment, including one or more of memory usage, processing speed, and integration with existing software infrastructure.

8. 8. The method of claim 1, further comprising the steps of displaying the translated code segments on a user interface, receiving user input for code editing, compiling the edited code, and executing the compiled code to achieve a transformation of a state in a machine, wherein the execution of the compiled code causes the machine to perform a series of operations that result in physical changes indicative of a function of the code in the target language.

9. 1. A system for efficiently translating program code from a source language to a target language, comprising: A processor operably coupled to a computer-readable storage medium, the processor comprising: Parsing the input source code into an intermediate representation (IR); Establishing a structural and semantic model of the input source code by applying static analysis to the IR; constructing a program skeleton of the target code from the IR, which includes a procedure for generating context-aware placeholders; Transforming the IR into a Single Static Assignment (SSA) form and constructing a System Dependence Graph (SDG) from the SSA form; traversing the SDG to order translation tasks and translating the ordered tasks into the target language using a large-scale language model (LLM); generating a translated program by integrating the translated code segments into a coherent program structure in the target language; The system is configured as follows:

10. The system of claim 9 , wherein the processor is further configured to receive as an input a mixed task batch of code segments for processing.

11. 10. The system of claim 9, wherein the applying the static analysis further comprises generating nodes and corresponding metadata for each input code segment.

12. The system of claim 9 , wherein the translating step comprises populating placeholders with contextually relevant translations obtained from the static analysis.

13. The system of claim 9 , wherein the processor is further configured to augment the IR with runtime dependency data from a source computing environment.

14. 10. The system of claim 9, wherein the translating of the ordered tasks into the target language includes adapting the translated code segments to comply with runtime constraints of a target computing environment to enable execution of the translated program within the target computing environment.

15. 10. The system of claim 9, wherein the translating of the ordered tasks into the target language includes adapting the translated code segments to meet specific performance metrics and resource constraints of a target computing environment, including one or more of memory usage, processing speed, and integration with existing software infrastructure.

16. 16. The system of claim 9, wherein the processor is further configured to display the translated code segments on a user interface, receive user input for editing code, compile the edited code, and execute the compiled code to achieve a transformation of a state in a machine, wherein the execution of the compiled code causes the machine to perform a series of operations that result in physical changes indicative of a function of the code in the target language.

17. 1. A computer program operatively coupled to a processor device for efficiently translating program code from a source language to a target language, the computer program, when executed on a computer, causes the computer to: parsing the input source code into an intermediate representation (IR) using a processor device; establishing a structural and semantic model of the input source code by applying static analysis to the IR; constructing a program skeleton of the target code from the IR, which includes generating context-aware placeholders; reducing the IR into a Single Static Assignment (SSA) form and constructing a System Dependence Graph (SDG) from the SSA form; traversing the SDG to order translation tasks and translating the ordered tasks into the target language using a large-scale language model (LLM); and generating a translated program by integrating the translated code segments into a coherent program structure in the target language; A computer program that executes the following:

18. 20. The computer program product of claim 17, wherein the step of applying the static analysis further comprises generating nodes and corresponding metadata for each input code segment.

19. 20. The computer program product of claim 17, wherein the step of translating comprises populating placeholders with contextually relevant translations obtained from the static analysis.

20. 20. The computer program product of claim 17, wherein translating the ordered tasks into the target language comprises adapting the translated code segments to meet particular performance metrics and resource constraints of a target computing environment, including one or more of memory usage, processing speed, and integration with existing software infrastructure.