Methods, apparatuses, devices, media, and products for generating code

CN122837835APending Publication Date: 2026-09-29CHENGDU HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510329786.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

而低资源编程语言,其代码资源相对匮乏,缺乏广泛且成熟的代码库支持,往往可能不容易获取到针对该编程语言的程序代码

Benefits of technology

[0028]根据本申请的第五方面,还提供了一种计算机程序产品,包括计算机可执行指令,其中所述计算机可执行指令被处理器执行时实现根据本申请的第一方面所述的方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122837835A_ABST
    Figure CN122837835A_ABST
Patent Text Reader

Abstract

The application provides a method, device, equipment, medium and product for generating code, and relates to the technical field of artificial intelligence. The method comprises determining a first program code of a first programming language corresponding to a first dialect of the programming language. The method further comprises generating a first intermediate representation for the first dialect based on the first program code. The method further comprises generating a second intermediate representation for a second dialect of the programming language based on the first intermediate representation, the first characteristic set of the first dialect comprising the second characteristic set of the second dialect. The method further comprises generating a second program code of a second programming language corresponding to the second dialect based on the second intermediate representation. Through the method, corresponding program codes can be quickly obtained for different programming languages, the efficiency of obtaining comparable corpora is improved, the number of comparable corpora is increased, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this application primarily relate to the field of artificial intelligence technology. More specifically, the embodiments of this application relate to methods, apparatus, devices, media, and products for generating code. Background Technology

[0002] A programming language is a formalized language used for communication between humans and computers. Through a specific set of syntax, semantic rules, and instructions, it allows programmers to express algorithms, data structures, and logical flows to achieve various functions. Programs generated by a programming language are sets of instructions written according to the language's rules, capable of running on a computer system and completing specific tasks. These tasks can cover a wide range of fields, from simple numerical calculations to complex artificial intelligence applications.

[0003] High-resource programming languages, with their abundant code resources, typically possess vast open-source code libraries, comprehensive standard libraries, and numerous ready-made code examples and tools. These rich code resources significantly improve development efficiency, allowing developers to quickly reuse existing code to build complex systems. For example, Python boasts a wealth of third-party libraries, enabling convenient access to various functionalities in fields such as data science and artificial intelligence. In contrast, low-resource programming languages ​​have relatively scarce code resources, lacking widespread and mature code library support, and often making it difficult to obtain code specific to that language. Furthermore, utilizing this code presents numerous challenges that need to be addressed. Summary of the Invention

[0004] The embodiments of this application provide a scheme for generating code.

[0005] According to a first aspect of this application, a method for generating code is provided. The method includes determining first program code for a first programming language corresponding to a first dialect of a programming language; generating a first intermediate representation for the first dialect based on the first program code; generating a second intermediate representation for a second dialect of the programming language based on the first intermediate representation, wherein a first feature set of the first dialect includes a second feature set of the second dialect; and generating second program code for a second programming language corresponding to the second dialect based on the second intermediate representation.

[0006] This method allows for the conversion of program code to an intermediate representation of the corresponding dialect, and then the conversion between intermediate representations of different dialects, ultimately yielding the program code corresponding to another programming language dialect. This enables the rapid acquisition of corresponding program code even for low-resource programming languages, improving the efficiency of obtaining comparable corpora, increasing the number of comparable corpora, and enhancing the user experience.

[0007] In some embodiments, generating a first intermediate representation for a first dialect based on first program code includes: determining a first conversion rule from a first programming language to a first dialect; and adjusting the first program code into a first intermediate representation for the first dialect based on the first conversion rule. This method allows for fast and accurate conversion from program code to an intermediate representation.

[0008] In some embodiments, generating a second intermediate representation for a second dialect of a programming language includes: determining partial features in a first feature set that correspond to a second feature set; and adjusting the first intermediate representation to the second intermediate representation based on the partial features. This method enables rapid conversion between code representations of different dialects, improving the efficiency and accuracy of code conversion.

[0009] In some embodiments, generating second program code corresponding to a second dialect in a second programming language based on a second intermediate representation includes: determining a second conversion rule between the second programming language and the second dialect; and adjusting the second intermediate representation into second program code based on the second conversion rule. This method allows for rapid conversion from intermediate representation to target program code, improving the efficiency of obtaining comparable corpora.

[0010] In some embodiments, the method further includes: training a machine learning model that can be used to implement code conversion between different programming languages, based on first-level code and second-level code. This approach can rapidly improve the accuracy of the machine learning model.

[0011] In some embodiments, the method further includes: generating a third intermediate representation for a third dialect of a programming language based on the second intermediate representation, wherein the second feature set of the second dialect includes the third feature set of the third dialect; and reverse compiling the third intermediate representation into third program code corresponding to a third programming language of the third dialect. This approach allows for the addition of program code for more programming languages, thereby increasing the number of comparable corpora.

[0012] In some embodiments, training a machine learning model for converting program code between different programming languages, based on first-level code and second-level code, includes: training a machine learning model for converting program code between different programming languages ​​based on first-level code, second-level code, and third-level code. This approach can rapidly improve the accuracy of the machine learning model in code conversion.

[0013] In some embodiments, the method further includes: in response to the addition of a fourth programming language for a third language, determining fourth program code corresponding to the third programming language of the third language; compiling the fourth program code into a fourth intermediate representation for the third language; and decompiling the fourth intermediate representation into fifth program code for the fourth programming language. This approach allows for the rapid addition of program code for new programming languages, thereby increasing the comparable corpus for the new programming languages.

[0014] In some embodiments, the method further includes adjusting the machine learning model based on the fourth and fifth program codes. This enables the machine learning model to accurately handle code translation for new programming languages.

[0015] In some embodiments, the method further includes: obtaining source code for a second programming language; and generating target code for a first programming language by inputting the source code into a trained machine learning model. This approach allows for rapid code conversion using a machine learning model.

[0016] According to a second aspect of this application, an apparatus for generating code is provided. The apparatus includes: a first program code determining unit configured to determine first program code for a first programming language corresponding to a first dialect of the programming language; a first intermediate representation generating unit configured to generate a first intermediate representation for the first dialect based on the first program code; a second intermediate representation generating unit configured to generate a second intermediate representation for a second dialect of the programming language based on the first intermediate representation, wherein a first feature set of the first dialect includes a second feature set of the second dialect; and a second program code generating unit configured to generate second program code for a second programming language corresponding to the second dialect based on the second intermediate representation.

[0017] In some embodiments, the first intermediate representation generation unit includes: a first rule determination unit configured to determine a first conversion rule from a first programming language to a first dialect; and a compilation unit configured to adjust the first program code into a first intermediate representation for the first dialect based on the first conversion rule.

[0018] In some embodiments, the second intermediate representation generation unit includes: a partial feature determination unit configured to determine partial features corresponding to the second feature set in the first feature set; and an adjustment unit configured to adjust the first intermediate representation to the second intermediate representation based on the partial features.

[0019] In some embodiments, the second program code generation unit includes: a second rule determination unit configured to determine a second conversion rule between a second programming language and a second dialect; and a code adjustment unit configured to adjust a second intermediate representation into second program code based on the second conversion rule.

[0020] In some embodiments, the apparatus further includes: a first training unit configured to train a machine learning model that can be used to implement program code conversion between different programming languages, based on a first level code and a second program code.

[0021] In some embodiments, the apparatus further includes: a third intermediate representation generation unit configured to generate a third intermediate representation for a third language of a programming language based on the second intermediate representation, wherein the second feature set of the second dialect includes the third feature set of the third language; and a third program code generation unit configured to reverse compile the third intermediate representation into third program code corresponding to the third programming language of the third language.

[0022] In some embodiments, the first training unit includes a second training unit configured to train a machine learning model that can be used to implement the conversion of program code in different programming languages, based on the first program code, the second program code, and the third program code.

[0023] In some embodiments, the apparatus further includes: a fourth program code determining unit configured to determine fourth program code corresponding to a third programming language in response to the addition of a fourth programming language for a third language; a fourth intermediate representation generating unit configured to compile the fourth program code into a fourth intermediate representation for the third language; and a fifth program code generating unit configured to reverse compile the fourth intermediate representation into fifth program code for the fourth programming language.

[0024] In some embodiments, the apparatus further includes a machine learning model adjustment unit configured to adjust the machine learning model based on a fourth program code and a fifth program code.

[0025] In some embodiments, the apparatus further includes: a source code acquisition unit configured to acquire source code for a second programming language; and a target code generation unit configured to generate target code for a first programming language by inputting the source code into a trained machine learning model.

[0026] According to a third aspect of this application, an electronic device is also provided, comprising: at least one computing unit; and at least one memory coupled to the at least one computing unit and storing instructions for execution by the at least one computing unit, the instructions, when executed by the at least one computing unit, causing the device to perform the method according to the first aspect of this application.

[0027] According to a fourth aspect of this application, a computer-readable storage medium is also provided, on which a computer program is stored, which, when executed by a processor, implements the method described according to a first aspect of this application.

[0028] According to a fifth aspect of this application, a computer program product is also provided, including computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method described according to a first aspect of this application.

[0029] Understandably, the apparatus of the second aspect, the electronic device of the third aspect, the computer storage medium of the fourth aspect, or the computer program product of the fifth aspect provided above are used to perform the method provided in the first aspect. Therefore, the explanations or descriptions regarding the first aspect also apply to the second, third, fourth, and fifth aspects. Furthermore, the beneficial effects achieved by the second, third, fourth, and fifth aspects can be referred to the beneficial effects in the corresponding methods, and will not be repeated here. Attached Figure Description

[0030] The above and other features, advantages and aspects of the embodiments of this application will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description.

[0031] In the accompanying drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0032] Figure 1 A schematic diagram illustrates an example environment in which several embodiments of this application can be implemented;

[0033] Figure 2 A schematic flowchart of a method for generating code according to some embodiments of this application is shown;

[0034] Figure 3 A schematic diagram illustrating the dialect hierarchy of programming languages ​​according to some embodiments of this application is shown;

[0035] Figure 4 A schematic diagram illustrating the entire process of generating code according to some embodiments of this application is shown;

[0036] Figure 5 A schematic diagram illustrating the generation of a low-level dialect from a high-level dialect according to some embodiments of this application is shown;

[0037] Figure 6 A schematic diagram illustrating a multilingual neural machine translation-based data synthesis process according to some embodiments of this application is shown;

[0038] Figure 7 A schematic diagram illustrating low-resource or resource-free programming language data synthesis according to some embodiments of this application is shown;

[0039] Figure 8 Block diagrams of apparatus for generating code according to some embodiments of this application are shown; and

[0040] Figure 9 A block diagram of a computing device capable of implementing several embodiments of the present application is shown. Detailed Implementation

[0041] Embodiments of this application will now be described in more detail with reference to the accompanying drawings. While some embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this application. It should be understood that the drawings and embodiments of this application are for illustrative purposes only and are not intended to limit the scope of protection of this application.

[0042] In the description of embodiments of this application, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0043] When using machine learning models to generate code in high-resource programming languages, performance typically prioritizes generating code in low-resource programming languages. This is related to the amount and proportion of each programming language in the training data. In natural language processing, different languages ​​within the same language family usually possess considerable transfer learning capabilities. However, strong transfer learning capabilities have not been observed between different programming languages.

[0044] One traditional approach is to use common compiler representations to enhance neural machine translation models. This approach combines program code with compiler representations to tune the machine learning model. While this method can achieve high accuracy when translating code from high-resource programming languages, its accuracy is lower when translating code from low-resource languages ​​due to limited data. Furthermore, the syntax of compiler representations can lead to the loss of structure and semantics after compilation. Another traditional approach is to fine-tune a pre-trained CodeLarge Language Model (CLM) using semi-synthetic training data to improve its performance in generating low-resource programming languages. This technique first uses the CLM to generate unit tests in high-resource programming languages ​​and then translates these unit tests into low-resource languages. Then, the CLM generates code in low-resource languages, and the translated unit tests are used to verify the correctness of the generated code. However, this technique only ensures code correctness and cannot guarantee structural or semantic fidelity; additionally, because it relies on the CLM to generate code, it does not support resource-free programming languages. While this technique includes translation unit tests to verify the correctness of low-resource code, it still relies on an underlying virtual machine to generate the correct low-resource code, resulting in extremely poor efficiency. Research shows that comparable corpus code data can significantly improve the pre-training and fine-tuning of machine learning models. However, in practical applications, it is difficult to quickly obtain low-resource programming language code data, thus hindering further improvements to neural machine translation models.

[0045] To address at least some of the aforementioned problems and other potential issues, in embodiments of this application, the computing device may first determine first program code in a first programming language corresponding to a first dialect of the programming language. Then, the computing device compiles the first program code into a first intermediate representation for the first dialect. Next, the computing device can use this first intermediate representation to generate a second intermediate representation for a second dialect of the programming language, wherein a first feature set of the first dialect includes a second feature set of the second dialect. Finally, the computing device can reverse-compile the second intermediate representation into second program code in a second programming language corresponding to the second dialect. In this way, by converting the program code to the intermediate representation of the corresponding dialect, and then through the conversion between intermediate representations of different dialects, the program code corresponding to the programming language dialect in another dialect can be finally obtained. This allows for the rapid acquisition of corresponding program code even for low-resource programming languages, improving the efficiency of obtaining comparable corpora, increasing the number of comparable corpora, and improving the user experience.

[0046] Figure 1 A schematic diagram of an example environment 100 in which various embodiments of this application can be implemented is shown. For example... Figure 1As shown, example environment 100 includes computing device 104, which can be used to generate corpora for training machine learning models, such as generating comparable corpora for training neural machine translation models, including program code for different programming languages. Computing device 104 includes, but is not limited to, personal computers, server computers, handheld or laptop devices, mobile devices (such as mobile phones, personal digital assistants (PDAs), media players, etc.), multiprocessor systems, consumer electronics, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices.

[0047] To enable code conversion between programming languages, multiple programming languages ​​can be categorized into different dialect families. Each dialect family has a corresponding dialect, which preserves the structure and semantics of the programming languages ​​within that dialect family and can be considered a normalized representation of all programming languages ​​within that dialect family. Program code in a programming language within a dialect family can be compiled into program code in that dialect, and conversely, program code for that dialect can be decompiled back into program code in the original programming language. Furthermore, different dialects within different dialect families have different levels of authority. For example, multiple programming languages ​​such as Rust, C++, C, Java, JavaScript, Julia, and Python can be categorized into different dialect families. Rust belongs to one dialect family, C++, C, and Java to another, and JavaScript, Julia, and Python to yet another. Additionally, each of these three dialect families can have a corresponding dialect, which can be obtained by normalizing all programming languages ​​within each dialect family. Furthermore, dialects within the Rust dialect family can be at a higher level than those within the C++ dialect family. A higher-level dialect's feature set includes the feature set of a lower-level dialect; that is, the feature set of a lower-level dialect is a subset of the feature set of a higher-level dialect. A dialect's feature set typically includes dialectal characteristic information that describes the dialect's rules and requirements, such as how to declare variables, write functions, declare structs and classes, and define modules.

[0048] The computing device 104 can utilize the dialects within each dialect family to perform program code conversion. Since some information is inevitably lost when converting from a high-level dialect to a low-level dialect, it is impossible to convert from a low-level dialect back to a high-level dialect. Therefore, it is generally only permissible to use the dialect to convert program code in the programming language corresponding to the high-level dialect into program code in the programming language corresponding to the low-level dialect.

[0049] For example, computing device 104 receives first program code 102, which is written in a first programming language belonging to a first dialect family, which includes a first dialect 108. In this case, computing device 104 can acquire code from a lower-level programming language corresponding to the first program code to form a comparable corpus. A comparable corpus refers to a set of corpora that are similar or comparable in some aspects. The obtained comparable corpus can be used to train a machine learning model that supports code conversion between different programming languages.

[0050] Next, the computing device 104 can use the first program code 102 to generate a first intermediate representation 106 for the first dialect 108, which is represented using statements in the first dialect. For example, the first program code 102 is compiled into the first intermediate representation 106. Then, the computing device 104 can convert the first intermediate representation 106 for the first dialect 108 into a second intermediate representation 112 for a lower-level second dialect 114. In order to achieve inter-dialect conversion, the structure and semantics of the dialects should be consistent. Therefore, the first intermediate representation 106 is converted into the second intermediate representation 112 by using pre-set conversion rules.

[0051] After obtaining the second intermediate representation 112, the computing device 104 can use the second intermediate representation 112 to generate second program code 110 for the second programming language. For example, the computing device 104 can reverse compile the second intermediate representation 112, which is a compilation that converts one code form into another. Therefore, through reverse compilation, the second intermediate representation 112 can be reverse compiled into the second program code 110, which is for the second programming language in the second dialect family to which the second dialect 114 belongs.

[0052] After obtaining the second program code 110, combined with the first program code 102, a corpus for training the machine learning model 116 can be formed. Using this corpus, the machine learning model 116 can be trained or fine-tuned. The trained or adjusted machine learning model can perform conversions from program code in the first programming language to program code in the second programming language, or vice versa. Furthermore, comparable corpora formed from program code in different programming languages ​​can be obtained in this way, enabling the machine learning model to achieve rapid conversion between program code in different programming languages. Additionally, this machine learning model can be a neural machine translation model.

[0053] This method allows for the conversion of program code to an intermediate representation of the corresponding dialect, and then the conversion between intermediate representations of different dialects, ultimately yielding the program code corresponding to another programming language dialect. This enables the rapid acquisition of corresponding program code even for low-resource programming languages, improving the efficiency of obtaining comparable corpora, increasing the number of comparable corpora, and enhancing the user experience.

[0054] The above combination Figure 1 A schematic diagram illustrating an example environment 100 in which embodiments of this application can be implemented is described. The following is in conjunction with... Figure 2 A schematic flowchart illustrating Example 200 of a method for generating code according to some embodiments of this application. Example 200 can be provided by... Figure 1 The computing device 104 and any suitable computing device in it shall execute.

[0055] At box 202, computing device 104 determines first program code in a first programming language corresponding to a first dialect of the programming language. To obtain machine learning models, such as neural machine translation models, that implement the conversion between program code in different programming languages ​​for training, comparable corpora generated from program code in different programming languages ​​are needed. To generate comparable corpora, computing device 104 may first obtain a program code, such as program code written in C. Since C can be classified into a dialect family, the dialect of the dialect family to which C belongs can be called the first dialect. The program code written in C conforms to the rules and requirements of C's syntax. In one example, the first program code is provided to computing device 104 by a user. In another example, the first program code is obtained by computing device 104 via a network, for example, from a predetermined program code library.

[0056] At box 204, computing device 104 generates a first intermediate representation for a first dialect based on the first program code. To facilitate the conversion of the first program code into program code of a lower-level programming language, the first program code needs to be compiled into the dialect code of the dialect family in which the programming language of the program code belongs.

[0057] In some embodiments, computing device 104 may first determine conversion rules from the first programming language dialect to the first dialect. Then, computing device 104 uses these conversion rules to adjust or compile the first program code into a first intermediate representation for the first dialect. For example, the rules define the correspondence between the syntax, semantics, and functions of the two languages, such as the correspondence and conversion methods between functions in the first programming language and functions in the first dialect, or how parameters in the first programming language are converted to parameters in the first dialect. Computing device 104 can use these rules to convert the first program code to the first intermediate representation for the first dialect. In some embodiments, a compiler for converting from the first programming language to the first dialect can be pre-generated, and then the compiler can be used to implement the conversion from the first program code to the first intermediate representation. In some embodiments, a pre-trained machine learning model can also be used to implement the conversion from the first program code to the first intermediate representation. The above examples are merely for describing this disclosure and are not intended to limit the scope of this disclosure.

[0058] At box 206, based on the first intermediate representation, computing device 104 generates a second intermediate representation for a second dialect of a programming language, wherein the first feature set of the first dialect includes the second feature set of the second dialect. After obtaining the first intermediate representation, computing device 104 can use it to convert to a second intermediate representation of a second dialect that is at a lower level than the first dialect. For example, if the first intermediate representation corresponds to a dialect of the C language family, the second dialect could be a dialect of the Julia language family. Alternatively, if the first intermediate representation corresponds to a dialect of the Rust language family, the second dialect could be a dialect of the C language family.

[0059] In some embodiments, when generating a second intermediate representation for a second dialect of a programming language using a first intermediate representation, the computing device 104 may first determine some features in the first feature set that correspond to the second feature set. Since the first feature set of the first dialect includes the second feature set of the second dialect, the first feature set also includes features not included in the second feature set. At this time, the computing device 104 may first determine the features that correspond to the first and second feature sets; for example, if both feature sets have an addition function implementation, these common corresponding features can be identified. Then, the computing device 104 uses these features to adjust the first intermediate representation to the second intermediate representation. If the first intermediate representation contains a part not present in the second feature set, that part is removed from the first intermediate representation. In some embodiments, a machine learning model for converting intermediate representations may also be pre-trained, and then this pre-trained machine learning model may be used to implement the conversion from the first intermediate representation to the second intermediate representation. The above examples are merely for describing this disclosure and are not intended to limit the scope of this disclosure.

[0060] At box 208, computing device 104 generates second program code in a second programming language corresponding to a second dialect based on the second intermediate representation. After obtaining the second intermediate representation, computing device 104 can perform decompilation to convert the second intermediate representation into second program code written in the second programming language corresponding to the second dialect.

[0061] In some embodiments, when decompiling a second intermediate representation into second program code corresponding to a second programming language of a second dialect, the computing device 104 may first determine a second conversion rule between the second programming language and the second dialect. Then, the computing device 104 further utilizes this second conversion rule to adjust the second intermediate representation into the second program code. For example, the second conversion rule defines the correspondence between the syntax rules of the second dialect and the semantics, syntax rules, functions, and other information of the second programming language (e.g., the Julia programming language). This correspondence is used to achieve the conversion from the second intermediate representation to the second program code. In some embodiments, a decompiler may be pre-generated, and the conversion from the second intermediate representation to the second program code may be implemented using the decompiler. The above examples are merely for describing this disclosure and are not intended to limit the scope of this disclosure.

[0062] After the computing device 104 obtains the second program code according to the above operations, it can combine it with the first program code to form a comparable corpus. Furthermore, the above method can obtain comparable corpora of program codes from different programming languages ​​within the dialect families of two different levels of dialects. The comparable corpus can be used to train a machine learning model for converting program codes from different programming languages. For example, the first program code can be input into a machine learning model to generate program code for a second programming language, and then the generated program code can be compared with the second program code to adjust the machine learning model; alternatively, the second program code can be used as input, and the machine learning model can be used to generate program code for a first programming language, thereby comparing the generated program code with the first program code to adjust the machine learning model. The first programming language can be any programming language within the dialect family of the first dialect, and the second programming language can be any programming language within the dialect family of the second dialect. Additionally, a first intermediate representation and a second intermediate representation can be added during training to train the machine learning model.

[0063] While the process of obtaining comparable corpora between two dialect groups has been described above, the process of obtaining comparable corpora between multiple dialect groups is similar. For example, there may also exist a third dialect at a lower level than the second dialect. In this case, the second intermediate representation can be further transformed. For example, computing device 104 can use the second intermediate representation to generate a third intermediate representation for a third dialect of a programming language, where the second feature set of the second dialect includes the third feature set of the third dialect. For example, computing device 104 transforms the second intermediate representation into the third intermediate representation according to the transformation rules between the second dialect and the third dialect. Then, the third intermediate representation can be reverse-compiled into third program code corresponding to the third programming language of the third dialect. At this point, comparable corpora formed by the first program code, the second program code, and the third program code can be obtained. Therefore, computing device 104 can further train a machine learning model for the transformation of program code for different programming languages. For example, using the first program code, the second program code, and the third program code, a machine learning model that can be used to implement the transformation of program code for different programming languages ​​can be trained. Additionally, during the training of this machine learning model, a first intermediate representation, a second intermediate representation, and a third intermediate representation can be further added for training.

[0064] If a new programming language is added to a dialect family, the lack of corpus data for this language may prevent machine learning models used for converting program code between different programming languages ​​from converting the code for this new language. Therefore, the method described above can be used to convert the code for a higher-level programming language into the code for this new programming language. Furthermore, the computing device 104 can also convert the code for other existing programming languages ​​within the same dialect family into the code for this new programming language. For example, if a new fourth programming language is added to a third dialect, the fourth program code written in the third programming language within the dialect family of that third dialect can be obtained first. Then, the fourth program code is compiled into a fourth intermediate representation for the third dialect. Next, the computing device 104 further decompiles the fourth intermediate representation into a fifth program code for the fourth programming language. Additionally, the computing device 104 can further obtain program code written in other programming languages ​​within the dialect family of the third dialect, compile it into the intermediate representation of the third dialect, and then further decompile it into program code written in the dialect of the fourth programming language. Once these comparable corpora are obtained, the fourth and fifth program codes can be used to train the machine learning model. Comparable corpora for new programming language dialects can also be generated in this way in other dialect groups.

[0065] However, if the newly added programming language belongs to a high-level dialect family, it may be impossible to convert program code from other dialect families to this new programming language. In this case, program code for the new programming language can be generated using program code from other programming languages ​​within the same dialect family. Then, the machine learning model can be trained using the corresponding program code from different programming languages ​​within the same dialect family. Additionally, the machine learning model can be further trained by combining program code from other programming languages ​​with program code from other dialect families.

[0066] In some embodiments, after the machine learning model for converting program code between different programming languages ​​has been trained, program code in any programming language can be received to generate program code in another programming language. For example, source code for a second programming language can be obtained. Then, the computing device 104 inputs the source code into the trained machine learning model to generate target program code for a first programming language.

[0067] This method allows for the conversion of program code into an intermediate representation of the corresponding dialect, and then the conversion between intermediate representations of different dialects, ultimately yielding the program code corresponding to another programming language dialect. This enables the rapid acquisition of corresponding program code even for low-resource programming languages, improving the efficiency of obtaining comparable corpora, increasing the number of comparable corpora, and enhancing the user experience.

[0068] The above combination Figure 2 A schematic flowchart illustrating a method for generating code according to some embodiments of this application is described below. Figure 3 A schematic diagram illustrating the dialect hierarchy of programming languages ​​according to some embodiments of this application.

[0069] In Example 300, computing device 104 uses nested circles to represent the dialect hierarchy of different programming languages. Dialect-A, Dialect-B, and Dialect-C in the diagram represent languages ​​at different levels, and the nesting relationship indicates that the feature set of an inner language is a subset of the feature set of an outer language. That is, each outer language contains features of its inner language, but outer languages ​​can be more extensive and flexible, and allow more features, while inner languages ​​may be more strict and standardized. The outermost nested circle containing Dialect-A represents the language with the most extensive features and fewest restrictions. It includes lifecycle management, features, and unsafe features, indicating that programming languages ​​within this scope may allow manual memory management or direct access to underlying resources, etc. It also includes features of other dialects, such as interfaces, immutability, object-oriented programming, and types.

[0070] The circle for dialect-A includes a circle for dialect-B. Dialect-B includes features such as interfaces and immutability. The addition of interfaces allows for modular development through abstraction, improving code extensibility, such as the interface mechanism between Java and C. Immutability, on the other hand, is a common programming paradigm that enhances concurrency safety and prevents accidental data modification. However, dialect-B does not possess features such as lifecycles found in dialect-A.

[0071] The innermost circle containing dialect-C is included within the circle containing dialect-B. It includes the characteristics of object-oriented programming and types, which means that this type of programming language emphasizes structured design, such as organizing code through classes, inheritance, and polymorphism. It also introduces a strict type system, such as Python and Julia. However, it does not include the interface and immutability characteristics of dialect-B, nor the lifecycle characteristics of dialect-A.

[0072] This approach, employing a nested structure, demonstrates the progressive process from dialect A to dialect C. By utilizing the similarity and equivalence relationships between programming languages ​​for normalization, a hierarchical relationship between programming languages ​​is established, facilitating the analysis of the similarities, equivalence structures, and differences in human factors engineering among different languages.

[0073] The above combination Figure 3 A schematic flowchart illustrating examples of dialect hierarchy relationships in programming languages ​​according to some embodiments of this application is provided below. Figure 4 A schematic diagram illustrating the entire process of generating code according to some embodiments of this application.

[0074] Example 400 illustrates multiple programming languages ​​corresponding to various dialect sets and demonstrates a compiler enhancement method based on intermediate representation (IR) (e.g., dialects). This method optimizes the code transformation process by defining dialects for different programming languages ​​and constructing a dialect hierarchy. The method divides multiple programming languages ​​into different dialect sets and establishes a hierarchical relationship within these sets. Dialect-A, Dialect-B, and Dialect-C represent sets of programming languages ​​with different complexities and feature sets, respectively. Dialect-A 402 is the highest-level dialect, possessing the richest feature set. For example, the Rust language, known for its strict memory safety, lifecycle management, and complex type system, is classified in Dialect-A 402. Dialect-B 404 is at an intermediate level, representing a set of dialects with relatively concise features but still strong expressive power, including C++, C, and Java. These languages ​​share some similarities with Rust in terms of performance, structure, and semantics, but their feature sets are slightly fewer than those in Dialect-A. Dialect-C406 represents the lowest level of language, with the most simplified features, and includes JavaScript, Julia, and Python. Although these languages ​​have simplified features, they still retain their core structure and semantics to facilitate recognition and processing by neural machine translation models during the translation process. Taking the Rust language as an example, Rust code is first compiled into dialect-A (402) using compiler facilities. This process preserves Rust's core structure and semantics, expressing it in dialect-A form. Then, dialect-A code can be further downgraded through a downgrade mechanism into dialect-B and dialect-C, gradually losing some information during the downgrade process. For example, dialect-A402 contains lifecycle-related syntax; after downgrading, this information is lost because dialect-B lacks this syntax.

[0075] In the pre-training process of the neural machine translation model 408, data from multiple programming languages ​​and their corresponding dialect representations were introduced. The pre-training phase not only learned the characteristics of each programming language but also achieved consistent modeling of language features and semantics through the hierarchical relationship between dialect-A, dialect-B, and dialect-C. To further improve model performance, this method introduced a large amount of comparable corpus during the model fine-tuning phase. Comparable corpus refers to parallel datasets formed by code snippets with the same function or similar structure implemented in different programming languages. For example, conversion pairs such as C language code to Python code and RUST code to Python code can form comparable corpus. This data effectively helps the neural machine translation model 408 identify structural alignment points and semantic correspondences between different languages, thereby further enhancing model performance and cross-language transfer capabilities. Using this neural machine translation model, high-level language code can be regenerated from low-level language code. The introduction of this mechanism largely compensates for the negative impact of information loss during the downgrading process, especially in helping to recover complex structures and high-level semantic features in high-level languages.

[0076] This method, combining intermediate representations and neural machine translation models, reduces the cost of multiple inference steps and improves the quality of code translation results. Simultaneously, by unifying the semantics and structure across multiple dialects, it significantly enhances support for low-resource and resource-free programming languages, reduces reliance on large amounts of manually intervened data, and provides a more efficient and accurate solution for code conversion in multilingual programming scenarios.

[0077] The above combination Figure 4 A schematic diagram illustrating the entire process of generating code according to some embodiments of this application is provided below. Figure 5 The illustration depicts the generation of a low-level dialect from a high-level dialect according to some embodiments of this application. Figure 5 Dialect-A 502, Dialect-B 504 and Dialect-C 506 are shown.

[0078] Example 500 illustrates the complete process of synthesizing a low-level dialect programming language from a high-level dialect programming language, specifically using the conversion from C to Julia as an example. It details how to achieve efficient and accurate data synthesis using dialect hierarchies and intermediate representation techniques. The process begins with existing C language data. Since C belongs to an intermediate dialect—dialect-B 504—the first step is to compile the original C code into dialect-B form. This compilation process, aided by compiler facilities, transforms the core logic, syntactic structure, and data flow of the C code into the general expression form of dialect-B. Because dialect-B is an intermediate dialect with a richer feature set, many complex semantic features of the original C language can be fully preserved during this process.

[0079] After compiling from C to dialect-B, the next step is to perform a downgrade operation, further converting dialect-B to dialect-C 506. During the downgrade process, because dialect-C506 is a lower-level dialect with a more streamlined feature set, some advanced features may be lost. Since dialect-C506 is closer to dynamic or scripting languages, it may not fully cover these complex mechanisms. Therefore, during the downgrade process, some complex elements of the original C code are appropriately simplified or replaced with equivalent structures to ensure that the converted dialect-C code retains its core semantics and has a more concise structure, allowing for smoother subsequent conversion to the target language.

[0080] After successfully downgrading the C language code to dialect-C, the final step is to use a decompilation mechanism to convert the dialect-C code into the Julia language. Furthermore, with this corpus available, a neural machine translation model 508 can be trained, allowing the model to decode the dialect-C code and reassemble it into an equivalent Julia implementation.

[0081] Ultimately, the generated Julia code and the original C code form a naturally comparable corpus. This naturally comparable corpus is crucial for further fine-tuning and training of the neural machine translation model. By introducing this data, the model can not only more accurately capture the structural transformation patterns between different programming languages, but also possess stronger transformational learning capabilities when facing low-resource languages ​​or newer programming languages. This method effectively improves the automation of data synthesis while reducing the cost of relying on extensive manual annotation and manual compilation of comparable data in traditional methods.

[0082] The above combination Figure 5 A schematic diagram illustrating the synthesis of a low-level dialect from a high-level dialect according to some embodiments of this application is described below. Figure 6 A schematic diagram illustrating a multilingual neural machine translation-based data synthesis process according to some embodiments of this application.

[0083] In Example 600, Rust, C, and Python belong to different dialect levels. Rust is in the highest-level dialect -A 602, C is in the mid-level dialect -B 604, and Python is in the lowest-level dialect -C 606. During the data transformation process, Rust code can be compiled to dialect -A and, if necessary, reverse-compiled back to the original Rust code. Similarly, C code can be compiled to dialect -B, and Python code to dialect -C. When dialect -A code needs to be downgraded, a degradation mechanism can demote it to dialect -B or dialect -C. This degradation operation will result in the loss of some high-level syntactic features, but it can be used to simplify code structure and verify the correctness of code logic.

[0084] During the training of the neural machine translation model 608, a comparable corpus is first constructed using multilingual code samples extracted from dialects A, B, and C. This corpus includes corresponding code snippets between languages ​​such as Rust, C, and Python. The neural machine translation model uses this data for pre-training to learn certain mapping relationships and common conversion rules between different programming languages. After pre-training, the model can be further fine-tuned using the comparable corpus 610 to improve the translation quality between specific language pairs.

[0085] This method enables a fully trained neural machine translation model to achieve mutual conversion between multiple programming languages, particularly supporting bidirectional conversion between C, Rust, and Python. Furthermore, by combining multilingual dialects, comparable corpora, and the neural machine translation model, it not only improves the accuracy of programming language translation but also offers significant advantages over traditional methods in supporting low-resource and resource-poor programming languages.

[0086] The above combination Figure 6 A schematic diagram illustrating a multilingual neural machine translation-based data synthesis process according to some embodiments of this application is provided below. Figure 7 A schematic diagram illustrating low-resource or resource-free programming language data synthesis according to some embodiments of this application.

[0087] Example 700 illustrates how to integrate low-resource or resource-free programming languages ​​into existing multilingual compilation and translation systems, reusing existing facilities to achieve more efficient programming language conversion and training. Example 700 includes dialects A 702, B 704, and C 706, with GDScript as an example, describing how to enhance support for new languages ​​through existing compiler frameworks and neural machine translation models 708.

[0088] First, by implementing compilation and reverse compilation capabilities from GDScript to dialect-C706, GDScript code can be represented and converted within the dialect-C framework. This mechanism allows GDScript code to be transformed into dialect-C form. With the addition of dialect-C, GDScript code can be converted to and from other dialect-C languages, such as Python, JavaScript, and Julia, generating a large amount of comparable corpus. This comparable corpus provides crucial data support for fine-tuning the neural machine translation model 708. By utilizing existing neural machine translation models and combining them with these newly synthesized corpora for fine-tuning, the translation performance of GDScript with other languages ​​can be further improved.

[0089] Meanwhile, since GDScript is located at the dialect-C level, a degradation mechanism can be used to gradually downgrade high-level languages ​​to dialect-C form and ultimately convert them into GDScript code. However, if the newly added low-resource or no-resource programming language belongs to the dialect-A702 level, in order to support the efficient translation of multiple programming languages ​​into this language, the neural machine translation model needs to undergo a more complex pre-training and fine-tuning process to ensure that the model can understand and correctly translate the syntax and features of the programming language.

[0090] This method, which incorporates new programming languages ​​into the dialect system, enables faster language expansion, makes full use of existing resources, and achieves efficient construction of multilingual programming environments.

[0091] Figure 8 A block diagram of an apparatus 800 for generating code according to an embodiment of this application is further shown. The apparatus 800 is applied to a computing device and may include multiple modules for performing tasks such as... Figure 2 The corresponding steps in example 200 of the method discussed herein. For example... Figure 8 As shown, the apparatus 800 includes: a first program code determination unit 802 configured to determine a first program code for a first programming language corresponding to a first dialect of the programming language; a first intermediate representation generation unit 804 configured to generate a first intermediate representation for the first dialect based on the first program code; a second intermediate representation generation unit 806 configured to generate a second intermediate representation for a second dialect of the programming language based on the first intermediate representation, wherein a first feature set of the first dialect includes a second feature set of the second dialect; and a second program code generation unit 808 configured to generate a second program code for a second programming language corresponding to the second dialect based on the second intermediate representation.

[0092] In some embodiments, the first intermediate representation generation unit 804 includes: a first rule determination unit configured to determine a first conversion rule from a first programming language to a first dialect; and a compilation unit configured to adjust the first program code into a first intermediate representation for the first dialect based on the first conversion rule.

[0093] In some embodiments, the second intermediate representation generation unit 806 includes: a partial characteristic determination unit configured to determine partial characteristics corresponding to the second characteristic set in the first characteristic set; and an adjustment unit configured to adjust the first intermediate representation to the second intermediate representation based on the partial characteristics.

[0094] In some embodiments, the second program code generation unit 808 includes: a second rule determination unit configured to determine a second conversion rule between a second programming language and a second dialect; and a code adjustment unit configured to adjust a second intermediate representation into second program code based on the second conversion rule.

[0095] In some embodiments, the device 800 further includes: a first training unit configured to train a machine learning model that can be used to implement program code conversion between different programming languages ​​based on a first level code and a second program code.

[0096] In some embodiments, the apparatus 800 further includes: a third intermediate representation generation unit configured to generate a third intermediate representation for a third language of a programming language based on the second intermediate representation, wherein the second feature set of the second dialect includes the third feature set of the third language; and a third program code generation unit configured to reverse compile the third intermediate representation into third program code corresponding to the third programming language of the third language.

[0097] In some embodiments, the first training unit includes a second training unit configured to train a machine learning model that can be used to implement the conversion of program code in different programming languages, based on the first program code, the second program code, and the third program code.

[0098] In some embodiments, the apparatus 800 further includes: a fourth program code determining unit configured to determine fourth program code corresponding to a third programming language in response to the addition of a fourth programming language for a third language; a fourth intermediate representation generating unit configured to compile the fourth program code into a fourth intermediate representation for the third language; and a fifth program code generating unit configured to reverse compile the fourth intermediate representation into fifth program code for the fourth programming language.

[0099] In some embodiments, the device 800 further includes a machine learning model adjustment unit configured to adjust a machine learning model based on a fourth program code and a fifth program code.

[0100] In some embodiments, the apparatus 800 further includes: a source code acquisition unit configured to acquire source code for a second programming language; and a target code generation unit configured to generate target code for a first programming language by inputting the source code into a trained machine learning model.

[0101] Figure 9 A schematic block diagram of an example device 900 that can be used to implement embodiments of the present application is shown. For example, embodiments of the present application... Figure 1 The computing device 104 can be implemented by the example device 900. As shown, device 900 includes a central processing unit (CPU) 901, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 902 or loaded from storage unit 908 into random access memory (RAM) 903. Various programs and data required for the operation of device 900 can also be stored in RAM 903. CPU 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904.

[0102] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of monitors, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0103] The various processes and handling described above, such as method example 200, can be executed by processing unit 901. For example, in some embodiments, method example 200 can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by CPU 901, one or more actions of method example 200 described above can be performed.

[0104] This application may be a method, apparatus, system, chip, and / or computer program product. A chip may include a processing unit and a communication interface, the processing unit being capable of processing program instructions received from the communication interface. A computer program product may include a computer-readable storage medium on which computer-readable program instructions for performing various aspects of this application are stored.

[0105] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0106] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0107] The computer program instructions used to perform the operations of this application may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuits, such as programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), are personalized by utilizing state information from the computer-readable program instructions. These electronic circuits can execute the computer-readable program instructions to implement various aspects of this application.

[0108] Various aspects of this application are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0109] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0110] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0111] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0112] The various embodiments of this application have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for generating code, characterized in that, The method includes: Determine the first program code of the first programming language corresponding to the first dialect of the programming language; Based on the first program code, a first intermediate representation for the first dialect is generated; Based on the first intermediate representation, a second intermediate representation for a second dialect of the programming language is generated, wherein the first feature set of the first dialect includes the second feature set of the second dialect; and Based on the second intermediate representation, second program code corresponding to the second dialect in the second programming language is generated.

2. The method according to claim 1, characterized in that, Based on the first program code, generating the first intermediate representation for the first dialect includes: Determine the first conversion rule from the first programming language to the first dialect; and Based on the first conversion rule, the first program code is adjusted to the first intermediate representation for the first dialect.

3. The method according to claim 1, characterized in that, Generating a second intermediate representation for a second dialect of a programming language includes: Determine the partial characteristics in the first characteristic set that correspond to the second characteristic set; and Based on the aforementioned characteristics, the first intermediate representation is adjusted to the second intermediate representation.

4. The method according to claim 1, characterized in that, Based on the second intermediate representation, generating second program code corresponding to the second dialect in the second programming language includes: Determine the second conversion rule between the second programming language and the second dialect; and Based on the second conversion rule, the second intermediate representation is adjusted to the second program code.

5. The method according to claim 1, characterized in that, Also includes: Based on the first level code and the second program code, a machine learning model is trained that can be used to convert program code between different programming languages.

6. The method according to claim 5, characterized in that, Also includes: Based on the second intermediate representation, a third intermediate representation for a third dialect of the programming language is generated, wherein the second feature set of the second dialect includes the third feature set of the third dialect; and The third intermediate representation is reverse-compiled into third program code corresponding to the third programming language.

7. The method of claim 6, wherein training a machine learning model for converting program code between different programming languages, based on the first level code and the second program code, comprises: Based on the first program code, the second program code, and the third program code, a machine learning model is trained that can be used to convert program code between different programming languages.

8. The method according to claim 6, characterized in that, Also includes: In response to the addition of a fourth programming language for a third-party language, a fourth program code corresponding to the third programming language of the third-party language is determined; The fourth program code is compiled into a fourth intermediate representation for the third language; The fourth intermediate representation is reverse-compiled into fifth program code for the fourth programming language.

9. The method according to claim 8, characterized in that, Also includes: The machine learning model is adjusted based on the fourth and fifth program codes.

10. The method according to claim 5, characterized in that, Also includes: Obtain the source code for the second programming language; as well as By inputting the source code into the trained machine learning model, target code for the first programming language is generated.

11. An apparatus for generating code, characterized in that, The device includes: The first program code determining unit is configured to determine the first program code of the first programming language corresponding to the first dialect of the programming language; The first intermediate representation generation unit is configured to generate a first intermediate representation for the first dialect based on the first program code. The second intermediate representation generation unit is configured to generate a second intermediate representation for a second dialect of a programming language based on the first intermediate representation, wherein a first feature set of the first dialect includes a second feature set of the second dialect; and The second program code generation unit is configured to generate second program code corresponding to the second dialect in a second programming language based on the second intermediate representation.

12. The apparatus according to claim 11, characterized in that, The first intermediate representation generation unit includes: The first rule-determining unit is configured to determine a first conversion rule from the first programming language to the first dialect; and The compilation unit is configured to adjust the first program code into the first intermediate representation for the first dialect based on the first conversion rule.

13. The apparatus according to claim 11, characterized in that, The second intermediate representation generation unit includes: A partial characteristic determination unit is configured to determine partial characteristics in the first characteristic set that correspond to the second characteristic set; and The adjustment unit is configured to adjust the first intermediate representation to the second intermediate representation based on the aforementioned partial characteristics.

14. The apparatus according to claim 11, characterized in that, The second program code generation unit includes: The second rule-determining unit is configured to determine a second conversion rule between the second programming language and the second dialect; and The code adjustment unit is configured to adjust the second intermediate representation into the second program code based on the second conversion rule.

15. The apparatus according to claim 11, characterized in that, Also includes: The first training unit is configured to train a machine learning model that can be used to implement program code conversion between different programming languages, based on the first level code and the second program code.

16. The apparatus according to claim 15, characterized in that, Also includes: The third intermediate representation generation unit is configured to generate a third intermediate representation for a third dialect of a programming language based on the second intermediate representation, wherein the second feature set of the second dialect includes the third feature set of the third dialect; as well as The third program code generation unit is configured to reverse compile the third intermediate representation into third program code corresponding to the third programming language.

17. The apparatus of claim 16, wherein the first training unit comprises: The second training unit is configured to train a machine learning model that can be used to convert program code between different programming languages, based on the first program code, the second program code, and the third program code.

18. The apparatus according to claim 16, characterized in that, Also includes: The fourth program code determining unit is configured to determine a fourth program code corresponding to a third programming language in response to the addition of a fourth programming language for the third language; The fourth intermediate representation generation unit is configured to compile the fourth program code into a fourth intermediate representation for the third language; The fifth program code generation unit is configured to reverse compile the fourth intermediate representation into fifth program code for the fourth programming language.

19. The apparatus according to claim 18, characterized in that, Also includes: The machine learning model adjustment unit is configured to adjust the machine learning model based on the fourth program code and the fifth program code.

20. The apparatus according to claim 15, characterized in that, Also includes: The source code acquisition unit is configured to acquire source code for the second programming language; as well as The target program code generation unit is configured to generate target program code for the first programming language by inputting the source program code into the trained machine learning model.

21. An electronic device, comprising: At least one computing unit; At least one memory coupled to the at least one computing unit and storing instructions for execution by the at least one computing unit, the instructions, when executed by the at least one computing unit, causing the device to perform the method according to any one of claims 1-10.

22. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method according to any one of claims 1-10.

23. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1-10.