Methods, apparatus, devices, storage media, and program products for code conversion.

By using machine learning models to perform code conversion based on mapping relationships and descriptive information, the problems of high cost and low accuracy in existing technologies are solved, and efficient and accurate code conversion is achieved.

CN122489078APending Publication Date: 2026-07-31QUANXIN INTELLIGENT MFG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QUANXIN INTELLIGENT MFG TECH CO LTD
Filing Date
2026-06-29
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing code conversion methods rely on manually maintaining rule bases, resulting in high costs and low accuracy, making it difficult to adapt to the code conversion needs of different users.

Method used

A machine learning model is used to perform code conversion based on the mapping relationship and descriptive information between source language instructions and target language instructions, reducing the dependence on rule bases and improving conversion accuracy.

Benefits of technology

Using machine learning models for code conversion reduces the difficulty and cost of code conversion, improves the accuracy of conversion, and reduces the need for maintaining the rule base.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122489078A_ABST
    Figure CN122489078A_ABST
Patent Text Reader

Abstract

According to embodiments of this disclosure, a method, apparatus, device, storage medium, and program product for code conversion are provided. The method includes: determining at least one source language instruction included in source language code to be converted; determining at least one target language instruction corresponding to each of the at least one source language instruction based on a mapping relationship between the source language instructions and target language instructions, the mapping relationship indicating the target language instruction corresponding to each source language instruction; obtaining description information for each of the at least one source language instruction and description information for each of the at least one target language instruction; and using a machine learning model, based on the mapping relationship, the description information for each of the at least one source language instruction, and the description information for each of the at least one target language instruction, performing code conversion on the source language code to determine target language code corresponding to the source language code, the target language code including at least one target language instruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of photolithography, and particularly to methods, apparatus, devices, computer-readable storage media, and computer program products for code conversion. Background Technology

[0002] Code conversion, also known as code translation, refers to the process of converting code from one language to another. The code to be converted is typically called the source language code, and the converted code is called the target language code. Code conversion offers several advantages, such as allowing the converted target language code to run in the target language environment, resolving interoperability issues between different systems and platforms, and improving code portability and cross-platform capabilities. Summary of the Invention

[0003] In a first aspect of this disclosure, a method for code conversion is provided. The method includes: determining at least one source language instruction included in source language code to be converted; determining at least one target language instruction corresponding to each of the at least one source language instruction based on a mapping relationship between the source language instructions and target language instructions, the mapping relationship indicating the target language instruction corresponding to each source language instruction; obtaining description information for each of the at least one source language instruction and description information for each of the at least one target language instruction; and using a machine learning model, based on the mapping relationship, the description information for each of the at least one source language instruction, and the description information for each of the at least one target language instruction, performing code conversion on the source language code to determine target language code corresponding to the source language code, the target language code including at least one target language instruction.

[0004] In a second aspect of this disclosure, an apparatus for code conversion is provided. The apparatus includes: a first instruction determination module configured to determine at least one source language instruction included in source language code to be converted; a second instruction determination module configured to determine at least one target language instruction corresponding to each of the at least one source language instruction based on a mapping relationship between source language instructions and target language instructions, the mapping relationship indicating the target language instruction corresponding to each source language instruction; a description information acquisition module configured to acquire description information for each of the at least one source language instruction and description information for each of the at least one target language instruction; and a code conversion module configured to use a machine learning model to perform code conversion on the source language code based on the mapping relationship, the description information for each of the at least one source language instruction, and the description information for each of the at least one target language instruction to determine target language code corresponding to the source language code, the target language code including at least one target language instruction.

[0005] In a third aspect of this disclosure, an electronic device is provided. The electronic device includes: at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to a first aspect of this disclosure when executed by the at least one processor.

[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, which, when executed by a processor, cause the processor to perform the method according to a first aspect of this disclosure.

[0007] In a fifth aspect of this disclosure, a computer program product is provided, which is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by an electronic device, cause the electronic device to perform the method according to a first aspect of this disclosure.

[0008] As will be understood from the following description, according to embodiments of this disclosure, at least one source language instruction included in the source language code to be converted can be determined, and at least one target language instruction corresponding to each of the at least one source language instruction can be determined based on the mapping relationship between the source language instructions and the target language instructions. A machine learning model can be used to perform code conversion on the source language code based on the description information of each of the at least one source language instruction, the description information of each of the at least one target language instruction, and the mapping relationship to determine the target language code corresponding to the source language code, which includes at least one target language instruction. In this way, code conversion can be performed based on the description information of language instructions using a machine learning model, thereby reducing the difficulty of code conversion and improving the accuracy of code conversion.

[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] In the following detailed description, in conjunction with the accompanying drawings, the above and other features, advantages, and aspects of the various implementations of this disclosure will become more apparent. In the accompanying drawings, the same or similar reference numerals denote the same or similar elements, wherein: Figure 1A and Figure 1B A schematic diagram of an example environment in which embodiments of the present disclosure can be implemented is shown; Figure 2A An example architecture for code conversion according to some embodiments of this disclosure is shown; Figure 2BAn example architecture for determining mapping relationships according to some embodiments of this disclosure is shown; Figure 2C Examples of descriptive information for source language instructions and target language instructions according to some embodiments of this disclosure are shown; Figure 2D An example architecture for model fine-tuning according to some embodiments of this disclosure is shown; Figure 3 A flowchart of a method for code conversion according to some embodiments of the present disclosure is shown; Figure 4 A schematic structural block diagram of an example device for code conversion, based on some examples, is shown; and Figure 5 A block diagram is shown in which one or more examples of an electronic device can be implemented. Detailed Implementation

[0011] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0012] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below. In this document, unless explicitly stated otherwise, performing a step in response to A does not mean that the step is performed immediately after "A", but may include one or more intermediate steps.

[0013] As used in this paper, the term "model" refers to a model that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs using multiple layers of processing units. A neural network model is an example of a deep learning-based model. In this paper, "model" may also be referred to as a "machine learning model," "learning model," "machine learning network," or "learning network," and these terms are used interchangeably.

[0014] Existing code conversion methods mostly employ engineering and rule bases for code transformation. Specifically, existing methods typically use rule bases, templates, etc., to perform semantic mapping on the source language code to determine the target language code. These methods are heavily reliant on the quality of the rule base, but they usually rely on manual selection of the rule base, which requires significant manpower and greatly increases the cost of code conversion. Furthermore, different users have varying code conversion abilities, which may lead to a higher error rate in the rule base, affecting the accuracy of the code conversion.

[0015] Furthermore, code transformation and semantic mapping typically require a large number of rules to fine-tune the quality of the final target language code. For example, mapping may include conditional matching, priority, template replacement, context constraints, and so on. Currently, dedicated personnel are needed to maintain and fine-tune multiple rules in the rule base to adapt to different application scenarios. This increases the cost of maintaining and updating the rule base, further increasing the cost of code transformation.

[0016] In view of this, this disclosure proposes an improved code conversion scheme. According to the scheme of this disclosure, at least one source language instruction included in the source language code to be converted is determined. Based on the mapping relationship between the source language instructions and the target language instructions, at least one target language instruction corresponding to each of the at least one source language instruction is determined. The mapping relationship indicates the target language instruction corresponding to each source language instruction. Description information of each of the at least one source language instruction and description information of each of the at least one target language instruction are obtained. Using a machine learning model, based on the mapping relationship, the description information of each of the at least one source language instruction, and the description information of each of the at least one target language instruction, the source language code is converted to determine the target language code corresponding to the source language code. The target language code includes at least one target language instruction.

[0017] In this way, machine learning models can be used to perform code conversion based on the descriptive information of language instructions, thereby reducing the difficulty of code conversion and improving its accuracy. Compared with traditional rule-based solutions, the descriptive information used in this embodiment includes more comprehensive content. Therefore, the target language code determined based on the descriptive information is more accurate, and users do not need to frequently maintain the rule base.

[0018] The following will describe in detail various example implementations of this scheme with reference to the accompanying drawings.

[0019] First see Figure 1A , Figure 1AA schematic diagram of an example environment 100A according to some examples is shown. In this example environment 100A, an application 115 is installed on a terminal device 110. A user 140 can interact with the application 115 via the terminal device 110 and / or an attached device to the terminal device 110. In some embodiments, the application 115 is capable of providing code conversion services to the user 140. The application 115 can be any suitable application with specific code conversion functionality.

[0020] exist Figure 1A In environment 100A, if application 115 is active, terminal device 110 can display the interface 150 of application 115. Interface 150 may include various interfaces that application 115 can provide, such as code acquisition interface, code conversion interface, code editing interface, etc.

[0021] In some embodiments, terminal device 110 communicates with server 120 to provide services to application 115. Terminal device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some examples, terminal device 110 may also support any type of user-facing interface (such as "wearable" circuitry).

[0022] Server 120 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Server 120 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in cloud environments, etc.

[0023] A communication connection can be established between server 120 and terminal device 110. This communication connection can be established via wired or wireless means. The communication connection can include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus (USB), and Wireless Fidelity (WiFi) connections. In some examples, server 120 and terminal device 110 can exchange signaling signals through their communication connection.

[0024] In some examples, application 115 can use machine learning model 130 to perform code transformation. Machine learning model 130 can be deployed locally on terminal device 110 or on other devices / systems (e.g., server 120). Application 115 can, for example, directly utilize the locally deployed machine learning model 130 or invoke the machine learning model 130 deployed on other devices / systems via communication connections to generate media content.

[0025] Machine learning model 130 can be based on any suitable model architecture, including but not limited to Transformer models, Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), Deep Neural Networks (DNNs), and so on. In some examples, machine learning model 130 can be based on a language model (LM), for example. Machine learning model 130 can include language models, Large Language Models (LLMs), Vision-Language Models (VLMs), Multimodal Large Language Models (MLLMs), and so on. Language models, by learning from large corpora, possess question-answering capabilities. Machine learning model 130 can also be based on other suitable models. Although only a single machine learning model 130 is shown in the figure, multiple machine learning models can exist. If machine learning model 130 includes multiple models, these multiple models can include models configured to perform the same or similar tasks, or models configured to perform different tasks.

[0026] Secondly, refer to Figure 1B , Figure 1BA schematic diagram of an example environment 100B in which embodiments of the present disclosure can be implemented is shown. As shown in example environment 100B, terminal device 110 can acquire source language code 102. Source language code 102 is also known as source language code. Terminal device 110 can receive source language code 102 from other devices or receive source language code 102 input by a user (e.g., user 140). By way of example only, terminal device 110 can present an input box and receive source language code 102 input by the user via the input box.

[0027] Terminal device 110 can perform code conversion on source language code 102 to determine the corresponding target language code 104. Target language code 104 is the code for the target language, which can be any suitable language different from the source language. Terminal device 110 can perform code conversion in any suitable manner, for example. As an example only, terminal device 110 can use machine learning model 130 to perform code conversion on source language code 102.

[0028] In some examples, terminal device 110 may also upload source language code 102 to server 120, whereby server 120 performs code conversion on source language code 102 to determine target language code 104. Server 120 may also perform code conversion in any suitable manner. As an example only, server 120 may use machine learning model 130 to perform code conversion on source language code 102. Terminal device 110 can then obtain target language code 104 from server 120 via a communication connection with server 120.

[0029] It should be understood that the structure and function of the various elements in environments 100A and 100B are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure. Some exemplary embodiments of this disclosure will continue to be described below with reference to the accompanying drawings.

[0030] The following will combine Figures 2A to 2C This disclosure describes the code conversion method provided. The following example illustrates this method implemented on terminal device 110. It should be noted that the operations performed by terminal device 110 in this disclosure can be performed by a related application (e.g., application 115) installed on terminal device 110 or by a client. Some operations described regarding terminal device 110 may require the assistance of server 120.

[0031] First refer to Figure 2A , Figure 2A An example architecture 200A for code conversion according to some embodiments of this disclosure is shown. For example... Figure 2AAs shown, the example architecture 200A includes a source language instruction determination unit 210, a target language instruction determination unit 220, an information acquisition unit 230, and a machine learning model 130.

[0032] In some embodiments, both the source language and the target language can be specialized scripting languages ​​in the field of Electronic Design Automation (EDA) in the integrated circuit manufacturing industry, such as the Recipe scripting language. The Recipe scripting language is a specialized language used to describe physical verification rules and control verification processes. It can be divided into two categories: tool-specific rule languages ​​and general process control languages.

[0033] Tool-specific rule languages ​​(Rule Files / Runsets) can directly define geometric and electrical rules for Design Rule Check (DRC) / Live Layout Verification (LVS), which are the core basis for final sign-off verification before chip tape-out. As an example, tool-specific rule languages ​​may include Standard Verification Rule Format (SVRF) language, Tcl Verification Format (TVF), etc. General flow control languages ​​(Flow Scripts) are used for batch execution, tool invocation, result parsing, and process automation. As an example, general flow control languages ​​may include tool command language (Tcl), Python, etc.

[0034] After obtaining the source language code 102, the terminal device 110 can provide it to the source language instruction determination unit 210. The source language instruction determination unit 210 can determine at least one source language instruction 218 included in the source language code 102 to be converted. The source language instruction determination unit 210 can use any suitable method to determine the at least one source language instruction 218 included in the source language code 102. In some embodiments, the source language instruction determination unit 210 can directly process the source language code 102 to determine the at least one source language instruction 218 included in the source language code 102. For example, the source language instruction determination unit 210 can directly utilize the trained machine learning model 130 to determine the at least one source language instruction 218 included in the source language code 102 based on the source language code 102.

[0035] In some embodiments, the source language instruction determination unit 210 may include a structured code determination subunit 212 and a source language instruction determination subunit 216. The structured code determination subunit 212 may acquire source language syntax information 202. The source language syntax information 202 may be used to describe the syntax rules of the source language. The source language syntax information 202 may be provided to the terminal device 110 by the user, may be obtained by the terminal device 110 from other electronic devices, or may be predetermined by the terminal device 110.

[0036] In some embodiments, the terminal device 110 may pre-obtain sample code in the source language and description information for each of the multiple source language instructions. The description information for each source language instruction may describe its name, definition, function, etc. The description information may also include parameter information for the source language instruction, which may include parameter name, parameter type, parameter explanation, etc. It is understood that each source language instruction may include at least one parameter.

[0037] In some embodiments, the terminal device 110 can use a machine learning model 130 to determine the source language syntax information 202 based on example code of the source language and the description information of each of the multiple source language instructions. It is understood that in some embodiments, after obtaining the source language syntax information 202, the terminal device 110 can store the source language syntax information 202 for subsequent direct application. In this case, the terminal device 110 does not need to obtain the source language syntax information 202 every time, which can improve the efficiency of information processing.

[0038] The structured code determination subunit 212 can parse the source language code 102 based on the source language syntax information 202 to determine the structured code 214 corresponding to the source language code 102. The source language instruction determination subunit 216 can then determine at least one source language instruction 218 based on the structured code 214. As an example only, the source language syntax information 202 can be formally expressed in Extended Backus-Naur Form (EBNF) to represent the grammatical rules of the source language. In some scenarios, the structured code determination subunit 212 can use an EBNF parsing tool to parse the source language code 102 based on the source language syntax information 202 to determine the structured code 214, and the source language instruction determination subunit 216 can then determine at least one source language instruction 218 based on the structured code 214. Therefore, by first determining the structured code 214 corresponding to the source language code 102 and then determining at least one source language instruction 218 based on the structured code 214, the accuracy of the determined at least one source language instruction 218 can be improved.

[0039] The target language instruction determination unit 220 can acquire the mapping relationship 204 between source language instructions and target language instructions, and determine at least one target language instruction 222 corresponding to at least one source language instruction 218 based on the mapping relationship 204. The mapping relationship 204 can indicate the target language instruction corresponding to each source language instruction. Similar to the source language syntax information 202, the mapping relationship 204 can be provided by the user to the terminal device 110, obtained by the terminal device 110 from other electronic devices, or predetermined by the terminal device 110.

[0040] Regarding the method of determining mapping relationship 204, in some embodiments, terminal device 110 may determine the prompting information for machine learning model 130 based at least on reference information for mapping relationship 204, example code and syntax documentation for the source language, and example code and syntax documentation for the target language. Terminal device 110 may then utilize the machine learning model to determine mapping relationship 204 by providing the prompting information to the machine learning model. As an example only, the reference information may indicate a manually determined reference mapping relationship between source language instructions and target language instructions.

[0041] The following is combined with Figure 2B This describes how the mapping relationship 204 is determined. Figure 2B An example architecture 200B for determining mapping relationships 204 according to some embodiments of the present disclosure is shown. Example architecture 200B can be implemented at terminal device 110. Example architecture 200B involves a preprocessing unit 250 and a machine learning model 130.

[0042] In some embodiments, the terminal device 110 may obtain source language development documentation 242 and target language development documentation 244. Source language development documentation 242 may include sample code and syntax descriptions for the source language. Target language development documentation 244 may include sample code and syntax descriptions for the target language. Of course, in some embodiments, the terminal device 110 may also obtain sample code and syntax descriptions for the source language and the target language, respectively. The following description uses the example of the terminal device 110 directly obtaining source language development documentation 242 and target language development documentation 244 as an example. In some embodiments, source language development documentation 242 and target language development documentation 244 may each include other content; for example, source language development documentation 242 and target language development documentation 244 may each include descriptions of source language instructions and descriptions of target language instructions.

[0043] The source language development document 242 and the target language development document 244 may include content in various forms such as text, images, links, and icons. To improve the quality of the final determined mapping relationship 204, the preprocessing unit 250 may perform preprocessing operations on the acquired documents (i.e., the source language development document 242 and the target language development document 244) to obtain a source language preprocessed document set 252 and a target language preprocessed document set 254. Preprocessing operations may include cleaning, deduplication, sorting, and organization. In some embodiments, the preprocessed source language preprocessed document set 252 and the target language preprocessed document set 254 may contain only text content. In some embodiments, the source language preprocessed document set 252 may include multiple source language preprocessed documents, each corresponding to a source language instruction. Similarly, in some embodiments, the target language preprocessed document set 254 may include multiple target language preprocessed documents, each corresponding to a target language instruction.

[0044] In some embodiments, the terminal device 110 may also acquire the prompt word template 256 and reference information 258 for the mapping relationship. As mentioned above, the reference information 258 may indicate a manually determined reference mapping relationship between source language instructions and target language instructions. Of course, the reference information 258 may also include other information that can be used to help determine the mapping relationship 204, and this disclosure does not limit the specific content of the reference information 258. The prompt word template 256 may be a template determined in real time by the user based on actual scenarios, historical experience, etc., or it may be a pre-defined template.

[0045] In some embodiments, the terminal device 110 may determine the prompt information for the machine learning model 130 based at least on the source language preprocessed document set 252, the target language preprocessed document set 254, the prompt word template 256, and the reference information 258 for mapping relationships. For example, the terminal device 110 may determine the prompt information by filling the prompt word template 256 with the source language preprocessed document set 252, the target language preprocessed document set 254, and the reference information 258 for mapping relationships. In some embodiments, the terminal device 110 may also obtain the description information of each of the multiple source language instructions and the description information of each of the multiple target language instructions, and may combine the description information of each of the multiple source language instructions and the description information of each of the multiple target language instructions to determine the prompt information for the machine learning model 130.

[0046] The description information for each source language instruction can describe the instruction's name, definition, function, and parameter information. The description information for each target language instruction can describe the instruction's name, definition, function, and parameter information. Parameter information may include parameter name, parameter type, and parameter explanation.

[0047] Terminal device 110 can provide prompt information to machine learning model 130 and obtain the output of machine learning model 130 in response to the prompt information. Terminal device 110 can determine mapping relationship 204 based on this output. It is understood that, similar to source language syntax information 202, in some embodiments, after obtaining mapping relationship 204, terminal device 110 can store mapping relationship 204 for subsequent direct application. In this case, terminal device 110 does not need to obtain mapping relationship 204 every time, which can improve the efficiency of information processing.

[0048] In some embodiments, the terminal device 110 may also receive feedback information regarding the mapping relationship 204. This feedback information may include feedback from a user, or feedback from other electronic devices or other models / systems. For example, the terminal device 110 may provide the mapping relationship 204 to a user and receive feedback from the user on the mapping relationship. Alternatively, the terminal device 110 may provide the mapping relationship 204 to a trained evaluation model and obtain feedback from the evaluation model on the mapping relationship. The feedback information may indicate the quality of the mapping relationship, errors, modification suggestions, etc. In some embodiments, the terminal device 110 may update the mapping relationship based on the feedback information received. Therefore, updating the mapping relationship based on the feedback information can improve the accuracy of the mapping relationship.

[0049] certainly, Figure 2B Only one example of determining the mapping relationship is shown. In embodiments of this disclosure, terminal device 110 may determine the mapping relationship in any suitable manner. As an example, terminal device 110 may directly provide source language development document 242 and target language development document 244 to machine learning model 130 to directly determine the mapping relationship between source language instructions and target language instructions using machine learning model 130.

[0050] Return to reference Figure 2AIn some embodiments, at least one source language instruction 218 and at least one target language instruction 222 are provided to the information acquisition unit 230. The information acquisition unit 230 can acquire description information 232 for each of the at least one source language instruction and description information 234 for each of the at least one target language instruction. As an example only, the information acquisition unit 230 can acquire description information 224, which includes description information 226 for the source language instruction and description information 228 for the target language instruction. It is understood that the description information 226 for the source language instruction may include description information for multiple source language instructions, and the description information 228 for the target language instruction may include description information 228 for multiple target language instructions. For example, the information acquisition unit 230 can search for description information 232 for each of the source language instruction and description information 234 for each of the target language instructions from the description information 224 based on at least one source language instruction 218 and at least one target language instruction 222.

[0051] Terminal device 110 can then utilize machine learning model 130 to perform code conversion on source language code 102 based on mapping relationship 204, description information 232 of at least one source language instruction, and description information 234 of at least one target language instruction to determine target language code 104 corresponding to source language code 102. Target language code 104 includes at least one target language instruction 222. As an example only, terminal device 110 can obtain a prompt word template for guiding machine learning model 130 to perform code conversion, and can obtain prompt word information for machine learning model 130 by filling the prompt word template with at least mapping relationship 204, description information 232 of at least one source language instruction, description information 234 of at least one target language instruction, and source language code 102. Terminal device 110 can provide this prompt word information to machine learning model 130 to enable machine learning model 130 to perform code conversion on source language code 102.

[0052] In some embodiments, the terminal device 110 can directly provide the mapping relationship 204, the description information 232 of each of at least one source language instruction, the description information 234 of each of at least one target language instruction, and the source language code 102 to the machine learning model 130 to generate the target language code 104 using the machine learning model 130. In some embodiments, the terminal device 110 can determine at least one source language code segment included in the source language code. In some scenarios, each source language code segment may include only a single source language instruction. In some scenarios, each source language code segment may also include only multiple source language instructions with an association relationship. The association relationship here can be any appropriate association relationship, which can be determined based on the actual scenario, user configuration, etc. It is understood that the terminal device 110 can use any appropriate method to divide the source language code 102 to determine at least one source language code segment. As an example only, the terminal device 110 can divide the source language code based on the source language instructions.

[0053] In this scenario, terminal device 110 can utilize machine learning model 130 to perform code conversion on each source language code segment in at least one source language code segment, thereby determining the target language code segment corresponding to each of the at least one source language code segment. For example, in conjunction with... Figure 2C , Figure 2C Example 200C illustrates description information for source language instructions and target language instructions according to some embodiments of the present disclosure. In example 200C, the source language can be Recipe language N, and the target language can be Recipe language M. The Recipe language can include the instruction "xor" (i.e., the source language instruction is xor), and the Recipe language M can include the instruction "^" (i.e., the target language instruction is ^). Example 200C includes description information 262 for the source language instruction xor and description information 264 for the target language instruction ^ corresponding to the source language instruction xor. If the source language code includes the source language code fragment xor(geo1, geo2, 1e-8), the terminal device 110 can obtain description information 262 and description information 264, and use machine learning model 130 to perform code conversion on the source language code fragment xor(geo1, geo2, 1e-8) based on description information 262 and description information 264 to determine the corresponding target language code fragment geo1^(geo2, 1e-8). The terminal device 110 can then determine the target language code 104 corresponding to the source language code 102 based on the target language code fragment corresponding to each of the at least one source language code fragment.

[0054] Therefore, machine learning models can be used to translate Recipe script code without relying on traditional rule bases. This can significantly reduce migration costs in different scenarios, improve translation quality, and reduce reliance on human intervention.

[0055] In some embodiments, the terminal device 110 can also detect whether anomalies occur in the code conversion process of each source language code segment. For example, in response to an anomaly in the code conversion process of a certain source language code segment, the terminal device 110 can receive natural language input for that source language code segment. This natural language input can be correction information or supplementary information manually entered by the user. The terminal device 110 can, for example, re-perform the code conversion of the source language code segment based on the natural language input. Thus, each source language code segment can be processed individually, and the source language code segments that have anomalies can be corrected independently. Compared to methods that convert the entire source language code and correct the entire target language code, this processing and correction method targets a smaller granularity and improves the efficiency of anomaly correction.

[0056] It should be noted that although the above example illustrates the code conversion process implemented on terminal device 110, in some scenarios, the code conversion process mentioned in this disclosure can also be implemented on server 120. Terminal device 110 can, in response to obtaining source language code 102, provide source language code 102 to server 120 and obtain the target language code 104 corresponding to source language code 102 from server 120. This disclosure does not limit this aspect.

[0057] In some embodiments, the terminal device 110 may also provide a code conversion interface. The code conversion interface includes source language code and target language code, which are presented in different areas of the interface. Both the area presenting the source language code and the area presenting the target language code can be located at any suitable position on the interface, and these two areas can have any suitable positional relationship. For example, the area presenting the source language code can be located in the left half of the interface, and the area presenting the target language code can be located in the right half of the interface. Alternatively, the area presenting the target language code can be overlaid on the area presenting the source language code as a floating window. This disclosure does not limit this.

[0058] In some embodiments, terminal device 110 may, in response to receiving user interaction via a code conversion interface regarding a source language instruction in the source language code, highlight the source language instruction and the corresponding target language instruction in the target language code. Terminal device 110 may highlight the interacted source language instruction and the corresponding target language instruction in any appropriate manner. As an example only, terminal device 110 may highlight the interacted source language instruction and the corresponding target language instruction in any appropriate manner, such as highlighting, bolding, italics, adjusting font size, adjusting color, or adding a background, and may display the remaining content in a regular font.

[0059] It is understood that in some embodiments, the terminal device 110 may also, in response to receiving user interaction with a target language instruction in the target language code via the code conversion interface, highlight the target language instruction and the corresponding source language instruction in the source language code. It is understood that the interaction here can be any appropriate interaction. For example, the terminal device 110 may, in response to detecting a click operation on a source language instruction or a target language instruction, highlight the source language instruction and the corresponding target language instruction. This allows the user to quickly understand the currently interacting instruction and facilitates comparison between the source language instruction and the target language instruction.

[0060] In some embodiments, the terminal device 110 can, in response to detecting a browsing operation on source language code or target language code, determine the content currently being browsed by the user and synchronize the currently displayed source language code and target language code. For example, if the user is currently browsing a second source language instruction in the source language code, the terminal device 110 will simultaneously present the second source language instruction and the corresponding target language instruction in the target language code. This ensures the synchronization of browsing between the source language code and the target language code, facilitating comparison between the two codes by the user.

[0061] In some embodiments, the terminal device 110 may also receive calibration information for a corresponding target language instruction in response to an interaction with a source language instruction in the code conversion interface. In some embodiments, the terminal device 110 may directly receive calibration information for a target language instruction in response to an interaction with it. It is understood that the interaction here can be any appropriate interaction. As an example only, the terminal device 110 may present an input box in response to detecting a hover operation on a source language instruction or a target language instruction, and receive calibration information via the input box.

[0062] Terminal device 110 may, for example, re-perform code conversion on the source language code based on the received calibration information to generate updated target language code, which includes calibrated target language instructions. In some embodiments, terminal device 110 may also, based on the received calibration information, only regenerate the target language instructions, keeping the rest of the target speech code unchanged. In some examples, terminal device 110 may also present an undo control. For example, in response to triggering the undo control, terminal device 110 may return to presenting the target language code before the update. Thus, supporting user interaction to correct the target language code can improve the accuracy of the target language code.

[0063] In some embodiments, the machine learning model 130 can also be fine-tuned based on historical conversion records of code conversions performed by the machine learning model 130. Historical conversion records are records of code conversions performed by the machine learning model 130 on historically received source language code to obtain historical target language code. As mentioned above, both the source language and the target language can be EDA-specific scripting languages, such as the Recipe scripting language. Fine-tuning the machine learning model 130 based on historical records of code conversions performed on EDA-specific scripting languages ​​allows the fine-tuned machine learning model 130 to be more suitable for the EDA domain, improving the performance of the machine learning model 130 in performing code conversions on EDA-specific scripting languages.

[0064] refer to Figure 2D , Figure 2D An example architecture 200D for model fine-tuning according to some embodiments of this disclosure is shown. It should be noted that the example architecture 200D can be implemented at the terminal device 110 or at other devices (i.e., fine-tuning of the machine learning model 130 can be performed by other electronic devices). This document only describes the example of model fine-tuning performed by the terminal device 110.

[0065] Example architecture 200D involves a training sample determination unit 280 and a model fine-tuning unit 290. The training sample determination unit 280 can obtain historical transformation records 270 of the machine learning model 130. The historical transformation records 270 can indicate the historical source language code 271, the historical target language code 272 corresponding to the historical source language code 271, the description information 273, the historical structured code 274 corresponding to the historical source language code 271, etc., involved in the historical code transformation process of the machine learning model 130.

[0066] It can be understood that the historical conversion record 270 is a record of the historical code conversion of the machine learning model 130. The historical source language code 271 may include one or more source language codes converted from the historical code conversion of the machine learning model 130. The historical target language code 272 may include one or more target language codes obtained from the historical code conversion of the machine learning model 130. The description information 273 may include description information of each instruction acquired historically. The historical structured code 274 may include the structured code determined during the historical code conversion process.

[0067] As an example only, historical source language code 271 may correspond to the aforementioned source language code 102, historical target language code 272 may correspond to the aforementioned target language code 104, description information 273 may correspond to the description information 232 of each of the aforementioned source language instructions and the description information 234 of each of the aforementioned target language instructions, and historical structured code 274 may correspond to the aforementioned structured code 214.

[0068] The training sample determination unit 280 can generate training samples 281 based at least on the historical transformation records of the machine learning model 130. It is understood that although only one training sample 281 is shown in the figure, the training sample determination unit 280 can actually generate one or more training samples 281. In some examples, the training sample determination unit 280 can also obtain the mapping relationship 204 and generate training samples 281 based on the mapping relationship 204 and the historical transformation records 270.

[0069] In some examples, the training sample determination unit 280 can generate at least one candidate training sample based on the mapping relationship 204 and the historical transformation record 270. For example, the training sample determination unit 280 can determine a candidate training sample by simply concatenating the single code transformation record in the historical transformation record 270 (which may include the source language code received by the machine learning model 130 in that instance, the target language code generated in that instance, the description information of at least one source language instruction, the description information and structured code of at least one target language instruction) and the mapping relationship 204. The training sample determination unit 280 can then perform preprocessing operations on the at least one candidate training sample to determine at least one training sample.

[0070] The preprocessing operations here may include at least one of data cleaning and data augmentation operations. The training sample determination unit 280 can remove candidate training samples with inaccurate target language codes, semantically identical candidate training samples, etc., from at least one candidate training sample by performing data cleaning operations. For example, the training sample determination unit 280 may determine that the target language code of a candidate training sample is inaccurate in response to one or more of the following problems: mismatch between target language code and source language code, missing target language code, or grammatical errors in the target language code.

[0071] The training sample determination unit 280 can improve the quality and diversity of at least one candidate training sample by performing data augmentation operations. Data augmentation operations can include task addition operations, sample construction operations, and so on. For example, the original cue words for each candidate training sample can be used to guide the machine learning model 130 to perform code translation on the source language code therein. The training sample determination unit 280 can adjust the cue words to guide the machine learning model 130 to perform other tasks (such as code interpretation tasks). The adjusted cue words can simultaneously indicate the code translation task and the newly added other tasks, or only indicate the newly added other tasks. This can help improve the machine learning model 130's ability to understand instructions.

[0072] For example, the training sample determination unit 280 can generate a negative sample that is extremely similar to, but still different from, a candidate training sample. The training sample determination unit 280 can then identify this negative sample as another candidate training sample. For instance, the training sample determination unit 280 can adjust the parameter order, parameter names, parameter values, etc., in the source language code, and generate corresponding target language code based on the adjusted source language code. Then, based on the current code conversion record, a new candidate training sample can be generated. By using such a pair of similar but different training samples to fine-tune the machine learning model 130, the machine learning model 130's ability to distinguish subtle differences can be improved.

[0073] The training sample determination unit 280 can determine one or more candidate training samples obtained through preprocessing operations as at least one training sample 281. Each training sample 281 may include an input sample 282 and an output sample 283. In some examples, the input sample 282 may include at least a source language code sample 284 and a description information sample 285. The source language code sample 284 includes at least one source language instruction sample, which may correspond to at least one target language instruction sample respectively. The description information sample 285 includes description information samples of each of the at least one source language instruction sample and description information samples of each of the at least one target language instruction sample. In some examples, the input sample 282 may also include a structured code sample 286 corresponding to the source language code sample 284.

[0074] Output sample 283 may include at least the target language code sample 288 corresponding to source language code sample 284. In some examples, training sample determination unit 280 may also determine thought chain sample 287 based on mapping relationship 204, input sample 282, and target language code sample 288. Thought chain sample 287 may indicate the reasoning process of machine learning model 130 from source language code sample 284 to target language code sample 288. For example, thought chain sample 287 may indicate the reasoning process of machine learning model 130 performing code conversion on a specialized scripting language in the EDA domain.

[0075] The model fine-tuning unit 290 can then use at least one training sample 281 to fine-tune the machine learning model. For example, the model fine-tuning unit 290 can use the machine learning model 130 to determine the predicted output 291 corresponding to the input sample 282, based on the input sample 282. The predicted output 291 may include, for example, a predicted thought chain and a predicted target language code. The model fine-tuning unit 290 can determine, for example, the difference between the predicted output 291 and the output sample 283, which may specifically include a first difference between the predicted thought chain and the thought chain sample 287 and a second difference between the predicted target language code and the target language code sample 288. The model fine-tuning unit 290 can determine a loss 292 for fine-tuning the machine learning model 130 based on this difference and use the loss 292 to fine-tune the machine learning model 130.

[0076] Therefore, during the fine-tuning phase, the machine learning model 130 does not need to determine its output based on the mapping relationship 204. Instead, it learns the external mapping relationship 204, allowing the fine-tuned model to perform code conversion without relying on it. This reduces the need for managing and maintaining the mapping relationship 204, thus decreasing the workload. Furthermore, constructing training samples based on historical code conversion records reduces the difficulty and cost of sample construction, solving the problem of sample construction difficulties in the EDA field. Additionally, using real-world records to fine-tune the model helps improve the accuracy of subsequent code conversions. Moreover, relying on samples with a thought chain, the model can autonomously master the complete reasoning logic of instruction recognition, mapping matching, and parameter arrangement.

[0077] It should be noted that the training sample determination unit 280 and the model fine-tuning unit 290 can also be implemented in different electronic devices. For example, the training sample determination unit 280 can be deployed locally on the terminal device 110. After the terminal device 110 generates training samples locally, it provides them to the electronic device where the model fine-tuning unit 290 is located so that fine-tuning can continue.

[0078] It should also be noted that, in some examples, to further improve the performance of the machine learning model 130, the electronic device (e.g., terminal device 110) that performs model fine-tuning may also use only a portion (e.g., 70%) of a large number of training samples to fine-tune the machine learning model 130, and use another portion (e.g., the remaining 30%) to perform validation and evaluation on the fine-tuned machine learning model 130.

[0079] Furthermore, in some examples, the above processing of training samples (or candidate training samples) can also be achieved manually, which is not limited in this paper.

[0080] In some examples, terminal device 110 may also acquire a fine-tuned machine learning model 130 and utilize the fine-tuned machine learning model 130 to perform subsequent code conversion tasks. In this case, terminal device 110 may perform code conversion on the source language code based solely on the description information of at least one source language instruction and at least one target language instruction included in the received source language code to determine the target language code corresponding to the source language code. Since examples of code conversion have already been described in detail above, they will not be repeated here.

[0081] Figure 3 A flowchart of a method 300 for code conversion according to some embodiments of the present disclosure is shown. Method 300 can be implemented at terminal device 110. Reference is made below. Figure 1A and Figure 1B Let's describe method 300.

[0082] In box 310, terminal device 110 determines at least one source language instruction included in the source language code to be converted.

[0083] In box 320, terminal device 110 determines at least one target language instruction corresponding to at least one source language instruction based on the mapping relationship between source language instructions and target language instructions. The mapping relationship indicates the target language instruction corresponding to each source language instruction.

[0084] In box 330, terminal device 110 acquires description information for at least one source language instruction and description information for at least one target language instruction.

[0085] In box 340, terminal device 110 uses a machine learning model to perform code conversion on source language code based on mapping relationships, description information of at least one source language instruction, and description information of at least one target language instruction to determine target language code corresponding to the source language code. The target language code includes at least one target language instruction.

[0086] In some embodiments, determining at least one source language instruction included in the source language code to be converted includes: obtaining source language syntax information, which describes the syntax rules of the source language; parsing the source language code based on the source language syntax information to determine the structured code corresponding to the source language code; and determining at least one source language instruction based on the structured code.

[0087] In some embodiments, source language syntax information is determined by a machine learning model based on example code of the source language and description information of each of the multiple source language instructions.

[0088] In some embodiments, the mapping relationship is determined based on: determining prompts for the machine learning model based at least on reference information for the mapping relationship, sample code and syntax documentation for the source language, and sample code and syntax documentation for the target language; providing prompts to the machine learning model to obtain the output of the machine learning model; and determining the mapping relationship based on the output of the machine learning model.

[0089] In some embodiments, method 300 further includes: updating the mapping relationship based on the feedback information received regarding the mapping relationship.

[0090] In some embodiments, determining the target language code corresponding to the source language code includes: determining at least one source language code segment included in the source language code, each source language code segment including a single source language instruction or multiple source language instructions with an association relationship; using a machine learning model to perform code conversion on each source language code segment in the at least one source language code segment to determine the target language code segment corresponding to each of the at least one source language code segments; and determining the target language code corresponding to the source language code based on the target language code segment corresponding to each of the at least one source language code segments.

[0091] In some embodiments, determining the target language code corresponding to the source language code includes: receiving natural language input for the source language code segment in response to an anomaly occurring in the code conversion process of at least one source language code segment; and re-performing code conversion on the source language code segment based on the natural language input.

[0092] In some embodiments, method 300 further includes: providing a code conversion interface, the code conversion interface including source language code and target language code, the source language code and target language code being presented in different areas of the code conversion interface; and in response to interaction with a first source language instruction in the source language code, highlighting the first source language instruction and a first target language instruction in the target language code corresponding to the first source language instruction.

[0093] In some embodiments, method 300 further includes: receiving calibration information for the first target language instruction in response to an interaction with a first source language instruction or a first target language instruction; and re-performing code conversion on the source language code based on the calibration information to generate updated target language code, the updated target language code including the calibrated first target language instruction.

[0094] In some embodiments, method 300 further includes: generating at least one training sample based at least on the historical conversion records of the machine learning model, the historical conversion records including historical source language code and historical target language code of the historical execution code conversion of the machine learning model; and fine-tuning the machine learning model using at least one training sample.

[0095] In some embodiments, generating at least one training sample includes generating at least one training sample based on mapping relationships and historical transformation records.

[0096] In some embodiments, generating at least one training sample includes: generating at least one candidate training sample based at least on historical transformation records of a machine learning model; and performing preprocessing operations on the at least one candidate training sample to determine at least one training sample, wherein the preprocessing operations include data cleaning operations and / or data augmentation operations.

[0097] In some embodiments, each training sample includes a sample input and a sample output, wherein the sample input includes at least a sample source language code and sample description information, the sample source language code includes at least one sample source language instruction, the sample description information includes sample description information for each of at least one sample source language instruction and description information for each of at least one sample target language instruction, and the at least one sample target language instruction corresponds to at least one sample source language instruction; and wherein the sample output includes a sample thought chain and sample target language code corresponding to the sample source language code, the sample thought chain indicating the reasoning process from the sample source language code to the sample target language code.

[0098] In some embodiments, the sample input may also include sample structured code corresponding to the sample source language code.

[0099] In some embodiments, samples are determined based on mapping relationships, sample inputs, and sample target language codes.

[0100] In some embodiments, the fine-tuning objective includes making the machine learning model so that it does not require code transformation based on mapping relationships.

[0101] In some embodiments, fine-tuning a machine learning model using at least one training sample includes: using the machine learning model to determine a predicted output corresponding to a sample input based on the sample input; and fine-tuning the machine learning model based on the difference between the predicted output and the sample output.

[0102] The various examples described herein also provide corresponding apparatus for implementing the methods or processes described above. Figure 4 A schematic structural block diagram of an example device 400 for code conversion, based on some examples, is shown. Device 400 may be implemented as or included in terminal device 110. The various modules / components in device 400 may be implemented by hardware, software, firmware, or any combination thereof.

[0103] like Figure 4As shown, the apparatus 400 includes: a first instruction determination module 410, configured to determine at least one source language instruction included in the source language code to be converted; a second instruction determination module 420, configured to determine at least one target language instruction corresponding to each of the at least one source language instruction based on the mapping relationship between the source language instructions and the target language instructions, wherein the mapping relationship indicates the target language instruction corresponding to each source language instruction; a description information acquisition module 430, configured to acquire description information of each of the at least one source language instruction and description information of each of the at least one target language instruction; and a code conversion module 440, configured to use a machine learning model to perform code conversion on the source language code based on the mapping relationship, the description information of each of the at least one source language instruction, and the description information of each of the at least one target language instruction to determine the target language code corresponding to the source language code, wherein the target language code includes at least one target language instruction.

[0104] In some embodiments, the first instruction determination module 410 is further configured to: obtain source language syntax information, which describes the syntax rules of the source language; parse the source language code based on the source language syntax information to determine the structured code corresponding to the source language code; and determine at least one source language instruction based on the structured code.

[0105] In some embodiments, source language syntax information is determined by a machine learning model based on example code of the source language and description information of each of the multiple source language instructions.

[0106] In some embodiments, the mapping relationship is determined based on: determining prompts for the machine learning model based at least on reference information for the mapping relationship, sample code and syntax documentation for the source language, and sample code and syntax documentation for the target language; providing prompts to the machine learning model to obtain the output of the machine learning model; and determining the mapping relationship based on the output of the machine learning model.

[0107] In some embodiments, the apparatus 400 further includes a mapping update module, configured to update the mapping relationship based on feedback information received regarding the mapping relationship.

[0108] In some embodiments, the code conversion module 440 is further configured to: determine at least one source language code segment included in the source language code, each source language code segment including a single source language instruction or multiple source language instructions having an association relationship; use a machine learning model to convert each source language code segment in the at least one source language code segment one by one to determine the target language code segment corresponding to each of the at least one source language code segments; and determine the target language code corresponding to the source language code based on the target language code segment corresponding to each of the at least one source language code segments.

[0109] In some embodiments, the code conversion module 440 is further configured to: receive natural language input for the source language code segment in response to an anomaly in the code conversion process of at least one source language code segment; and re-perform code conversion on the source language code segment based on the natural language input.

[0110] In some embodiments, the apparatus 400 further includes: an interface providing module configured to provide a code conversion interface, the code conversion interface including source language code and target language code, the source language code and target language code being presented in different areas of the code conversion interface; and a highlighting module configured to, in response to interaction with a first source language instruction in the source language code, highlight the first source language instruction and a first target language instruction in the target language code corresponding to the first source language instruction.

[0111] In some embodiments, the apparatus 400 further includes: a calibration information receiving module configured to receive calibration information for a first target language instruction in response to an interaction with a first source language instruction or a first target language instruction; and a re-execution module configured to re-perform code conversion on the source language code based on the calibration information to generate updated target language code, the updated target language code including the calibrated first target language instruction.

[0112] In some embodiments, the apparatus 400 further includes: a sample generation module configured to generate at least one training sample based on at least a historical conversion record of a machine learning model, the historical conversion record including historical source language code and historical target language code converted by historical execution code of the machine learning model; and a model fine-tuning module configured to fine-tune the machine learning model using at least one training sample.

[0113] In some embodiments, the sample generation module is further configured to generate at least one training sample based on the mapping relationship and historical transformation records.

[0114] In some embodiments, the sample generation module is further configured to: generate at least one candidate training sample based at least on historical transformation records of the machine learning model; and perform preprocessing operations on the at least one candidate training sample to determine at least one training sample, wherein the preprocessing operations include data cleaning operations and / or data augmentation operations.

[0115] In some embodiments, each training sample includes a sample input and a sample output, wherein the sample input includes at least a sample source language code and sample description information, the sample source language code includes at least one sample source language instruction, the sample description information includes sample description information for each of at least one sample source language instruction and description information for each of at least one sample target language instruction, and the at least one sample target language instruction corresponds to at least one sample source language instruction; and wherein the sample output includes a sample thought chain and sample target language code corresponding to the sample source language code, the sample thought chain indicating the reasoning process from the sample source language code to the sample target language code.

[0116] In some embodiments, the sample input may also include sample structured code corresponding to the sample source language code.

[0117] In some embodiments, the sample thought chain is determined based on mapping relationships, sample inputs, and sample target language codes.

[0118] In some embodiments, the fine-tuning objective includes making the machine learning model so that it does not require code transformation based on mapping relationships.

[0119] In some embodiments, the model fine-tuning module is further configured to: use a machine learning model to determine the predicted output corresponding to the sample input based on the sample input; and fine-tune the machine learning model based on the difference between the predicted output and the sample output.

[0120] The modules included in device 400 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some examples, one or more modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units in device 400 can be implemented at least partially by one or more hardware logic components. By way of example, and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard parts (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.

[0121] Figure 5 A block diagram of an electronic device 500 in which one or more examples may be implemented is shown. It should be understood that... Figure 5 The electronic device 500 shown is merely exemplary and should not be construed as limiting the functionality and scope of the examples described herein. Figure 5 The electronic device 500 shown can be used to implement the terminal device 110 discussed above.

[0122] like Figure 5 As shown, electronic device 500 is in the form of a general-purpose electronic device. Components of electronic device 500 may include, but are not limited to, one or more processing units or processors 510, memory 520, storage devices 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processor 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 500.

[0123] Electronic device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof). Storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 500.

[0124] Electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 5 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various examples.

[0125] The communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 500 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 500 can operate in a networked environment using logical connections to one or more other servers, networked personal computers, or another network node.

[0126] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device that enables electronic device 500 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).

[0127] A computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. A computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0128] The flowcharts and / or block diagrams of the methods, apparatus, devices, and computer program products referred to herein describe various aspects. It should be understood that each block of the flowcharts and / or block diagrams, as well as combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0129] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0130] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0131] The flowcharts and block diagrams in the accompanying figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to some examples. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the figures. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0132] Various examples have been described above. The foregoing descriptions are exemplary and not exhaustive, nor are they limited to the proposed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the proposed implementations.

Claims

1. A code conversion method, comprising: Identify at least one source language instruction included in the source language code to be converted; Based on the mapping relationship between source language instructions and target language instructions, at least one target language instruction corresponding to each of the at least one source language instructions is determined, wherein the mapping relationship indicates the target language instruction corresponding to each source language instruction; Obtain the description information of each of the at least one source language instruction and the description information of each of the at least one target language instruction; as well as Using a machine learning model, based on the mapping relationship, the description information of each of the at least one source language instruction and the description information of each of the at least one target language instruction, the source language code is converted to determine the target language code corresponding to the source language code, wherein the target language code includes the at least one target language instruction.

2. The method of claim 1, wherein determining at least one source language instruction included in the source language code to be converted comprises: Obtain source language syntax information, which is used to describe the syntax rules of the source language; The source language code is parsed based on the source language syntax information to determine the structured code corresponding to the source language code; as well as Based on the structured code, the at least one source language instruction is determined.

3. The method according to claim 2, wherein the source language syntax information is determined by the machine learning model based on example code of the source language and description information of each of the multiple source language instructions.

4. The method of claim 1, wherein the mapping relationship is determined based on the following: Based at least on reference information for the mapping relationship, sample code and syntax documentation for the source language, and sample code and syntax documentation for the target language, prompts for the machine learning model are determined. The prompting information is provided to the machine learning model to obtain the output of the machine learning model; as well as The mapping relationship is determined based on the output of the machine learning model.

5. The method according to claim 4, further comprising: In response to receiving feedback information regarding the mapping relationship, the mapping relationship is updated based on the feedback information.

6. The method according to claim 1, wherein determining the target language code corresponding to the source language code comprises: The source language code is determined to include at least one source language code segment, and each source language code segment in the at least one source language code segment includes a single source language instruction, or multiple source language instructions that are related. Using the machine learning model, each source language code segment in the at least one source language code segment is converted to determine the target language code segment corresponding to each of the at least one source language code segments; as well as Based on the target language code segment corresponding to each of the at least one source language code segment, the target language code corresponding to the source language code is determined.

7. The method of claim 6, wherein determining the target language code corresponding to the source language code comprises: In response to an anomaly occurring in the code conversion process of the source language code segment in the at least one source language code segment, natural language input is received for the source language code segment; as well as Based on the natural language input, the source language code fragment is re-transformed.

8. The method according to claim 1, further comprising: A code conversion interface is provided, which includes the source language code and the target language code, and the source language code and the target language code are presented in different areas of the code conversion interface; In response to interaction with a first source language instruction in the source language code, the first source language instruction and the first target language instruction in the target language code corresponding to the first source language instruction are highlighted.

9. The method of claim 8, further comprising: In response to an interaction with a first source language instruction or a first target language instruction, receive calibration information for the first target language instruction; as well as Based on the calibration information, the source language code is re-transformed to generate the updated target language code, which includes the calibrated first target language instructions.

10. The method of claim 1, further comprising: At least one training sample is generated based on the historical transformation record of the machine learning model, wherein the historical transformation record includes the historical source language code and historical target language code transformed by the historical execution code of the machine learning model; and The machine learning model is fine-tuned using the at least one training sample.

11. The method of claim 10, wherein generating at least one training sample comprises: Based on the mapping relationship and the historical transformation record, at least one training sample is generated.

12. The method of claim 10, wherein generating at least one training sample comprises: At least one candidate training sample is generated based on the historical transformation records of the machine learning model. Preprocessing operations are performed on the at least one candidate training sample to determine the at least one training sample, wherein the preprocessing operations include data cleaning operations and / or data augmentation operations.

13. The method of claim 10, wherein each training sample comprises an input sample and an output sample. The input sample includes at least a source language code sample and a description information sample. The source language code sample includes at least one source language instruction sample. The description information sample includes description information samples of each of the at least one source language instruction sample and description information samples of each of the at least one target language instruction sample. The at least one target language instruction sample corresponds to the at least one source language instruction sample. Furthermore, the output sample includes a thought chain sample and a target language code sample corresponding to the source language code sample, wherein the thought chain sample indicates the reasoning process from the source language code sample to the target language code sample.

14. The method according to claim 13, wherein the input sample further includes a structured code sample corresponding to the source language code sample.

15. The method of claim 13, wherein the thought chain sample is determined based on the mapping relationship, the input sample, and the target language code sample.

16. The method of claim 13, wherein fine-tuning the machine learning model using the at least one training sample comprises: Using the machine learning model, the predicted output corresponding to the input sample is determined based on the input sample; as well as The machine learning model is fine-tuned based on the difference between the predicted output and the output sample.

17. An apparatus for code conversion, comprising: The first instruction determination module is configured to determine at least one source language instruction included in the source language code to be converted; The second instruction determination module is configured to determine at least one target language instruction corresponding to each of the at least one source language instruction based on the mapping relationship between the source language instruction and the target language instruction, wherein the mapping relationship indicates the target language instruction corresponding to each source language instruction. The description information acquisition module is configured to acquire description information of each of the at least one source language instruction and description information of each of the at least one target language instruction; as well as The code conversion module is configured to use a machine learning model to convert the source language code based on the mapping relationship, the description information of each of the at least one source language instruction and the description information of each of the at least one target language instruction to determine the target language code corresponding to the source language code, wherein the target language code includes the at least one target language instruction.

18. An electronic device, characterized in that, include: At least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 16 when executed by the at least one processor.

19. A computer-readable storage medium, characterized in that, It stores computer-executable instructions that can be executed by a processor to implement the method according to any one of claims 1 to 16.

20. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 16.