Warehouse level code translation method and device based on large model
By constructing a self-evolving knowledge base of target language code samples, dependency usage examples, and successful translation functions, this approach addresses the insufficient dependency handling in existing repository-level code translation methods, thereby improving the accuracy and adaptability of repository-level code translation.
Patent Information
- Application Number
- CN202511068634.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-07
AI Technical Summary
Existing repository-level code translation methods only involve function translation pairs and ignore dependency knowledge, resulting in poor performance of code translation tasks in repository-level contexts.
By acquiring the repository to be translated, open source projects, and historically successful translated function pairs, a pre-built large model is used to generate the repository architecture. Based on a tree parser, a self-evolving knowledge base of target language code samples, dependency usage examples, and successful translated function pairs is constructed to determine the target triple translation knowledge of the function to be translated. The implementation code is then output and embedded into the repository architecture.
It enhances the large model's ability to handle dependencies, improves the performance of code translation tasks in repository-level contexts, and enhances the accuracy and adaptability of translation.
Smart Images

Figure CN120909589A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of code translation, and in particular to a warehouse-level code translation method and device based on a large model. BACKGROUND
[0002] Code translation refers to the migration of code projects from one programming language to another for the purpose of adapting to different running environments, improving running speed, improving software security, etc. Warehouse-level code translation refers to the complete warehouse as the object to be translated, which is closer to the real translation demand scene.
[0003] Repository-Level Code Translation is a new software engineering method that aims to solve the limitations of traditional single-file code translation. Its core idea is to analyze the context information of the entire code repository (including project structure, dependency relationship, design pattern, etc.) to achieve more accurate and complete cross-language code conversion. It is particularly suitable for large project migration or legacy system modernization, which can significantly reduce manual intervention and improve the maintainability and functional integrity of the translated code. Key technologies include cross-file semantic analysis, dependency graph construction, and deep learning-based code pattern matching.
[0004] Most existing warehouse-level code translation methods extract a large number of function-level code pairs from open source projects, and then input these isolated function pairs into a large model. However, this method only involves function translation pairs, and ignores dependency knowledge, making the existing method insufficient in handling dependencies, resulting in poor performance in warehouse-level context code translation tasks. SUMMARY
[0005] The present application provides a warehouse-level code translation method and device based on a large model, which solves the technical problem that the existing warehouse-level code translation method only involves function translation pairs, resulting in poor performance in warehouse-level context code translation tasks.
[0006] The first aspect of the present application provides a warehouse-level code translation method based on a large model, comprising:
[0007] Obtaining a warehouse to be translated, an open source project and a plurality of historical successful translation function pairs, and using a pre-installed large model to perform architecture translation according to the warehouse to be translated to generate a warehouse architecture;
[0008] constructing, based on the tree-shaped parser and the preset large model, a target language code sample self-evolution knowledge base, a dependency usage example self-evolution knowledge base, and a successful translation function pair self-evolution knowledge base according to the open source project, the plurality of to-be-translated functions in the to-be-translated repository, and the plurality of historical successful translation function pairs;
[0009] determining target triple translation knowledge corresponding to each of the plurality of to-be-translated functions according to the target language code sample self-evolution knowledge base, the dependency usage example self-evolution knowledge base, the successful translation function pair self-evolution knowledge base, and the plurality of to-be-translated functions;
[0010] inputting the target triple translation knowledge corresponding to each of the plurality of to-be-translated functions and each of the plurality of to-be-translated functions as inputs of the preset large model, and outputting implementation code corresponding to each of the plurality of to-be-translated functions;
[0011] embedding the implementation code corresponding to each of the plurality of to-be-translated functions into the repository architecture to generate a repository-level code translation result.
[0012] Optionally, the generating of the repository architecture by using the preset large model according to the to-be-translated repository includes:
[0013] extracting import statements in a plurality of files in the to-be-translated repository, and constructing a file-level calling dependency graph;
[0014] determining a file architecture translation order based on the file-level calling dependency graph;
[0015] performing a function body removal operation on the plurality of to-be-translated functions in the to-be-translated repository to determine a function header and an empty function body corresponding to each of the plurality of to-be-translated functions;
[0016] inputting the function header and the empty function body corresponding to each of the plurality of to-be-translated functions as a file architecture and sequentially inputting each of the file architectures into the preset large model for translation in the file architecture translation order to determine a file architecture under a plurality of target language versions;
[0017] constructing the repository architecture according to the file architecture under the plurality of target language versions.
[0018] Optionally, the constructing of the target language code sample self-evolution knowledge base, the dependency usage example self-evolution knowledge base, and the successful translation function pair self-evolution knowledge base based on the tree-shaped parser and the preset large model according to the open source project, the plurality of to-be-translated functions in the to-be-translated repository, and the plurality of historical successful translation function pairs includes:
[0019] constructing the target language code sample self-evolution knowledge base according to the open source project by using the tree-shaped parser;
[0020] The codes of each of the functions to be translated are input into the preset large model for translation, and target functions corresponding to each of the functions to be translated in a target language version are output;
[0021] The tree-shaped parser is used to construct a dependent use example self-evolution knowledge base according to the target functions in each of the target language versions;
[0022] A successful translation function pair self-evolution knowledge base is constructed according to a plurality of successful translation function pairs.
[0023] Optionally, the tree-shaped parser is used to construct a target language code sample self-evolution knowledge base according to the open source project, and the method comprises the following steps of:
[0024] The tree-shaped parser is used to extract functions in the open source project as target language code samples;
[0025] A target language code sample self-evolution knowledge base is constructed according to the target language code samples.
[0026] Optionally, the tree-shaped parser is used to construct a dependent use example self-evolution knowledge base according to the target functions in each of the target language versions, and the method comprises the following steps of:
[0027] The tree-shaped parser is used to identify calling nodes in the target functions in each of the target language versions;
[0028] Function dependency extraction is performed on each of the calling nodes to determine a plurality of function call statements;
[0029] The plurality of function call statements are screened according to function dependency names corresponding to the target functions in each of the target language versions to determine a plurality of target function call statements;
[0030] A plurality of code execution statements are extracted from the target functions in each of the target language versions;
[0031] The plurality of code execution statements are screened according to variable dependency names corresponding to the target functions in each of the target language versions to determine a plurality of target variable dependency call statements;
[0032] A plurality of dependent use examples are generated according to the plurality of target function call statements and the plurality of target variable dependency call statements;
[0033] A dependent use example self-evolution knowledge base is constructed according to the plurality of dependent use examples.
[0034] Optionally, the determining the target triple translation knowledge corresponding to each of the to-be-translated functions according to the target language code sample self-evolution knowledge base, the dependent usage example self-evolution knowledge base, the successful translation function pair self-evolution knowledge base and the plurality of to-be-translated functions comprises:
[0035] preprocessing the plurality of to-be-translated functions, the target language code sample self-evolution knowledge base, the dependent usage example self-evolution knowledge base and the successful translation function pair self-evolution knowledge base to determine a plurality of query functions, a preprocessed target language code sample self-evolution knowledge base, a preprocessed dependent usage example self-evolution knowledge base and a preprocessed successful translation function pair self-evolution knowledge base;
[0036] performing best matching similarity calculation on each of the to-be-translated functions and a plurality of target language code samples in the preprocessed target language code sample self-evolution knowledge base to determine a first best matching similarity between each of the to-be-translated functions and each of the target language code samples;
[0037] selecting target language code samples corresponding to the first best matching similarities in the first preset number of positions as intermediate target language code samples;
[0038] performing cosine similarity calculation on each of the to-be-translated functions and each of the intermediate target language code samples to determine a first cosine similarity between each of the to-be-translated functions and each of the intermediate target language code samples;
[0039] selecting intermediate target language code samples corresponding to the first cosine similarities in the second preset number of positions as final target target language code samples;
[0040] performing best matching similarity calculation on each of the to-be-translated functions and a plurality of dependent usage examples in the preprocessed dependent usage example self-evolution knowledge base to determine a second best matching similarity between each of the to-be-translated functions and each of the dependent usage examples;
[0041] selecting dependent usage examples corresponding to the second best matching similarities in the first preset number of positions as intermediate dependent usage examples;
[0042] performing cosine similarity calculation on each of the to-be-translated functions and each of the intermediate dependent usage examples to determine a second cosine similarity between each of the to-be-translated functions and each of the intermediate dependent usage examples;
[0043] selecting intermediate dependent usage examples corresponding to the second cosine similarities in the second preset number of positions as final dependent usage examples;
[0044] The third optimal matching similarity between each of the to-be-translated functions and each of the successful translation function pairs is determined by performing optimal matching similarity calculation on each of the to-be-translated functions and the preprocessed successful translation function pairs in the self-evolution knowledge base.
[0045] The successful translation function pair corresponding to the third optimal matching similarity of the first preset number of bits is selected as an intermediate successful translation function pair.
[0046] The third cosine similarity between each of the to-be-translated functions and each of the intermediate successful translation function pairs is determined by performing cosine similarity calculation on each of the to-be-translated functions and each of the intermediate successful translation function pairs.
[0047] The intermediate successful translation function pair corresponding to the third cosine similarity of the second preset number of bits is selected as a final successful translation function pair.
[0048] The final target target language code sample, the final dependent usage example, and the final successful translation function pair corresponding to each of the to-be-translated functions are used as target triple translation knowledge.
[0049] The second aspect of the present application provides a warehouse-level code translation device based on a large model, which comprises:
[0050] An acquisition module is configured to acquire a to-be-translated warehouse, an open source project, and a plurality of historical successful translation function pairs, and use a preset large model to perform architecture translation based on the to-be-translated warehouse to generate a warehouse architecture.
[0051] A construction module is configured to use the preset large model to construct a target language code sample self-evolution knowledge base, a dependent usage example self-evolution knowledge base, and a successful translation function pair self-evolution knowledge base based on a tree parser and according to the open source project, a plurality of to-be-translated functions in the to-be-translated warehouse, and a plurality of historical successful translation function pairs.
[0052] A determination module is configured to determine target triple translation knowledge corresponding to each of the to-be-translated functions based on the target language code sample self-evolution knowledge base, the dependent usage example self-evolution knowledge base, the successful translation function pair self-evolution knowledge base, and a plurality of to-be-translated functions.
[0053] An output module is configured to use the target triple translation knowledge corresponding to each of the to-be-translated functions and each of the to-be-translated functions as input of the preset large model, and output implementation code corresponding to each of the to-be-translated functions.
[0054] A generation module is configured to embed the implementation code corresponding to each of the to-be-translated functions into the warehouse architecture to generate a warehouse-level code translation result.
[0055] The third aspect of the present application provides a computer device, comprising a memory and a processor, the memory stores a computer program, and the computer program is executed by the processor to make the processor execute the steps of the warehouse-level code translation method based on a large model according to any one of the above.
[0056] The fourth aspect of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed to implement the steps of the warehouse-level code translation method based on a large model according to any one of the above.
[0057] The fifth aspect of the present application provides a computer program product, which comprises a computer program stored on a non-transitory computer readable storage medium, and the computer program comprises program instructions, wherein when the program instructions are executed by a computer, the computer executes the steps of the warehouse-level code translation method based on a large model according to any one of the above.
[0058] From the above technical solutions, the present application has the following advantages:
[0059] The above-mentioned scheme of the present application provides a warehouse-level code translation method based on a large model. First, the to-be-translated warehouse, the open source project and the plurality of historical successful translation function pairs are obtained, and the preset large model is used to perform architecture translation according to the to-be-translated warehouse to generate a warehouse architecture. Then, based on a tree-shaped parser, the preset large model is used to construct a target language code sample self-evolution knowledge base, a dependency usage example self-evolution knowledge base and a successful translation function pair self-evolution knowledge base according to the open source project, the plurality of to-be-translated functions in the to-be-translated warehouse and the plurality of historical successful translation function pairs. According to the target language code sample self-evolution knowledge base, the dependency usage example self-evolution knowledge base, the successful translation function pair self-evolution knowledge base and the plurality of to-be-translated functions, the target triple translation knowledge corresponding to each to-be-translated function is determined. The target triple translation knowledge corresponding to each to-be-translated function and each to-be-translated function are taken as inputs of the preset large model, and the implementation code corresponding to each to-be-translated function is output. Finally, the implementation code corresponding to each to-be-translated function is embedded into the warehouse architecture to generate a warehouse-level code translation result. Based on the above-mentioned scheme, the present application uses a tree-shaped parser and a preset large model to construct a target language code sample self-evolution knowledge base, a dependency usage example self-evolution knowledge base and a successful translation function pair self-evolution knowledge base according to the obtained open source project, to-be-translated warehouse and plurality of historical successful translation function pairs, so as to output the target triple translation knowledge corresponding to each to-be-translated function. In combination with the obtained implementation code corresponding to each to-be-translated function and the generated warehouse architecture, the process of outputting the warehouse-level code translation result, the present application takes the dependency into account, which can effectively improve the processing capability of the large model on the dependency, thereby enhancing the performance of the code translation task in the warehouse-level context. BRIEF DESCRIPTION OF DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the accompanying drawings in the following description only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0061] Figure 1 A step flow chart of a warehouse-level code translation method based on a large model provided for the first embodiment of the present application;
[0062] Figure 2 An example diagram of a file-level call dependency graph provided for the first embodiment of the present application;
[0063] Figure 3 An example diagram of function architecture provided for the first embodiment of the present application;
[0064] Figure 4 A framework schematic diagram of a warehouse-level code translation method based on a large model provided for the first embodiment of the present application;
[0065] Figure 5 A structural block diagram of a warehouse-level code translation device based on a large model provided for the second embodiment of the present application. DETAILED DESCRIPTION
[0066] The embodiments of the present application provide a warehouse-level code translation method and device based on a large model, which are used to solve the technical problem that the existing warehouse-level code translation method only involves function translation pairs, resulting in poor performance of code translation tasks in the warehouse-level context.
[0067] Term explanation:
[0068] Large model: Large Language Model,
[0069] Knowledge-driven: Since the knowledge source of the large model is pre-training corpus, there is a fatal flaw in the knowledge boundary. At the same time, the large model performs poorly on the less resource problem with a small proportion in the training corpus. The above problems can be attributed to the lack of relevant knowledge of the large model. In view of the deficiency of the large model, the performance of the model on specific problems can be enhanced through the knowledge-driven way of external knowledge base.
[0070] Warehouse-level code translation: Code translation refers to the migration of code projects from one programming language to another programming language for the purpose of adapting to different running environments, improving running speed, improving software security, etc. Warehouse-level code translation refers to that the object to be translated is a complete warehouse, which is closer to the real translation demand scene.
[0071] In order to make the application purposes, features and advantages of the present application more obvious and easy to understand, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings of the embodiments of the present application. Obviously, the embodiments described below are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0072] Please refer to Figure 1 , Figure 1 A step flowchart of a warehouse-level code translation method based on a large model provided for the first embodiment of the present application.
[0073] The warehouse-level code translation method based on a large model provided by the present application comprises:
[0074] Step 101, obtaining a warehouse to be translated, an open source project and a plurality of historical successful translation function pairs, and using a preset large model to perform architecture translation according to the warehouse to be translated to generate a warehouse architecture;
[0075] It should be noted that for the warehouse to be translated, this part translates the warehouse architecture in two stages: 1) performing file-level dependency inference construction on the warehouse to be translated to construct a file-level call dependency graph; 2) performing file-level architecture translation according to the topological order of the call dependency graph.
[0076] Specifically, step 101 can include the following sub-steps S11-S15:
[0077] Step S11, extracting import statements in a plurality of files in the warehouse to be translated, and constructing a file-level call dependency graph;
[0078] Step S12, determining a file architecture translation order based on the file-level call dependency graph;
[0079] Step S13, performing a function body removal operation on a plurality of functions to be translated in the warehouse to be translated to determine the function header and the empty function body corresponding to each function to be translated;
[0080] Step S14, taking the function header and the empty function body corresponding to each function to be translated as a file architecture and sequentially inputting each file architecture to the preset large model for translation according to the file architecture translation order to determine the file architecture under a plurality of target language versions;
[0081] Step S15, constructing a warehouse architecture according to the file architecture under a plurality of target language versions.
[0082] The warehouse to be translated is a warehouse of an implemented source language version. The warehouse to be translated is a warehouse of an implemented source language version.
[0083] The files in the to-be-translated repository are code files in the implemented source language version.
[0084] The file architecture in the target language version is the code file architecture in the to-be-generated target language version of the repository.
[0085] The to-be-translated functions in the to-be-translated repository are functions in the implemented source language version of the repository.
[0086] The preset large model is a black box model.
[0087] It should be noted that, please refer to Figure 2 For each file in the to-be-translated repository, the import statement in the file is extracted to determine the dependency relationship of the current file to the remaining files in the repository. Then the file is taken as a node, and the dependency relationship is taken as a directed edge. A directed edge from node C to node A represents that file C has a dependency relationship with file A, indicating that the implementation of file C calls the variables, data types or functions implemented in file A. Since circular dependencies are not allowed, the file-level call dependency graph constructed is a directed acyclic graph.
[0088] Further, there may be a dependency relationship between different files in the same repository, for example, file C depends on file A, which means that the implementation of file C depends on file A, and the implementation of file A will have a certain influence and guiding effect on the implementation of file C. Therefore, the translation effect of translating file A first and then translating file C will be better than that of translating C first and then A. Therefore, based on the constructed file-level call dependency graph, a topological order can be obtained. The specific method is as follows: for the constructed file-level call dependency graph, first obtain the nodes with an out-degree of 0 (indicating that there is no unprocessed dependent file), output these nodes and eliminate the edges originally pointing to these nodes, indicating that the corresponding dependency relationship has been processed. At this time, update the graph and find the nodes with an out-degree of 0 again. Since the constructed file-level call dependency graph is a directed acyclic graph, an effective topological order can be finally obtained and used as the order of file architecture translation.
[0089] Further, please refer to Figure 3 For each file, the following is an example in python: remove the function body of each function, i.e. the specific implementation code of the function, and only leave the function header and empty function body as the file architecture. Then, according to the obtained topological order, translate each file architecture using the large model to obtain the corresponding file architecture in the target language version, thereby constructing the repository architecture and completing the repository architecture translation.
[0090] Step 102, based on the tree parser, using the preset large model to construct a target language code sample self-evolution knowledge base, a dependency usage example self-evolution knowledge base, and a successful translation function pair self-evolution knowledge base according to the open source project, the plurality of to-be-translated functions in the to-be-translated repository, and the plurality of historical successful translation function pairs;
[0091] It should be noted that the application enhances the function translation effect of the large model in the repository level context through triple knowledge enhancement. The three key knowledge sources include: (1) target language code samples from open source projects, (2) dependency usage examples from the repository being translated (to-be-translated repository), and (3) successful translation function pairs from code translation history. Figure 4 The framework of the application is shown. When a repository-level context function that needs to be translated (a to-be-translated function in a to-be-translated repository) is given, the application generates target language code (i.e., the implementation code corresponding to the to-be-translated function) using triple knowledge. This part is divided into three stages: Stage 1 (translation knowledge base construction) is performed offline, Stage 2 (translation knowledge retrieval) and Stage 3 (knowledge-enhanced code translation) are performed online. The application also proposes a self-evolution framework to continuously improve translation quality as the target language code base, repository structure, and translation history continue to develop. By dynamically incorporating the latest translation knowledge, the application improves the adaptability and accuracy of code translation. The translation knowledge base consists of three parts: a target language code sample self-evolution knowledge base, a dependency usage example self-evolution knowledge base, and a successful translation function pair self-evolution knowledge base. These parts aim to enhance the large model's ability to understand target language grammar, correctly identify and call dependencies, and identify grammar differences. Figure 4 The existing repository context in the application is the to-be-translated repository. The target language repository is the open source project.
[0092] Specifically, step 102 can include the following sub-steps S21-S24:
[0093] Step S21, using a tree parser to construct a target language code sample self-evolution knowledge base according to the open source project;
[0094] Further, step S21 can include the following sub-steps S211-S212:
[0095] Step S211, using a tree parser to extract functions in the open source project as target language code samples;
[0096] Step S212, constructing a target language code sample self-evolution knowledge base according to the target language code samples.
[0097] The tree parser is tree-sitter, an open-source parser generation tool.
[0098] It should be noted that, since the codes implementing the same function usually use similar function names and variable names, resulting in high text similarity, the application can effectively identify code samples implementing similar functions from projects in similar fields by searching for target language codes most similar in text to the source function. These code samples help improve translation effectiveness. In order to simulate the most comprehensive target language reference code available at present, the application obtains open source projects from the open source community, extracts functions in the open source projects using tree-sitter, and uses all extracted functions as target language code samples, and uses all target language code samples to build a target language code sample self-evolution knowledge base. This type of knowledge is used to enhance the code generated by the large model to better meet the grammar requirements of the target language. This component updates the target language code sample self-evolution knowledge base by automatically detecting and downloading newly added target language projects from the open source community and updating according to the preconfigured update interval.
[0099] Step S22, input the code of each function to be translated into the preset large model for translation, and output the target function corresponding to each function to be translated in the target language version;
[0100] The target function in the target language version is the function to be translated in the target language version, wherein the target language version refers to the programming language version as the translation target.
[0101] Step S23, using a tree parser, constructing a dependent usage example self-evolution knowledge base according to the target function in each target language version;
[0102] Specifically, step S23 can include the following sub-steps S231-S237:
[0103] Step S231, using a tree parser to identify the call nodes in the target function in each target language version;
[0104] Step S232, function dependency extraction is performed on each call node to determine a plurality of function call statements;
[0105] Step S233, filtering the plurality of function call statements according to the function dependency names corresponding to the target function in each target language version to determine a plurality of target function call statements;
[0106] Step S234, extracting a plurality of code execution statements in the target function in each target language version;
[0107] Step S235, filtering the plurality of code execution statements according to the variable dependency names corresponding to the target function in each target language version to determine a plurality of target variable dependency call statements;
[0108] Step S236, according to the plurality of target function call statements and the plurality of target variable dependency call statements, generate a plurality of dependency usage examples;
[0109] Step S237, according to the plurality of dependency usage examples, construct a dependency usage example self-evolution knowledge base.
[0110] It should be noted that the same dependency is usually called multiple times in the same project, so how the dependency is called elsewhere in the project can be used as a useful reference for using the dependency in the target function. Specifically, when the scope is the same, the call path and input parameter type of the dependency are the same; when the scope is different, the call path is still similar, and the input parameter type is usually the same. Therefore, for each involved dependency, the present application extracts the corresponding dependency call statement from the current project that shares the same scope as the target function. If multiple call statements are found, only the first occurrence is kept as an example. Specifically, using tree-sitter, the function dependency is extracted by identifying the call node in the target function under the target language version, so as to obtain all function call statements. Then, according to the name of the target function dependency, all function call statements are matched to obtain the required dependency call statement (target function call statement). For variable dependencies, the present application extracts code execution statements in the target function and filters the code execution statements according to the name of the target variable dependency corresponding to the target function, so as to obtain the corresponding variable dependency call statement (target variable dependency call statement). These extracted statements (i.e., target function call statements, target variable dependency call statements) constitute dependency usage examples, and further constitute a dependency usage example self-evolution knowledge base, enhancing the ability of the pre-trained large model to correctly identify and call dependencies. This component constitutes a dependency usage example self-evolution knowledge base by continuously performing the above dependency extraction operation during translation.
[0111] Step S24, according to the plurality of historical successful translation function pairs, construct a successful translation function pair self-evolution knowledge base.
[0112] It should be noted that providing the model with information of the same category as the target task can maximize its performance on the current task. However, due to the scarcity of existing warehouse-level context code translation datasets, it is difficult to obtain a large number of high-quality function-level equivalent pairs to construct a successful translation function pair library for warehouse-level context. Therefore, the present application collects successful translation function pairs during translation and adds them to the knowledge base to construct a successful translation function pair self-evolution knowledge base, which is used to enhance the model's ability to identify syntactic differences when performing code translation tasks. This component constitutes a successful translation function pair self-evolution knowledge base by continuously extracting successful translation function pairs verified by test cases from previous translation history.
[0113] Step 103, determining the target triple translation knowledge corresponding to each of the to-be-translated functions according to the target language code sample self-evolution knowledge base, the dependency usage example self-evolution knowledge base, the successful translation function pair self-evolution knowledge base and the plurality of to-be-translated functions;
[0114] The target triple translation knowledge is composed of the final target target language code sample, the final dependency usage example and the final successful translation function pair corresponding to the to-be-translated function.
[0115] It should be noted that the input of the warehouse-level context code translation task includes three parts: the source code of the to-be-translated function, the function signature of the target function and the dependency of the target function. For a given warehouse-level context code translation task, the present application retrieves relevant translation knowledge from the constructed translation knowledge base and constructs a translation prompt library through a three-step retrieval process: target function and dependency extraction, candidate knowledge retrieval and candidate knowledge reordering. Among them, the target function refers to the to-be-translated function in the target language version. The function signature of the target function refers to the function signature in the target language obtained by translating the function (to-be-translated function) in the source language in the architecture translation stage, and the dependency of the target function refers to the dependency needed to implement the to-be-translated function in the target language version.
[0116] Further, for each warehouse-level context code translation task, the present application extracts the source code of the to-be-translated function and its corresponding dependency through pattern matching. The source code of the to-be-translated function is used to retrieve target language code examples and successful translation function pairs. At the same time, the dependency is matched with the source code in the dependency usage example, and the corresponding usage example is extracted as the dependency usage example of the dependency.
[0117] Specifically, step 103 can include the following sub-steps S31-S314:
[0118] Step S31, preprocessing the plurality of to-be-translated functions, the target language code sample self-evolution knowledge base, the dependency usage example self-evolution knowledge base and the successful translation function pair self-evolution knowledge base to determine the plurality of query functions, the preprocessed target language code sample self-evolution knowledge base, the preprocessed dependency usage example self-evolution knowledge base and the preprocessed successful translation function pair self-evolution knowledge base;
[0119] Step S32, calculating the first best matching similarity between each to-be-translated function and each target language code sample by respectively best matching each to-be-translated function with the plurality of target language code samples in the preprocessed target language code sample self-evolution knowledge base;
[0120] Step S33, selecting the target language code samples corresponding to the first pre-set number of first best matching similarities as intermediate target language code samples;
[0121] Step S34, cosine similarity calculation is performed between each function to be translated and each intermediate target language code sample to determine the first cosine similarity between each function to be translated and each intermediate target language code sample;
[0122] Step S35, the intermediate target language code samples corresponding to the first cosine similarity of the first preset second number of bits are selected as the final target target language code samples;
[0123] Step S36, best matching similarity calculation is performed between each function to be translated and the preprocessed multiple dependent use examples in the self-evolution knowledge base to determine the second best matching similarity between each function to be translated and each dependent use example;
[0124] Step S37, the dependent use examples corresponding to the second best matching similarity of the first preset first number of bits are selected as intermediate dependent use examples;
[0125] Step S38, cosine similarity calculation is performed between each function to be translated and each intermediate dependent use example to determine the second cosine similarity between each function to be translated and each intermediate dependent use example;
[0126] Step S39, the intermediate dependent use examples corresponding to the second cosine similarity of the first preset second number of bits are selected as the final dependent use examples;
[0127] Step S310, best matching similarity calculation is performed between each function to be translated and the preprocessed multiple successful translation function pairs in the self-evolution knowledge base to determine the third best matching similarity between each function to be translated and each successful translation function pair;
[0128] Step S311, the successful translation function pairs corresponding to the third best matching similarity of the first preset first number of bits are selected as intermediate successful translation function pairs;
[0129] Step S312, cosine similarity calculation is performed between each function to be translated and each intermediate successful translation function pair to determine the third cosine similarity between each function to be translated and each intermediate successful translation function pair;
[0130] Step S313, the intermediate successful translation function pairs corresponding to the third cosine similarity of the first preset second number of bits are selected as the final successful translation function pairs;
[0131] Step S314, the final target target language code samples, the final dependent use examples, and the final successful translation function pairs corresponding to each function to be translated are used as target triple translation knowledge.
[0132] BM25 is the Best - Matching 25. It is a term weighting algorithm used in the field of information retrieval to calculate the relevance score between a document and a query.
[0133] The query function is a preprocessed function to be translated.
[0134] It should be noted that for each function to be translated, the present application retrieves the top N target language code samples, dependency usage examples, and successful translation function pair knowledge items of each query function using BM25. Before calculating the BM25 similarity, the function to be translated, the retrieval document (i.e., the target language code sample self-evolution knowledge base, the dependency usage example self-evolution knowledge base, and the successful translation function pair self-evolution knowledge base) all need to go through necessary preprocessing steps, including word segmentation, morphological restoration, and stop word removal.
[0135] Further, after the function to be translated, the target language code sample self-evolution knowledge base, the dependency usage example self-evolution knowledge base, and the successful translation function pair self-evolution knowledge base are preprocessed, the best matching similarity between each query function and each target language code sample in the preprocessed target language code sample self-evolution knowledge base, each dependency usage example in the preprocessed dependency usage example self-evolution knowledge base, and each successful translation function pair in the preprocessed successful translation function pair self-evolution knowledge base is calculated, obtaining a plurality of first best matching similarities, a plurality of second best matching similarities, and a plurality of third best matching similarities corresponding to each query function; for example, assuming that the first preset number is 3, the second preset number is 2, the number of query functions is 3, the number of target language code samples in the preprocessed target language code sample self-evolution knowledge base is 5, the number of dependency usage examples in the preprocessed dependency usage example self-evolution knowledge base is 5, and the number of successful translation function pairs in the preprocessed successful translation function pair self-evolution knowledge base is 5, then the number of first best matching similarities corresponding to each query function is 5, the number of second best matching similarities is 5, and the number of third best matching similarities is 5. The best matching similarities (first best matching similarities, second best matching similarities, and third best matching similarities) are sorted in descending order, and the top 3 first best matching similarities corresponding to the target language code samples are selected as intermediate target language code samples. Then, the first cosine similarity between each query function and the three intermediate target language code samples is calculated, and the three first cosine similarities are sorted in descending order. The top 2 first cosine similarities corresponding to the intermediate target language code samples are selected as the final target target language code samples. The processing principles of the final dependency usage examples and the final successful translation function pairs are consistent with the principles described above, and the present application will not be described in more detail.
[0136] Further, the calculation formula of the best matching similarity is specifically:
[0137]
[0138] wherein, is the best matching similarity, the best matching similarity calculation includes the best matching similarity calculation between the function to be translated and the target language code sample, the best matching similarity calculation between the function to be translated and the dependent use example, and the best matching similarity calculation between the function to be translated and the successful translation function pair, D is a candidate knowledge in the knowledge base, is the function to be translated, represents the i-th word in the function to be translated; is the inverse document frequency of the i-th word of the function to be translated; is the word frequency of the i-th word in the candidate knowledge D, the candidate knowledge is the preprocessed target language code sample in the target language code sample self-evolution knowledge base, the preprocessed dependent use example in the dependent use example self-evolution knowledge base, and the preprocessed successful translation function pair in the successful translation function pair self-evolution knowledge base; is a hyperparameter, k takes a value of 1.2, and b takes a value of 0.75; is the number of words of the candidate knowledge; is the average length of the candidate knowledge in the knowledge base; n is the number of words of the function to be translated.
[0139] It is worth mentioning that although BM25 performs well in text-based similarity calculation, it is difficult to effectively identify noise such as comments and other key information in code such as Abstract Syntax Tree (AST). Therefore, the present application re-ranks the retrieved knowledge items (i.e., intermediate dependent use examples, intermediate successful translation function pairs, and intermediate target language code samples) by using a unified cross-modal pre-training programming language model UniXcoder to calculate the cosine similarity. In the experiment, the top N' target language code samples, dependent use examples, and successful translation function pairs with the highest UniXcoder scores among the retrieved knowledge items will be selected as the final knowledge items provided to the preset large model base code translation.
[0140] Step 104, taking each target triple translation knowledge corresponding to each function to be translated and each function to be translated as input of the preset large model, outputting the implementation code corresponding to each function to be translated;
[0141] It should be noted that the pre-set large model is used to perform translation based on the retrieved triple translation knowledge, to obtain the result (the implementation code corresponding to each function to be translated), the code repair based on the large model is applied to correct any identified problems, and finally the obtained implementation code corresponding to each function to be translated is embedded in the warehouse architecture to obtain the warehouse-level code translation result.
[0142] It is worth mentioning that based on the retrieved translation knowledge items, the present application translates the code from the source language to the target language using a large model. The prompt word design adopts a format suitable for the large model: markdown (markdown format), Chain-of-Thought strategy (Chain-of-Thought strategy), and guides the large model to translate the function step by step. First, the large model is required to confirm the function to be implemented by the current function. Second, in order to understand the differences between the source language and the target language, such as dependency relationships, syntax, and available local variables, the present application uses a paraphrasing technique, which requires the large model to list all the dependencies and local variables used, and to distinguish the syntax differences. Finally, the large model translates based on the function to be implemented, the dependencies used, and the syntax differences.
[0143] Step 105, embedding the implementation code corresponding to each function to be translated into the warehouse architecture to generate a warehouse-level code translation result.
[0144] It should be noted that for simple syntax errors that may be contained in the code translated by the large model, the present application uses error information to guide the large model to iteratively improve the translation results that fail the test due to compilation or functional errors, further improving the accuracy of the translation results.
[0145] As a comparison of technical effects, in combination with the prior art, 1) the neglect of warehouse-level problems of the existing method: the problem scenario of the existing method is in the case of a single function or a single file, without considering the warehouse-level context, nor considering the complete warehouse as the translation subject. But in the real code translation demand scenario, the task to be translated is a huge project with complex context and dependency information, and the task object handled in the existing method is at the single file level, without involving cross-file dependencies and warehouse-level context, and cannot be effectively applied to the complex warehouse scenario of the real translation demand scenario. 2) Insufficient processing ability for dependencies: the content provided to the large model in the existing method only involves function translation pairs, ignoring dependency knowledge, making the existing method insufficient in processing dependencies and performing poorly in warehouse-level context code translation tasks, and cannot be effectively applied to real translation demand scenarios. The present application enhances the large model's ability to handle dependencies by extracting dependency call examples from existing warehouse contexts to build a dependency knowledge base, improving the large model's ability to correctly identify and call dependencies, and enhancing the large model's performance in warehouse-level context translation tasks. 3) Knowledge scale cannot be expanded: the knowledge base built in the existing method is static and cannot be expanded after construction, so the range of knowledge base that can support the large model is fixed. As the translation process continues, the types of code translation tasks handled gradually increase, and the knowledge base of the existing method cannot continuously improve the code translation ability of the large model. At the same time, previous translation success tasks cannot be fully utilized, while the present application builds a self-evolving code translation knowledge base that extracts triple knowledge from existing warehouses and translation results in an automated manner and updates it to the warehouse, continuously increasing the knowledge scale, expanding the knowledge base coverage area, and enhancing the knowledge base's ability to improve the large model's performance in warehouse-level context code translation tasks.
[0146] To solve the above problems, the present application provides a warehouse-level code translation method based on a large model, a warehouse architecture translation based on a large model, and a self-evolving knowledge-driven warehouse-level context function translation based on a large model. The present application first obtains a translated and compilable warehouse architecture through the first part for the complete warehouse to be translated, and then translates each function based on the warehouse architecture obtained by translation. The compilable warehouse architecture obtained by the first part allows each translated function of the second part to be separately verified for correctness. By collecting multiple correctly translated code pairs in the target language as a corpus of knowledge, the construction of the knowledge base is completed. During retrieval, the most similar k examples are dynamically retrieved based on cosine similarity. Then the retrieved translation pairs are sent into the large model together with the original query for translation, thereby improving the accuracy and efficiency of code translation.
[0147] Based on the above, the present application takes the warehouse level context and the complete warehouse into account in the code translation technology, proposes a warehouse level code translation technology based on a large model, effectively reduces the distance between existing tools and real translation demand scenarios. The extraction and use of dependency call knowledge effectively improve the ability of the large model to handle dependencies, enhance the ability of the large model to correctly identify and call dependencies, and improve the performance of the code translation method based on the large model in handling warehouse level code translation tasks. At the same time, the present application extracts target language syntax knowledge, dependency call knowledge and syntax difference knowledge from the target language code samples of existing projects, the dependency usage examples in the warehouse being translated, and the successful translation function pairs of code translation history respectively by an automatic way, and constructs a knowledge base, reduces the model's target language syntax misunderstanding, dependency misuse and syntax difference confusion, and improves the model's code translation ability from various aspects and dimensions. In addition, the present application automatically increases the knowledge scale of the existing knowledge base, expands the coverage field of the knowledge base, improves the supporting role and range of the knowledge base for the large model; through the self-repairing way of the large model, the code translation effect based on the large model is further improved, the syntax errors and function inconsistency problems existing in the translation results of the large model are effectively reduced, and the code translation result quality of the large model is improved.
[0148] Compared with the prior art, the task scene of the existing method does not consider the warehouse level, while in the real development scene, there are usually rich dependencies and context information, and the translated object should be a complete warehouse. The present application first considers the warehouse level in the code translation technology, proposes a warehouse level code translation technology based on a large model, and effectively reduces the distance between the existing tool and the real translation demand scene. The existing method ignores the dependency knowledge, and the present application solves this problem by extracting dependency calling examples from the existing warehouse context to construct a dependency knowledge base, and effectively enhances the processing ability of the large model to the dependency, improves the ability of the large model to correctly identify and call the dependency, and enhances the performance of the large model in the translation task of the warehouse level context. The existing method only involves a single knowledge, and cannot improve the code translation ability of the model from various aspects and dimensions. The present application extracts target language syntax knowledge, dependency calling knowledge and syntax difference knowledge from the target language code samples of the existing project, the dependency usage examples in the warehouse being translated, and the successful translation function pairs of the code translation history in an automated manner, and constructs a knowledge base, reduces the misunderstanding of the model in the target language syntax, the misuse of the dependency, and the confusion of the syntax difference, and improves the code translation ability of the model from various aspects and dimensions. The existing method knowledge base is not expandable, and the present application solves this problem by constructing a self-evolving code translation knowledge base. The triple knowledge is extracted from the existing warehouse and translation result in an automated manner, and is updated to the warehouse, continuously increases the knowledge scale, expands the knowledge base coverage field, and enhances the improvement ability of the knowledge base to the large model in the warehouse level context code translation task. The existing method does not have any post-processing method to further improve the translation effect after translation. The present application further improves the code translation effect based on the large model through the self-repairing of the large model, effectively reduces the syntax error and function inconsistency problem existing in the translation result of the large model, and improves the code translation result quality of the large model.
[0149] In the embodiment of the application, the warehouse-level code translation method based on a large model is provided. First, a to-be-translated warehouse, an open source project and a plurality of historical successful translation function pairs are acquired, and a preset large model is used to perform architecture translation according to the to-be-translated warehouse to generate a warehouse architecture. Then, based on a tree-shaped parser, the preset large model is used to construct a target language code sample self-evolution knowledge base, a dependency usage example self-evolution knowledge base and a successful translation function pair self-evolution knowledge base according to the open source project, a plurality of to-be-translated functions in the to-be-translated warehouse and the plurality of historical successful translation function pairs. The target triple translation knowledge corresponding to each to-be-translated function is determined according to the target language code sample self-evolution knowledge base, the dependency usage example self-evolution knowledge base, the successful translation function pair self-evolution knowledge base and the plurality of to-be-translated functions. The target triple translation knowledge corresponding to each to-be-translated function and each to-be-translated function are used as inputs of the preset large model, and the implementation code corresponding to each to-be-translated function is output. Finally, the implementation code corresponding to each to-be-translated function is embedded into the warehouse architecture to generate a warehouse-level code translation result. Based on the above scheme, the tree-shaped parser and the preset large model are used to construct the target language code sample self-evolution knowledge base, the dependency usage example self-evolution knowledge base and the successful translation function pair self-evolution knowledge base according to the acquired open source project, to-be-translated warehouse and plurality of historical successful translation function pairs, so as to output the target triple translation knowledge corresponding to each to-be-translated function. In combination with the obtained implementation code corresponding to each to-be-translated function and the generated warehouse architecture, the process of outputting the warehouse-level code translation result, the application takes the dependency into account, can effectively improve the processing capability of the large model on the dependency, and thus enhances the performance of the code translation task in the warehouse-level context.
[0150] Please refer to Figure 5 , Figure 5 The structure block diagram of the warehouse-level code translation device based on a large model provided in the second embodiment of the application is shown in FIG. 2.
[0151] The warehouse-level code translation device based on a large model provided in the application comprises:
[0152] The acquisition module 501 is configured to acquire a to-be-translated warehouse, an open source project and a plurality of historical successful translation function pairs, and use a preset large model to perform architecture translation according to the to-be-translated warehouse to generate a warehouse architecture.
[0153] The construction module 502 is configured to use the preset large model to construct a target language code sample self-evolution knowledge base, a dependency usage example self-evolution knowledge base and a successful translation function pair self-evolution knowledge base according to the open source project, a plurality of to-be-translated functions in the to-be-translated warehouse and the plurality of historical successful translation function pairs based on a tree-shaped parser.
[0154] The determining module 503 is configured to determine target triple translation knowledge corresponding to each function to be translated according to the target language code sample self-evolution knowledge base, the dependency usage example self-evolution knowledge base, the successful translation function pair self-evolution knowledge base, and the plurality of functions to be translated.
[0155] The output module 504 is configured to output the implementation code corresponding to each function to be translated by taking the target triple translation knowledge corresponding to each function to be translated and each function to be translated as inputs of the preset large model.
[0156] The generating module 505 is configured to embed the implementation code corresponding to each function to be translated into the warehouse architecture to generate a warehouse-level code translation result.
[0157] Further, the obtaining module 501 is specifically configured to:
[0158] extract import statements in a plurality of files in the warehouse to be translated, and construct a file-level call dependency graph;
[0159] determine a file architecture translation order based on the file-level call dependency graph;
[0160] perform a function body removal operation on a plurality of functions to be translated in the warehouse to be translated to determine a function header and an empty function body corresponding to each function to be translated;
[0161] take the function header and the empty function body corresponding to each function to be translated as a file architecture and input the file architecture into the preset large model in sequence according to the file architecture translation order to determine a file architecture under a plurality of target language versions;
[0162] construct a warehouse architecture according to the file architecture under the plurality of target language versions.
[0163] Further, the constructing module 502 includes:
[0164] The first submodule is configured to construct a target language code sample self-evolution knowledge base according to an open source project by using a tree-shaped parser;
[0165] The second submodule is configured to input the code of each function to be translated into the preset large model to translate and output a target function under a target language version corresponding to each function to be translated;
[0166] The third submodule is configured to construct a dependency usage example self-evolution knowledge base according to the target function under each target language version by using a tree-shaped parser;
[0167] The fourth submodule is configured to construct a successful translation function pair self-evolution knowledge base according to a plurality of historical successful translation function pairs.
[0168] Further, the first submodule is specifically configured to:
[0169] extracting functions in the open source project by using a tree parser and taking the target language code samples;
[0170] According to the target language code samples, a target language code sample self-evolution knowledge base is constructed.
[0171] Further, the second submodule is specifically configured to:
[0172] The calling nodes in the target functions under each target language version are identified by using a tree parser;
[0173] The function dependencies of each calling node are extracted to determine a plurality of function calling statements;
[0174] The plurality of function calling statements are screened according to the function dependency names corresponding to the target functions under each target language version to determine a plurality of target function calling statements;
[0175] A plurality of code execution statements are extracted from the target functions under each target language version;
[0176] The plurality of code execution statements are screened according to the variable dependency names corresponding to the target functions under each target language version to determine a plurality of target variable dependency calling statements;
[0177] According to the plurality of target function calling statements and the plurality of target variable dependency calling statements, a plurality of dependency usage examples are generated;
[0178] According to the plurality of dependency usage examples, a dependency usage example self-evolution knowledge base is constructed.
[0179] Further, the determining module 503 is specifically configured to:
[0180] The plurality of functions to be translated, the target language code sample self-evolution knowledge base, the dependency usage example self-evolution knowledge base, and the successfully translated function pair self-evolution knowledge base are preprocessed to determine a plurality of query functions, a preprocessed target language code sample self-evolution knowledge base, a preprocessed dependency usage example self-evolution knowledge base, and a preprocessed successfully translated function pair self-evolution knowledge base;
[0181] The first best matching similarity between each function to be translated and each target language code sample is determined by performing best matching similarity calculation on each function to be translated and the plurality of target language code samples in the preprocessed target language code sample self-evolution knowledge base;
[0182] The target language code samples corresponding to the first best matching similarities of the first preset number of positions are selected as intermediate target language code samples;
[0183] The cosine similarity between each function to be translated and each intermediate target language code sample is calculated, to determine the first cosine similarity between each function to be translated and each intermediate target language code sample;
[0184] The intermediate target language code sample corresponding to the first cosine similarity of the first preset number of bits is selected as the final target target language code sample;
[0185] The best matching similarity between each function to be translated and the plurality of dependent use examples in the preprocessed dependent use example self-evolution knowledge base is calculated, to determine the second best matching similarity between each function to be translated and each dependent use example;
[0186] The dependent use example corresponding to the second best matching similarity of the first preset number of bits is selected as the intermediate dependent use example;
[0187] The cosine similarity between each function to be translated and each intermediate dependent use example is calculated, to determine the second cosine similarity between each function to be translated and each intermediate dependent use example;
[0188] The intermediate dependent use example corresponding to the second cosine similarity of the first preset number of bits is selected as the final dependent use example;
[0189] The best matching similarity between each function to be translated and the plurality of successful translation function pairs in the preprocessed successful translation function pair self-evolution knowledge base is calculated, to determine the third best matching similarity between each function to be translated and each successful translation function pair;
[0190] The successful translation function pair corresponding to the third best matching similarity of the first preset number of bits is selected as the intermediate successful translation function pair;
[0191] The cosine similarity between each function to be translated and each intermediate successful translation function pair is calculated, to determine the third cosine similarity between each function to be translated and each intermediate successful translation function pair;
[0192] The intermediate successful translation function pair corresponding to the third cosine similarity of the first preset number of bits is selected as the final successful translation function pair;
[0193] The final target target language code sample, the final dependent use example, and the final successful translation function pair corresponding to each function to be translated are used as target triple translation knowledge.
[0194] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described device, module and sub-module can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.
[0195] The embodiment of the present application also provides a computer device, comprising a memory and a processor, the memory stores a computer program; the computer program is executed by the processor, so that the processor executes the steps of the warehouse-level code translation method based on a large model according to any one of the above embodiments.
[0196] The embodiment of the present application also provides a computer readable storage medium, which stores a computer program / instruction, and the computer program / instruction is executed by a processor to implement the steps of the warehouse-level code translation method based on a large model according to any one of the above embodiments.
[0197] The embodiment of the present application also provides a computer program product, comprising a computer program / instruction, and the computer program / instruction is executed by a processor to implement the steps of the warehouse-level code translation method based on a large model according to any one of the above embodiments.
[0198] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented by other manners. For example, the apparatus embodiment described above is only schematic, for example, the division of units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, apparatuses or units, and can be electrical, mechanical or other forms.
[0199] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0200] The above description and the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A warehouse-level code translation method based on a large model, characterized by, The method comprises the following steps: acquiring a to-be-translated repository, an open source project, and a plurality of historical successful translation function pairs, and using a preset large model to perform architecture translation on the to-be-translated repository to generate a repository architecture; based on a tree-like parser, using the preset large model to construct a target language code sample self-evolution knowledge base, a dependency usage example self-evolution knowledge base, and a successful translation function pair self-evolution knowledge base according to the open source project, a plurality of to-be-translated functions in the to-be-translated repository, and a plurality of historical successful translation function pairs; determining target triple translation knowledge corresponding to each to-be-translated function according to the target language code sample self-evolution knowledge base, the dependency usage example self-evolution knowledge base, the successful translation function pair self-evolution knowledge base, and a plurality of to-be-translated functions; taking the target triple translation knowledge corresponding to each to-be-translated function and each to-be-translated function as input of the preset large model, and outputting implementation code corresponding to each to-be-translated function; embedding the implementation code corresponding to each to-be-translated function into the repository architecture to generate a repository-level code translation result.
2. The large model-based warehouse-level code translation method of claim 1, wherein, The method comprises the following steps: extracting import statements in a plurality of files in the to-be-translated repository and constructing a file-level calling dependency graph; determining a file architecture translation order based on the file-level calling dependency graph; performing a function body removal operation on a plurality of to-be-translated functions in the to-be-translated repository to determine a function header and an empty function body corresponding to each to-be-translated function; taking the function header and the empty function body corresponding to each to-be-translated function as a file architecture and inputting each file architecture into a preset large model in sequence according to the file architecture translation order to determine a file architecture under a plurality of target language versions; constructing a repository architecture according to the file architecture under a plurality of target language versions.
3. The large model-based warehouse-level code translation method of claim 1, wherein, The method comprises the following steps: using the tree-like parser to extract functions in the open source project as target language code samples; constructing a target language code sample self-evolution knowledge base according to the target language code samples; inputting the code of each to-be-translated function into the preset large model to translate and output a target function under a target language version corresponding to each to-be-translated function; using the tree-like parser to construct a dependency usage example self-evolution knowledge base according to the target function under each target language version; 4. The warehouse-level code translation method based on a large model according to claim 3, characterized by, constructing a successful translation function pair self-evolution knowledge base according to a plurality of historical successful translation function pairs. The method comprises the following steps: using the tree-like parser to extract functions in the open source project as target language code samples; constructing a target language code sample self-evolution knowledge base according to the target language code samples.
5. The large model-based warehouse-level code translation method of claim 3, wherein, The tree parser is used to construct a dependent use example self-evolution knowledge base according to the target functions under each target language version, including: The tree parser is used to identify the call nodes in the target functions under each target language version; The function dependencies of each call node are extracted to determine a plurality of function call statements; The function call statements are screened according to the function dependency names corresponding to the target functions under each target language version to determine a plurality of target function call statements; A plurality of code execution statements are extracted from the target functions under each target language version; The code execution statements are screened according to the variable dependency names corresponding to the target functions under each target language version to determine a plurality of target variable dependency call statements; A plurality of dependent use examples are generated according to the target function call statements and the target variable dependency call statements; A dependent use example self-evolution knowledge base is constructed according to the dependent use examples.
6. The large model-based warehouse-level code translation method of claim 1, wherein, The target triple translation knowledge corresponding to each to-be-translated function is determined according to the target language code sample self-evolution knowledge base, the dependent use example self-evolution knowledge base, the successful translation function pair self-evolution knowledge base, and the plurality of to-be-translated functions, including: The to-be-translated functions, the target language code sample self-evolution knowledge base, the dependent use example self-evolution knowledge base, and the successful translation function pair self-evolution knowledge base are preprocessed to determine a plurality of query functions, a preprocessed target language code sample self-evolution knowledge base, a preprocessed dependent use example self-evolution knowledge base, and a preprocessed successful translation function pair self-evolution knowledge base; The to-be-translated functions are respectively matched with the target language code samples in the preprocessed target language code sample self-evolution knowledge base to determine the first best matching similarity between the to-be-translated functions and the target language code samples; The target language code samples corresponding to the first best matching similarity of the first preset number of positions are selected as intermediate target language code samples; The to-be-translated functions are respectively matched with the intermediate target language code samples to determine the first cosine similarity between the to-be-translated functions and the intermediate target language code samples; The intermediate target language code samples corresponding to the first cosine similarity of the second preset number of positions are selected as final target target language code samples; The to-be-translated functions are respectively matched with the dependent use examples in the preprocessed dependent use example self-evolution knowledge base to determine the second best matching similarity between the to-be-translated functions and the dependent use examples; The dependent use examples corresponding to the second best matching similarity of the first preset number of positions are selected as intermediate dependent use examples; The to-be-translated functions are respectively matched with the intermediate dependent use examples to determine the second cosine similarity between the to-be-translated functions and the intermediate dependent use examples; Select the intermediate dependency usage example corresponding to the second cosine similarity of the first preset number of bits as the final dependency usage example; Each of the functions to be translated is respectively matched with the pre-processed successful translation function pair in the self-evolution knowledge base to calculate the third best matching similarity between each of the functions to be translated and each of the successful translation function pairs; Select the successful translation function pair corresponding to the third best matching similarity of the first preset number of bits as the intermediate successful translation function pair; Each of the functions to be translated is respectively matched with each of the intermediate successful translation function pairs to calculate the third cosine similarity between each of the functions to be translated and each of the intermediate successful translation function pairs; Select the intermediate successful translation function pair corresponding to the third cosine similarity of the first preset number of bits as the final successful translation function pair; The final target target language code sample, the final dependency usage example, and the final successful translation function pair corresponding to each of the functions to be translated are used as target triple translation knowledge.
7. A warehouse-level code translation apparatus based on a large model, characterized by, It includes: An acquisition module is configured to acquire a translation warehouse, an open source project, and a plurality of historical successful translation function pairs, and use a preset large model to perform architecture translation based on the translation warehouse to generate a warehouse architecture; A construction module is configured to use the preset large model to construct a target language code sample self-evolution knowledge base, a dependency usage example self-evolution knowledge base, and a successful translation function pair self-evolution knowledge base based on a tree parser and based on the open source project, a plurality of functions to be translated in the translation warehouse, and a plurality of historical successful translation function pairs; A determination module is configured to determine target triple translation knowledge corresponding to each of the functions to be translated based on the target language code sample self-evolution knowledge base, the dependency usage example self-evolution knowledge base, the successful translation function pair self-evolution knowledge base, and a plurality of functions to be translated; An output module is configured to use the target triple translation knowledge corresponding to each of the functions to be translated and each of the functions to be translated as input of the preset large model, and output implementation code corresponding to each of the functions to be translated; A generation module is configured to embed the implementation code corresponding to each of the functions to be translated into the warehouse architecture to generate a warehouse-level code translation result.
8. A computer device, comprising: The computer program is executed to implement the warehouse-level code translation method based on the large model.
9. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed to implement the warehouse-level code translation method based on the large model.
10. A computer program product, characterised in that, The computer program product includes a computer program stored on a non-transitory computer readable storage medium, the computer program including program instructions, wherein when the program instructions are executed by a computer, the computer executes the warehouse-level code translation method based on the large model.
Citation Information
Cited By
Warehouse-level code translation agent method and device based on code graph structure
CN122331909A