Code generation method and related equipment
By extracting project-level contexts from multiple dimensions and inputting them into the code generation model, the problem of poor code generation effect of intelligent programming assistants in real application scenarios is solved, and more efficient and accurate code generation is achieved.
Patent Information
- Application Number
- CN202410291045.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-07
- Filing Date
- 2024-03-13
- Publication Date
- 2025-05-09
AI Technical Summary
In code development projects for real application scenarios, intelligent programming assistants have poor code generation effects and low code acceptance rate, making it difficult to meet business needs.
By starting from multiple dimensions such as the project's static structure, developer behavior, code warehouse evolution history, etc., we perceive and extract the in-file context and cross-file context related to the code generation task in the project, and use multi-dimensional project-level context and user input as inputs to the code generation model, improving the ability of the code generation model to utilize cross-file context in the code generation process.
It improves the code generation effect in actual development scenarios, improves the code generation model's perception and utilization ability of project-level context, and enhances the accuracy and efficiency of code generation.
Smart Images

Figure CN119960823A_ABST
Abstract
Description
[0001] This application claims the priority of the Chinese patent application filed with the State Intellectual Property Office on November 7, 2023, with application number 202311473710.1 and invention name “A code generation method and related equipment”, the entire contents of which are incorporated by reference in this application. Technical Field
[0002] The present application relates to the field of artificial intelligence (AI) technology, and in particular to a code generation method, a code development platform, a computing device cluster, a computer-readable storage medium, and a computer program product. Background Art
[0003] As software scale and complexity increase, more and more developers are trying to use code generation technology for software development. Code generation technology is committed to reducing the manual programming workload of developers and improving the efficiency of code development, and has received widespread attention from the academic and industrial circles of software engineering (SE). In recent years, thanks to the progress of artificial intelligence (AI) research in natural language processing and breakthroughs in large language models, the relevant results of large language models have prompted code generation technology to gradually move from the academic research stage to the practical application stage, and various intelligent programming assistant products based on large language models have continued to emerge.
[0004] Intelligent programming assistant products mainly focus on text-to-code conversion (Text2Code) scenarios, which are used to generate code that implements the above requirements based on the requirements described by developers in natural language. Specifically, developers complete the writing of code function annotations during the process of writing code and trigger code generation. Then, the intelligent programming assistant product can use the generative pre-trained transformation model (Generative Pre-trained Transformer, GPT) to generate code snippets that implement the functions described in the annotations based on the annotations and related information provided by the developer. Intelligent programming assistant products can present code snippets to developers in the form of recommendations, allowing developers to decide to accept or reject their recommendations, or accept them and then make further modifications.
[0005] However, in code development projects targeting real application scenarios, the code generation effect of intelligent programming assistants is poor, the code acceptance rate is low, and it is difficult to meet business needs. Summary of the invention
[0006] The present application provides a code generation method, which starts from multiple dimensions such as the static structure of the project, developer behavior, and the evolution history of the code repository, perceives and extracts the in-file context and cross-file context related to the code generation task in the project, and uses the above multi-dimensional project-level context together with user input (such as task description information and input code in natural language form) as the input of the code generation model, improves the ability of the code generation model to utilize cross-file context during the code generation process, and improves the code generation effect in actual development scenarios. The present application also provides a code development platform, a computing device cluster, a computer-readable storage medium, and a computer program product corresponding to the above method.
[0007] In the first aspect, the present application provides a code generation method. The method can be performed by a code development platform. The code development platform can be software, which can be independently running software, or integrated into other software, such as a functional module in an integrated development environment (integrated development environment) or a plug-in integrated into an IDE, or a code editor. The software can be deployed on a computing device cluster, and the computing device cluster executes the program code of the software system, thereby executing the code generation method of the present application. In some possible implementations, the code development platform can also be hardware, such as a computing device cluster that provides code generation capabilities, and when the computing device cluster is running, the code generation method of the present application is executed.
[0008] Specifically, the code development platform receives the input information of the first code file in the project from the user, the input information includes the task description information of the code generation task or at least one of the input codes, and then the code development platform obtains the in-file context and the first cross-file context according to the static structure of the project, obtains the second cross-file context according to the behavioral characteristics of the user in the development project, and obtains the third cross-file context according to the evolutionary coupling degree of at least one second code file and the first code file in the code repository of the project. Among them, the in-file context is the context in the first code file, and the cross-file context is the context in the code files other than the first code file in the project. Then the code development platform can generate prompt information according to the input information, the in-file context, the first cross-file context, the second cross-file context and the third cross-file context. The code development platform inputs the prompt information into the code generation model for reasoning to obtain at least one set of generated code. The code development platform can display at least one set of generated code to the user.
[0009] This method starts from multiple dimensions such as the static structure of the project, developer behavior, and the evolution history of the code repository, and perceives and extracts the in-file context and cross-file context related to the code generation task in the project. The multi-dimensional project-level context and user input (such as task description information and input code in natural language form) are used as the input of the code generation model, improving the code generation model's ability to utilize cross-file context during the code generation process and enhancing the code generation effect in actual development scenarios.
[0010] In some possible implementations, the code development platform may obtain a second cross-file context based on at least one of the metadata of an opened code file, the editing popularity of a code file in a project, or a search record of a user in a development project, wherein the metadata of an opened code file includes at least one of a relative position relationship of a class, a member, a method in the opened code file, or an editor of the opened code file.
[0011] This method perceives the user's focus in the development process (or programming process) by analyzing behavioral characteristics such as opened code files, editing heat, and search history. It can thus extract context that is helpful for code generation from the user behavior dimension and improve the quality of code generation.
[0012] In some possible implementations, the code development platform can obtain the behavior characteristics of users in the development project through the behavior perception interface. This can realize the analysis of user behavior during the development process and provide assistance for subsequent code generation.
[0013] In some possible implementations, the code development platform may also obtain the submission record of the code repository. The code development platform may perform an evolutionary coupling analysis on at least one second code file and the first code file in the code repository of the project based on the submission record of the code repository, and obtain the evolutionary coupling degree of at least one second code file and the first code file. The evolutionary coupling degree represents the degree of evolutionary coupling or the degree of evolutionary correlation, which can be quantified by the probability that the first code file and the second code file are changed and submitted at the same time in all submission histories.
[0014] For code files with high evolutionary coupling, the context within the code file can provide a reference for code generation in the current code file, which helps to improve the quality of code generation in the current code file.
[0015] In some possible implementations, the code development platform can construct a project structure diagram based on the static structure of the project. The static structure includes the hierarchical structure of the project, the hierarchical structure includes the hierarchical relationship of modules, packages, classes or code blocks in the project, and the project structure diagram includes the hierarchical relationship and dependency information. The code development platform can obtain the subgraph corresponding to the code generation task according to the position of the code generation task in the project structure diagram, and obtain the file context and the first cross-file context according to the subgraph corresponding to the code generation task.
[0016] In this method, by combining the static structure of the project to obtain the subgraph corresponding to the code generation task, the scope of determining the intra-file context and cross-file context can be narrowed. Based on the subgraph, the intra-file context and the first cross-file context can be accurately extracted to avoid interference of other contexts on code generation and improve the quality of code generation.
[0017] In some possible implementations, the code development platform can obtain the internal import statements and file context of the project based on the subgraph corresponding to the code generation task, and the file context includes at least one of the library file import statement, the class to which the code generation task belongs, the context of the first code file, or the context of the first code file. The code development platform can obtain the dependent class of the belonging class based on the internal import statement of the project, and obtain the first cross-file context based on the dependent class, and the first cross-file context includes at least one of the member variable name, method signature, constant, and access control keyword in the dependent class.
[0018] This method uses subgraphs to identify the project's internal import statements, library file import statements, the class to which the code generation task belongs, and the context of the first code file. Based on the above internal import statements, the dependent class of the class can be identified, and then the member variable names, method signatures, constants, access control keywords and other contexts of the dependent class can be identified, providing rich reference information for code generation.
[0019] In some possible implementations, the code development platform may abstract at least one of the intra-file context, the first cross-file context, the second cross-file context, or the third cross-file context into a grammatically correct interface declaration, and then generate prompt information through a prompt project based on the input information and the interface declaration.
[0020] This method can achieve the purpose of compressing context information by abstracting the context into an interface declaration. On the other hand, by abstracting the context code into a grammatically correct interface declaration form, the purpose of reusing the programming language grammar knowledge learned by the model in unsupervised pre-training can be achieved.
[0021] In some possible implementations, the code development platform may also sort the cross-file contexts according to at least one of access rights, topological distance, editing heat, semantic similarity, or evolutionary coupling. Accordingly, the code development platform may assemble the prompt information according to the input information, the context within the file, the first cross-file context, the second cross-file context, and the third cross-file context in combination with the sorting results of the cross-file contexts.
[0022] Taking into account the situation where the overall length of the context is too long, by sorting the context, relatively important contexts can be input as prompt information without being truncated, thus ensuring the quality of code generation.
[0023] In some possible implementations, the code development platform may also add start marks to different types of information in the input information, the in-file context, the first cross-file context, the second cross-file context, and the third cross-file context to generate prompt information.
[0024] In this way, it is possible to display information related to the code generation task, improve the ability of the code generation model to utilize cross-file context during the code generation process, and enhance the code generation effect in actual development scenarios.
[0025] In some possible implementations, the code generation model is obtained as follows:
[0026] Obtain training data including cross-file context;
[0027] Based on the training data, the base model is directly pre-trained, multi-stage pre-trained, or fine-tuned to obtain a code generation model.
[0028] Compared with adding cross-file context only in the reasoning stage, this application also supports aligning training tasks with reasoning tasks, thereby focusing on improving the code generation model's ability to perceive and utilize context, and achieving better results and experience in actual use.
[0029] In a second aspect, the present application provides a code development platform. The code development platform includes:
[0030] An interaction module, configured to receive input information of a first code file in a project from a user, the input information including at least one of task description information of a code generation task or input code;
[0031] A context extraction module is used to obtain an intra-file context and a first cross-file context according to a static structure of a project, obtain a second cross-file context according to a behavioral characteristic of a user in developing a project, and obtain a third cross-file context according to an evolutionary coupling degree between at least one second code file and a first code file in a code repository of the project, wherein the intra-file context is a context within the first code file, and the cross-file context is a context within code files other than the first code file in the project;
[0032] A prompt module, used for generating prompt information according to input information, an intra-file context, a first cross-file context, a second cross-file context and a third cross-file context;
[0033] A generation module, used for inputting prompt information into a code generation model for reasoning to obtain at least one set of generated codes;
[0034] The interactive module is further used to display at least one set of generated codes to the user.
[0035] In some possible implementations, the context extraction module is specifically used to:
[0036] A second cross-file context is obtained based on at least one of the metadata of an opened code file, the editing popularity of a code file in a project, or a search record of a user in a development project, wherein the metadata of the opened code file includes at least one of a relative position relationship of a class, a member, a method in the opened code file, or an editor of the opened code file.
[0037] In some possible implementations, the context extraction module is further used to:
[0038] The behavior characteristics of users in the development project are obtained through the behavior perception interface.
[0039] In some possible implementations, the context extraction module is further used to:
[0040] Get the commit record of the code repository;
[0041] According to the submission record of the code repository, an evolutionary coupling analysis is performed on at least one second code file and a first code file in the code repository of the project to obtain an evolutionary coupling degree between the at least one second code file and the first code file.
[0042] In some possible implementations, the context extraction module is specifically used to:
[0043] According to the static structure of the project, build a project structure diagram. The static structure includes the hierarchical structure of the project. The hierarchical structure includes the hierarchical relationship of modules, packages, classes or code blocks in the project. The project structure diagram includes the hierarchical relationship and dependency information.
[0044] According to the position of the code generation task in the project structure diagram, obtain the sub-diagram corresponding to the code generation task;
[0045] According to the subgraph corresponding to the code generation task, the intra-file context and the first cross-file context are obtained.
[0046] In some possible implementations, the context extraction module is specifically used to:
[0047] According to the subgraph corresponding to the code generation task, an internal import statement and an in-file context of the project are obtained, where the in-file context includes at least one of a library file import statement, a class to which the code generation task belongs, a context above the first code file, or a context below the first code file;
[0048] According to the internal import statement of the project, the dependent class of the belonging class is obtained, and the first cross-file context is obtained according to the dependent class. The first cross-file context includes at least one of the member variable name, method signature, constant, and access control keyword in the dependent class.
[0049] In some possible implementations, the prompt module is specifically used to:
[0050] Abstracting at least one of the in-file context, the first cross-file context, the second cross-file context, or the third cross-file context into a grammatically correct interface declaration;
[0051] Generate prompt information through prompt project according to input information and interface declaration.
[0052] In some possible implementations, the prompt module is further used to:
[0053] sorting the cross-file contexts according to at least one of access rights, topological distance, editing heat, semantic similarity, or evolutionary coupling;
[0054] The prompt module is specifically used for:
[0055] According to the input information, the context within the file, the first cross-file context, the second cross-file context and the third cross-file context, the prompt information is obtained by assembling the sorting results of the cross-file context.
[0056] In some possible implementations, the prompt module is specifically used to:
[0057] Start marks are added to different types of information in input information, in-file context, first cross-file context, second cross-file context and third cross-file context to generate prompt information.
[0058] In some possible implementations, the code development platform further includes:
[0059] The training module is used to obtain training data including cross-file contexts, and to perform direct pre-training, multi-stage pre-training or fine-tuning on the base model according to the training data to obtain a code generation model.
[0060] In a third aspect, the present application provides a computing device cluster. The computing device cluster includes at least one computing device, and the at least one computing device includes at least one processor and at least one memory. The at least one processor and the at least one memory communicate with each other. The at least one processor is used to execute instructions stored in the at least one memory, so that the computing device or the computing device cluster executes the code generation method described in the first aspect or any implementation of the first aspect.
[0061] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, wherein the instructions instruct a computing device or a computing device cluster to execute the code generation method described in the above-mentioned first aspect or any one of the implementations of the first aspect.
[0062] In a fifth aspect, the present application provides a computer program product comprising instructions, which, when executed on a computing device or a computing device cluster, enables the computing device or the computing device cluster to execute the code generation method described in the first aspect or any one of the implementations of the first aspect.
[0063] Based on the implementations provided in the above aspects, this application can also be further combined to provide more implementations. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical method of the embodiments of the present application, the drawings required for use in the embodiments are briefly introduced below.
[0065] Figure 1 A schematic diagram of the architecture of a code development platform provided for this application;
[0066] Figure 2 A flowchart of a code generation method provided for this application;
[0067] Figure 3 A schematic diagram of a code editing interface provided for this application;
[0068] Figure 4 A schematic diagram of a flow chart of code generation in the reasoning phase provided for this application;
[0069] Figure 5 A schematic diagram of a code editing interface provided for this application;
[0070] Figure 6 A flowchart of context processing in the reasoning phase provided by this application;
[0071] Figure 7 A schematic diagram of a model architecture based on a transformer decoder provided for this application;
[0072] Figure 8 A schematic diagram of the structure of training data including file-level context provided for this application;
[0073] Fig. 9 A schematic diagram of a structure of training data including cross-file context provided for this application;
[0074] Fig.10 A schematic diagram of the front-end interface of a code generation plug-in provided in this application;
[0075] Fig.11 A schematic diagram of the front-end interface of a code generation plug-in provided in this application;
[0076] Fig.12 A schematic diagram of the comparison of code generation results provided by this application;
[0077] Fig.13 A schematic diagram of triggering code generation through a human-computer interaction interface provided in this application;
[0078] Fig.14 A schematic diagram of constructing prompt information and generating code based on the prompt information provided by the present application;
[0079] Fig.15 A schematic diagram of the structure of a code development platform provided for this application;
[0080] Fig.16 A schematic diagram of the structure of a computing device provided for this application;
[0081] Fig.17 A schematic diagram of the structure of a computing device cluster provided for this application;
[0082] Fig.18 A schematic diagram of the structure of another computing device cluster provided for this application;
[0083] Fig.19 A schematic diagram of the structure of another computing device cluster provided in this application. DETAILED DESCRIPTION
[0084] The terms "first" and "second" in the embodiments of the present application are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features.
[0085] First, some technical terms involved in the embodiments of the present application are introduced.
[0086] Code generation refers to the automatic generation of code based on user input information, such as incomplete code or natural language description, through automated tools or technologies to make the code complete or implement the functions described by the natural language description. Depending on the granularity of the generated code, code generation can include line-level code generation and method-level code (or function-level code) generation.
[0087] A Large Language Model (LLM) is a language model that consists of an artificial neural network with many parameters (usually billions of weights or more) and is trained on a large amount of unlabeled text using self-supervised learning or semi-supervised learning. Large language models can be divided into different types according to the model architecture. Large language models that are widely used in the field of code generation include Generative Pre-trained Transformer (GPT).
[0088] Intelligent programming assistant products based on large language models such as GPT support the generation of code snippets that implement the functions described in the comments based on the comments and related information provided by the developer. User feedback and large-scale survey results during the public beta phase show that the above-mentioned intelligent programming assistant products can effectively reduce the cost of developers frequently switching between actual code writing and searching for knowledge, consulting documents, and finding reusable components, thereby improving the efficiency of software development. However, in code development projects for real application scenarios, the business logic code generation effect is poor, which is specifically manifested in the presence of completely non-existent methods or variables in the generated code (usually referred to as hallucinations in the AI field), or the failure to make good use of the existing cross-file code in the project for code generation, resulting in repeated development and increased code complexity.
[0089] After research, the inventors found that the code generation model in the intelligent programming assistant product is usually obtained based on LLM training such as GPT. The training and reasoning methods of the code generation model are derived from natural language processing (NLP) technology, which does not fully consider the logical structure of the code project and the common calling relationship between the code file content. Specifically:
[0090] During the model training phase, LLM processes the training corpus in a natural language order, which is inconsistent with the logic of actual code development. Moreover, LLM is trained based on file-level context in text form, making it difficult to learn and utilize global information such as the static structure of the project and code call relationships.
[0091] During the model reasoning phase, the context scope that LLM can perceive is limited to the current code file. The generated code is likely to include calls to code that does not exist in the project, but cannot correctly call the code that already exists in the project. If the context scope is expanded and still in text form, the private data in the user code is very easy to leak.
[0092] It can be seen that compared with other software development tools, intelligent programming assistant products based on LLM are still in their infancy. To this end, the industry has proposed a solution to use project-level context as an input prefix to enhance the input of the code generation model in the reasoning stage to improve the code generation effect. The solution first proposes a series of rules to define project-level context, including the current class (Current), parent class (Parent Class), reference class (Import Class), class under the same package (Sibling), class with the same name as the current class in the project (Similar Name), child class (Child Class), reference class of parent class (Import of Parent Class), reference class in class of the same package (Import of Sibling), etc. At the same time, in order to solve the problem of too long context, the solution also specifically classifies the context information in the code to select a context of appropriate length.
[0093] Based on the above rules, the solution designs a rule selector, such as a multi-label classifier, which selects contexts of different levels and contents through a combination of different rules and inputs them into the intelligent programming assistant product. The rule selector is optimized based on the comparison between the generated code of the intelligent programming assistant product and the real code label. In the inference stage, the optimized rule selector extracts contexts of different levels for different code scenarios, and splices them with the code near the generation point in proportion as the input of code generation tools such as intelligent programming assistant products.
[0094] However, the dimensions considered in the above solution are relatively single, mainly considering the static dimensions of the project, which makes it difficult to accurately locate the context scope required for the code generation task. In addition, considering the flexibility of programming languages and the unpredictability of project codes, defining the context scope through predefined rules cannot cover various situations in actual development scenarios, which can easily lead to omissions or redundancies of context. Based on this, the above solution has limited improvement or gain in the effect of code generation, and the code acceptance rate of the generated code is still difficult to meet business needs.
[0095] In view of this, the present application provides a code generation method. This method aims to alleviate the limitations of current intelligent programming assistants in code generation from the perspective of context awareness, and proposes an intelligent perception and dynamic construction technology for project-level context (including in-file context and cross-file context) in object-oriented code generation. Through multi-dimensional cross-file context perception and extraction technology, a context that helps code generation is constructed. Furthermore, the project-level context awareness and utilization capabilities of the code generation model are improved through the joint optimization of the model training and reasoning stages, thereby alleviating the problem of poor code generation effect caused by the above-mentioned deficiencies.
[0096] Wherein, the code generation method can be executed by a code development platform. The code development platform can be software, which can be independently running software, or integrated into other software, such as a functional module in an integrated development environment (integrated development environment) or a plug-in integrated into an IDE, or a code editor. The software can be deployed in a computing device cluster, and the computing device cluster executes the program code of the software system, thereby executing the code generation method of the present application. Wherein, the software can be provided to the user in the form of a software package, for example, the software can be a new function of a client code editor or IDE, provided to the user with the version update iteration, or a new feature of a code generation plug-in based on a pre-trained language model, provided to the user with the version update iteration. The user can run the software package in a local data center or a private cloud to deploy the software. Alternatively, the software can be provided to the user in the form of a cloud service, such as software as a service (SaaS). For example, the software can appear as an auxiliary coding function of a cloud code editor or development environment, and expose the functional interface to the outside in the form of a cloud service, such as an application programming interface (API), and other tools can use the above-mentioned code generation function or capability by calling the interface. In some possible implementations, the code development platform may also be hardware, such as a computing device cluster that provides code generation capabilities. When the computing device cluster is running, it executes the code generation method of the present application.
[0097] Specifically, the code development platform can receive input information of the first code file in the project from the user, the input information includes task description information of the code generation task or at least one of the input codes, and then the code development platform obtains the in-file context and the first cross-file context according to the static structure of the project, obtains the second cross-file context according to the behavioral characteristics of the user in the development project, and obtains the third cross-file context according to the evolutionary coupling degree of at least one second code file and the first code file in the code repository of the project. The code development platform generates prompt information based on the input information, the in-file context, the first cross-file context, the second cross-file context, and the third cross-file context, inputs the prompt information into the code generation model for reasoning, obtains at least one set of generated code, and displays at least one set of generated code to the user.
[0098] This method starts from multiple dimensions such as the static structure of the project, developer behavior, and the evolution history of the code repository, and perceives and extracts the in-file context and cross-file context related to the code generation task in the project. The multi-dimensional project-level context and user input (such as task description information and input code in natural language form) are used as the input of the code generation model, improving the code generation model's ability to utilize cross-file context during the code generation process and enhancing the code generation effect in actual development scenarios.
[0099] Moreover, this method introduces the above-mentioned multi-dimensional cross-file context into the training expectations, aligns the task form and data distribution in the training phase with the task form and data distribution in the reasoning phase, so as to enhance the code generation model's ability to perceive and utilize cross-file context, and further improve the code generation capability of the code generation model.
[0100] In order to make the technical solution of the present application clearer and easier to understand, the system architecture of the code development platform of the present application is first introduced below.
[0101] See also Figure 1 The schematic diagram of the architecture of a code development platform shown in the figure, the code development platform 10 includes an inference platform 100, and the inference platform 100 is used to perform inference based on a trained code generation model to generate code. Further, the code development platform 10 can also include a training platform 200, and the training platform 200 is used to train the base model based on a training corpus including a multi-dimensional cross-file context to obtain a code generation model. The inference process is performed in the inference phase, and the training process is performed in the training phase.
[0102] The reasoning platform 100 is used to obtain multi-dimensional cross-file context from the project through multi-dimensional project-level cross-file context awareness technology for code generation tasks, and to extract in-file context through code generation key information extraction technology for code generation tasks. Among them, the in-file context is also called the current file context. Specifically, the reasoning platform 100 can receive input information of the first code file in the project from the user, and the input information includes at least one of the task description information or input code of the code generation task, and then obtain the in-file context and the first cross-file context according to the static structure of the project, obtain the second cross-file context according to the behavioral characteristics of the user in the development project, and obtain the third cross-file context according to the evolutionary coupling degree of at least one second code file and the first code file in the code repository of the project. The multi-dimensional cross-file context may include a combination of the above-mentioned first cross-file context, the second cross-file context and the third cross-file context. The reasoning platform 100 is used to generate prompt information prompt according to the above input information (such as the task description information and input code of the code generation task), the context within the file and the multi-dimensional cross-file context (such as the first cross-file context, the second cross-file context and the third cross-file context), input the prompt information into the code generation model for reasoning, and obtain at least one set of generated code. The generated code can be, for example, a generated line-level code snippet, or a method-level code snippet. The reasoning platform 100 is used to display at least one set of generated code to the user.
[0103] In some possible implementations, the reasoning platform 100 is also used to sort the multi-dimensional cross-file context by context priority to obtain the cross-file context priority. Accordingly, the reasoning platform 100 is used to combine the cross-file context priority to construct the input information (such as the task description information and input code of the code generation task), the context in the file, and the multi-dimensional cross-file context as prompt information, for example, a context-enhanced prompt is constructed through a context-aware prompt template. Among them, when constructing the prompt, the location of the code generation task is also combined, also known as the task point. It should be noted that the reasoning platform 100 takes into account the situation where the overall length of the context is too long, and sorts the context so that the relatively important context can be input into the prompt without being truncated. In some cases, such as when the overall length of the context is within a controllable range, the reasoning platform 100 may not perform the above sorting process.
[0104] The code generation model can be obtained by training on the training platform 200. The training platform 200 is used to obtain training data including cross-file context, for example, training data including the above-mentioned multi-dimensional cross-file context, and then directly pre-train, multi-stage pre-train or fine-tune the base model according to the training data to obtain the code generation model. Figure 1This is an example of how the training platform 200 fine-tunes instructions on the base model.
[0105] In specific implementation, the training platform 200 is used to obtain code files of a large number of projects from a massive software code repository (also referred to as a code warehouse for short), and pre-process the code files of the projects, such as extraction, screening or deduplication, to form a code warehouse data set. The training platform 200 is used to perform multi-dimensional project-level cross-file context perception on the data in the code warehouse data set, thereby extracting the multi-dimensional project-level cross-file context for the function (or method), and constructing training data based on the multi-dimensional project-level cross-file context. Among them, the training platform 200 can segment, encode, and split the function's annotations, declarations, and the function's in-file context, multi-dimensional project-level cross-file context, in-file context, and function body, and then form instructions (for example, prompts in the form of instructions). The training platform 200 is used to fine-tune the instructions recorded as a model through the constructed instructions, so as to obtain a code generation model.
[0106] based on Figure 1 The code development platform 10 shown in the figure, the present application provides a code generation method. The code generation method of the present application is introduced below in conjunction with the accompanying drawings.
[0107] See also Figure 2 A flowchart of a code generation method is shown, the method comprising the following steps:
[0108] S202: The code development platform 10 receives input information of a first code file in a project from a user.
[0109] Project is the abbreviation of Software Project. It is an engineering file created by developers for the software to be developed. A project can include multiple code files, such as code files used to implement different functions or features of the software. The code files in a project can be developed independently by one user or collaboratively by multiple users. When developing, users can use the code generation capability to automatically generate code.
[0110] Taking the first code file as an example, the user's input information in the first code file may include at least one of the task description information of the code generation task or the input code. The task description information may be a task description in natural language. For example, when the code generation task is used to generate the code of the target method or target function, the task description information may be a requirement description for creating the target method or target function, and the task description information may generally be input in the form of comments. The input code may include a function declaration. The function declaration may include a function name and a parameter name. Accordingly, the code generation task may be a code snippet that generates a function body.
[0111] Specifically, the code development platform 10 may present a code editing interface to the user, wherein the code editing interface may be a user interface of a code editor, and the code editing interface may be a graphical user interface (GUI) or a command user interface (CUI).
[0112] For ease of description, this application uses the code editing interface as a GUI example. Figure 3 As shown, the code editing interface 300 may include an editing window 302 for the first code file, and the user may enter task description information 304 in the editing window 302, and the task description information 304 is used to describe the code to be generated. In this example, the task description information 304 may be "Create a http sever instance and start it", and input method declaration 306, such as "public void init(){}". The user may trigger the code generation control 308 of the code editing interface 300, thereby triggering the code generation operation. Accordingly, the code development platform 10 may receive the task description information 304 and method declaration 306 (input code) input by the user, and then generate code based on the task description information 304 and method declaration 306.
[0113] Need to explain, Figure 3 It is only to illustrate the example that the user's input information includes task description information and input code. In other possible implementations of the embodiment of the present application, the user's input information may also include task description information, or include input code. For example, when editing the code of some methods, the user can directly write the method without adding task description information in the form of comments. For another example, when editing the code of some methods, the user can add task description information in the form of comments without entering the code. Among them, when entering the task description information 304 in the form of comments, the user can first enter the keyword of the comment, such as the " / * / " character, the @ character or the # character, and then enter the task description information, so that the execution of the above-mentioned task description information 304 can be avoided when executing the code file.
[0114] also, Figure 3 The example of the user triggering the code generation control 308 to trigger the code generation operation is used for illustration. In actual application, the code development platform 10 also supports triggering the code generation operation by other means. For example, the code development platform 10 also supports a shortcut key or menu (such as a right-click menu) to trigger the code generation operation. Alternatively, see Figure 4When editing the code in the first code file, the user can also call the language model, such as LLM, to input the task description information and / or input the code of the code generation task in the form of instructions / dialogue. Figure 4 In the example, the task description information can be "Based on the current project, add a course selection method for the Student class", and the LLM interactive interface can also include a selection control. When the selection control is selected, LLM can refer to the current project to generate code. This application triggers code generation or inputs code generation task description information and inputs code in any way, which can be selected according to actual conditions.
[0115] S204 . The code development platform 10 obtains the intra-file context and the first cross-file context according to the static structure of the project.
[0116] The static structure of a project may include a hierarchical structure of the project. The hierarchical structure includes the hierarchical relationship of modules, packages, classes, or code blocks in the project. The hierarchical structure can be represented by a tree structure (referred to as a tree). Specifically, the tree takes the project as the root node and the method as the leaf node, and is obtained by expanding the modules, packages, classes, and code blocks under the project layer by layer.
[0117] The code development platform 10 can construct a project structure diagram according to the static structure of the project, and the project structure diagram includes hierarchical relationships and dependency information. Among them, the dependency information can be dependency information between classes. Then the code development platform 10 can obtain the subgraph corresponding to the code generation task according to the position of the code generation task in the project structure diagram, for example, the position of the class to which the code generation task belongs in the project structure diagram. The subgraph can be a local graph related to the project structure diagram and the code generation task.
[0118] For a folder corresponding to a project / code warehouse, the code development platform 10 can construct a graph-formatted data structure through preprocessing, which is called a project structure diagram. The project structure diagram can be generated based on a project hierarchy tree, where the project hierarchy tree can be a tree that models the hierarchical structure of a project. Specifically, the code development platform 10 can convert the project hierarchy tree into a project structure diagram that models the hierarchical and reference relationships between classes within a project by analyzing the dependencies between classes and adding edges based on the project hierarchy tree.
[0119] For a code generation task, the code development platform 10 can determine the scope of code that the method has the authority to access or call based on the location of the code generation task in the project structure diagram, the access control permissions of the class (such as public / private / default, etc.), and the hierarchy and reference relationships involved in the class, which is reflected as a sub-diagram of the project structure diagram.
[0120] The code development platform 10 can obtain the in-file context and the first cross-file context according to the subgraph corresponding to the code generation task. Among them, the in-file context refers to the context within the first code file. The context before the input information (or target method) is called the preceding context (also called the foreword), and the context after the target method is called the following context (also called the following context). The preceding context generally includes package statements and import statements, the signature of the belonging class (the class where the target method is located), partial declarations of the belonging class (such as member variables, constructors, partial methods), etc.; and the following context generally includes other method declarations of the belonging class. The first cross-file context refers to the context from other code files in the code warehouse obtained from the static structure dimension. In specific implementation, the code development platform 10 can obtain code snippets from other code files in the project used in the first code file based on the edges in the project structure diagram and the import statements of the first code file (current file). The first cross-file context may include code snippets from other code files in the project used in the first code file obtained according to the above-mentioned import statements and the project structure diagram. As Figure 4 As shown, the first cross-file context may include at least one of the cross-file classes, class members, or class methods in the imported project, wherein the class method may include a constructor or other methods. The first cross-file context may also include code in the same package, parent class members, and parent class methods.
[0121] Among them, the code development platform 10 can obtain the first cross-file context such as cross-file classes and methods in the project that the current file depends on through analysis tools, or access the intermediate data obtained by the IDE analyzing the project through the API provided by the IDE, such as the Program Structure Interface (PSI), to obtain the first cross-file context.
[0122] S206 : The code development platform 10 obtains a second cross-file context according to the behavior characteristics of the user in the development project.
[0123] Behavior features refer to the characteristic representation of the opening code file, editing code file or searching behavior triggered when developing a project. Among them, the behavior features may include at least one of the metadata of the opened code file, the editing heat of the code file in the project, or the search record of the user in the development project. The metadata of the opened code file may include at least one of the relative position relationship of the class, member, method or editor of the opened code file in the opened code file. The relative position relationship of the editor of the opened code file may include the group where the editor is located and the relative position of each editor. The editing heat of the code file can be characterized by the x files that have been modified recently, the modification frequency or the modification amount. The search record can be the search record of the user in the search box of the code development platform 10.
[0124] The code development platform 10 can perceive the user's focus in the development process (or programming process) based on the user's behavioral characteristics in the development project (such as real-time behavioral characteristics). The focus is the code file in the project that the user focuses on (for example, the code file in the project other than the first code file). The code development platform 10 can extract cross-file code from the code file in the project that the user focuses on, thereby obtaining a second cross-file context.
[0125] The code development platform 10 may provide a behavior-aware interface. For example, code development platforms such as IDEs provide APIs for perceiving user behavior. The code development platform 10 may obtain the user's behavior characteristics in the development project through the behavior-aware interface, thereby obtaining another dimension of cross-file context.
[0126] S208. The code development platform 10 obtains a third cross-file context according to the evolution coupling degree between at least one second code file and the first code file in the code repository of the project.
[0127] Evolutionary coupling refers to the tendency of code entity pairs (such as two code entities) to change together (covariate) in the software revision history. The evolutionary coupling relationship between code entities X and Y can be represented by X→Y. If code entity X changes, entity Y also tends to change. Among them, X and Y can be granularities such as source files, classes, modules, methods, variables, etc. The degree of evolutionary coupling indicates the degree of evolutionary coupling or evolutionary correlation.
[0128] Specifically, the code development platform 10 can analyze the evolutionary coupling between at least one second code file and the first code file (the file where the code generation task is located) in the project based on the evolutionary history of the code warehouse using an evolutionary correlation analysis method, and extract the third cross-file context from the file with high coupling. Among them, the second code file can be a code file other than the first code file in the project, and the high coupling can mean that the coupling is higher than a set value, or it is ranked top k when the coupling is sorted from high to low. In some examples, the third cross-file context can be a code snippet in a code file that has been modified at the same time in the submission history.
[0129] Among them, the code development platform 10 can obtain the submission record of the code repository, such as the git submission record, and then the code development platform 10 can perform an evolutionary coupling analysis on at least one second code file and the first code file in the code repository of the project according to the submission record of the code repository to obtain the evolutionary coupling degree of at least one second code file and the first code file.
[0130] Among them, the above S204, S206, and S208 can be executed in parallel, or executed successively in a set order, and the embodiment of the present application does not limit this.
[0131] S210 , the code development platform 10 generates prompt information according to the input information, the context within the file, the first cross-file context, the second cross-file context, and the third cross-file context.
[0132] Specifically, the code development platform 10 can fill the input information, the file context, the first cross-file context, the second cross-file context, and the third cross-file context into the corresponding positions of the prompt template, thereby generating a prompt. Figure 4 As shown, the prompt template may include instruction filling indication information and context filling indication information. In the reasoning stage, the instruction filling indication information may indicate filling of the user's input information, such as task description information and input code, and the context filling indication information may indicate filling of the intra-file context, the first cross-file context, the second cross-file context and the third cross-file context. Furthermore, the context filling indication information may also be divided into intra-file context filling indication information and cross-file context filling indication information. Among them, the intra-file context filling indication information may include the following filling indication information and the preceding filling indication information.
[0133] In some possible implementations, the prompt template may also include a system prompt, such as Figure 4 The system prompt in can be used to prompt that the task is a programming task based on the project-level context. In order to achieve task alignment between the reasoning phase and the training phase, the training phase and the reasoning phase can share the prompt template. Based on this, the prompt template can also include generated code filling indication information, which is used to indicate the filling position of the expected generated code.
[0134] Since the input length of the code generation model (usually a language model, such as LLM) has certain restrictions (called window size, usually 1024 to 8192), when the input length is limited, the length of the project-level context may exceed the input length limit of the code generation model. Based on this, the present application can also support sorting cross-file contexts to reduce the probability of important contexts being deleted by truncation strategies when the input is too long. Among them, the code development platform 10 can sort cross-file contexts according to at least one of access rights, topological distance, editing heat, semantic similarity, or evolutionary coupling (or evolutionary relevance, evolutionary relevance). Among them, the topological distance is the probability relative to the physical distance, the physical distance can be the distance between the target method and the context in the code file, and the topological distance can be the distance between the target method and the context in the project structure diagram, which can be characterized by the number of hops (or the number of nodes in the interval). Accordingly, the code development platform 10 can be assembled to obtain prompt information based on the input information, the upper and lower parts of the file, the first cross-file context, the second cross-file context, and the third cross-file context, combined with the sorting results of the cross-file context. The code development platform 10 quantifies the importance of different cross-file contexts to the code generation task by comprehensively using indicators such as access rights, topological distance, editing heat, semantic similarity, and evolutionary correlation, and sorts the cross-file contexts according to their importance (importance) to avoid truncation of important cross-file contexts due to excessively long input.
[0135] Furthermore, if the original code is used directly as the context, redundant information can be introduced, which limits the amount of context information available to the code generation model, thereby reducing the efficiency of context utilization. To this end, the code development platform 10 also supports direct context abstraction, such as abstracting at least one of the context within the file, the first cross-file context, the second cross-file context, or the third cross-file context into a grammatically compliant interface declaration. Among them, the code development platform 10 can retain the hierarchy and identifier (Identifier, ID) by removing at least one of the comments, variable assignments, and method bodies, and the retained ID may include method names and parameter names. Accordingly, the code development platform 10 can generate prompt information through a prompt project based on input information and interface declarations. Among them, the code development platform 10 can filter out the context (or interface declaration) related to the current method (such as the target method) in the abstracted context (such as an interface declaration) through identifier similarity, such as the part that is most likely to be related to the current method, thereby realizing context intelligent recommendation and filtering. Furthermore, the code development platform 10 can adjust the filtered context or interface declaration to the position before the user input (such as task description information), and organize the position of each context in the Prompt in reverse order of context importance, thereby reducing the probability of important context being truncated and increasing the weight of important context for code generation (such as method body generation of the target method).
[0136] Among them, the code development platform 10 can achieve the purpose of compressing context information by abstracting the context into an interface declaration. On the other hand, the code development platform 10 can achieve the purpose of reusing the programming language grammar knowledge learned by the model in unsupervised pre-training by abstracting the context code into a grammatical interface declaration form.
[0137] In some possible implementations, the code development platform 10 can add a start mark to different types of information in the input information, the context within the file, the first cross-file context, the second cross-file context, and the third cross-file context to generate prompt information. In this way, it is possible to display information related to the code generation task, improve the ability of the code generation model to use the cross-file context during the code generation process, and enhance the code generation effect in the actual development scenario.
[0138] S212: The code development platform 10 inputs the prompt information into the code generation model for reasoning to obtain at least one set of generated codes.
[0139] The code generation model is a language model used to generate code based on user input information. Considering the code generation effect, the language model can be a causal language model based on the Transformer decoder (Decoder) architecture, such as the LLM of the GPT architecture. This type of model can accept a code sequence as input, autoregressively predict the next word (token) in the code as output, and use the output as the next input (i.e., perform the Next Token Prediction task).
[0140] The code development platform 10 inputs the prompt information into the code generation model, and the code generation model can combine the multi-dimensional project-level context in the prompt information to infer the subsequent code of the input code, thereby obtaining at least one set of generated code.
[0141] S214. The code development platform 10 displays at least one set of generated codes to the user.
[0142] Specifically, the code development platform 10 can display at least one set of generated codes to the user in the code editing interface of the first code file. When the code development platform 10 infers multiple sets of generated codes for the code generation task, the code development platform 10 can display the generated codes in combination with probability.
[0143] In some possible implementations, the code development platform 10 can sort the generated codes according to probability, and the code development platform 10 can display the generated codes with high probability. When the user triggers the operation of viewing the next set of generated codes, the code development platform 10 displays the generated codes whose probability sorting results are after the current generated codes.
[0144] In some other possible implementations, the code development platform 10 may also display multiple groups of generated codes at one time. Specifically, the code development platform 10 may sort the generated codes according to probability, and display the top n generated codes at one time according to the sorting result.
[0145] Furthermore, the code development platform 10 can also receive user feedback on the generated code, wherein the feedback can include acceptance, rejection or revision of the generated code. When the user determines that the generated code is usable, the user can accept the second code snippet; when the user determines that the generated code is not usable, the user can reject the second code snippet; when the user determines that the generated code is partially usable, the user can revise the generated code.
[0146] like Figure 5As shown, the code development platform 10 can present a code editing interface 500 to the user, and the code editing interface 500 includes a generated code 502, and a feedback control corresponding to the generated code 502, wherein the feedback control may include an accept control 504, a reject control 506, and a revision control 508. The user can trigger different types of feedback controls to perform different types of feedback such as accepting, rejecting, or revising the generated code 502. Further, the code editing interface 500 can also include a coding process analysis 503 for the generated code 502, so that the user can determine whether the generated code 502 meets the requirements according to the coding process analysis 503, and then decide the type of feedback for the generated code 502.
[0147] In some possible implementations, when the feedback is rejection or revision, the code development platform 10 can also update the code generation model according to the user's feedback on the generated code 502. For example, the code development platform 10 can construct the revised code and related input information and context as training data, and update it to the training data set for subsequent updating of the code generation model. In this way, the accuracy of the code generation model reasoning can be improved, and the output precision can be improved.
[0148] Compared with file-level context awareness, the code generation method of the present application takes into account the diversity of project-level context scope and types, expands the context scope to the entire project, and is closer to the background knowledge range required by human developers in actual programming. Moreover, by extracting context from multiple dimensions for the entire project, it can accurately perceive and obtain project-level context that is helpful for the current code generation task for code generation, thereby improving the code generation effect of the code generation model in actual development scenarios, especially improving the generation ability of object-oriented programs and codes that rely on custom types / methods, and significantly improving the generation effect of object-oriented code.
[0149] Furthermore, compared to uploading the context in plain text form to the server and inputting it into the code generation model, the present application supports preprocessing the context locally on the user side, such as abstracting it into an interface declaration. On the one hand, it compresses the context length and allows a wider range of context to be included with the same input length. On the other hand, it also reduces the risk of privacy data leakage in user code to a certain extent.
[0150] Figure 2 The embodiment of the invention introduces the code generation method. The following introduces the multi-dimensional project-level context perception and extraction, context importance quantification and intelligent recommendation, and context reorganization process in the code generation method in combination with an example.
[0151] See also Figure 6 The flowchart of context processing in the reasoning phase shown in the figure may specifically include the following phases:
[0152] Phase 1: The code development platform 10 constructs a project structure diagram and determines a sub-diagram from the project structure diagram.
[0153] Specifically, for a project, the code development platform 10 (such as an IDE) usually provides a more powerful and accurate project-level analysis tool, which can analyze the opened project (such as the current project) as a whole, and the analysis results can be used through the plug-in API as a source of context for the reasoning stage. Among them, the analysis results may include the hierarchy and dependency information (such as reference relationships) between various types within the project. The code development platform 10 can use the above analysis results as a project index and cache them.
[0154] The code development platform 10 can obtain a project hierarchy tree (e.g. Figure 6 In the project structure in the code generation task, the code development platform 10 adds edges to the project hierarchy tree based on the project index or cache, thereby converting the project hierarchy tree into a project structure diagram that models the hierarchy and reference relationships between the various classes in the project. For code generation tasks, the code development platform 10 can determine the code scope that the method has the right to access or call based on the location of the code generation task (e.g., the target method requested to be generated) in the project structure diagram, the access control permissions of the class (e.g., public / private / default, etc.), and the hierarchy and reference relationships involved in the class. Figure 6 As shown in Figure 1, the code scope can be represented by a sub-graph of the project structure graph. Based on this, the sub-graph can also be called a context scope graph.
[0155] Compared with using a syntax tree parser, this application can effectively use the project index or cache provided by the IDE, which can provide more accurate and complete analysis results on the one hand, thereby improving the accuracy of code generation. On the other hand, it will bring additional analysis costs (considering that code generation is usually an activity with high real-time requirements, and the model itself requires time reasoning, this method uses project indexes, does not require additional analysis, shortens response time, and avoids the cost increase caused by additional analysis.
[0156] At this stage, the code development platform 10 can also obtain code submission records for subsequent evolutionary coupling analysis.
[0157] Phase 2: The code development platform 10 performs code partitioning within the file.
[0158] The code development platform 10 divides the source code file (such as the first code file) where the code generation task is located, and the part before the target method is called the context, and the part after the target method is called the context. The context generally includes package and import statements, the signature of the class, and some declarations of the class (such as member variables, constructors, and some methods), etc.; and the context generally includes other method declarations of the class.
[0159] The import statements can be further divided into the internal import statements of the project and the library file import statements. The internal import statements of the project are used to import other code files in the project. The library file import statements are used to import standard libraries or third-party libraries. Since the internal import statements of the project can be used to determine cross-file contexts, the code development platform 10 can partition or divide the internal import statements and library file import statements of the project. Figure 6 As shown, the code development platform 10 can partition the context, input information, and context in the first code file, wherein the context can also be divided into a partition of package statements and internal import statements of the project, a partition of library file import statements, and a partition of partial declarations of the class. Different partitions can be distinguished by different styles, such as different fill colors.
[0160] Phase 3: The code development platform 10 can obtain multi-dimensional project-level context through multi-dimensional project-level context perception and extraction.
[0161] Specifically, the code development platform 10 can obtain the code content of other code files in the project used in the first file based on the edges in the project structure diagram and the import statements in the first code file, such as the internal import statements of the project, thereby obtaining the first cross-file context.
[0162] Among them, the code development platform 10 can obtain the internal import statement and file context of the project according to the subgraph corresponding to the code generation task, and the file context includes at least one of the library file import statement, the class to which the code generation task belongs, the previous context of the first code file, or the following context of the first code file. Then the code development platform 10 obtains the dependent class of the belonging class according to the internal import statement of the project, and obtains the first cross-file context according to the dependent class. The first cross-file context may include at least one of the member variable name, method signature, constant, and access control keyword in the dependent class.
[0163] The code development platform 10 can also rely on the IDE API to perceive the code files in the project that the developer focuses on and extract the second cross-file context. In addition, the code development platform 10 can also obtain the code files with high evolutionary coupling with the file where the current code generation task is located through evolution correlation analysis and extract the third cross-file context.
[0164] Phase 4: The code development platform 10 performs context screening and sorting.
[0165] The project-level context may exceed the context limit of the language model. Therefore, the code development platform 10 can also filter out the part of the project-level context that is most likely to be related to the current method based on the description, method name, return type, parameter type and other information of the method to be generated by the similarity of identifiers. The code development platform 10 can sort the filtered contexts (for example, cross-file contexts with high relevance). Among them, the code development platform 10 can also uniformly adjust the filtered contexts to the position before the input information to reduce the probability of being truncated when entering the code generation model, and increase the weight of the filtered context on the generation of the method body by reducing the distance.
[0166] Phase 5: The code development platform 10 performs context reorganization.
[0167] The code development platform 10 abstracts the cross-file context of the file, for example, removing comments, variable assignments, method bodies and other information from the cross-file code, retaining only the hierarchy and identifier information, so as to achieve the purpose of compressing the context information. In addition, the code development platform 10 abstracts the context code into a grammatical interface declaration form, so as to achieve the purpose of reusing the programming language grammar knowledge learned by the model in unsupervised pre-training. Then, the code development platform 10 can organize the position of each context in the prompt in reverse order of context importance.
[0168] Compared to adding cross-file context only in the reasoning stage, this solution also supports aligning training tasks with reasoning tasks, thereby focusing on improving the code generation model's ability to perceive and utilize context, and achieving better results and experience in actual use. Among them, the code development platform 10 can obtain training data including cross-file context, for example, training data including multi-dimensional cross-file context, and perform direct pre-training, multi-stage pre-training or fine-tuning on the base model according to the training data to obtain a code generation model. The model training process in the training stage is described below in conjunction with the embodiments.
[0169] During the training phase, this solution is positioned as a common training data processing solution and format among models. Therefore, it can be used to pre-train code generation models from random initialization, and can also be used to perform targeted tuning training on existing code generation models.
[0170] The core of this stage is multi-dimensional project-level context awareness. Different from the typical data processing method that takes source files as the basic unit, the solution of this application takes projects as the basic unit, and specifically includes the following steps:
[0171] Step 1: Analyze the project structure of the entire project and represent the project structure as a tree structure with functions / methods as leaf nodes;
[0172] Step 2: Filter out leaf nodes that meet the criteria.
[0173] Specifically, from the project's methods, filter out complete methods that meet the set conditions. The set conditions may include but are not limited to the method body not being empty and not being a special method. Special methods can be configured by the user, for example, special methods may include get / set / constructor / toString / hashCode and other methods.
[0174] Step 3: For each method, determine the context scope and extract information from the context and perform desensitization processing.
[0175] Step 4: Take the project-level context, method annotation (if any), and method signature of each leaf node as input, and the method body code as the expected output to form a training data.
[0176] Step 5: After the above processing, the methods in multiple software code repositories together constitute a training data set, which is used for model pre-training or tuning of existing models.
[0177] The following describes possible implementation solutions for this application in the training phase from the aspects of model architecture, training method, data processing, etc.
[0178] Considering that the causal language model based on the Transformer decoder architecture has a good effect in the field of code generation, this solution uses the GPT model as the base model to perform model training and obtain the code generation model. Figure 7 A schematic diagram of a model architecture based on a transformer decoder is shown. The model may include a transformer decoder and a token embedding layer and a position embedding layer. The decoder may include a feedforward neural network (FNN) and a masked multi-head attention layer. The training data may be encoded by the token embedding layer and the position embedding layer, and then input into the transformer decoder for decoding.
[0179] In order to adapt to the input format requirements of GPT, it is usually necessary to process the context information into a sequence form, while taking into account the characteristics of the language model and trying to conform to the grammatical rules of the programming language itself.
[0180] In the training phase, the current training method generally uses a single source code file as a unit and generates training data / samples through a sliding window algorithm. Figure 8 A structure of training data including file-level context is shown, and the above data only includes context information within a certain range at the file level. In order to specifically improve the model's ability to perceive and utilize project-level context information, this solution processes the training data through the technology introduced in the previous article, expands the context range and compresses the context content. After the context extraction is completed, different types of information are spliced in the order of project-level context, file-level context, class-level context, method comments, and method code snippets. Furthermore, the code generation platform 10 can use special tokens (such as <--xl_start-->, <comment> 、 <java>etc.) to mark the starting position of different information, so that different loss update strategies can be formulated for different types of information during training. The format of a processed training data can be as follows Fig. 9 shown.
[0181] The process of converting training data into a computable tensor format is similar to other models. Specifically, a specially constructed tokenizer and vocab are used to convert each word (Token) in the training data into its corresponding index in the vocab, forming a sequence of the corresponding indexes of each word in the data in the vocab, and then the vector representation of the corresponding word in the word embedding model is obtained through the index. The vector representations of multiple samples are combined into a word vector matrix and sent to the base model for training.
[0182] Code generation models generally use the pretrain-finetune paradigm for model training and optimization. In the pretraining stage, the model needs to learn the grammar and patterns of the language through unsupervised training on a large amount of corpus; in the fine-tuning stage, it is necessary to process the data in a targeted manner according to the downstream task objectives and perform targeted optimization through supervised learning.
[0183] Depending on the optimization objectives, this solution can be applied to different stages:
[0184] Direct pretraining: Start training directly from a randomly initialized model.
[0185] Multi-stage pretraining: Based on a pretrained model, replace the data, modify the hyperparameters and continue training.
[0186] Fine-tuning: Based on a pre-trained model, fine-tune through training methods such as instruction tuning and prompt tuning.
[0187] This solution trains the model by predicting the next word based on the existing words (Next Token Prediction), but only calculates the target code part to be predicted (for example Fig. 9 middle <java>The loss value of this part is used to update the model weights, thereby specifically optimizing the model's prediction ability for the code implementation part under the premise of known project-level context information and the current method function description.
[0188] Next, the solution of this application is introduced from the perspective of front-end interface and human-computer interaction.
[0189] Similar to similar code generation tools and intelligent programming assistant products, the main implementation form of this application is as an extension function or plug-in of a code editor or IDE, so the front-end interface is embodied as a code generation, completion or auxiliary coding tool embedded in the IDE. An example of generating method-level code for an IDE for the Java language is shown.
[0190] Fig.10 A schematic diagram of the front-end interface of a code generation plug-in is shown. After the user inputs comments and method declarations to trigger code generation, the code generation plug-in can execute the aforementioned code generation method to obtain generated code. The generated code can be the method body code. The front-end interface can also carry acceptance controls, view the next set of generated code controls, and recommend more generated code controls. The user can trigger the corresponding control to achieve the corresponding function.
[0191] It should be noted that most code generation tools can provide Fig.10 The experience shown in the figure is different from other tools in that this solution provides the following options, allowing users to explicitly select and control the cross-file context involved in code generation. Fig.11 As shown, the front-end interface can support the user to display and select the cross-file context for code generation in the form of a dialogue.
[0192] How to make intelligent programming assistants cooperate with existing code development and management tools, how to make the human-computer interaction logic of programming assistants more in line with developers' development habits, and how to present code generation technology and the product form of programming assistants still need to be continuously explored in practice.
[0193] Next, an example of human-computer interaction engineering is given.
[0194] like Fig.12 As shown in the figure, current similar tools generally only send the context within a certain range near the code generation point in the current file as a text request to the backend model for reasoning. For example, the trigger position in the figure is after the signature of the init() method (that is, the cursor is at the beginning of line 19 at this time). If the backend model input window is 1024, the plug-in will input a maximum of 1024 tokens before and after the cursor position into the model, including the method declaration, comments, member variable declarations above init(), class declaration statements, import statements, etc. of init(). Some models (such as Meta's InCoder) also allow other content below init() to be used as input. Under this input, since the information obtained is at most the current file content, the model tends to generate code implementation solutions that appear frequently in the training data, such as directly using ServerSocket, ServerHandler and other classes to implement from scratch. This result is too low-level and may contain code from other projects that are not imported and dependent, which does not meet user expectations.
[0195] Applying the project-level context-aware mechanism proposed in this solution, if Figure 7 The code generation plug-in (plug-in side) will first analyze the location of the current task point, and then perform the following steps respectively:
[0196] In-file context analysis: First, the current file content is segmented into the project's internal import statements, standard library import statements, third-party library import statements, the class where the task point is located, and the current file context. For the project's internal import statements, use the out-of-file context analysis to perform dependency expansion and extraction; for standard library and third-party library import statements, directly retain them as part of the input; for the current file context and context, directly retain them as part of the input.
[0197] Out-of-file context analysis: Analyze other classes (or interfaces, enumeration types, etc.) in the project that the current class depends on through package and import statements, expand their contents, analyze the member variable names, method signatures, constants and their access control keywords (such as public / private / default / friendly, etc.) defined therein, and remove specific assignment statements, initialization blocks, method bodies, etc., so as to achieve the purpose of code interface abstraction. For context filtering and sorting, first filter the information accessible to the generation point based on the relationship between the package where the generated task is located and the package where the current class is located; then sort the accessible information based on the similarity between the identifier and the method name / comment where the generated task is located.
[0198] against Fig.12 For example, after the above steps, the context information finally obtained includes: (1) the createServer() method declaration in the util.Helper class; (2) the Java standard library imported by the current file; (3) the member variable ip declared by the current class Server; (4) the method start() defined after the generation point of the current class Server. This context information will be attached to the prompt information and input into the code generation model, making it easier for the code generation model to generate code implementations that are relevant to the current project context and reuse existing encapsulations as much as possible, so as to better meet user expectations.
[0199] like Fig.13 As shown in the figure, the user can trigger the cross-file context extraction by right-clicking in the editing area to open the menu bar and clicking on cross-file code generation. Then, the extracted cross-file context and the input information such as the task description information of the code generation task are spliced into the prompt, and the spliced prompt is sent to the code generation model for reasoning. The generated result is as follows Fig.14 As shown, if the user accepts the generated code, the code generation task is completed.
[0200] It should be noted that the core idea of this application is to achieve a certain balance between the model's requirements for contextual information and the security of the user's local code / privacy data through selective perception of project-level context information, under the premise that the code generation model based on the causal language model has a limited acceptable input window, so as to help developers better use code generation tools to improve the efficiency of software development. With the recent rapid advancement of code generation tools based on large models in technology and business, more and more users have begun to pay attention to privacy and security issues during use. This application proposes a better solution to privacy and security issues, which can not only be implemented as a local IDE capability or plug-in, but also as a cloud IDE function, deployed and used in the form of cloud services.
[0201] Based on the aforementioned code generation method, the present application also provides a code development platform 10. Fig.15 As shown, the code development platform 10 includes:
[0202] Interaction module 1502, used to receive input information of a first code file in a project from a user, the input information including at least one of task description information of a code generation task or input code;
[0203] A context extraction module 1504 is used to obtain an intra-file context and a first cross-file context according to a static structure of the project, obtain a second cross-file context according to a behavioral feature of a user in developing the project, and obtain a third cross-file context according to an evolutionary coupling degree between at least one second code file and a first code file in a code repository of the project, wherein the intra-file context is a context within the first code file, and the cross-file context is a context within a code file other than the first code file in the project;
[0204] A prompt module 1506, configured to generate prompt information according to the input information, the intra-file context, the first cross-file context, the second cross-file context, and the third cross-file context;
[0205] A generation module 1508, configured to input the prompt information into a code generation model for reasoning to obtain at least one set of generated codes;
[0206] The interaction module 1502 is further configured to display at least one set of generated codes to the user.
[0207] Among them, the interaction module 1502, the context extraction module 1504, the prompt module 1506, and the generation module 1508 can be modules in the aforementioned reasoning platform 100. Exemplarily, the interaction module 1502, the context extraction module 1504, the prompt module 1506, and the generation module 1508 can be implemented by hardware, or can be implemented by software.
[0208] When implemented by software, the interaction module 1502, the context extraction module 1504, the prompt module 1506, and the generation module 1508 can be applications running on a computing device, such as a computing engine. Among them, the application can be virtualized through a virtualization service and provided to the user for use. Virtualization services may include virtual machine (VM) services, bare metal server (BMS) services, and container services. Among them, the VM service can be a service that virtualizes a virtual machine (VM) resource pool on multiple physical hosts through virtualization technology to provide VMs for users to use on demand. The BMS service is a service that virtualizes a BMS resource pool on multiple physical hosts to provide BMS for users to use on demand. The container service is a service that virtualizes a container resource pool on multiple physical hosts to provide containers for users to use on demand. VM is a simulated virtual computer, that is, a logical computer. BMS is a high-performance computing service that can be elastically scalable, and its computing performance is no different from that of a traditional physical machine, and it has the characteristics of secure physical isolation. Container is a kernel virtualization technology that can provide lightweight virtualization to achieve the purpose of isolating user space, processes and resources. It should be understood that the VM service, BMS service and container service in the above virtualization services are only used as specific examples. In actual applications, virtualization services can also be other lightweight or heavyweight virtualization services, which are not specifically limited here.
[0209] When implemented by hardware, the interaction module 1502, the context extraction module 1504, the prompt module 1506, and the generation module 1508 may include at least one computing device, such as a server, etc. Alternatively, the interaction module 1502, the context extraction module 1504, the prompt module 1506, and the generation module 1508 may also be implemented by using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0210] In some possible implementations, the context extraction module 1504 is specifically configured to:
[0211] A second cross-file context is obtained based on at least one of the metadata of an opened code file, the editing popularity of a code file in a project, or a search record of a user in a development project, wherein the metadata of the opened code file includes at least one of a relative position relationship of a class, a member, a method in the opened code file, or an editor of the opened code file.
[0212] In some possible implementations, the context extraction module 1504 is further configured to:
[0213] The behavior characteristics of users in the development project are obtained through the behavior perception interface.
[0214] In some possible implementations, the context extraction module 1504 is further configured to:
[0215] Get the commit record of the code repository;
[0216] According to the submission record of the code repository, an evolutionary coupling analysis is performed on at least one second code file and a first code file in the code repository of the project to obtain an evolutionary coupling degree between the at least one second code file and the first code file.
[0217] In some possible implementations, the context extraction module 1504 is specifically configured to:
[0218] According to the static structure of the project, build a project structure diagram. The static structure includes the hierarchical structure of the project. The hierarchical structure includes the hierarchical relationship of modules, packages, classes or code blocks in the project. The project structure diagram includes the hierarchical relationship and dependency information.
[0219] According to the position of the code generation task in the project structure diagram, obtain the sub-diagram corresponding to the code generation task;
[0220] According to the subgraph corresponding to the code generation task, the intra-file context and the first cross-file context are obtained.
[0221] In some possible implementations, the context extraction module 1504 is specifically configured to:
[0222] According to the subgraph corresponding to the code generation task, an internal import statement and an in-file context of the project are obtained, where the in-file context includes at least one of a library file import statement, a class to which the code generation task belongs, a context above the first code file, or a context below the first code file;
[0223] According to the internal import statement of the project, the dependent class of the belonging class is obtained, and the first cross-file context is obtained according to the dependent class. The first cross-file context includes at least one of the member variable name, method signature, constant, and access control keyword in the dependent class.
[0224] In some possible implementations, the prompt module 1506 is specifically configured to:
[0225] Abstracting at least one of the in-file context, the first cross-file context, the second cross-file context, or the third cross-file context into a grammatically correct interface declaration;
[0226] Generate prompt information through prompt project according to input information and interface declaration.
[0227] In some possible implementations, the prompt module 1506 is further configured to:
[0228] sorting the cross-file contexts according to at least one of access rights, topological distance, editing heat, semantic similarity, or evolutionary coupling;
[0229] The prompt module 1506 is specifically used for:
[0230] According to the input information, the context within the file, the first cross-file context, the second cross-file context and the third cross-file context, the prompt information is obtained by assembling the sorting results of the cross-file context.
[0231] In some possible implementations, the prompt module 1506 is specifically configured to:
[0232] Start marks are added to different types of information in input information, in-file context, first cross-file context, second cross-file context and third cross-file context to generate prompt information.
[0233] In some possible implementations, the code development platform 10 further includes:
[0234] The training module 1509 is used to obtain training data including cross-file contexts, and perform direct pre-training, multi-stage pre-training or fine-tuning on the base model according to the training data to obtain a code generation model.
[0235] The training module 1509 may be a module in the aforementioned training platform 200. Exemplarily, the training module 1509 may be implemented by hardware or software.
[0236] When implemented by software, the training module 1509 may be an application program running on a computing device, such as a VM service, a BMS service, or a container service. When implemented by hardware, the training module 1509 may include at least one computing device, such as a server, etc. Alternatively, the training module 1509 may also be a device implemented by ASIC or PLD, etc.
[0237] The present application also provides a computing device 1600. Fig.16 As shown, computing device 1600 includes: bus 1602, processor 1604, memory 1606 and communication interface 1608. Processor 1604, memory 1606 and communication interface 1608 communicate through bus 1602. Computing device 1600 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in computing device 1600.
[0238] The bus 1602 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.16 The bus 1602 may include a path for transmitting information between various components of the computing device 1600 (eg, the memory 1606, the processor 1604, and the communication interface 1608).
[0239] The processor 1604 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0240] The memory 1606 may include a volatile memory, such as a random access memory (RAM). The memory 1606 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD) or a solid state drive (SSD). The memory 1606 stores executable program code, and the processor 1604 executes the executable program code to implement the aforementioned code generation method. Specifically, the memory 1606 stores instructions for the code development platform 10 to execute the code generation method.
[0241] The communication interface 1608 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1600 and other devices or communication networks.
[0242] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.
[0243] like Fig.17 As shown, the computing device cluster includes at least one computing device 1600. The memory 1606 in one or more computing devices 1600 in the computing device cluster may store the same code development platform 10, which is used to execute instructions of the code generation method.
[0244] In some possible implementations, one or more computing devices 1600 in the computing device cluster may also be used to execute some instructions of the code development platform 10 for executing the code generation method. In other words, a combination of one or more computing devices 1600 may jointly execute instructions of the code development platform 10 for executing the code generation method.
[0245] It should be noted that the memory 1606 in different computing devices 1600 in the computing device cluster can store different instructions for executing part of the functions of the code development platform 10 .
[0246] Fig.18 A possible implementation is shown. Fig.18 As shown, two computing devices 1600A and 1600B are connected via a communication interface 1608. The memory in computing device 1600A stores instructions for executing the functions of interaction module 1502, context extraction module 1504, and prompt module 1506. The memory in computing device 1600B stores instructions for executing the functions of generation module 1508. Optionally, the memory of computing device 1600B also stores instructions for executing the functions of training module 1509. In other words, the memory 1606 of computing devices 1600A and 1600B jointly stores instructions for the code development platform 10 to execute the code generation method.
[0247] Fig.18 The connection mode between the computing device clusters shown may be considered to be that the code generation method provided by the present application requires a lot of computing power for model reasoning. Therefore, it is considered that the functions implemented by the generation module 1508 for model reasoning are handed over to the computing device 1600B for execution.
[0248] It should be understood that Fig.18 The functions of the computing device 1600A shown in FIG. 1 may also be completed by multiple computing devices 1600. Similarly, the functions of the computing device 1600B may also be completed by multiple computing devices 1600.
[0249] In some possible implementations, one or more computing devices in the computing device cluster may be connected via a network, which may be a wide area network or a local area network. Fig.19 A possible implementation is shown. Fig.19 As shown, two computing devices 1600C and 1600D are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 1606 in the computing device 1600C stores instructions for executing the functions of the interaction module 1502, the context extraction module 1504, and the prompt module 1506. At the same time, the memory 1606 in the computing device 1600D stores instructions for executing the functions of the generation module 1508.
[0250] Fig.19 The connection method between the computing device clusters shown may be that considering that the code generation method provided in the present application requires a large amount of computing power for model reasoning to generate code, it is considered that the functions implemented by the generation module 1508 are handed over to the computing device 1600D for execution.
[0251] It should be understood that Fig.19 The functions of the computing device 1600C shown in FIG. 1600A may also be completed by multiple computing devices 1600. Similarly, the functions of the computing device 1600D may also be completed by multiple computing devices 1600.
[0252] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes instructions that instruct the computing device to execute the above-mentioned code development platform 10 for executing the code generation method.
[0253] The embodiment of the present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the above-mentioned code generation method.
[0254] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.< / java> < / java> < / comment>
Claims
1. A code generation method, characterized in that: The method comprises: The code development platform receives input information of a first code file in a project from a user, wherein the input information includes at least one of task description information of a code generation task or input code; The code development platform obtains an in-file context and a first cross-file context according to the static structure of the project, obtains a second cross-file context according to the behavioral characteristics of the user in developing the project, and obtains a third cross-file context according to the evolutionary coupling degree between at least one second code file in the code repository of the project and the first code file, wherein the in-file context is a context within the first code file, and the cross-file context is a context within a code file in the project other than the first code file; The code development platform generates prompt information according to the input information, the intra-file context, the first cross-file context, the second cross-file context and the third cross-file context; The code development platform inputs the prompt information into a code generation model for reasoning to obtain at least one set of generated codes; The code development platform displays the at least one set of generated codes to the user.
2. The method according to claim 1, characterized in that The acquiring a second cross-file context according to the behavior characteristics of the user in developing the project includes: The code development platform obtains a second cross-file context based on at least one of the metadata of an opened code file, the editing popularity of the code file in the project, or the search record of the user in developing the project, wherein the metadata of the opened code file includes at least one of the relative position relationship of the class, member, method in the opened code file, or the editor of the opened code file.
3. The method according to claim 2, characterized in that The method further comprises: The code development platform obtains the behavior characteristics of the user in developing the project through a behavior perception interface.
4. The method according to claim 1, characterized in that The method further comprises: The code development platform obtains the submission record of the code repository; The code development platform performs an evolutionary coupling analysis on at least one second code file in the code repository of the project and the first code file according to the submission record of the code repository to obtain an evolutionary coupling degree between the at least one second code file and the first code file.
5. The method according to claim 1, characterized in that The code development platform obtains the intra-file context and the first cross-file context according to the static structure of the project, including: The code development platform constructs a project structure diagram according to the static structure of the project, wherein the static structure includes a hierarchical structure of the project, the hierarchical structure includes a hierarchical relationship among modules, packages, classes or code blocks in the project, and the project structure diagram includes the hierarchical relationship and dependency information; The code development platform obtains a subgraph corresponding to the code generation task according to the position of the code generation task in the project structure diagram; The code development platform obtains the intra-file context and the first cross-file context according to the subgraph corresponding to the code generation task.
6. The method according to claim 5, characterized in that The code development platform obtains the intra-file context and the first cross-file context according to the subgraph corresponding to the code generation task, including: The code development platform obtains, according to the subgraph corresponding to the code generation task, an internal import statement of the project and the in-file context, wherein the in-file context includes at least one of a library file import statement, a category of the code generation task, a context above the first code file, or a context below the first code file; The code development platform obtains the dependent class of the belonging class according to the internal import statement of the project, and obtains the first cross-file context according to the dependent class, wherein the first cross-file context includes at least one of the member variable name, method signature, constant, and access control keyword in the dependent class.
7. The method according to any one of claims 1 to 6, characterized in that: The code development platform generates prompt information according to the input information, the intra-file context, the first cross-file context, the second cross-file context, and the third cross-file context, including: Abstracting at least one of the in-file context, the first cross-file context, the second cross-file context, or the third cross-file context into a grammatically correct interface declaration; Prompt information is generated through a prompt project according to the input information and the interface declaration.
8. The method according to any one of claims 1 to 7, characterized in that: The method further comprises: The code development platform sorts the cross-file context according to at least one of access rights, topological distance, editing heat, semantic similarity, or evolutionary coupling; The code development platform generates prompt information according to the input information, the intra-file context, the first cross-file context, the second cross-file context, and the third cross-file context, including: The code development platform assembles prompt information according to the input information, the context within the file, the first cross-file context, the second cross-file context and the third cross-file context in combination with the sorting result of the cross-file context.
9. The method according to any one of claims 1 to 8, characterized in that: The code development platform generates prompt information according to the input information, the intra-file context, the first cross-file context, the second cross-file context, and the third cross-file context, including: The code development platform adds a start mark to different types of information in the input information, the in-file context, the first cross-file context, the second cross-file context and the third cross-file context to generate prompt information.
10. The method according to any one of claims 1 to 9, characterized in that: The code generation model is obtained in the following way: Obtain training data including cross-file context; According to the training data, the base model is directly pre-trained, multi-stage pre-trained or fine-tuned to obtain the code generation model.
11. A code development platform, characterized in that: The code development platform includes: An interaction module, configured to receive input information of a first code file in a project from a user, wherein the input information includes at least one of task description information of a code generation task or input code; a context extraction module, configured to obtain an intra-file context and a first cross-file context according to a static structure of the project, obtain a second cross-file context according to a behavioral characteristic of the user in developing the project, and obtain a third cross-file context according to an evolutionary coupling degree between at least one second code file in a code repository of the project and the first code file, wherein the intra-file context is a context within the first code file, and the cross-file context is a context within a code file in the project other than the first code file; a prompt module, configured to generate prompt information according to the input information, the intra-file context, the first cross-file context, the second cross-file context and the third cross-file context; A generation module, used for inputting the prompt information into a code generation model for reasoning to obtain at least one set of generated codes; The interaction module is further used to display the at least one set of generated codes to the user.
12. The code development platform according to claim 11, characterized in that: The context extraction module is specifically used for: A second cross-file context is obtained based on at least one of metadata of an opened code file, editing popularity of the code file in the project, or a search record of the user in developing the project, wherein the metadata of the opened code file includes at least one of a relative position relationship of a class, a member, a method in the opened code file, or an editor of the opened code file.
13. The code development platform according to claim 12, characterized in that: The context extraction module is also used to: The behavior characteristics of the user in developing the project are obtained through a behavior perception interface.
14. The code development platform according to claim 11, characterized in that: The context extraction module is also used for: Obtain the submission record of the code repository; According to the submission record of the code repository, an evolutionary coupling analysis is performed on at least one second code file in the code repository of the project and the first code file to obtain an evolutionary coupling degree between the at least one second code file and the first code file.
15. The code development platform according to claim 11, characterized in that: The context extraction module is specifically used for: According to the static structure of the project, a project structure diagram is constructed, wherein the static structure includes a hierarchical structure of the project, the hierarchical structure includes a hierarchical relationship among modules, packages, classes or code blocks in the project, and the project structure diagram includes the hierarchical relationship and dependency information; According to the position of the code generation task in the project structure diagram, obtaining a subgraph corresponding to the code generation task; According to the subgraph corresponding to the code generation task, an intra-file context and a first cross-file context are obtained.
16. The code development platform according to claim 15, characterized in that: The context extraction module is specifically used for: According to the subgraph corresponding to the code generation task, an internal import statement of the project and the in-file context are obtained, wherein the in-file context includes at least one of a library file import statement, a category of the code generation task, a context above the first code file, or a context below the first code file; According to the internal import statement of the project, the dependent class of the belonging class is obtained, and the first cross-file context is obtained according to the dependent class, wherein the first cross-file context includes at least one of the member variable name, method signature, constant, and access control keyword in the dependent class.
17. The code development platform according to any one of claims 11 to 16, characterized in that: The prompt module is specifically used for: Abstracting at least one of the in-file context, the first cross-file context, the second cross-file context, or the third cross-file context into a grammatically correct interface declaration; Prompt information is generated through a prompt project according to the input information and the interface declaration.
18. The code development platform according to any one of claims 11 to 17, characterized in that: The prompt module is also used for: sorting the cross-file contexts according to at least one of access rights, topological distance, editing heat, semantic similarity, or evolutionary coupling; The prompt module is specifically used for: The prompt information is obtained by assembling the input information, the context in the file, the first cross-file context, the second cross-file context and the third cross-file context in combination with the sorting result of the cross-file context.
19. The code development platform according to any one of claims 11 to 18, characterized in that: The prompt module is specifically used for: A start mark is added to different types of information in the input information, the in-file context, the first cross-file context, the second cross-file context, and the third cross-file context to generate prompt information.
20. The code development platform according to any one of claims 11 to 19, characterized in that: The code development platform also includes: The training module is used to obtain training data including cross-file contexts, and perform direct pre-training, multi-stage pre-training or fine-tuning on the base model according to the training data to obtain the code generation model.
21. A computing device cluster, characterized in that: The computing device cluster includes at least one computing device, and the at least one computing device includes at least one processor and at least one memory, wherein the at least one memory stores computer-readable instructions; the at least one processor executes the computer-readable instructions so that the computing device cluster executes the code generation method according to any one of claims 1 to 10.
22. A computer-readable storage medium, characterized in that: The method comprises computer-readable instructions; the computer-readable instructions are used to implement the code generation method according to any one of claims 1 to 10.
23. A computer program product, characterized in that The method comprises computer-readable instructions; the computer-readable instructions are used to implement the code generation method according to any one of claims 1 to 10.
Citation Information
Cited By
Code adoption condition determination method and device, electronic equipment and storage medium
CN120508837A
Code pre-training and generating method and system oriented to function call relation
CN120803428A
A method and system for function call relationship-oriented code pre-training and generation
CN120803428B
Code generation method and related device
EP4797080A1
Code generation method and related device
WO2025097689A1