Code generation method and related equipment
By using prompt information from the same data source and unified prompt templates in the professional field for LLM code generation, the problem of low code acceptance rate in the professional field is solved, and efficient and accurate code generation is achieved.
Patent Information
- Application Number
- CN202410013304.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-04
- Publication Date
- 2025-07-04
AI Technical Summary
Existing code generation tools based on large language model (LLM) have low code acceptance rates in the professional field and are difficult to meet business needs.
By using prompt information from the same data source in different processes to fully share and align prompt information, and assemble context information and business knowledge in combination with a unified prompt template, the accuracy of LLM generation code is improved.
It has achieved high code acceptance rate in the professional field, met business needs, and improved the accuracy and efficiency of LLM generation code.
Smart Images

Figure CN120255895A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of code generation, and particularly to a code generation method, a code development platform, a computing device cluster, a computer-readable storage medium, and a computer program product. Background Art
[0002] In recent years, as the digitization of information technology (IT) has become a development trend, more and more industries have started to develop software to achieve digitization. Software development is a product development process of building a software system or the software part in a system according to user requirements, and usually includes stages such as requirement acquisition, development planning, requirement analysis and design, programming implementation, software testing, and version control. Among them, programming implementation and software testing usually involve code development.
[0003] To improve code development efficiency, the industry has provided various technologies to assist code development, including but not limited to code continuation (also known as code completion) or code generation. With the breakthroughs of artificial intelligence (AI) technology in various tasks, especially the extremely excellent code generation ability and semantic understanding ability of large language models (LLMs) in code programming, more and more developers use LLMs for code development.
[0004] However, code development tools based on LLMs mainly demonstrate the ability or potential of code generation in general application software, but for professional fields, the code acceptance rate is low and it is basically in an unusable state. Summary of the Invention
[0005] This application provides a code generation method. This method uses prompt information from the same data source to prompt the LLM in different processes, realizes the complete sharing and complete alignment of prompt information, improves the accuracy of code generated in a single time, and moreover, this method assembles context information and business knowledge through a unified prompt template, provides a unified service, realizes the efficient collaboration between R & D assets and the LLM, and makes the code snippets generated by the LLM more accurate. Even in professional fields, a relatively high code acceptance rate can be obtained, which can meet business requirements. This application also provides a code development platform, a computing device cluster, a computer-readable storage medium, and a computer program product corresponding to the above method.
[0006] In a first aspect, the present application provides a code generation method. This method can be executed by a code development platform. The code development platform supports code generation. Among them, the code development platform can be software, which can be an independent running software, or integrated into other software, such as a functional module within an integrated development environment or a plug-in integrated into an IDE. The software can be deployed in a computing device cluster, and the computing device cluster executes the program code of the software system to execute the code generation method of the present application. Among them, the software can be provided to users in the form of a software package, and users can run the software package in a local data center or a private cloud to deploy the software. Or, the software can be provided to users in the form of a cloud service, such as software as a service (SaaS). The code development platform can also be hardware, such as a computing device cluster that provides code generation capabilities, and this computing device cluster can be used as an AI infrastructure. Among them, each computing device in the computing device cluster can include a display card, thus forming a display card cluster, and the display card resources in the display card cluster can form a display card resource pool. Among them, a display card is also called a graphics card, a video card, a graphics adapter, or a video adapter, and it is an expansion card with a Graphics Processing Unit (GPU) as the core. Its purpose is to provide a microprocessor other than the central processing unit to assist in calculating image information, and convert the display information required by the computing device and provide progressive or interlaced scanning signals to the display device.
[0007] Specifically, the code development platform receives the input information of the user in the first code file, extracts the context information of the input information according to the data warehouse to which the first code file belongs, and retrieves the corresponding business knowledge base of the data warehouse according to the input information to obtain the target business knowledge. Then the code development platform splices the input information, the context information, and the target business knowledge according to the prompt template to obtain the prompt information. Next, the code development platform inputs the prompt information into the large language model LLM for reasoning and presents the code snippet generated by the LLM reasoning to the user.
[0008] In this method, the prompt information (or prompt, prompt knowledge) of the LLM in different processes (such as training, inference, running the RAG retrieval process) comes from the same data source, specifically the data warehouse to which the code files belong. The prompt information is fully shared and fully aligned. Through the prompt information at the data warehouse level, the accuracy of the code generated by the LLM once can be improved. Moreover, this method assembles context information and business knowledge through a unified prompt template (such as a unified three-in-one prompt template for training, inference, and RAG), provides unified services, realizes the efficient collaboration between R & D assets and the LLM, and makes the code snippets generated by the LLM more accurate. In this way, in the professional field (vertical business scenarios), a relatively high code acceptance rate can also be obtained, which can meet business requirements.
[0009] In some possible implementation manners, the importance of different context information and different target business knowledge to the input information of the user may be different. For example, the context information or business knowledge with a relatively high relevance to the input information may be more important and have a higher weight. The code snippets generated by splicing different context information and different target business knowledge into the prompt information may be different. Based on this, the code development platform can sort the context information and target business knowledge according to importance. Correspondingly, the code development platform can splice the input information with the context information and target business knowledge according to the sorting result according to the prompt template to obtain the prompt information. For example, the code development platform can splice the input information with the context information and target business knowledge with high importance according to the sorting result according to the prompt template to obtain the prompt information. By preferentially splicing the context information and business knowledge with high importance into the prompt information and providing it to the LLM for inference, this method can improve the efficiency of generating code snippets that meet requirements and avoid unnecessary resource waste caused by unimportant context information and business knowledge being input to the LLM for inference first.
[0010] In some possible implementation manners, the code development platform can extract the context information of the input information according to the dependency graph corresponding to the data warehouse to which the first code file belongs. The dependency graph includes at least one of the dependency relationships between folders in the data warehouse, the dependency relationships between files, the dependencies of files, or the dependency relationships of components. The components include at least one of functions, structures, or macro definitions. The input information may include component identifiers, such as function names, function signatures, etc. The code development platform can use a graph search algorithm to search the dependency graph according to the above component identifiers such as function names and function signatures, so as to extract the context information of the input information.
[0011] By leveraging the dependency graph corresponding to the data warehouse, this method can efficiently extract context information and visualize the extracted context information through the dependency graph, providing developers with richer information and enhancing the user experience.
[0012] In some possible implementation manners, the code development platform can also obtain the full-source code of the data warehouse to which the first code file belongs, the compilation information of the full-source code, or the product design document. Then, the code development platform analyzes the full-source code, compilation information, or product design document to obtain dependency information. Next, the code development platform constructs a dependency graph based on the dependency information.
[0013] By constructing a dependency graph, this method can provide developers with more convenient code viewing and analysis capabilities. Moreover, during the training and inference processes, the dependency graph can be used to extract context information, improving the efficiency of context information extraction.
[0014] In some possible implementation manners, the LLM is trained as follows: According to the data warehouse and the business knowledge base corresponding to the data warehouse, training corpus is constructed according to the prompt template. The training corpus includes input samples, context samples obtained from the data warehouse, retrieval result samples obtained from the business knowledge base, and code samples in the data warehouse. Then, based on the training corpus, the base model is trained through supervised fine-tuning (SFT).
[0015] By constructing training corpus in the professional field and performing SFT on the base model, this method can improve the accuracy of the model in generating code in the professional field, and even in the professional field, a relatively high code acceptance rate can be obtained.
[0016] In some possible implementation manners, the code development platform can construct training corpus according to the dependency graph corresponding to the data warehouse and the business knowledge base corresponding to the data warehouse. The dependency graph includes at least one of the dependency relationships between folders in the data warehouse, the dependency relationships between files, the dependencies of files, or the dependency relationships of components. The components include at least one of functions, structures, or macro definitions.
[0017] By using the constructed dependency graph to construct training corpus, this method can improve the efficiency of constructing training corpus, thereby improving the training efficiency. Moreover, based on the dependency graph, training corpus from the same data source can be constructed, enabling complete sharing and alignment of prompt information and improving the accuracy of code generation in a single instance.
[0018] In some possible implementation manners, the code development platform may verify the training corpus and divide the verified training corpus into a training set (or referred to as a training data set) and an evaluation set. Among them, the evaluation set may include an objective evaluation set and a subjective evaluation set. The above evaluation set may be used to evaluate the trained model. In this way, high-performance models can be selected for code generation, improving the accuracy of the generated code.
[0019] In some possible implementation manners, the code development platform may also receive feedback from the user on the code snippet generated by the LLM inference. Among them, the feedback from the user on the generated code snippet may include acceptance, rejection, or revision of the code snippet.
[0020] This method supports human intervention to accept, reject, or revise the code, combining automatic code generation with human assistance to improve usability.
[0021] In some possible implementation manners, when the feedback is rejection or revision, the code development platform may also update the business knowledge base and / or the LLM according to the feedback from the user on the code snippet.
[0022] Among them, the code development platform updates the business knowledge base according to the rejection or revision of the code snippet by the user, enabling the business knowledge base to update knowledge in a timely manner and correct itself in a timely manner, providing help for improving the accuracy of the LLM output. The code development platform updates the LLM according to the rejection or revision of the second code snippet by the user, improving the accuracy of the LLM inference.
[0023] In some possible implementation manners, the code development platform may receive the first code snippet input by the user in the first code file. The first code snippet is an example code snippet or a code snippet to be completed. Among them, the code snippet to be completed may be a code snippet already input by the user, for example, a partial code snippet in a function. Or, the code development platform may also receive the requirement description input by the user in the first code file. The requirement description is used to describe the second code snippet to be generated. Or, the input information may be a combination of the first code snippet and the requirement description. For example, the code development platform may receive the first code snippet and the requirement description input by the user.
[0024] This method supports the user to input different types of information, such as code snippets or natural language, to generate code snippets, with high usability.
[0025] In a second aspect, the present application provides a code development platform. The code development platform includes:
[0026] An interaction module, configured to receive the input information of the user in the first code file;
[0027] An extraction module, configured to extract context information of the input information according to the data warehouse to which the first code file belongs;
[0028] A retrieval module, configured to retrieve a business knowledge base corresponding to the data warehouse according to the input information to obtain target business knowledge;
[0029] A prompt module, configured to splice the input information with the context information and the target business knowledge according to a prompt template to obtain prompt information;
[0030] An inference module, configured to input the prompt information into a large language model (LLM) for inference, and present code snippets generated by the LLM inference to the user.
[0031] In some possible implementation manners, the prompt module is further configured to:
[0032] Sort the context information and the target business knowledge according to importance;
[0033] Specifically, the prompt module is configured to:
[0034] According to the sorting result, splice the input information with the context information and the target business knowledge according to the prompt template to obtain prompt information.
[0035] In some possible implementation manners, specifically, the extraction module is configured to:
[0036] Extract context information of the input information according to a dependency graph corresponding to the data warehouse to which the first code file belongs, where the dependency graph includes at least one of dependencies between folders in the data warehouse, dependencies between files, dependencies of files, or dependencies of components, and the components include at least one of functions, structures, or macro definitions.
[0037] In some possible implementation manners, the code development platform further includes:
[0038] A graph construction module, configured to obtain full-source code of the data warehouse to which the first code file belongs, compilation information of the full-source code, or product design documents, analyze the full-source code, the compilation information, or the product design documents to obtain dependency information, and construct the dependency graph according to the dependency information.
[0039] In some possible implementation manners, the code development platform further includes:
[0040] A corpus construction module, configured to construct training corpus according to the data warehouse and the business knowledge base corresponding to the data warehouse, according to the prompt template, where the training corpus includes input samples, context samples obtained from the data warehouse, retrieval result samples obtained from the business knowledge base, and code samples in the data warehouse;
[0041] A training module, configured to train a base model according to the training corpus through supervised fine-tuning (SFT).
[0042] In some possible implementation manners, the training module is specifically configured to:
[0043] Construct training corpus according to the dependency graph corresponding to the data warehouse and the business knowledge base corresponding to the data warehouse, according to the prompt template, where the dependency graph includes at least one of the dependency relationships of folders in the data warehouse, the dependency relationships of files, the dependencies of files, or the dependency relationships of components, and the components include at least one of functions, structures, or macro definitions.
[0044] In some possible implementation manners, the interaction module is further configured to:
[0045] Receive the user's feedback on the code snippet generated by the LLM inference, where the feedback includes acceptance, rejection, or revision of the code snippet.
[0046] In some possible implementation manners, the code development platform further includes:
[0047] An update module, configured to update the business knowledge base and / or the LLM according to the user's feedback on the code snippet when the feedback is rejection or revision.
[0048] In some possible implementation manners, the interaction module is specifically configured to:
[0049] Receive a first code snippet input by the user in a first code file, where the first code snippet is an example code snippet or a code snippet to be completed; or,
[0050] Receive a requirement description input by the user in the first code file, where the requirement description is used to describe a second code snippet to be generated.
[0051] In a third aspect, the present application provides a computing device cluster. The computing device cluster includes at least one computing device, and the at least one computing device includes at least one processor and at least one memory. The at least one processor and the at least one memory communicate with each other. The at least one processor is configured to execute instructions stored in the at least one memory, so that the computing device or the computing device cluster executes the code generation method as described in the first aspect or any implementation manner of the first aspect.
[0052] In a fourth aspect, the present application provides a computer-readable storage medium, in which instructions are stored, and the instructions direct a computing device or a computing device cluster to execute the code generation method as described in the first aspect or any implementation manner of the first aspect.
[0053] In a fifth aspect, the present application provides a computer program product including instructions, which, when running on a computing device or a computing device cluster, causes the computing device or the computing device cluster to execute the code generation method as described in the first aspect or any implementation manner of the first aspect.
[0054] Based on the implementation manners provided in the above aspects of the present application, further combinations can be made to provide more implementation manners. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] To more clearly illustrate the technical solutions of the embodiments of the present application, the drawings used in the embodiments will be briefly introduced below.
[0056] Figure 1 It is a schematic diagram of the architecture of a code development platform provided by the present application;
[0057] Figure 2 It is a schematic diagram of constructing a dependency graph provided by the present application;
[0058] Figure 3 It is a flowchart of a code generation method provided by the present application;
[0059] Figure 4 It is a schematic diagram of the interface of a code editing interface provided by the present application;
[0060] Figure 5 It is a schematic diagram of the interface of another code editing interface provided by the present application;
[0061] Figure 6 It is a schematic diagram of constructing a dependency graph and training a model based on the dependency graph to construct training corpus provided by the present application;
[0062] Figure 7 It is a schematic diagram of the process of training, inference, and retrieval-augmented generation provided by the present application;
[0063] Figure 8 A schematic diagram of an application scenario of a code generation method provided in this application;
[0064] Figure 9 A schematic diagram of the structure of a code development platform provided for this application;
[0065] Figure 10 A schematic diagram of the structure of a computing device provided for this application;
[0066] Figure 11 A schematic diagram of the structure of a computing device cluster provided for this application;
[0067] Figure 12 A schematic diagram of the structure of another computing device cluster provided for this application;
[0068] Figure 13 A schematic diagram of the structure of another computing device cluster provided in this application. DETAILED DESCRIPTION
[0069] The terms "first" and "second" in the embodiments of the present application are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features.
[0070] First, some technical terms involved in the embodiments of the present application are introduced.
[0071] Code generation refers to the process of automatically generating source code based on predefined rules, templates, and data through automated tools or technologies. Depending on the user input, code generation can usually be divided into code continuation (such as code completion) scenarios, code generation based on sample code, or code generation based on natural language. Among them, code continuation refers to the continuation of new code snippets based on partial code snippets input by the user, code generation based on sample code refers to the generation of new code based on sample code input by the user, and code generation based on natural language generates code snippets corresponding to the requirements described by the user in natural language. Code generation can greatly improve development efficiency, reduce human errors, and ensure code consistency.
[0072] Code generation methods can be classified into the following types: code generation methods based on template engines, code generation methods based on code generation libraries, code generation methods based on models, code generation methods based on AI (such as machine learning), or code generation methods based on code conversion. Among them, the code generation method based on a template engine uses predefined templates and generates code by replacing placeholders in the templates. The code generation method based on a code generation library uses libraries or frameworks provided by programming languages to write code to generate other code. The code generation method based on a model defines a domain-specific language (DSL) model or a Unified Modeling Language (UML) model to automatically generate code. The code generation method based on AI uses machine learning technologies such as deep learning and natural language processing to automatically generate code according to the input requirements or example code. The code generation based on code conversion converts the code of one programming language into the code of another programming language.
[0073] With the continuous development of deep learning, extremely large deep learning models pre-trained on a large amount of data, namely large language models (LLMs), have made breakthrough progress in various tasks. Especially in the field of code generation, LLMs have shown extraordinary code generation capabilities.
[0074] However, LLMs are usually trained through basic training corpora and do not include professional domain corpora. As a result, when LLMs are used for code generation in professional domains (such as vertical business domains like finance, education, law, and healthcare), the code acceptance rate is low and they are basically unusable. Fine-tuning LLMs with new corpora can be both expensive and time-consuming. For this reason, the industry has introduced Retrieval-Augmented Generation (RAG). Instead of using a new corpus to fine-tune the entire LLM, RAG uses the ability of retrieval to access relevant information on demand and enhances the prompt information input into the LLM, enabling it to reference knowledge bases outside the training data sources before generating responses to optimize the output of the LLM.
[0075] When developing LLMs for different business domains, there are obvious differences in development requirements. Different templates are used for the training, inference, or RAG of LLMs, resulting in an extremely large development workload and a high maintenance cost.
[0076] In view of this, the present application provides a code generation method. This method can be executed by a code development platform. The code development platform supports code generation. Among them, the code development platform can be software, which can be an independent running software, or integrated into other software, such as a functional module within an integrated development environment or a plugin integrated into an IDE. The software can be deployed in a computing device cluster, and the computing device cluster executes the program code of the software system to execute the code generation method of the present application. Among them, the software can be provided to users in the form of a software package, and users can run the software package in a local data center or a private cloud to deploy the software. Or, the software can be provided to users in the form of a cloud service, such as software as a service (SaaS). The code development platform can also be hardware, such as a computing device cluster that provides code generation capabilities, and this computing device cluster can be used as an AI infrastructure. Among them, each computing device in the computing device cluster can include a display card, thereby forming a display card cluster, and the display card resources in the display card cluster can form a display card resource pool. Among them, a display card is also called a graphics card, a video card, a graphics adapter, or a video adapter, and is an expansion card with a Graphics Processing Unit (GPU) as the core. Its purpose is to provide a microprocessor other than the central processing unit to assist in calculating image information, and convert the display information required by the computing device and provide progressive or interlaced scanning signals to the display device.
[0077] Specifically, the code development platform can receive the input information of the user in the first code file. The input information can be the first code snippet or a requirement description. Among them, the first code snippet is an example code snippet or a code snippet to be completed, and the requirement description is used to describe the second code snippet to be generated. Then, the code development platform extracts the context information of the input information according to the data warehouse to which the first code file belongs, and retrieves the corresponding business knowledge base of the data warehouse according to the input information to obtain the target business knowledge. Then, the code development platform splices the input information with the context information and the target business knowledge according to the prompt template to obtain the prompt information. The code development platform inputs the prompt information into the large language model LLM for reasoning, and presents the code snippet generated by the LLM reasoning to the user, such as the second code snippet.
[0078] In this method, the prompt information (or prompt, prompt knowledge) of the LLM in different processes (such as training, inference, and running the RAG retrieval process) comes from the same data source, specifically the data warehouse to which the code files belong. The prompt information is fully shared and fully aligned, and the accuracy of the code generated by the LLM once can be improved through the prompt information at the data warehouse level. Moreover, this method assembles context information and business knowledge through a unified prompt template (such as a unified three-in-one prompt template for training, inference, and RAG), provides unified services, realizes the efficient collaboration between R & D assets and the LLM, and makes the code snippets generated by the LLM more accurate. In this way, in the professional field (vertical business scenarios), a relatively high code acceptance rate can also be obtained, which can meet the business requirements.
[0079] To make the technical solution of this application clearer and easier to understand, the system architecture of the code development platform of this application will be introduced below with reference to the accompanying drawings.
[0080] See Figure 1 As shown in the schematic diagram of the architecture of a code development platform, the code development platform 100 is deployed on the AI infrastructure. Among them, the AI infrastructure includes a computing device cluster and a data warehouse. Figure 1 Taking the computing device cluster as an example of a graphics card cluster, in other possible implementation manners of the embodiments of this application, it may also be other forms of computing device clusters. The resources provided by the graphics cards in the graphics card cluster can be pooled to form a graphics card resource pool for unified scheduling and management. The graphics card cluster is used to provide computing power infrastructure. The data warehouse includes a code warehouse and product design documents (also called R & D documents in some cases). The data warehouse is used to provide data support.
[0081] The code development platform 100 includes an inference platform 102, a RAG platform 104, and a prompt template 106. Further, the code development platform 100 may also include a training platform 101. Among them, the prompt template 106 can be used as a basic dependency library, for example, a three-in-one (training, inference, RAG three-in-one) basic dependency library, and is deployed in the common basic service. The training platform 101 can call the common prompt template 106 to splice training corpora to train the LLM. The RAG platform 104 can call the common prompt template 106 to splice retrieval corpora to retrieve the business knowledge base. The inference platform 102 can call the common prompt template 106 to splice context information for inference. Among them, the inference platform 102 and the RAG platform 104 can be used as interfaces for business applications to provide task scheduling services to generate the code of business applications.
[0082] Specifically, the inference platform 102 is used to receive the input information of the user in the first code file, such as the first code snippet or requirement description input by the user in the first code file. Among them, the first code snippet can be an example code snippet or an incomplete code snippet to be completed. The incomplete code snippet is the code snippet that the user has already input, for example, part of the code of a function. The requirement description is used to describe the second code snippet to be generated, and the requirement description can be descriptive information based on natural language. In some possible implementation manners, the input information can also be a combination of the first code snippet and the requirement description, such as a combination of the example code snippet and the requirement description. The inference platform 102 is also used to extract the context information of the input information according to the data warehouse to which the first code file belongs.
[0083] The RAG platform 104 is used to retrieve the business knowledge base corresponding to the data warehouse according to the input information of the user in the first code file to obtain the target business knowledge.
[0084] The inference platform 102 is also used to splice the input information of the user in the first code file with the context information and the target business knowledge according to the prompt template 106 to obtain the prompt information, input the prompt information into the LLM for inference, and present the code snippet generated by the LLM inference to the user.
[0085] Among them, the LLM can be obtained by the training platform 101 constructing training corpus according to the data warehouse and the business knowledge base corresponding to the data warehouse according to the prompt template 106, and training the base model through supervised fine-tuning (SFT) according to the training corpus. Among them, the training corpus includes input samples, context samples obtained from the data warehouse, retrieval result samples obtained from the business knowledge base, and code samples in the data warehouse.
[0086] Supervised fine-tuning means pre-training a neural network model, that is, the source model, on the source dataset. Then create a new neural network model, that is, the target model. The target model replicates all the model designs and their parameters of the source model except for the output layer. These model parameters contain the knowledge learned on the source dataset, and this knowledge is also applicable to the target dataset. The output layer of the source model is closely related to the labels of the source dataset, so it is not adopted in the target model. During fine-tuning, an output layer with an output size equal to the number of categories of the target dataset is added to the target model, and the model parameters of this layer are randomly initialized. When training the target model on the target dataset, the output layer will be trained from scratch, and the parameters of the remaining layers are fine-tuned based on the parameters of the source model. Supervised fine-tuning can utilize the parameters and structure of the pre-trained model (such as the base model), avoid training the model from scratch, thereby accelerating the model training process, and can improve the performance of the model on the target task (such as code generation tasks in professional fields).
[0087] Considering that context information at the data warehouse level will be used during the training or inference process, the code development platform 100 also supports constructing a dependency graph corresponding to the data warehouse. For example, the code development platform 100 includes a graph construction service 108. The graph construction service 108 is used to obtain the full source code of the code warehouse to which the first code file belongs, the compilation information of the full source code, or product design documents (such as the interface document of the product), and then analyze the full source code, compilation information, or product design documents to obtain dependency information, and then construct a dependency graph based on the above dependency information. Among them, the dependency graph refers to the dependency information stored in a graph, such as the dependency information stored in a tree graph. In some examples, the dependency graph can also be called a code map. The dependency graph includes at least one of the dependency relationships between folders in the data warehouse, the dependency relationships between files, the dependencies of files, or the dependency relationships between components.
[0088] Among them, the graph construction service 108 can be a knowledge graph service. As Figure 2 shown, the dependency graph can present the above dependency relationships through a visual relationship graph. For example, the dependency graph can include a folder dependency graph, a file dependency graph, a dependency graph of file dependencies or component dependencies. Among them, the component includes at least one of a function, a structure, or a macro definition. Based on this, the component dependency graph can include a function dependency graph. It should be noted that the function dependency graph can include at least one of the dependency relationships between functions, the dependency relationships between functions and structures, or the dependency relationships between functions and macros.
[0089] By extracting various dependency information, such as the dependency relationships between folders, the dependency relationships between files, the dependencies of files, the dependency relationships between components, etc., from the full source code of the project's data warehouse, the compilation information of the full source code, or product design documents, and constructing a dependency graph based on this, it can provide developers with more convenient code viewing and analysis capabilities.
[0090] Based on the foregoing code development platform 100, the present application also provides a code generation method. The code generation method of the present application will be introduced below with reference to the accompanying drawings.
[0091] See Figure 3 the flowchart of a code generation method shown. This method can be executed by the code development platform 100, and specifically includes the following steps:
[0092] S302. The code development platform 100 receives the input information of the user in the first code file.
[0093] The first code file can be one or more code files in a project. When developing software, a user can create a project, and for the project, multiple code files can be created. Different code files can be used to implement different functions of the software. Among them, the code files in the project can be developed independently by a single user or collaboratively by multiple users. When developing, the user can use the code generation ability to automatically generate code.
[0094] Taking the first code file as an example, the input information of the user in the first code file can include the first code snippet or a requirement description. Specifically, the user can input the first code snippet in the first code file. The first code snippet can be an example code snippet or a code snippet to be completed, and then trigger code generation to generate the second code snippet. Alternatively, the user can input a requirement description in the first code file, for example, input the requirement description through natural language. The requirement description is used to describe the second code snippet to be generated, and then trigger code generation to generate the second code snippet. In some examples, the input information of the user in the first code file can also be a combination of the first code snippet and the requirement description.
[0095] Specifically, the code development platform 100 can present a code editing interface to the user. The code editing interface can be a graphical user interface (GUI) or a command user interface (CUI). For the sake of description, this application takes the code editing interface as a GUI example. As Figure 4 shown, the code editing interface 400 can include an editing window 402 for the first code file. The user can input a requirement description 404 in the editing window 402. The requirement description 404 is used to describe the second code snippet to be generated. In this example, the requirement description 404 can be "create a function named CreatAccount", and then the user can trigger the code generation control 406 of the code editing interface 400, thereby triggering a code generation operation. Correspondingly, the code development platform 100 can receive the requirement description 404 input by the user and then generate code based on the requirement description 404.
[0096] It should be noted that Figure 4It is only an example of generating code examples based on requirement descriptions. In other possible implementation manners, the user can also input code snippets to be completed or example code snippets, so as to generate code based on the code snippets to be completed or example code snippets. Among them, when inputting a requirement description 404 or an example code snippet, it can be input in a comment manner. For example, the user can first input the keyword of the comment, such as the @ character or the # character, and then input the requirement description 404 or the example code snippet, so as to avoid executing the above requirement description 404 or example code snippet when executing the code file.
[0097] In addition, Figure 4 It is an example of the user triggering the code generation control 406 to trigger the code generation operation. In actual applications, the code development platform 100 also supports triggering the code generation operation in other ways. For example, the code development platform 100 also supports triggering the code generation operation through shortcut keys or menus (such as right-click menus). This application does not limit this.
[0098] S304. The code development platform 100 extracts the context information of the input information according to the data warehouse to which the first code file belongs.
[0099] To improve the accuracy of code generation, the code development platform 100 can extract the context information at the data warehouse level for code generation. Among them, the context information at the data warehouse level can include cross-file context information. Specifically, the data warehouse to which the first code file belongs not only includes the first code file, but may also include a second code file, and the second code file may be a file other than the first code file in the data warehouse. The code development platform 100 can extract the context information of the input information in the first code file and the context information of the input information in the second code file according to the data warehouse to which the first code file belongs, and obtain the context information at the data warehouse level. This context information may include dependency information related to the first code snippet or requirement description.
[0100] In some possible implementation manners, the code development platform 100 can extract the context information of the input information according to the dependency graph (code map) corresponding to the data warehouse to which the first code file belongs. Among them, the dependency graph can be constructed when training the LLM. In the inference stage, the code development platform 100 can reuse this dependency graph to extract the context information of the input information.
[0101] Specifically, the dependency graph includes at least one of the dependencies of folders in the data warehouse, the dependencies of files, the dependencies of file dependencies or components. The components include at least one of functions, structures, or macro definitions. The code development platform 100 can query the dependency graph based on the component identifier in the first code snippet or requirement description, such as the component name or component signature (function signature), to obtain dependency information related to the component. Among them, the dependency information may include at least one of the dependent components of the component, the code file to which the component belongs, the dependent files of the code file to which the component belongs, the dependencies of the code file and the dependent files to which it belongs, and the dependent folders of the code file to which the component belongs.
[0102] S306. The code development platform 100 retrieves the business knowledge base corresponding to the data warehouse according to the input information of the user to obtain the target business knowledge.
[0103] The business knowledge base stores knowledge related to the business. For example, in the financial business field, the business knowledge base can store knowledge related to finance. In the education field, the business knowledge base can store knowledge related to education. Among them, the business knowledge base can be a vector library, and the knowledge in the business knowledge base can be stored in vector form. Accordingly, the code development platform 100 can vectorize according to the input information of the user, such as the first code snippet or requirement description input by the user in the first code file, to obtain the input vector. Then, the code development platform 100 can match according to the input vector and the knowledge vector in the business knowledge base to retrieve the business knowledge base corresponding to the data warehouse and obtain the target business knowledge.
[0104] Among them, the code development platform 100 can calculate the distance between the input vector and the knowledge vector. The distance can include, but is not limited to, Euclidean distance and cosine distance. When the distance between the input vector and the knowledge vector is less than the threshold, it means that the matching is successful, and the code development platform 100 can determine the knowledge vector as the target business knowledge. Or, the code development platform 100 can sort the knowledge vectors in the business knowledge base according to the distance between the input vector and the knowledge vector, for example, sort them in ascending order, and determine the top n knowledge vectors as the target business knowledge. Among them, n is greater than or equal to 1, and n can be set according to empirical values. This embodiment does not limit this.
[0105] It should be noted that the above S304 and S306 can be executed in parallel or sequentially according to a set order. This embodiment does not limit this.
[0106] S308. The code development platform 100 splices the input information of the user in the first code file with the context information and the target business knowledge according to the prompt template to obtain the prompt information.
[0107] Among them, the prompt template can be a unified template, such as a unified three-in-one prompt template for training, inference, or RAG. The prompt template can include indications of the user input content, context information, and target business knowledge, which are used to indicate filling the corresponding content into the specified positions to achieve prompt splicing.
[0108] For ease of understanding, this application provides an example of the prompt template. As follows:
[0109] You are an expert in c / c++ code development. Your task is to complete code development based on the xx design document:
[0110] Interface declaration: {statement}
[0111] Code example: {API example}
[0112] Struct: {struct descript}
[0113] Code map: {code}
[0114] Business knowledge: {Business Knowledge}
[0115] Requirement: {requirement}
[0116] Please output the code and a detailed explanation of the coding process in markdown format: code:ˋˋˋc++ˋˋˋ
[0117] Analyze the coding process: ""
[0118] Among them, the requirement description input by the user can be filled in the {requirement} position of the above template, the example code snippet input by the user can be filled into {API example}, and the context information can be filled into the {statement}, {structdescript}, and / or {code} positions. The generated code snippet can be filled in the code: "c++" position, and the analysis process of the generated code can be filled in the reserved position for analyzing the coding process.
[0119] It should be noted that the prompt template can provide indications of the user input content, context information, and target business knowledge, but when splicing, some content can also be default. For example, some context information or target business knowledge can be default.
[0120] In some possible implementation manners, the code development platform 100 can also sort the context information and the target business knowledge according to importance, and then, according to the sorting result, the code development platform splices the user's input information with the context information and the target business knowledge according to the prompt template to obtain the prompt information.
[0121] Specifically, there can be multiple pieces of context information, and there can also be multiple pieces of target business knowledge. The importance of different context information and different business knowledge to the user's input information can be different. For example, the context information or business knowledge with a greater relevance to the input information can be more important and have a higher weight. The code development platform 100 can, according to the sorting result, splice the first code snippet or requirement description input by the user with the context information and target business knowledge with high importance according to the prompt template to obtain the prompt information.
[0122] S310. The code development platform 100 inputs the prompt information into the LLM for inference and presents the code snippet generated by the LLM inference to the user.
[0123] During inference, the prompt does not modify the user's input. It is mainly used to provide additional information for the LLM, including but not limited to context information and target business knowledge, so that the LLM can better understand the user's input and question. For example, in a question-and-answer task, the prompt can generate a prompt related to the question to help the LLM better understand the question and the user's input and generate a more accurate answer.
[0124] Among them, the context information can include information related to the user's input. For example, it is the dependency information of the function generated by the user's requirement. This information can help the LLM better understand the user's input, so as to generate an output that meets the requirements. The target business knowledge can be information related to the task. For example, it can be the question type, answer type, or restriction conditions. This information can help the LLM better understand the task and the input, so as to generate a more accurate output. In addition, the prompt information can also include templates for specific tasks to help the LLM generate outputs that meet the requirements. For example, in a question-and-answer task, the LLM can use the template to guide the LLM to generate an answer that meets the requirements.
[0125] The code development platform 100 inputs the prompt information into the LLM for inference. The inference process of the LLM model can be divided into prompt processing and the autoregressive calculation of subsequent output tokens. Among them, prompt processing can include vectorizing the prompt. When the prompt includes text, it is usually necessary to tokenize the text, specifically by splitting the text into words to break it into a sequence of discrete tokens. Each token can be represented as a word embedding representation, i.e., a word vector, through the embedding layer of the LLM, thus achieving vectorization. Autoregressive calculation means predicting the next token each time, concatenating the predicted token to the currently generated sentence, and then predicting the next token based on the concatenated sentence, repeating this process until it ends. In this example, when the LLM predicts the next token, it can combine the context information, target business knowledge, or prompt template in the prompt information for prediction to improve accuracy. When the LLM infers and generates a code snippet, in order to facilitate the distinction from the first code snippet input by the user, this application refers to the code snippet generated by the LLM inference as the second code snippet. The LLM can output or return this code snippet. Correspondingly, the code development platform 100 can present the second code snippet generated by inference to the user. Among them, the LLM can infer and generate multiple candidate code snippets, and then evaluate the multiple candidate code snippets. The second code snippet is determined from the multiple candidate code snippets according to the evaluation results. Among them, the evaluation results can be calculated according to the set evaluation metrics. The evaluation metrics can be set according to empirical values. For example, the evaluation metric can be accuracy.
[0126] Furthermore, the code development platform 100 can also receive the user's feedback on the second code snippet. Among them, the feedback can include acceptance, rejection, or revision of the second code snippet. When the user determines that the second code snippet is available, the user can accept the second code snippet; when the user determines that the second code snippet is not available, the user can reject the second code snippet; when the user determines that the second code snippet is partially available, the user can revise the second code snippet.
[0127] Such as Figure 5As shown, the code development platform 100 can present a code editing interface 500 to the user. The code editing interface 500 includes a second code snippet 502 and feedback controls corresponding to the second code snippet. Among them, the feedback controls can include an acceptance control 504, a rejection control 506, and a revision control 508. The user can perform different types of feedback such as acceptance, rejection, or revision on the second code snippet by triggering different types of feedback controls. Further, the code editing interface 500 can also include a coding process analysis 503 for the second code snippet 502. In this way, the user can determine whether the currently generated second code snippet meets the requirements based on the coding process analysis 503, and then decide the feedback type for the second code snippet.
[0128] In some possible implementation manners, when the feedback is rejection or revision, the code development platform 100 can also update the business knowledge base and / or the LLM according to the user's feedback on the second code snippet. Among them, the code development platform 100 updating the business knowledge base according to the user's rejection or revision of the second code snippet can enable the business knowledge base to update knowledge in a timely manner, correct itself in a timely manner, and provide help for improving the accuracy of the LLM output. The code development platform 100 updating the LLM according to the user's rejection or revision of the second code snippet can improve the accuracy of the LLM inference.
[0129] Based on the foregoing description, the code generation method of the present application uses a single data source, such as a data warehouse to which the code file belongs, to obtain prompt information for different processes. The prompt information is fully shared and fully aligned. Through the prompt information at the data warehouse level, the accuracy of the LLM generating code once can be improved. Moreover, this method assembles context information and business knowledge through a unified prompt template (such as a three-in-one prompt template), provides unified services, realizes the efficient collaboration between R & D assets and the LLM, and makes the code snippets generated by the LLM more accurate. In this way, in the professional field (vertical business scenarios), a relatively high code acceptance rate can also be obtained, which can meet the business requirements.
[0130] Figure 3 The LLM in the illustrated embodiment can be obtained from a base model. Among them, the base model can be a pre-trained model trained through general training corpora. For a professional field (or a specific business field, vertical business scenario), training corpora for this field (such as SFT training corpora) can also be constructed to perform SFT on the base model, so as to obtain the LLM for code generation. The following takes the code development platform 100 constructing training corpora and performing SFT on the base model as an example for illustration.
[0131] In specific implementation, the code development platform 100 can construct training corpus according to the data warehouse and the corresponding business knowledge base of the data warehouse, in accordance with the prompt template. Among them, the training corpus includes input samples, context samples obtained from the data warehouse, retrieval result samples obtained from the business knowledge base, and code samples in the data warehouse. It should be noted that in some cases, some samples can be absent. For example, when the retrieval is unsuccessful, the retrieval result sample can be absent. Then, the code development platform 100 can perform SFT on the base model according to the above training corpus.
[0132] To improve efficiency, the code development platform 100 can also construct a dependency graph corresponding to the data warehouse (such as a code map). Accordingly, the code platform 100 can construct training corpus according to the dependency graph corresponding to the data warehouse, and perform SFT on the base model based on this training corpus.
[0133] As Figure 6 shown, the code development platform 100 can determine the data warehouse of the business, specifically the data warehouse to which the first code file belongs, and then obtain the full-source code, the compilation information of the full-source code, or the product design file, and analyze them according to the full-source code, the compilation information of the full-source code, or the product design file to obtain dependency information, such as the dependency relationship between functions. Among them, the dependency relationship can include the dependency relationship between functions, the dependency relationship between functions and structures, or the dependency relationship between functions and macros. The dependency relationship can be represented by a tree structure. For example, a node of the tree represents a function declaration, and this function declaration node includes multiple child nodes, namely an identifier (ID) child node and a block statement child node. Among them, the block statement child node includes multiple child nodes, such as a variable declaration child node and a return statement child node. The code development platform 100 can construct a dependency graph according to the above dependency information.
[0134] Accordingly, the code development platform 100 can construct training corpus according to the dependency graph. As Figure 6 shown, for the functions in the dependency graph, the code development platform 100 can obtain the function signature, construct input samples according to the function signature, and obtain the function body, and construct code samples according to the function body. In addition, the code development platform 100 can also obtain the function dependencies of the function (such as external methods), structure definitions, context information of the function, obtain context samples, and retrieve in the business knowledge base according to the function signature to obtain retrieval result samples. The code development platform 100 arranges the above input samples, code samples, context samples, and retrieval result samples according toFigure 6 Concatenate according to the instructions in the prompt template to form a training corpus.
[0135] The code development platform 100 inputs the training corpus into the base model and fine-tunes the base model through SFT. As Figure 7 shown, the prompt templates in the training corpus (such as the SFT training corpus) are consistent with the inference state and the running state. The same set of prompt templates is not only used to assemble the SFT training corpus (usually the stock corpus), but also used to assemble the RAG corpus (usually the incremental corpus), and is also used to supplement context information during inference, achieving the integration of three "codes".
[0136] To make the technical solution of this application clearer and easier to understand, the code generation method will be described in detail below in combination with specific application scenarios.
[0137] See Figure 8 the application scenario schematic diagram of a code generation method shown in the figure. This method can be divided into a data preparation node, a training and evaluation phase, and an inference phase. The specific implementation of different phases will be described below.
[0138] In the data preparation phase, first, it is necessary to evaluate the customer's code repository and select high-quality code repositories. Then, clean the code in the code repository according to the data annotation and cleaning specifications developed during the R & D process of the vertical business scenario. Then, based on the "Training & Inference Corpus Hierarchy Table", construct a code map, and verify the code map through manual spot checks or automated script verification. Finally, construct a training corpus according to the code map, unify the format of the training corpus, and generate an objective evaluation set and a subjective evaluation set.
[0139] In the training and evaluation phase, for the training corpus prepared in the data preparation phase, perform training iterations through the checkpoints generated every x epochs. On this basis, batch model evaluation can be performed to select available LLMs. In specific implementation, the model can be deployed in the Alpha environment for users to evaluate and use.
[0140] In the inference stage, the LLM selected in the training and evaluation stages can be deployed in the production environment. The IDE implements two processes: context extraction (inference state) and RAG retrieval (runtime state) at the code repository level through plugins. Among them, the IDE can perform cross-file context analysis at the project level to extract cross-file context information that the user is concerned about in the current editing area and assemble it in the prompt engineering. At the same time, the IDE can retrieve the business knowledge vector library according to the user's input and intention. In the prompt engineering, the context information and business knowledge will be sorted according to importance and sent to the large language model LLM for inference after splicing. The code generated by the LLM inference is post-processed and finally presented on the IDE user interface.
[0141] In this method, the base model is subjected to SFT using training corpora in the professional field, which improves the inference ability of the LLM in the professional field. Moreover, this method takes into account the problem of code singularity, and the code correlation between products is small. Therefore, context information at the code repository level is extracted to improve the accuracy of the generated code. Moreover, this method provides prompt information through a three-in-one prompt template, enabling the LLM model to understand both business and code. Even for complex business, it can easily participate in business code development. And this method introduces the RAG mechanism to solve the problems that the knowledge of the LLM is prone to expiration, the cost of retraining the LLM is too high, the training cycle is long, and the maintenance cost is high.
[0142] Based on the foregoing code generation method, the present application also provides a code generation platform 100. The code generation platform 100 will be introduced from the perspective of functional modularization below.
[0143] As Figure 9 shown, the code generation platform 100 may include:
[0144] An interaction module 902, configured to receive input information of the user in the first code file;
[0145] An extraction module 904, configured to extract context information of the input information according to the data warehouse to which the first code file belongs;
[0146] A retrieval module 906, configured to retrieve a business knowledge base corresponding to the data warehouse according to the input information to obtain target business knowledge;
[0147] A prompt module 908, configured to splice the input information with the context information and the target business knowledge according to a prompt template to obtain prompt information;
[0148] An inference module 909, configured to input the prompt information into a large language model LLM for inference and present code snippets generated by the LLM inference to the user.
[0149] Among them, the interaction module 902 can Figure 1 modules in the inference platform 102 and / or the RAG platform 104 shown. The extraction module 904 can be a module in the inference platform 102, and the retrieval module 906 can be Figure 1 a module in the RAG platform 104 shown. The prompting module 908 and the inference module 909 can be modules in the inference platform 102.
[0150] Exemplarily, the above interaction module 902, extraction module 904, retrieval module 906, prompting module 908, and inference module 909 can be implemented by hardware or can be implemented by software.
[0151] When implemented by software, the interaction module 902, extraction module 904, retrieval module 906, prompting module 908, and inference module 909 can be application programs running on a computing device, such as a computing engine, etc. The application program can be provided in the form of a virtualization service. The virtualization service can include virtual machine (VM) service, bare metal server (BMS) service, and container service. Among them, the VM service can be a service that virtualizes a virtual machine resource pool on multiple physical hosts to provide VMs for users on demand. The BMS service is a service that virtualizes a BMS resource pool on multiple physical hosts to provide BMS for users on demand. The container service is a service that virtualizes a container resource pool on multiple physical hosts to provide containers for users on demand. A VM is a simulated virtual computer, that is, a computer logically. A BMS is an elastic and scalable high-performance computing service, with computing performance no different from that of a traditional physical machine and having the characteristic of secure physical isolation. A container is a kernel virtualization technology that can provide lightweight virtualization to achieve the purpose of isolating user space, processes, and resources. It should be understood that the VM service, BMS service, and container service in the above virtualization service are only specific examples. In actual applications, the virtualization service can also be other lightweight or heavyweight virtualization services, which are not specifically limited here.
[0152] When implemented by hardware, at least one computing device, such as a server, etc., may be included in the interaction module 902, the extraction module 904, the retrieval module 906, the prompting module 908, and the inference module 909. Alternatively, the interaction module 902, the extraction module 904, the retrieval module 906, the prompting module 908, and the inference module 909 may also be devices implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). Among them, the above PLD may be implemented by a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0153] In some possible implementation manners, the prompting module 908 is further configured to:
[0154] Sort the context information and the target service knowledge according to importance;
[0155] Specifically, the prompting module 908 is configured to:
[0156] According to the sorting result, splice the input information with the context information and the target service knowledge according to a prompting template to obtain prompting information.
[0157] In some possible implementation manners, the extraction module 904 is specifically configured to:
[0158] Extract the context information of the input information according to the dependency graph corresponding to the data warehouse to which the first code file belongs. The dependency graph includes at least one of the dependency relationships of folders in the data warehouse, the dependency relationships of files, the dependencies of files, or the dependency relationships of components. The components include at least one of functions, structures, or macro definitions.
[0159] In some possible implementation manners, the code development platform 100 further includes:
[0160] A graph construction module 901, configured to obtain the full-source code of the data warehouse to which the first code file belongs, the compilation information of the full-source code, or the product design document, analyze the full-source code, the compilation information, or the product design document to obtain dependency information, and construct the dependency graph according to the dependency information.
[0161] Among them, the graph construction module 901 can be Figure 1 a module in the graph construction service 108 in the figure. Similar to the above-mentioned inference module 909, the graph construction module 901 can be implemented by hardware or can be implemented by software. When implemented by software, the graph construction module 901 can be an application program running on a computing device, such as a computing engine, etc. The application program can be provided in the form of a virtualization service, for example, provided through a VM or container service. When implemented by hardware, the graph construction module 901 can include at least one computing device, such as a server, etc. Alternatively, the graph construction module 901 can also be a device implemented using an application-specific integrated circuit ASIC or a programmable logic device PLD, etc.
[0162] In some possible implementation manners, the code development platform 100 further includes:
[0163] a corpus construction module 903, configured to construct a training corpus according to the data warehouse and the business knowledge base corresponding to the data warehouse, according to the prompt template, where the training corpus includes input samples, context samples obtained from the data warehouse, retrieval result samples obtained from the business knowledge base, and code samples in the data warehouse;
[0164] a training module 905, configured to train a base model according to the training corpus through supervised fine-tuning SFT.
[0165] Among them, the corpus construction module 903 and the training module 905 can be implemented by hardware or can be implemented by software. When implemented by software, the corpus construction module 903 and the training module 905 can be application programs running on a computing device. The application program can be provided in the form of a virtualization service, for example, provided through a VM or container service. When implemented by hardware, the corpus construction module 903 and the training module 905 can include at least one computing device, such as a server, etc. Alternatively, the corpus construction module 903 and the training module 905 can also be devices implemented using an application-specific integrated circuit ASIC or a programmable logic device PLD, etc.
[0166] In some possible implementation manners, the training module 905 is specifically configured to:
[0167] Construct a training corpus according to the dependency graph corresponding to the data warehouse and the business knowledge base corresponding to the data warehouse, where the dependency graph includes at least one of the dependency relationships of folders in the data warehouse, the dependency relationships of files, the dependencies of files, or the dependency relationships of components, and the components include at least one of functions, structures, or macro definitions.
[0168] In some possible implementations, the interaction module 902 is further configured to:
[0169] Receive feedback from the user on the code snippet generated by the LLM inference, where the feedback includes acceptance, rejection, or revision of the code snippet.
[0170] In some possible implementations, the code development platform 100 further includes:
[0171] An update module 907, configured to update the business knowledge base and / or the LLM according to the user's feedback on the code snippet when the feedback is rejection or revision.
[0172] Wherein, the update module 907 can be implemented by hardware or by software. When implemented by software, the update module 907 can be an application running on a computing device. The application can be provided in the form of a virtualization service, such as provided through a VM or container service. When implemented by hardware, the update module 907 can include at least one computing device, such as a server, etc. Or, the update module 907 can also be a device implemented using an application-specific integrated circuit ASIC or a programmable logic device PLD, etc.
[0173] In some possible implementations, the interaction module 902 is specifically configured to:
[0174] Receive the first code snippet input by the user in the first code file, where the first code snippet is an example code snippet or a code snippet to be completed; or,
[0175] Receive the requirement description input by the user in the first code file, where the requirement description is used to describe the second code snippet to be generated.
[0176] This application also provides a computing device 1000. As Figure 10 shown, the computing device 1000 includes: a bus 1002, a processor 1004, a memory 1006, and a communication interface 1008. The processor 1004, the memory 1006, and the communication interface 1008 communicate with each other through the bus 1002. The computing device 1000 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 1000.
[0177] The bus 1002 can be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 10 only one line is used in Figure 10 , but this does not mean that there is only one bus or one type of bus. The bus 1002 can include a path for transmitting information between various components of the computing device 1000 (e.g., the memory 1006, the processor 1004, the communication interface 1008).
[0178] The processor 1004 can include any one or more of processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0179] The memory 1006 can include volatile memory, such as random access memory (RAM). The memory 1006 can also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD). Executable program code is stored in the memory 1006, and the processor 1004 executes the executable program code to implement the foregoing code generation method. Specifically, instructions for the code development platform 100 to execute the code generation method are stored on the memory 1006.
[0180] The communication interface 1008 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1000 and other devices or a communication network.
[0181] The embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.
[0182] As shown Figure 11 in the figure, the computing device cluster includes at least one computing device 1000. Instructions for executing the code generation method of the same code development platform 100 may be stored in the memory 1006 of one or more of the computing devices 1000 in the computing device cluster.
[0183] In some possible implementation manners, one or more of the computing devices 1000 in the computing device cluster may also be used to execute some instructions of the code development platform 100 for executing the code generation method. In other words, a combination of one or more computing devices 1000 may jointly execute the instructions of the code development platform 100 for executing the code generation method.
[0184] It should be noted that the memories 1006 in different computing devices 1000 in the computing device cluster may store different instructions for executing some functions of the code development platform 100.
[0185] Figure 12 A possible implementation manner is shown. As shown Figure 12 in the figure, two computing devices 1000A and 1000B are connected through a communication interface 1008. Instructions for executing the functions of the interaction module 902, extraction module 904, prompt module 908, and inference module 909 are stored in the memory of the computing device 1000A. Instructions for executing the function of the retrieval module 906 are stored in the memory of the computing device 1000B. In other words, the memories 1006 of the computing devices 1000A and 1000B jointly store the instructions of the code development platform 100 for executing the code generation method. Further, the memory 1006 of the computing device 1000B may also store instructions for executing the functions of the graph construction module 901, corpus construction module 903, training module 905, and update module 907.
[0186] Figure 12 The connection manner between the computing device clusters shown in the figure may be considered in view of the need to update the business knowledge base in a timely manner for the code generation method provided in this application. Therefore, it is considered to hand over the function implemented by the retrieval module 906 to the computing device 1000B for execution.
[0187] It should be understood that Figure 12 the functions of the computing device 1000A shown in the figure may also be completed by multiple computing devices 1000. Similarly, the functions of the computing device 1000B may also be completed by multiple computing devices 1000.
[0188] In some possible implementation manners, one or more of the computing devices in the computing device cluster may be connected through a network. Among them, the network may be a wide area network or a local area network, etc. Figure 13 A possible implementation manner is shown. As shown Figure 13As shown, two computing devices 1000C and 1000D are connected via a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this possible implementation, the memory 1006 in computing device 1000C stores instructions for implementing the functions of the interaction module 902, extraction module 904, prompt module 908, and inference module 909. At the same time, the memory 1006 in computing device 1000D stores instructions for implementing the function of the retrieval module 906. Further, the memory 1006 of computing device 1000D can also store instructions for implementing the functions of the graph construction module 901, corpus construction module 903, training module 905, and update module 907.
[0189] Figure 13 The connection method between the computing device clusters shown can be considered in view of the need to update the business knowledge base in a timely manner for the code generation method provided in this application. Therefore, it is considered to hand over the function implemented by the retrieval module 906 to computing device 1000B for execution.
[0190] It should be understood that Figure 13 the functions of computing device 1000C shown in can also be completed by multiple computing devices 1000. Similarly, the functions of computing device 1000D can also be completed by multiple computing devices 1000.
[0191] This application embodiment also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to execute the above-mentioned code generation method applied to the code development platform 100.
[0192] This application embodiment also provides a computer program product containing instructions. The computer program product can be software or a program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, it causes at least one computing device to execute the above-mentioned code generation method.
[0193] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.
Claims
1. A code generation method, characterized in that, The method includes: The code development platform receives the input information of the user in the first code file; The code development platform extracts the context information of the input information according to the data warehouse to which the first code file belongs, and retrieves the business knowledge base corresponding to the data warehouse according to the input information to obtain the target business knowledge; The code development platform splices the input information with the context information and the target business knowledge according to the prompt template to obtain the prompt information; The code development platform inputs the prompt information into the large language model LLM for reasoning and presents the code snippet generated by the LLM reasoning to the user.
2. The method according to claim 1, wherein The method further includes: The code development platform sorts the context information and the target business knowledge according to importance; The code development platform splices the input information with the context information and the target business knowledge according to the prompt template to obtain the prompt information, including: The code development platform splices the input information with the context information and the target business knowledge according to the sorting result and the prompt template to obtain the prompt information.
3. The method according to claim 1 or 2, characterized in that, The code development platform extracts the context information of the input information according to the data warehouse to which the first code file belongs, including: The code development platform extracts the context information of the input information according to the dependency graph corresponding to the data warehouse to which the first code file belongs. The dependency graph includes at least one of the dependency relationships of folders in the data warehouse, the dependency relationships of files, the dependencies of files, or the dependency relationships of components. The components include at least one of functions, structures, or macro definitions.
4. The method according to claim 3, wherein The method further includes: The code development platform obtains the full-source code of the data warehouse to which the first code file belongs, the compilation information of the full-source code, or the product design document; The code development platform analyzes the full-source code, the compilation information, or the product design document to obtain dependency information; The code development platform constructs the dependency graph according to the dependency information.
5. The method according to any one of claims 1 to 4, characterized in that, The LLM is trained in the following manner: According to the data warehouse and the business knowledge base corresponding to the data warehouse, training corpus is constructed according to the prompt template. The training corpus includes input samples, context samples obtained from the data warehouse, retrieval result samples obtained from the business knowledge base, and code samples in the data warehouse; According to the training corpus, the base model is trained through supervised fine-tuning SFT.
6. The method according to claim 5, wherein The constructing the training corpus according to the data warehouse and the business knowledge base corresponding to the data warehouse according to the prompt template includes: According to the dependency graph corresponding to the data warehouse and the business knowledge base corresponding to the data warehouse, training corpus is constructed according to the prompt template. The dependency graph includes at least one of the dependency relationships of folders in the data warehouse, the dependency relationships of files, the dependencies of files, or the dependency relationships of components. The components include at least one of functions, structures, or macro definitions.
7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: The code development platform receives the user's feedback on the code snippets generated by the LLM inference, and the feedback includes acceptance, rejection, or revision of the code snippets.
8. The method according to claim 7, characterized in that, The method further includes: When the feedback is rejection or revision, the code development platform updates the business knowledge base and / or the LLM according to the user's feedback on the code snippets.
9. The method according to any one of claims 1 to 8, characterized in that, The code development platform receives the input information of the user in the first code file, including: The code development platform receives the first code snippet input by the user in the first code file, and the first code snippet is an example code snippet or a code snippet to be completed; or, The code development platform receives the requirement description input by the user in the first code file, and the requirement description is used to describe the second code snippet to be generated.
10. A code development platform, characterized in that, The code development platform includes: An interaction module, which is used to receive the input information of the user in the first code file; An extraction module, which is used to extract the context information of the input information according to the data warehouse to which the first code file belongs; A retrieval module, which is used to retrieve the business knowledge base corresponding to the data warehouse according to the input information to obtain the target business knowledge; A prompt module, which is used to splice the input information with the context information and the target business knowledge according to the prompt template to obtain prompt information; An inference module, which is used to input the prompt information into a large language model LLM for inference and present the code snippets generated by the LLM inference to the user.
11. The code development platform according to claim 10, characterized in that, The prompt module is further used for: Sorting the context information and the target business knowledge according to importance; Specifically, the prompt module is used for: According to the sorting result, splicing the input information with the context information and the target business knowledge according to the prompt template to obtain prompt information.
12. The code development platform according to claim 10 or 11, characterized in that Specifically, the extraction module is used for: According to the dependency graph corresponding to the data warehouse to which the first code file belongs, extracting the context information of the input information, and the dependency graph includes at least one of the dependency relationships between folders in the data warehouse, the dependency relationships between files, the dependencies of files, or the dependency relationships between components, and the components include at least one of functions, structures, or macro definitions.
13. The code development platform according to claim 12, wherein The code development platform further includes: A graph construction module, which is used to obtain the full-source code of the data warehouse to which the first code file belongs, the compilation information of the full-source code, or the product design document, analyze the full-source code, the compilation information, or the product design document to obtain dependency information, and construct the dependency graph according to the dependency information.
14. The code development platform according to any one of claims 10 to 13, characterized in that The code development platform further includes: A corpus construction module, which is used to construct training corpus according to the data warehouse and the business knowledge base corresponding to the data warehouse according to the prompt template, and the training corpus includes input samples, context samples obtained from the data warehouse, retrieval result samples obtained from the business knowledge base, and code samples in the data warehouse; A training module, which is used to train the base model through supervised fine-tuning SFT according to the training corpus.
15. The code development platform according to claim 14, characterized in that, Specifically, the training module is used for: Construct training corpus according to the dependency graph corresponding to the data warehouse and the business knowledge base corresponding to the data warehouse, where the dependency graph includes at least one of the dependencies of folders in the data warehouse, the dependencies of files, the dependencies of file dependencies or components, and the components include at least one of functions, structures or macro definitions.
16. The code development platform according to any one of claims 10 to 15, characterized in that, The interaction module is further configured to: Receive the user's feedback on the code snippet generated by the LLM inference, where the feedback includes acceptance, rejection or revision of the code snippet.
17. The code development platform according to claim 16, wherein The code development platform further includes: An update module, configured to update the business knowledge base and / or the LLM according to the user's feedback on the code snippet when the feedback is rejection or revision.
18. The code development platform according to any one of claims 10 to 17, characterized in that Specifically, the interaction module is configured to: Receive the first code snippet input by the user in the first code file, where the first code snippet is an example code snippet or a code snippet to be completed; or, Receive the requirement description input by the user in the first code file, where the requirement description is used to describe the second code snippet to be generated.
19. A cluster of computing devices, characterized in that, The computing device cluster includes at least one computing device, and the at least one computing device includes at least one processor and at least one memory, and computer-readable instructions are stored in the at least one memory; the at least one processor executes the computer-readable instructions to cause the computing device cluster to execute the code generation method according to any one of claims 1 to 9.
20. A computer-readable storage medium, characterized in that, Include computer-readable instructions; the computer-readable instructions are used to implement the code generation method according to any one of claims 1 to 9.
21. A computer program product, characterized in that, Include computer-readable instructions; the computer-readable instructions are used to implement the code generation method according to any one of claims 1 to 9.
Citation Information
Cited By
Code generation method and system based on artificial intelligence
CN120803432A
Business code generation method and device fusing code specification detection
CN120803465A
Business program code generation method and device, medium and electronic equipment
CN121614122A