Code generation method and related device

By using prompt information from the same data source and unified template in the code development platform to share and align LLM's prompts, combined with dependency diagrams and manual feedback, the problem of low code generation accuracy in the professional field is solved, and efficient and accurate code generation is achieved.

WO2025145584A1PCT designated stage expired Publication Date: 2025-07-10HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD

Patent Information

Application Number
PCT/CN2024/109791
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-04
Filing Date
2024-08-05
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Existing code generation tools based on large language model (LLM) have low code acceptance rates in the professional field, are difficult to meet business needs, and are costly to develop and maintain.

Method used

By using prompt information from the same data source for complete sharing and alignment in different processes, combining unified prompt templates for assembling context information and business knowledge, building a dependency diagram, improving the accuracy of LLM generated code, and supporting manual feedback and knowledge base updates.

Benefits of technology

Achieve high code acceptance rate in the professional field, improve the accuracy and efficiency of code generation, and reduce development and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024109791_10072025_PF_FP_ABST
    Figure CN2024109791_10072025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present application is a code generation method, comprising: a code development platform receives input information by a user in a first code file, extracts context information of the input information according to a data warehouse to which the first code file belongs, and retrieves a business knowledge base corresponding to the data warehouse according to the input information so as to obtain target business knowledge; then the code development platform combines the input information, the context information and the target business knowledge according to a prompt template, so as to obtain prompt information; and then the code development platform inputs the prompt information into a large language model (LLM) for reasoning, and presents to the user a code segment reasoned by the LLM. Thus, using the prompt information from the same data source in different processes for prompting the LLM can achieve complete sharing and complete alignment of the prompt information, thereby improving the accuracy of generating code in single time. Furthermore, using a unified prompt template to combine the context information and the business knowledge can make a code segment generated by the LLM more accurate, thereby improving the code acceptance rate in professional domains.
Need to check novelty before this filing date? Find Prior Art

Description

A code generation method and related device

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on January 4, 2024, with application number 202410013304.5 and invention name “A code generation method and related equipment”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of code generation technology, and in particular to a code generation method, a code development platform, a computing device cluster, a computer-readable storage medium, and a computer program product. Background Art

[0003] In recent years, as the digitization of information technology (IT) has become a growing trend, more and more industries have begun developing software to achieve digitalization. Software development is the process of building a software system or the software component of a system according to user requirements. It typically includes phases such as requirements acquisition, development planning, requirements analysis and design, programming implementation, software testing, and version control. Programming implementation and software testing typically involve code development.

[0004] To improve code development efficiency, the industry offers a variety of technologies to assist with code development, including but not limited to code continuation (also known as code completion) and code generation. With breakthroughs in artificial intelligence (AI) technology across various tasks, particularly large language models (LLMs), which have demonstrated exceptional code generation and semantic understanding capabilities in code programming, an increasing number of developers are using LLMs for code development.

[0005] However, LLM-based code development tools mainly demonstrate the ability or potential of code generation on general application software, but for professional fields, the code acceptance rate is low and is basically unusable.

[0006] Summary of the Invention

[0007] The present application provides a code generation method that uses prompt information from the same data source to prompt LLM in different processes, thereby achieving complete sharing and alignment of prompt information, improving the accuracy of single-time code generation, and assembling context information and business knowledge through a unified prompt template, providing a unified service, and achieving efficient collaboration between R&D assets and LLM, making the code snippets generated by LLM more accurate. Even in professional fields, a high code acceptance rate can be achieved to meet business needs. The present application also provides a code development platform, a computing device cluster, a computer-readable storage medium, and a computer program product corresponding to the above method.

[0008] In a first aspect, the present application provides a code generation method. The method can be executed by a code development platform. The code development platform supports code generation. The code development platform can be software, which can be independently running software or integrated into other software, such as a functional module within an integrated development environment (IDE) or a plug-in integrated into an IDE. The software can be deployed on a computing device cluster, which executes the program code of the software system, thereby executing the code generation method of the present application. The software can be provided to users in the form of a software package, and users can run the software package in a local data center or a private cloud to deploy the software. Alternatively, the software can be provided to users in the form of a cloud service, such as software as a service (SaaS). The code development platform can also be hardware, such as a computing device cluster that provides code generation capabilities, which can serve as AI infrastructure. Each computing device in the computing device cluster can include a display card, thereby forming a graphics card cluster, and the graphics card resources in the graphics card cluster can form a graphics card resource pool. A graphics card, also known as a graphics card, video card, graphics adapter, or video adapter, is an expansion card with a graphics processing unit (GPU) as its core. Its purpose is to provide a microprocessor other than the central processing unit to help calculate image information, convert the display information required by the computing device, and provide progressive or interlaced scanning signals to the display device.

[0009] Specifically, the code development platform receives user input from a first code file, extracts contextual information from the data repository to which the first code file belongs, and searches the business knowledge base corresponding to the data repository based on the input information to obtain target business knowledge. The code development platform then combines the input information with the contextual information and target business knowledge according to a prompt template to generate prompt information. The code development platform then inputs the prompt information into a large language model (LLM) for inference, presenting the code snippet generated by the LLM inference to the user.

[0010] In this method, the prompt information (or called prompt, prompt knowledge) of LLM in different processes (such as training, reasoning, and running RAG retrieval process) comes from the same data source, specifically the data warehouse to which the code file belongs. The prompt information is fully shared and fully aligned. The prompt information at the data warehouse level can improve the accuracy of LLM's single code generation. Moreover, this method assembles contextual information and business knowledge through a unified prompt template (such as a unified three-in-one prompt template for training, reasoning, and RAG), provides a unified service, and realizes efficient collaboration between R&D assets and LLM, making the code snippets generated by LLM more accurate. In this way, in professional fields (vertical business scenarios), a higher code acceptance rate can also be obtained to meet business needs.

[0011] In some possible implementations, the importance of different context information and different target business knowledge to the user's input information may be different. For example, context information or business knowledge with a greater relevance to the input information may be more important and have a higher weight. The code snippets generated by splicing different context information and different target business knowledge into prompt information may be different. Based on this, the code development platform can sort the context information and target business knowledge according to importance. Accordingly, the code development platform can splice the input information with the context information and target business knowledge according to the prompt template based on the sorting results to obtain prompt information. For example, the code development platform can splice the input information with the context information and target business knowledge of high importance according to the prompt template based on the sorting results to obtain prompt information. This method can improve the efficiency of generating code snippets that meet the requirements by splicing the context information and business knowledge of high importance into the prompt information first and providing it to the LLM for reasoning, thereby avoiding unnecessary waste of resources caused by unimportant context information and business knowledge being input into the LLM for reasoning first.

[0012] In some possible implementations, the code development platform can extract context information of the input information based on the dependency graph corresponding to the data warehouse to which the first code file belongs. The dependency graph includes at least one of the dependencies of folders, files, file dependencies, or components in the data warehouse. The component includes at least one of a function, a structure, or a macro definition. The input information may include component identifiers, such as function names, function signatures, etc. The code development platform can use a graph search algorithm to search the dependency graph based on the component identifiers such as the function names and function signatures, thereby extracting the context information of the input information.

[0013] This method can efficiently extract context information by utilizing the dependency graph corresponding to the data warehouse, and can visualize the extracted context information through the dependency graph, providing developers with richer information and improving user experience.

[0014] In some possible implementations, the code development platform may also obtain the full source code, compilation information of the full source code, or product design documents from the data warehouse to which the first code file belongs. The code development platform then analyzes the full source code, compilation information, or product design documents to obtain dependency information. The code development platform then constructs a dependency graph based on the dependency information.

[0015] This method builds a dependency graph, providing developers with more convenient code review and analysis capabilities. This dependency graph can also be used to extract contextual information during both training and inference, improving the efficiency of contextual information extraction.

[0016] In some possible implementations, the LLM is trained as follows: a training corpus is constructed according to a data warehouse and the business knowledge base corresponding to the data warehouse according to a prompt template. The training corpus includes input samples, context samples obtained from the data warehouse, retrieval result samples obtained from the business knowledge base, and code samples in the data warehouse. The base model is then trained through supervised fine-tuning (SFT) based on the training corpus.

[0017] This method constructs training corpus in professional fields and performs SFT on the base model, which can improve the accuracy of the model in generating code in professional fields and achieve a higher code acceptance rate even in professional fields.

[0018] In some possible implementations, the code development platform can construct training corpus according to the dependency graph corresponding to the data warehouse and the business knowledge base corresponding to the data warehouse, according to the prompt template. The dependency graph includes at least one of the dependencies of folders, files, file dependencies or component dependencies in the data warehouse, and the components include at least one of functions, structures or macro definitions.

[0019] This method uses a pre-built dependency graph to construct training corpora, improving the efficiency of constructing training corpora and, consequently, training efficiency. Furthermore, the dependency graph allows for constructing training corpora from the same data source, enabling complete sharing and alignment of prompt information, thus improving the accuracy of single-pass code generation.

[0020] In some possible implementations, the code development platform can validate the training corpus and divide the validated training corpus into a training set (or training dataset) and an evaluation set. The evaluation set can include both objective and subjective evaluation sets, which can be used to evaluate the trained model. This allows for the selection of high-performance models for code generation, improving the accuracy of the generated code.

[0021] In some possible implementations, the code development platform may further receive user feedback on the code snippet generated by LLM reasoning, wherein the user feedback on the generated code snippet may include acceptance, rejection, or revision of the code snippet.

[0022] This method improves usability by supporting human intervention to accept, reject or revise the code, combining automatic code generation with human assistance.

[0023] In some possible implementations, when the feedback is rejection or revision, the code development platform may also update the business knowledge base and / or LLM based on the user's feedback on the code snippet.

[0024] The code development platform updates the business knowledge base based on user rejections or revisions of code snippets, enabling timely knowledge updates and self-correction, helping to improve the accuracy of LLM output. The code development platform also updates the LLM based on user rejections or revisions of the second code snippet, improving the accuracy of LLM reasoning.

[0025] In some possible implementations, the code development platform may receive a first code snippet input by a user in a first code file, where the first code snippet is a sample code snippet or a code snippet to be completed, wherein the code snippet to be completed may be a code snippet already input by the user, such as a partial code snippet within a function. Alternatively, the code development platform may receive a requirement description input by the user in the first code file, where the requirement description is used to describe the second code snippet to be generated. Alternatively, the input information may be a combination of the first code snippet and the requirement description, for example, the code development platform may receive the first code snippet and the requirement description input by the user.

[0026] This method supports users to input different types of information, such as code snippets or natural language, to generate code snippets, and has high usability.

[0027] In a second aspect, the present application provides a code development platform. The code development platform includes:

[0028] An interactive module, configured to receive user input information in a first code file;

[0029] An extraction module, configured to extract context information of the input information according to a data warehouse to which the first code file belongs;

[0030] A retrieval module is used to retrieve the business knowledge base corresponding to the data warehouse according to the input information to obtain target business knowledge;

[0031] A prompt module, configured to combine the input information, the context information, and the target business knowledge according to a prompt template to obtain prompt information;

[0032] The inference module is used to input the prompt information into a large language model (LLM) for inference, and present the code snippet generated by the LLM inference to the user.

[0033] In some possible implementations, the prompt module is further configured to:

[0034] sorting the context information and the target business knowledge according to importance;

[0035] The prompt module is specifically used for:

[0036] According to the sorting result, the input information, the context information and the target business knowledge are spliced ​​according to a prompt template to obtain prompt information.

[0037] In some possible implementations, the extraction module is specifically configured to:

[0038] According to the dependency graph corresponding to the data warehouse to which the first code file belongs, the context information of the input information is extracted, the dependency graph includes at least one of the dependencies of folders, files, file dependencies or components in the data warehouse, and the component includes at least one of a function, a structure or a macro definition.

[0039] In some possible implementations, the code development platform further includes:

[0040] A graph construction module is used to obtain the full source code of the data warehouse to which the first code file belongs, the compilation information of the full source code, or the product design document, analyze the full source code, the compilation information, or the product design document, obtain dependency information, and construct the dependency graph based on the dependency information.

[0041] In some possible implementations, the code development platform further includes:

[0042] a corpus construction module, configured to construct training corpus according to the prompt template based on the data warehouse and the business knowledge base corresponding to the data warehouse, wherein the training corpus includes input samples, context samples obtained from the data warehouse, search result samples obtained from the business knowledge base, and code samples in the data warehouse;

[0043] The training module is used to train the base model through supervised fine-tuning SFT according to the training corpus.

[0044] In some possible implementations, the training module is specifically configured to:

[0045] According to the dependency graph corresponding to the data warehouse and the business knowledge base corresponding to the data warehouse, the training corpus is constructed according to the prompt template, the dependency graph includes at least one of the dependencies of folders, files, file dependencies or components in the data warehouse, and the component includes at least one of a function, a structure or a macro definition.

[0046] In some possible implementations, the interaction module is further configured to:

[0047] Receive feedback from the user on the code snippet generated by the LLM reasoning, where the feedback includes acceptance, rejection, or revision of the code snippet.

[0048] In some possible implementations, the code development platform further includes:

[0049] An updating module is configured to update the business knowledge base and / or the LLM according to the user's feedback on the code snippet when the feedback is rejection or revision.

[0050] In some possible implementations, the interaction module is specifically configured to:

[0051] receiving a first code snippet input by the user in a first code file, where the first code snippet is a sample code snippet or a code snippet to be completed; or

[0052] A requirement description input by the user in the first code file is received, where the requirement description is used to describe a second code fragment to be generated.

[0053] In a third aspect, the present application provides a computing device cluster. The computing device cluster includes at least one computing device, wherein the at least one computing device includes at least one processor and at least one memory. The at least one processor and the at least one memory communicate with each other. The at least one processor is configured to execute instructions stored in the at least one memory, so that the computing device or computing device cluster performs the code generation method described in the first aspect or any implementation of the first aspect.

[0054] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, wherein the instructions instruct a computing device or a computing device cluster to execute the code generation method described in the first aspect or any implementation of the first aspect.

[0055] In a fifth aspect, the present application provides a computer program product comprising instructions, which, when executed on a computing device or a computing device cluster, enables the computing device or computing device cluster to execute the code generation method described in the first aspect or any one of the implementations of the first aspect.

[0056] Based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical methods of the embodiments of the present application, the following is a brief introduction to the drawings used in the embodiments.

[0058] FIG1 is a schematic diagram of the architecture of a code development platform provided by this application;

[0059] FIG2 is a schematic diagram of a dependency graph provided by the present application;

[0060] FIG3 is a flow chart of a code generation method provided by the present application;

[0061] FIG4 is a schematic diagram of a code editing interface provided by the present application;

[0062] FIG5 is a schematic diagram of another code editing interface provided by the present application;

[0063] FIG6 is a schematic diagram of constructing a dependency graph and constructing a training corpus training model based on the dependency graph provided by the present application;

[0064] FIG7 is a schematic diagram of a process of training, reasoning, and retrieval enhancement generation provided by this application;

[0065] FIG8 is a schematic diagram of an application scenario of a code generation method provided by this application;

[0066] FIG9 is a schematic diagram of the structure of a code development platform provided by this application;

[0067] FIG10 is a schematic diagram of the structure of a computing device provided by the present application;

[0068] FIG11 is a schematic diagram of the structure of a computing device cluster provided by this application;

[0069] FIG12 is a schematic diagram of the structure of another computing device cluster provided by the present application;

[0070] FIG13 is a schematic diagram of the structure of another computing device cluster provided in this application. DETAILED DESCRIPTION

[0071] The terms "first" and "second" in the embodiments of this application are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Therefore, features specified as "first" or "second" may explicitly or implicitly include one or more of the features.

[0072] First, some technical terms involved in the embodiments of this application are introduced.

[0073] Code generation refers to the process of automatically generating source code based on predefined rules, templates, and data through automated tools or technologies. Depending on the user input, code generation can generally be divided into code continuation scenarios (such as code completion), code generation based on sample code, or code generation based on natural language. Among them, code continuation refers to the continuation of new code snippets based on partial code snippets input by the user, code generation based on sample code refers to the generation of new code based on sample code input by the user, and code generation based on natural language generates code snippets corresponding to the requirements described by the user in natural language. Code generation can greatly improve development efficiency, reduce human errors, and ensure code consistency.

[0074] Code generation methods can be divided into the following types: code generation methods based on template engines, code generation methods based on code generation libraries, model-based code generation methods, AI-based (such as machine learning) code generation methods, or code generation methods based on code conversion. Among them, the code generation method based on template engines is to use predefined templates to generate code by replacing placeholders in the templates. The code generation method based on code generation libraries is to use libraries or frameworks provided by programming languages ​​to write code to generate other codes. The model-based code generation method is to automatically generate code by defining a domain-specific language (DSL) model or a unified modeling language (UML) model. The AI-based code generation method is to use machine learning technologies such as deep learning and natural language processing to automatically generate code based on input requirements or sample code. The generated code based on code conversion is to convert code in one programming language into code in another programming language.

[0075] With the continuous development of deep learning, large language models (LLMs), ultra-large deep learning models pre-trained based on large amounts of data, have made breakthrough progress in various tasks, especially in the field of code generation, where LLMs have demonstrated extraordinary code generation capabilities.

[0076] However, LLM is usually trained with basic training corpus and does not include professional domain corpus. As a result, when LLM is faced with code generation in professional fields (such as vertical business fields such as finance, education, law, and healthcare), the code acceptance rate is low and it is basically unusable. Fine-tuning LLM with a new corpus can be expensive and time-consuming. For this reason, the industry introduced Retrieval-Augmented Generation (RAG). Instead of fine-tuning the entire LLM with a new corpus, RAG uses the power of retrieval to access relevant information on demand and enhance the prompt information input to LLM, enabling it to reference knowledge bases outside the training data source before generating a response to optimize the output of LLM.

[0077] When developing LLMs for different business fields, the development requirements vary significantly. Different templates are used for LLM training, reasoning, or RAG, resulting in an extremely large development workload and high maintenance costs.

[0078] In view of this, the present application provides a code generation method. The method can be executed by a code development platform. The code development platform supports code generation. The code development platform can be software, which can be independently running software or integrated into other software, such as a functional module in an integrated development environment (IDE) or a plug-in integrated into an IDE. The software can be deployed in a computing device cluster, and the computing device cluster executes the program code of the software system, thereby executing the code generation method of the present application. The software can be provided to the user in the form of a software package, and the user can run the software package in a local data center or a private cloud to deploy the software. Alternatively, the software can be provided to the user in the form of a cloud service, such as software as a service (SaaS). The code development platform can also be hardware, such as a computing device cluster that provides code generation capabilities, and the computing device cluster can serve as an AI infrastructure. Each computing device in the computing device cluster can include a display card, thereby forming a graphics card cluster, and the graphics card resources in the graphics card cluster can form a graphics card resource pool. A graphics card, also known as a graphics card, video card, graphics adapter, or video adapter, is an expansion card with a graphics processing unit (GPU) as its core. Its purpose is to provide a microprocessor other than the central processing unit to help calculate image information, convert the display information required by the computing device, and provide progressive or interlaced scanning signals to the display device.

[0079] Specifically, the code development platform can receive input information from the user in the first code file, and the input information can be a first code snippet or a requirement description, wherein the first code snippet is a sample code snippet or a code snippet to be completed, and the requirement description is used to describe the second code snippet to be generated. Then the code development platform extracts the context information of the input information based on the data warehouse to which the first code file belongs, and retrieves the business knowledge base corresponding to the data warehouse based on the input information to obtain the target business knowledge. Then the code development platform splices the input information with the context information and the target business knowledge according to the prompt template to obtain prompt information. The code development platform inputs the prompt information into the large language model LLM for reasoning, and presents the code snippet generated by the LLM reasoning to the user, such as the second code snippet.

[0080] In this method, the prompt information (or called prompt, prompt knowledge) of LLM in different processes (such as training, reasoning, and running RAG retrieval process) comes from the same data source, specifically the data warehouse to which the code file belongs. The prompt information is fully shared and fully aligned. The prompt information at the data warehouse level can improve the accuracy of LLM's single code generation. Moreover, this method assembles contextual information and business knowledge through a unified prompt template (such as a unified three-in-one prompt template for training, reasoning, and RAG), provides a unified service, and realizes efficient collaboration between R&D assets and LLM, making the code snippets generated by LLM more accurate. In this way, in professional fields (vertical business scenarios), a higher code acceptance rate can also be obtained to meet business needs.

[0081] In order to make the technical solution of the present application clearer and easier to understand, the system architecture of the code development platform of the present application is introduced below with reference to the accompanying drawings.

[0082] Referring to the architectural diagram of a code development platform shown in FIG1 , the code development platform 100 is deployed on an AI infrastructure. Among them, the AI ​​infrastructure includes a computing device cluster and a data warehouse. FIG1 uses the computing device cluster as an example of a graphics card cluster. In other possible implementations of the embodiment of the present application, other forms of computing device clusters may also be used. The resources provided by the graphics cards in the graphics card cluster can be pooled to form a graphics card resource pool for unified scheduling and management. The graphics card cluster is used to provide computing power infrastructure. The data warehouse includes a code warehouse and product design documents (sometimes also referred to as R&D documents). The data warehouse is used to provide data support.

[0083] The code development platform 100 includes a reasoning platform 102, a RAG platform 104, and a prompt template 106. Furthermore, the code development platform 100 may also include a training platform 101. Among them, the prompt template 106 can be used as a basic dependency library, for example, a three-in-one (training, reasoning, RAG three-in-one) basic dependency library, which is deployed in a public basic service. The training platform 101 can call the public prompt template 106 to perform training corpus splicing to train and obtain LLM. The RAG platform 104 can call the public prompt template 106 to perform retrieval corpus splicing to retrieve the business knowledge base. The reasoning platform 102 can call the public prompt template 106 to splice context information for reasoning. Among them, the reasoning platform 102 and the RAG platform 104 can serve as interfaces for business applications, providing task scheduling services to generate code for business applications.

[0084] Specifically, the reasoning platform 102 is used to receive input information from the user in the first code file, such as the first code snippet or requirement description entered by the user in the first code file. The first code snippet can be an example code snippet or a code snippet to be completed, and the code snippet to be completed is a code snippet that the user has entered, such as a partial code of a function. The requirement description is used to describe the second code snippet to be generated, and the requirement description can be a description information based on natural language. In some possible implementations, the input information can also be a combination of the first code snippet and the requirement description, such as a combination of the example code snippet and the requirement description. The reasoning platform 102 is also used to extract context information of the input information based on the data warehouse to which the first code file belongs.

[0085] The RAG platform 104 is used to search the business knowledge base corresponding to the data warehouse according to the user's input information in the first code file to obtain target business knowledge.

[0086] The reasoning platform 102 is also used to combine the user's input information in the first code file with the context information and target business knowledge according to the prompt template 106 to obtain prompt information, input the prompt information into the LLM for reasoning, and present the code snippet generated by the LLM reasoning to the user.

[0087] The LLM can be obtained by constructing training corpus according to the prompt template 106 by the training platform 101 based on the data warehouse and the business knowledge base corresponding to the data warehouse, and training the base model through supervised fine-tuning (SFT) based on the training corpus. The training corpus includes input samples, context samples obtained from the data warehouse, search result samples obtained from the business knowledge base, and code samples in the data warehouse.

[0088] Supervised fine-tuning involves pre-training a neural network model on a source dataset, known as the source model. A new neural network model, known as the target model, is then created. The target model replicates the entire model design and parameters of the source model, except for the output layer. These model parameters incorporate knowledge learned on the source dataset, and this knowledge is also applicable to the target dataset. The output layer of the source model is closely tied to the labels of the source dataset and is therefore not used in the target model. During fine-tuning, an output layer with an output size equal to the number of categories in the target dataset is added to the target model, and the model parameters of this layer are randomly initialized. When training the target model on the target dataset, the output layer is trained from scratch, while the parameters of the remaining layers are fine-tuned based on the parameters of the source model. Supervised fine-tuning can leverage the parameters and structure of a pre-trained model (such as a base model) to avoid training the model from scratch, thereby accelerating the model training process and improving the model's performance on target tasks (such as code generation tasks in specialized domains).

[0089] Taking into account that contextual information at the data warehouse level will be used during training or inference, the code development platform 100 also supports the construction of a dependency graph corresponding to the data warehouse. For example, the code development platform 100 includes a graph construction service 108, which is used to obtain the full source code of the code warehouse to which the first code file belongs, the compilation information of the full source code, or the product design document (such as the interface document of the product), and then analyze the full source code, compilation information or product design document to obtain dependency information, and then construct a dependency graph based on the above dependency information. Among them, the dependency graph refers to dependency information stored in a graph, such as dependency information stored in a tree graph. In some examples, the dependency graph can also be called a code map. The dependency graph includes at least one of the dependencies of folders, files, file dependencies, or components in the data warehouse.

[0090] Among them, the graph construction service 108 can be a knowledge graph service. As shown in Figure 2, the dependency graph can present the above-mentioned dependency relationships through a visual relationship graph. For example, the dependency graph can include a folder dependency graph, a file dependency graph, a file dependency graph, or a component dependency graph. Among them, the component includes at least one of a function, a structure, or a macro definition. Based on this, the component dependency graph can include a function dependency graph. It should be noted that the function dependency graph can include at least one of the dependency relationships between functions, the dependency relationship between functions and structures, or the dependency relationship between functions and macros.

[0091] By extracting various dependency information, such as folder dependencies, file dependencies, file dependencies, component dependencies, etc., from the full source code of the project-based data warehouse, the compilation information of the full source code, or product design documents, and constructing a dependency graph based on this, developers can be provided with more convenient code viewing and analysis capabilities.

[0092] Based on the aforementioned code development platform 100, the present application further provides a code generation method. The code generation method of the present application is described below with reference to the accompanying drawings.

[0093] 3 shows a flowchart of a code generation method. The method can be executed by the code development platform 100 and specifically includes the following steps:

[0094] S302: The code development platform 100 receives user input information in a first code file.

[0095] The first code file can be one or more code files in a project. During software development, users can create a project and, for each project, multiple code files. Different code files can be used to implement different software functions. Code files in a project can be developed independently by a single user or collaboratively by multiple users. During development, users can use code generation capabilities to automatically generate code.

[0096] Taking the first code file as an example, the user's input information in the first code file may include a first code snippet or a requirement description. Specifically, the user can input the first code snippet in the first code file, which can be an example code snippet or a code snippet to be completed, and then trigger code generation to generate a second code snippet. Alternatively, the user can input a requirement description in the first code file, for example, input the requirement description in natural language, which is used to describe the second code snippet to be generated, and then trigger code generation to generate the second code snippet. In some examples, the user's input information in the first code file may also be a combination of the first code snippet and the requirement description.

[0097] Specifically, the code development platform 100 can present a code editing interface to the user, which can be a graphical user interface (GUI) or a command user interface (CUI). For ease of description, this application uses the code editing interface as a GUI example. As shown in Figure 4, the code editing interface 400 can include an editing window 402 for the first code file. The user can enter a requirement description 404 in the editing window 402. The requirement description 404 is used to describe the second code snippet to be generated. In this example, the requirement description 404 can be "Create a function named CreateAccount", and then the user can trigger the code generation control 406 of the code editing interface 400, thereby triggering the code generation operation. Accordingly, the code development platform 100 can receive the requirement description 404 input by the user, and then generate code based on the requirement description 404.

[0098] It should be noted that FIG4 is merely an example of generating code based on a requirement description. In other possible implementations, the user may also input a code snippet to be completed or an example code snippet so as to generate code based on the code snippet to be completed or the example code snippet. When inputting the requirement description 404 or the example code snippet, the input may be in the form of comments. For example, the user may first input a comment keyword, such as an @ character or a # character, and then input the requirement description 404 or the example code snippet. This may avoid executing the requirement description 404 or the example code snippet when executing the code file.

[0099] In addition, FIG4 illustrates an example of a user triggering a code generation control 406 to trigger a code generation operation. In actual applications, the code development platform 100 also supports triggering a code generation operation through other means. For example, the code development platform 100 also supports triggering a code generation operation through a shortcut key or menu (such as a right-click menu). This application does not limit this.

[0100] S304 : The code development platform 100 extracts context information of the input information according to the data warehouse to which the first code file belongs.

[0101] In order to improve the accuracy of code generation, the code development platform 100 can extract context information at the data warehouse level for code generation. The context information at the data warehouse level may include cross-file context information. Specifically, the data warehouse to which the first code file belongs includes not only the first code file, but also the second code file, and the second code file may be a file in the data warehouse other than the first code file. The code development platform 100 can extract the context information of the input information in the first code file and the context information of the input information in the second code file according to the data warehouse to which the first code file belongs, and obtain the context information at the data warehouse level. The context information may include dependency information related to the first code snippet or the requirement description.

[0102] In some possible implementations, the code development platform 100 can extract contextual information about the input information based on a dependency graph (code map) corresponding to the data repository to which the first code file belongs. The dependency graph can be constructed during LLM training. During the inference phase, the code development platform 100 can reuse the dependency graph to extract contextual information about the input information.

[0103] Specifically, the dependency graph includes at least one of the dependencies of folders, files, file dependencies, or components in the data warehouse, and the components include at least one of functions, structures, or macro definitions. The code development platform 100 can query the dependency graph based on the component identifier in the first code snippet or the requirement description, such as the component name, component signature (function signature), to obtain dependency information related to the component. The dependency information may include at least one of the dependent components of the component, the code file to which the component belongs, the dependent files of the code file to which the component belongs, the dependencies of the owned code file and the dependent files, and the dependent folder of the code file to which the component belongs.

[0104] S306 , the code development platform 100 searches the business knowledge base corresponding to the data warehouse according to the user's input information to obtain target business knowledge.

[0105] The business knowledge base stores business-related knowledge. For example, in the financial business field, the business knowledge base can store knowledge related to finance, and in the education field, the business knowledge base can store knowledge related to education. The business knowledge base can be a vector base, and the knowledge in the business knowledge base can be stored in vector form. Accordingly, the code development platform 100 can vectorize the user's input information, such as the first code snippet or requirement description entered by the user in the first code file, to obtain an input vector. The code development platform 100 can then match the input vector with the knowledge vector in the business knowledge base to retrieve the business knowledge base corresponding to the data warehouse and obtain the target business knowledge.

[0106] Among them, the code development platform 100 can calculate the distance between the input vector and the knowledge vector, and the distance may include but is not limited to Euclidean distance and cosine distance. When the distance between the input vector or the knowledge vector is less than the threshold, it means that the match is successful, and the code development platform 100 can determine the knowledge vector as the target business knowledge. Alternatively, the code development platform 100 can sort the knowledge vectors in the business knowledge base according to the distance between the input vector and the knowledge vector, for example, sorting them in order from small to large, and determining the top n knowledge vectors as the target business knowledge. Among them, n is greater than or equal to 1, and n can be set according to an empirical value, which is not limited in this embodiment.

[0107] It should be noted that the above S304 and S306 can be executed in parallel or in sequence according to a set order, and this embodiment does not limit this.

[0108] S308 , the code development platform 100 combines the user's input information in the first code file with the context information and the target business knowledge according to the prompt template to obtain prompt information.

[0109] The prompt template can be a unified template, such as a unified three-in-one prompt template for training, inference, or RAG. The prompt template can include an indication of user input content, context information, and target business knowledge. This indication is used to instruct the corresponding content to be filled into the specified location to achieve prompt splicing.

[0110] For ease of understanding, this application provides an example of a prompt template as shown below:

[0111] You are a C / C++ code development expert. Your task is to complete code development based on the xx design document:

[0112] Interface declaration: {statement}

[0113] Code example: {API example}

[0114] Structure: {struct descript}

[0115] Code map: {code}

[0116] Business Knowledge

[0117] Requirement: {requirement}

[0118] Please output the code and detailed explanation of the coding process in markdown format: code:```c++```

[0119] Analysis of the encoding process: ""

[0120] The user-entered requirement description can be placed in the {requirement} section of the template. The user-entered sample code snippet can be placed in {API example}. Context information can be placed in {statement}, {struct descript}, and / or {code}. The generated code snippet can be placed in the "code:c++" section. The analysis process for the generated code can be placed in the reserved section for the analysis and coding process.

[0121] It should be noted that the prompt template can provide an indication of the user input content, context information, and target business knowledge, but some content may be omitted during splicing. For example, some context information or target business knowledge may be omitted.

[0122] In some possible implementations, the code development platform 100 can also sort the context information and target business knowledge according to importance, and then the code development platform splices the user's input information with the context information and target business knowledge according to the prompt template based on the sorting results to obtain prompt information.

[0123] Specifically, context information may include multiple items, and target business knowledge may also include multiple items. Different pieces of context information and business knowledge may have different importance to the user's input information. For example, context information or business knowledge that is more relevant to the input information may be more important and have a higher weight. Based on the sorting results, the code development platform 100 may combine the first code snippet or requirement description entered by the user with the highly important context information and target business knowledge according to a prompt template to obtain prompt information.

[0124] S310 , the code development platform 100 inputs the prompt information into the LLM for reasoning, and presents the code snippet generated by the LLM reasoning to the user.

[0125] During reasoning, prompts are not used to modify user input. Instead, they provide the LLM with additional information, including but not limited to contextual information and target business knowledge, to better understand user input and questions. For example, in question-answering tasks, prompts can generate a relevant prompt, helping the LLM better understand the question and user input and generate more accurate answers.

[0126] Contextual information can include information related to user input, such as dependency information of functions generated by user requirements. This information can help the LLM better understand user input and generate output that meets the requirements. Target business knowledge can be task-related information, such as question types, answer types, or constraints. This information can help the LLM better understand the task and input, thereby generating more accurate output. Furthermore, prompt information can also include templates for specific tasks to help the LLM generate output that meets the requirements. For example, in a question-answering task, the LLM can use templates to guide the LLM in generating answers that meet the requirements.

[0127] The code development platform 100 inputs the prompt information into the LLM for reasoning. The process of LLM model reasoning can be divided into prompt processing and subsequent autoregressive calculation of the output token. Among them, prompt processing can include vectorizing the prompt. When the prompt includes text, it is usually necessary to tokenize the text, specifically to segment the text into discrete word units (token sequences). Each token can be represented as a word embedding representation, i.e., a word vector, through the embedding layer of the LLM, thereby achieving vectorization. Autoregressive calculation refers to each time the next token is predicted, the predicted token is spliced ​​into the currently generated sentence, and then the next token prediction is made based on the spliced ​​sentence, and the process is repeated until the end. In this example, when predicting the next token, the LLM can combine the contextual information in the prompt information, the target business knowledge or the prompt template to make a prediction to improve accuracy. When LLM generates a code snippet by reasoning, in order to distinguish it from the first code snippet input by the user, this application refers to the code snippet generated by LLM reasoning as a second code snippet. LLM can output or return the code snippet. Accordingly, the code development platform 100 can present the second code snippet generated by reasoning to the user. Among them, LLM can generate multiple candidate code snippets by reasoning, and then evaluate the multiple candidate code snippets, and determine the second code snippet from the multiple candidate code snippets based on the evaluation results. Among them, the evaluation result can be calculated according to the set evaluation index. The evaluation index can be set according to the empirical value, for example, the evaluation index can be the accuracy rate.

[0128] Furthermore, the code development platform 100 may also receive user feedback on the second code snippet, wherein the feedback may include acceptance, rejection, or revision of the second code snippet. If the user determines that the second code snippet is usable, the user may accept the second code snippet; if the user determines that the second code snippet is not usable, the user may reject the second code snippet; if the user determines that the second code snippet is partially usable, the user may revise the second code snippet.

[0129] As shown in FIG5 , the code development platform 100 can present a code editing interface 500 to the user. The code editing interface 500 includes a second code snippet 502 and feedback controls corresponding to the second code snippet, wherein the feedback controls may include an accept control 504, a reject control 506, and a revision control 508. The user can trigger different types of feedback controls to provide different types of feedback, such as accepting, rejecting, or revising the second code snippet. Furthermore, the code editing interface 500 may also include a coding process analysis 503 for the second code snippet 502. In this way, the user can determine whether the currently generated second code snippet meets the requirements based on the coding process analysis 503, and then decide the type of feedback to provide to the second code snippet.

[0130] In some possible implementations, when the feedback is rejection or revision, the code development platform 100 may also update the business knowledge base and / or LLM based on the user's feedback on the second code snippet. This updating of the business knowledge base by the code development platform 100 based on the user's rejection or revision of the second code snippet allows the business knowledge base to update knowledge and self-correct in a timely manner, thereby helping to improve the accuracy of LLM output. This updating of the LLM by the code development platform 100 based on the user's rejection or revision of the second code snippet can also improve the accuracy of LLM reasoning.

[0131] Based on the foregoing description, the code generation method of the present application uses prompt information from the same data source, such as the data warehouse to which the code file belongs, to obtain prompt information for different processes. The prompt information is fully shared and fully aligned. The prompt information at the data warehouse level can improve the accuracy of LLM single-time code generation. Moreover, the method assembles contextual information and business knowledge through a unified prompt template (such as a three-in-one prompt template), provides a unified service, and realizes efficient collaboration between R&D assets and LLM, making the code snippets generated by LLM more accurate. In this way, in professional fields (vertical business scenarios), a higher code acceptance rate can also be obtained to meet business needs.

[0132] The LLM in the embodiment shown in FIG3 can be obtained from a base model. The base model can be a pre-trained model obtained by training with general training corpus. For a professional field (or a specific business field, vertical business scenario), a training corpus (such as SFT training corpus) in that field can also be constructed, and the base model can be subjected to SFT to obtain an LLM for code generation. The following example illustrates the construction of training corpus using the code development platform 100 and the SFT of the base model.

[0133] In specific implementations, the code development platform 100 can construct training corpus according to the prompt template based on the data warehouse and the business knowledge base corresponding to the data warehouse. The training corpus includes input samples, context samples obtained from the data warehouse, search result samples obtained from the business knowledge base, and code samples in the data warehouse. It should be noted that in some cases, some samples may be omitted. For example, when a search is unsuccessful, the search result samples may be omitted. The code development platform 100 can then perform SFT on the base model based on the aforementioned training corpus.

[0134] To improve efficiency, the code development platform 100 can also construct a dependency graph (such as a code map) corresponding to the data warehouse. Accordingly, the code platform 100 can construct training corpus according to the prompt template based on the dependency graph corresponding to the data warehouse, and perform SFT on the base model based on the training corpus.

[0135] As shown in Figure 6, the code development platform 100 can determine the data warehouse of the business, specifically the data warehouse to which the first code file belongs, and then obtain the full source code, the compilation information of the full source code or the product design file, and analyze the full source code, the compilation information of the full source code or the product design file to obtain dependency information, such as the dependency relationship of the function. Among them, the dependency relationship may include the dependency relationship between functions, the dependency relationship between functions and structures, or the dependency relationship between functions and macros. The dependency relationship can be represented by a tree structure, for example, a node of the tree represents a function declaration (Function Declaration), and the function declaration node includes multiple child nodes, namely, an identifier (ID) child node and a block statement (Block Statement) child node. Among them, the block statement child node includes multiple child nodes, such as a variable declaration child node and a return statement child node. The code development platform 100 can construct a dependency graph based on the above dependency information.

[0136] Accordingly, the code development platform 100 can construct a training corpus based on the dependency graph. As shown in Figure 6, for the function in the dependency graph, the code development platform 100 can obtain the function signature, construct an input sample based on the function signature, and obtain the function body, and construct a code sample based on the function body. In addition, the code development platform 100 can also obtain the function dependency (for example, external methods), structure definition, and context information of the function to obtain context samples, and search the business knowledge base based on the function signature to obtain retrieval result samples. The code development platform 100 splices the above-mentioned input samples, code samples, context samples, and retrieval result samples according to the instructions of the prompt template in Figure 6 to form a training corpus.

[0137] The code development platform 100 inputs training data into the base model and fine-tunes the base model through SFT. As shown in Figure 7, the prompt templates in the training data (such as SFT training data) remain consistent with the inference state and the running state. The same set of prompt templates is used not only to assemble SFT training data (usually existing data), but also to assemble RAG data (usually incremental data), and to supplement context information during inference, thus achieving the integration of three "codes".

[0138] In order to make the technical solution of this application clearer and easier to understand, the code generation method is described in detail below in combination with specific application scenarios.

[0139] Referring to the schematic diagram of an application scenario of a code generation method shown in FIG8 , the method can be divided into a data preparation node, a training and evaluation phase, and an inference phase. The specific implementation of each phase is described below.

[0140] During the data preparation phase, we first need to evaluate the customer's code repository and select high-quality code repositories. Next, we clean the code in the repository according to the data labeling and cleaning specifications developed during the R&D process for each vertical business scenario. Then, based on the "Training & Inference Corpus Hierarchy Table," we construct a code map and verify it through manual spot checks or automated script verification. Finally, we construct a training corpus based on the code map, unify the format of the training corpus, and generate both objective and subjective evaluation sets.

[0141] During the training and evaluation phases, training iterations are performed using checkpoints generated every x epochs based on the training data prepared during the data preparation phase. Based on this, batch model evaluations are performed to select viable LLMs. In practice, the models can be deployed in the Alpha environment for user evaluation and use.

[0142] During the reasoning phase, the LLM selected in the training and evaluation phases can be deployed in the production environment. The IDE implements two processes: context extraction (inference state) and RAG retrieval (running state) at the code repository level through plug-ins. Among them, the IDE can perform context analysis across files at the project level, thereby extracting cross-file context information that the user is concerned about in the current editing area, and assembling it in the prompt project. At the same time, the IDE can retrieve the business knowledge vector library based on the user's input and intention. In the prompt project, context information and business knowledge will be sorted according to importance, and after splicing, they will be sent to the large language model LLM for reasoning. The code generated by LLM reasoning is post-processed and finally presented on the IDE user interface.

[0143] In this method, SFT is performed on the base model using training corpus from professional domains, improving the reasoning ability of the LLM in professional domains. Furthermore, this method takes into account the uniqueness of the code, as the code correlation between products is low. Therefore, contextual information at the code repository level is extracted to improve the accuracy of the generated code. Furthermore, this method provides prompt information through a three-in-one prompt template, enabling the LLM model to understand both the business and the code, making it easier to participate in business code development even for complex businesses. Furthermore, this method introduces the RAG mechanism to address the issues of LLM knowledge easily becoming outdated, the high cost of retraining the LLM, the long training cycle, and the high maintenance cost.

[0144] Based on the aforementioned code generation method, the present application further provides a code generation platform 100. The code generation platform 100 is introduced below from the perspective of functional modularization.

[0145] As shown in FIG9 , the code generation platform 100 may include:

[0146] Interaction module 902, for receiving user input information in the first code file;

[0147] An extraction module 904 is configured to extract context information of the input information according to a data warehouse to which the first code file belongs;

[0148] Retrieval module 906, configured to retrieve the business knowledge base corresponding to the data warehouse according to the input information to obtain target business knowledge;

[0149] The prompt module 908 is configured to combine the input information with the context information and the target business knowledge according to a prompt template to obtain prompt information;

[0150] The reasoning module 909 is configured to input the prompt information into a large language model (LLM) for reasoning, and present the code snippet generated by the LLM reasoning to the user.

[0151] Among them, the interaction module 902 can be a module in the reasoning platform 102 and / or the RAG platform 104 shown in Figure 1, the extraction module 904 can be a module in the reasoning platform 102, the retrieval module 906 can be a module in the RAG platform 104 shown in Figure 1, and the prompt module 908 and the reasoning module 909 can be modules in the reasoning platform 102.

[0152] Exemplarily, the interaction module 902 , the extraction module 904 , the retrieval module 906 , the prompt module 908 , and the reasoning module 909 may be implemented by hardware or software.

[0153] When implemented via software, the interaction module 902, extraction module 904, retrieval module 906, prompt module 908, and inference module 909 can be applications running on a computing device, such as a computing engine. Applications can be provided as virtualization services. Virtualization services can include virtual machine (VM) services, bare metal server (BMS) services, and container services. VM services use virtualization technology to create a virtual machine resource pool across multiple physical hosts, providing users with VMs on demand. BMS services create a virtual BMS resource pool across multiple physical hosts, providing users with BMSs on demand. Container services create a virtual container resource pool across multiple physical hosts, providing users with containers on demand. A VM is a simulated virtual computer, essentially a single logical computer. BMS is a scalable, high-performance computing service with computing performance comparable to traditional physical machines and secure physical isolation. Containers are a kernel virtualization technology that provides lightweight virtualization to isolate user space, processes, and resources. It should be understood that the VM service, BMS service and container service in the above-mentioned virtualization services are only specific examples. In actual applications, virtualization services can also be other lightweight or heavyweight virtualization services, which are not specifically limited here.

[0154] When implemented through hardware, the interaction module 902, the extraction module 904, the retrieval module 906, the prompt module 908, and the reasoning module 909 may include at least one computing device, such as a server. Alternatively, the interaction module 902, the extraction module 904, the retrieval module 906, the prompt module 908, and the reasoning module 909 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0155] In some possible implementations, the prompt module 908 is further configured to:

[0156] sorting the context information and the target business knowledge according to importance;

[0157] The prompt module 908 is specifically used to:

[0158] According to the sorting result, the input information, the context information and the target business knowledge are spliced ​​according to a prompt template to obtain prompt information.

[0159] In some possible implementations, the extraction module 904 is specifically configured to:

[0160] According to the dependency graph corresponding to the data warehouse to which the first code file belongs, the context information of the input information is extracted, the dependency graph includes at least one of the dependencies of folders, files, file dependencies or components in the data warehouse, and the component includes at least one of a function, a structure or a macro definition.

[0161] In some possible implementations, the code development platform 100 further includes:

[0162] The graph construction module 901 is used to obtain the full source code of the data warehouse to which the first code file belongs, the compilation information of the full source code, or the product design document, analyze the full source code, the compilation information, or the product design document, obtain dependency information, and construct the dependency graph based on the dependency information.

[0163] Among them, the graph construction module 901 can be a module in the graph construction service 108 in Figure 1. Similar to the above-mentioned reasoning module 909, the graph construction module 901 can be implemented by hardware, or can be implemented by software. When implemented by software, the graph construction module 901 can be an application running on a computing device, such as a computing engine. The application can be provided in the form of a virtualization service, for example, through a VM or container service. When implemented by hardware, the graph construction module 901 can include at least one computing device, such as a server. Alternatively, the graph construction module 901 can also be a device implemented using an application-specific integrated circuit ASIC or a programmable logic device PLD.

[0164] In some possible implementations, the code development platform 100 further includes:

[0165] Corpus construction module 903, configured to construct training corpus according to the prompt template based on the data warehouse and the business knowledge base corresponding to the data warehouse, wherein the training corpus includes input samples, context samples obtained from the data warehouse, search result samples obtained from the business knowledge base, and code samples in the data warehouse;

[0166] The training module 905 is used to train the base model through supervised fine-tuning SFT based on the training corpus.

[0167] The corpus construction module 903 and the training module 905 can be implemented via hardware or software. When implemented via software, the corpus construction module 903 and the training module 905 can be applications running on a computing device. Applications can be provided as virtualized services, such as VMs or container services. When implemented via hardware, the corpus construction module 903 and the training module 905 can include at least one computing device, such as a server. Alternatively, the corpus construction module 903 and the training module 905 can be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD).

[0168] In some possible implementations, the training module 905 is specifically configured to:

[0169] According to the dependency graph corresponding to the data warehouse and the business knowledge base corresponding to the data warehouse, the training corpus is constructed according to the prompt template, the dependency graph includes at least one of the dependencies of folders, files, file dependencies or components in the data warehouse, and the component includes at least one of a function, a structure or a macro definition.

[0170] In some possible implementations, the interaction module 902 is further configured to:

[0171] Receive feedback from the user on the code snippet generated by the LLM reasoning, where the feedback includes acceptance, rejection, or revision of the code snippet.

[0172] In some possible implementations, the code development platform 100 further includes:

[0173] The updating module 907 is configured to update the business knowledge base and / or the LLM according to the user's feedback on the code snippet when the feedback is rejection or revision.

[0174] The update module 907 can be implemented via hardware or software. When implemented via software, the update module 907 can be an application running on a computing device. The application can be provided as a virtualized service, such as a VM or container service. When implemented via hardware, the update module 907 can include at least one computing device, such as a server. Alternatively, the update module 907 can be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD).

[0175] In some possible implementations, the interaction module 902 is specifically configured to:

[0176] receiving a first code snippet input by the user in a first code file, where the first code snippet is a sample code snippet or a code snippet to be completed; or

[0177] A requirement description input by the user in the first code file is received, where the requirement description is used to describe a second code fragment to be generated.

[0178] This application also provides a computing device 1000. As shown in Figure 10, computing device 1000 includes a bus 1002, a processor 1004, a memory 1006, and a communication interface 1008. Processor 1004, memory 1006, and communication interface 1008 communicate with each other via bus 1002. Computing device 1000 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 1000.

[0179] Bus 1002 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG10 illustrates a single bus line, but this does not imply a single bus or type of bus. Bus 1002 may include a path for transmitting information between various components of computing device 1000 (e.g., memory 1006, processor 1004, and communication interface 1008).

[0180] The processor 1004 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0181] The memory 1006 may include a volatile memory, such as a random access memory (RAM). The memory 1006 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD). The memory 1006 stores an executable program code, and the processor 1004 executes the executable program code to implement the aforementioned code generation method. Specifically, the memory 1006 stores instructions for the code development platform 100 to execute the code generation method.

[0182] The communication interface 1008 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1000 and other devices or a communication network.

[0183] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0184] As shown in Figure 11, the computing device cluster includes at least one computing device 1000. The memory 1006 of one or more computing devices 1000 in the computing device cluster may store instructions for executing the code generation method using the same code development platform 100.

[0185] In some possible implementations, one or more computing devices 1000 in the computing device cluster may also be used to execute some of the instructions of the code development platform 100 for executing the code generation method. In other words, a combination of one or more computing devices 1000 may jointly execute the instructions of the code development platform 100 for executing the code generation method.

[0186] It should be noted that the memory 1006 in different computing devices 1000 in the computing device cluster may store different instructions for executing part of the functions of the code development platform 100 .

[0187] FIG12 illustrates a possible implementation. As shown in FIG12 , two computing devices 1000A and 1000B are connected via a communication interface 1008. The memory in computing device 1000A stores instructions for executing the functions of interaction module 902, extraction module 904, prompt module 908, and reasoning module 909. The memory in computing device 1000B stores instructions for executing the functions of retrieval module 906. In other words, the memories 1006 of computing devices 1000A and 1000B jointly store instructions for the code development platform 100 to execute the code generation method. Furthermore, the memory 1006 of computing device 1000B may also store instructions for executing the functions of graph construction module 901, corpus construction module 903, training module 905, and update module 907.

[0188] The connection mode between the computing device clusters shown in Figure 12 may be considered to be based on the fact that the code generation method provided by this application requires timely updating of the business knowledge base. Therefore, it is considered to delegate the functions implemented by the retrieval module 906 to the computing device 1000B.

[0189] It should be understood that the functionality of the computing device 1000A shown in FIG12 may also be accomplished by multiple computing devices 1000. Similarly, the functionality of the computing device 1000B may also be accomplished by multiple computing devices 1000.

[0190] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network. The network may be a wide area network or a local area network, etc. FIG13 shows a possible implementation. As shown in FIG13 , two computing devices 1000C and 1000D are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 1006 in the computing device 1000C stores instructions for executing the functions of the interaction module 902, the extraction module 904, the prompt module 908, and the reasoning module 909. At the same time, the memory 1006 in the computing device 1000D stores instructions for executing the functions of the retrieval module 906. Furthermore, the memory 1006 of the computing device 1000D may also store instructions for executing the functions of the graph construction module 901, the corpus construction module 903, the training module 905, and the update module 907.

[0191] The connection mode between the computing device clusters shown in Figure 13 may be considered to be based on the fact that the code generation method provided by this application requires timely updating of the business knowledge base. Therefore, it is considered to delegate the functions implemented by the retrieval module 906 to the computing device 1000B.

[0192] It should be understood that the functions of the computing device 1000C shown in FIG13 may also be completed by multiple computing devices 1000. Similarly, the functions of the computing device 1000D may also be completed by multiple computing devices 1000.

[0193] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the above-mentioned code development platform 100 for executing the code generation method.

[0194] The present application also provides a computer program product comprising instructions. The computer program product may be software or a program product comprising instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the aforementioned code generation method.

[0195] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A code generation method, characterized in that, The method includes: The code development platform receives the input information of the user in the first code file; The code development platform extracts the context information of the input information according to the data warehouse to which the first code file belongs, and retrieves the business knowledge base corresponding to the data warehouse according to the input information to obtain the target business knowledge; The code development platform splices the input information, the context information, and the target business knowledge according to the prompt template to obtain the prompt information; The code development platform inputs the prompt information into the large language model LLM for reasoning, and presents the code snippet generated by the LLM reasoning to the user.

2. The method according to claim 1, characterized in that, The method further includes: The code development platform sorts the context information and the target business knowledge according to importance; The code development platform splices the input information, the context information, and the target business knowledge according to the prompt template to obtain the prompt information, including: The code development platform splices the input information, the context information, and the target business knowledge according to the sorting result and the prompt template to obtain the prompt information.

3. The method according to claim 1 or 2, characterized in that, The code development platform extracts the context information of the input information according to the data warehouse to which the first code file belongs, including: The code development platform extracts the context information of the input information according to the dependency graph corresponding to the data warehouse to which the first code file belongs. The dependency graph includes at least one of the dependency relationships of folders in the data warehouse, the dependency relationships of files, the dependencies of files, or the dependency relationships of components. The components include at least one of functions, structures, or macro definitions.

4. The method according to claim 3, characterized in that, The method further includes: The code development platform obtains the full-source code of the data warehouse to which the first code file belongs, the compilation information of the full-source code, or the product design document; The code development platform analyzes the full-source code, the compilation information, or the product design document to obtain dependency information; The code development platform constructs the dependency graph according to the dependency information.

5. The method according to any one of claims 1 to 4, characterized in that The LLM is trained in the following manner: According to the data warehouse and the business knowledge base corresponding to the data warehouse, training corpus is constructed according to the prompt template. The training corpus includes input samples, context samples obtained from the data warehouse, retrieval result samples obtained from the business knowledge base, and code samples in the data warehouse; According to the training corpus, the base model is trained through supervised fine-tuning SFT.

6. The method according to claim 5, wherein The constructing the training corpus according to the data warehouse and the business knowledge base corresponding to the data warehouse according to the prompt template includes: According to the dependency graph corresponding to the data warehouse and the business knowledge base corresponding to the data warehouse, training corpus is constructed according to the prompt template. The dependency graph includes at least one of the dependency relationships of folders in the data warehouse, the dependency relationships of files, the dependencies of files, or the dependency relationships of components. The components include at least one of functions, structures, or macro definitions.

7. The method according to any one of claims 1 to 6, characterized in that The method further includes: The code development platform receives the user's feedback on the code snippets generated by the LLM inference, and the feedback includes acceptance, rejection, or revision of the code snippets.

8. The method according to claim 7, wherein The method further includes: When the feedback is rejection or revision, the code development platform updates the business knowledge base and / or the LLM according to the user's feedback on the code snippets.

9. The method according to any one of claims 1 to 8, characterized in that, The code development platform receives the input information of the user in the first code file, including: The code development platform receives the first code snippet input by the user in the first code file, and the first code snippet is an example code snippet or a code snippet to be completed; or, The code development platform receives the requirement description input by the user in the first code file, and the requirement description is used to describe the second code snippet to be generated.

10. A code development platform, characterized in that, The code development platform includes: An interaction module, configured to receive the input information of the user in the first code file; An extraction module, configured to extract the context information of the input information according to the data warehouse to which the first code file belongs; A retrieval module, configured to retrieve the business knowledge base corresponding to the data warehouse according to the input information to obtain the target business knowledge; A prompt module, configured to splice the input information with the context information and the target business knowledge according to a prompt template to obtain prompt information; An inference module, configured to input the prompt information into a large language model (LLM) for inference, and present the code snippets generated by the LLM inference to the user.

11. The code development platform according to claim 10, wherein The prompt module is further configured to: Sort the context information and the target business knowledge according to importance; Specifically, the prompt module is configured to: According to the sorting result, splice the input information with the context information and the target business knowledge according to the prompt template to obtain prompt information.

12. The code development platform according to claim 10 or 11, characterized in that Specifically, the extraction module is configured to: Extract the context information of the input information according to the dependency graph corresponding to the data warehouse to which the first code file belongs, and the dependency graph includes at least one of the dependency relationships of folders in the data warehouse, the dependency relationships of files, the dependencies of files, or the dependency relationships of components, and the components include at least one of functions, structures, or macro definitions.

13. The code development platform according to claim 12, wherein The code development platform further includes: A graph construction module, configured to obtain the full-source code of the data warehouse to which the first code file belongs, the compilation information of the full-source code, or the product design document, analyze the full-source code, the compilation information, or the product design document to obtain dependency information, and construct the dependency graph according to the dependency information.

14. The code development platform according to any one of claims 10 to 13, characterized in that The code development platform further includes: A corpus construction module, configured to construct training corpus according to the data warehouse and the business knowledge base corresponding to the data warehouse according to the prompt template, and the training corpus includes input samples, context samples obtained from the data warehouse, retrieval result samples obtained from the business knowledge base, and code samples in the data warehouse; A training module, configured to train the base model through supervised fine-tuning (SFT) according to the training corpus.

15. The code development platform according to claim 14, characterized in that, Specifically, the training module is configured to: Construct training corpus according to the dependency graph corresponding to the data warehouse and the business knowledge base corresponding to the data warehouse. The dependency graph includes at least one of the dependencies of folders in the data warehouse, the dependencies of files, the dependencies of file dependencies or components. The components include at least one of functions, structures or macro definitions.

16. The code development platform according to any one of claims 10 to 15, characterized in that The interaction module is further configured to: Receive the user's feedback on the code snippet generated by the LLM inference. The feedback includes acceptance, rejection or revision of the code snippet.

17. The code development platform according to claim 16, wherein The code development platform further includes: An update module, configured to update the business knowledge base and / or the LLM according to the user's feedback on the code snippet when the feedback is rejection or revision.

18. The code development platform according to any one of claims 10 to 17, characterized in that The interaction module is specifically configured to: Receive a first code snippet input by the user in a first code file. The first code snippet is an example code snippet or a code snippet to be completed; or, Receive a requirement description input by the user in the first code file. The requirement description is used to describe a second code snippet to be generated.

19. A cluster of computing devices, characterized in that, The computing device cluster includes at least one computing device. The at least one computing device includes at least one processor and at least one memory. Computer-readable instructions are stored in the at least one memory. The at least one processor executes the computer-readable instructions to cause the computing device cluster to execute the code generation method according to any one of claims 1 to 9.

20. A computer-readable storage medium, characterized in that, Includes computer-readable instructions; the computer-readable instructions are used to implement the code generation method according to any one of claims 1 to 9.

21. A computer program product, characterized in that, Includes computer-readable instructions; the computer-readable instructions are used to implement the code generation method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Code generation method and device

    CN116719520A

  • Program code generation method and device and model training method and device

    CN116841506A

  • Code completion method and device based on generative artificial intelligence, and medium

    CN117111952A

  • Code processing method and system and electronic equipment

    CN117130593A

  • Coding activity task (CAT) evaluation for source code generators

    US20230342116A1

Cited By

  • Target embedded code generation method and device, electronic equipment and storage medium

    CN120704694A

  • Multi-source heterogeneous code conversion method and device for cross-language development

    CN121143797A

  • Large code model training method and electronic equipment

    CN121187569A

  • Method for training a code large model and electronic device

    CN121187569B

  • Method and device for processing computer algorithm

    CN121478390A