Code generation method, device and system
By recalling and analyzing information from multiple data sources in parallel in the code generation system, generating complete context information and inputting it into a large language model, the problem of poor code quality in large-scale modifications and large-scale historical code bases is solved, and higher quality and efficient code generation is achieved.
Patent Information
- Application Number
- CN202510109885.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-01-21
AI Technical Summary
In the prior art, large models are difficult to deal with complex modifications and large-scale historical code bases when generating code, resulting in poor quality of generated code.
By obtaining user demand information, multiple first data are recalled in parallel in at least two databases, a target language file collection is generated, and complete context information is generated through the summary result set, thinking chain CoT analysis and the second data recall, and finally input it into a large language model to generate the target code.
It improves the quality and efficiency of code generation in large-scale models, can handle complex modifications and large-scale historical code bases, and the generated code is more in line with project requirements and code style.
Smart Images

Figure CN120085871A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and more specifically, to a method, apparatus, and system for code generation. Background Art
[0002] In the process of software development, artificial intelligence (AI) large models are often used for AI-assisted development of applications. For example, high-frequency requirements of small and medium-sized applications are used to generate code through large models.
[0003] In the prior art, there are limitations in using large models to generate code. They can only handle simple requirements and small-scale code libraries. When dealing with complex modifications and large-scale historical code libraries, the quality of the generated code is poor.
[0004] In summary, how to improve the quality of code generated by large models is a problem that needs to be solved currently. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a method, apparatus, and system for code generation, which can improve the quality of code generated by large models.
[0006] In a first aspect, an embodiment of the present invention provides a method for code generation, the method comprising:
[0007] Obtain user requirement information;
[0008] Parallelly recall a plurality of first data from at least two databases according to the user requirement information to generate a target language file set, wherein the target language file set includes the plurality of first data;
[0009] Obtain files corresponding to each of the first data in the target language file set to generate a summary result set, wherein the summary result set includes summaries corresponding to each of the files;
[0010] Group the summary result set, perform pre-set chain of thought CoT analysis on each group of summaries, and screen out selected files in each group of summaries to generate a selected file set;
[0011] Perform a second data recall according to each selected file in the selected file set to generate a recall result set, wherein the recall result set includes a plurality of second data;
[0012] Generate complete context information according to the recall result set;
[0013] Input the complete context information into a first large language model to generate target code.
[0014] Optionally, the parallel retrieval of multiple first data in at least two databases according to the user requirement information to generate a target language file set specifically includes:
[0015] Parallelly retrieve multiple first data from the code library, the change request history data set, and the interface library according to the user requirement information, where the first data is a first code snippet, change request data, or implementation class data;
[0016] Form the multiple first data into a target language file set.
[0017] Optionally, the obtaining of the files corresponding to each of the first data in the target language file set to generate a summary result set specifically includes:
[0018] Obtain the files corresponding to each of the first data in the target language file set;
[0019] Parallelly compress each file to generate multiple summaries;
[0020] Form the multiple summaries into the summary result set.
[0021] Optionally, the performing of the preset chain of thought CoT analysis on each group of summaries to screen out the selected files in each group of summaries specifically includes:
[0022] Input each group of summaries into a second large language model, and screen out the selected files in each group of summaries according to the prompt words corresponding to the preset chain of thought CoT analysis in the second large language model.
[0023] Optionally, the preset chain of thought CoT analysis specifically includes:
[0024] One or more of requirement analysis, historical change analysis, call analysis, and plan analysis.
[0025] Optionally, the performing of a second data retrieval according to each selected file in the selected file set to generate a retrieval result set specifically includes:
[0026] Parallelly retrieve multiple second data from the code library for each selected file, where the second data is a second code snippet;
[0027] Form the multiple second data into a retrieval result set.
[0028] Optionally, the generating of the complete context information according to the retrieval result set specifically includes:
[0029] Generate the complete context information according to the retrieval result set and a preset prompt word template.
[0030] Optionally, according to the recall result set and a preset prompt template, generate complete context information, specifically including:
[0031] Input the recall result set and the preset prompt template into a third large language model;
[0032] According to the third large language model, perform requirement analysis, historical analysis, call analysis, incremental analysis, and execution plan formulation on the recall result set and the preset prompt template, and generate complete context information.
[0033] In a second aspect, an embodiment of the present invention provides a method for code generation, the method including:
[0034] Obtain user requirement information;
[0035] Parallelly recall multiple first data in at least two databases according to the user requirement information, and generate a target language file set, where the target language file set includes the multiple first data;
[0036] Obtain the file corresponding to each first data in the target language file set, and generate a summary result set, where the summary result set includes the summary corresponding to each file;
[0037] Group the summary result set, perform a preset chain of thought CoT analysis on each group of summaries, screen out the selected files in each group of summaries, and generate a selected file set;
[0038] Input the selected file set into a first large language model to generate target code.
[0039] In a third aspect, an embodiment of the present invention provides a method for code generation, the method including:
[0040] Obtain user requirement information;
[0041] Sequentially recall multiple first data in at least two databases according to the user requirement information, and generate a target language file set, where the target language file set includes the multiple first data;
[0042] Obtain the file corresponding to each first data in the target language file set, and generate a summary result set, where the summary result set includes the summary corresponding to each file;
[0043] Group the summary result set, perform a preset chain of thought CoT analysis on each group of summaries, screen out the selected files in each group of summaries, and generate a selected file set;
[0044] Perform second data recall for each selected file in the selected file set to generate a recall result set, where the recall result set includes multiple second data;
[0045] Generate complete context information based on the recall result set;
[0046] Input the complete context information into the first large language model to generate target code.
[0047] In a fourth aspect, an embodiment of the present invention provides a code generation device, the device includes:
[0048] An acquisition unit for acquiring user requirement information;
[0049] A recall unit for parallelly recalling multiple first data in at least two databases according to the user requirement information to generate a target language file set, where the target language file set includes the multiple first data;
[0050] A generation unit for obtaining the file corresponding to each first data in the target language file set to generate a summary result set, where the summary result set includes the summary corresponding to each file;
[0051] The generation unit is further configured to group the summary result set, perform a pre-set chain of thought CoT analysis on each group of summaries, screen out the selected files in each group of summaries, and generate a selected file set;
[0052] The recall unit is further configured to perform second data recall for each selected file in the selected file set to generate a recall result set, where the recall result set includes multiple second data;
[0053] The generation unit is further configured to generate complete context information based on the recall result set;
[0054] The generation unit is further configured to input the complete context information into the first large language model to generate target code.
[0055] In a fifth aspect, an embodiment of the present invention provides a code generation device, the device includes:
[0056] An acquisition unit for acquiring user requirement information;
[0057] A recall unit for parallelly recalling multiple first data in at least two databases according to the user requirement information to generate a target language file set, where the target language file set includes the multiple first data;
[0058] A generation unit, configured to obtain files corresponding to each of the first data in the target language file set, and generate a summary result set, where the summary result set includes summaries corresponding to each of the files;
[0059] The generation unit is further configured to group the summary result set, perform a preset Chain of Thought (CoT) analysis on each group of summaries, screen out selected files in each group of summaries, and generate a selected file set;
[0060] The generation unit is further configured to input the selected file set into a first large language model to generate target code.
[0061] In a sixth aspect, an embodiment of the present invention provides a code generation device, where the device includes:
[0062] An acquisition unit, configured to acquire user requirement information;
[0063] A recall unit, configured to serially recall multiple first data in at least two databases in sequence according to the user requirement information, and generate a target language file set, where the target language file set includes the multiple first data;
[0064] A generation unit, configured to obtain files corresponding to each of the first data in the target language file set, and generate a summary result set, where the summary result set includes summaries corresponding to each of the files;
[0065] The generation unit is further configured to group the summary result set, perform a preset Chain of Thought (CoT) analysis on each group of summaries, screen out selected files in each group of summaries, and generate a selected file set;
[0066] The recall unit is further configured to perform a second data recall according to each selected file in the selected file set, and generate a recall result set, where the recall result set includes multiple second data;
[0067] The generation unit is further configured to generate complete context information according to the recall result set;
[0068] The generation unit is further configured to input the complete context information into a first large language model to generate target code.
[0069] In a seventh aspect, an embodiment of the present invention provides a code generation system, where the system includes:
[0070] A terminal side and a cloud server;
[0071] Among them, the terminal side is used to receive user requirement information and send the user requirement information to the cloud server;
[0072] The cloud server is used to execute the method described in any one of the first aspect, any possible one of the first aspect, the second aspect or the third aspect, and send the generated target code to the terminal side.
[0073] In a seventh aspect, an embodiment of the present invention provides an electronic device, including a memory and a processor. The memory is used to store one or more computer program instructions. Among them, the one or more computer program instructions are executed by the processor to implement the method described in any one of the first aspect, any possible one of the first aspect, the second aspect or the third aspect.
[0074] In an eighth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which computer program instructions are stored. The computer program instructions, when executed by a processor, implement the method described in any one of the first aspect, any possible one of the first aspect, the second aspect or the third aspect.
[0075] In the embodiment of the present invention, by obtaining user requirement information; recalling a plurality of first data in parallel in at least two databases according to the user requirement information to generate a target language file set, where the target language file set includes the plurality of first data; obtaining files corresponding to each of the first data in the target language file set to generate a summary result set, where the summary result set includes summaries corresponding to each of the files; grouping the summary result set, performing a pre-set chain of thought CoT analysis on each group of summaries, screening out selected files in each group of summaries to generate a selected file set; recalling second data according to each selected file in the selected file set to generate a recall result set, where the recall result set includes a plurality of second data; generating complete context information according to the recall result set; inputting the complete context information into a first large language model to generate target code. Through the above method, the quality of the code generated by the large model can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Through the following description of the embodiments of the present invention with reference to the drawings, the above and other objects, features and advantages of the present invention will become clearer. In the drawings:
[0077] Figure 1 is a flowchart of a method for code generation in an embodiment of the present invention;
[0078] Figure 2 is a flowchart of a method for code generation in an embodiment of the present invention;
[0079] Figure 3 It is a flowchart of another method for code generation in an embodiment of the present invention;
[0080] Figure 4 It is a flowchart of yet another method for code generation in an embodiment of the present invention;
[0081] Figure 5 It is a flowchart of still another method for code generation in an embodiment of the present invention;
[0082] Figure 6 It is a flowchart of a method for code generation in an embodiment of the present invention;
[0083] Figure 7 It is a flowchart of yet another method for code generation in an embodiment of the present invention;
[0084] Figure 8 It is a schematic diagram of a device for code generation in an embodiment of the present invention;
[0085] Figure 9 It is a schematic diagram of an electronic device in an embodiment of the present invention. Detailed implementation manners
[0086] The following describes the present application based on embodiments, but the present application is not limited to these embodiments. In the following detailed description of the present application, some specific details are described in detail. Those skilled in the art can fully understand the present application without the description of these details. In order to avoid obscuring the essence of the present application, well-known methods, processes, flows, elements, and circuits are not described in detail.
[0087] In addition, those of ordinary skill in the art should understand that the drawings provided herein are for illustrative purposes only, and the drawings are not necessarily drawn to scale.
[0088] Unless the context clearly requires otherwise, the words such as "including" and "comprising" in the entire application document should be interpreted as the meaning of including rather than exclusive or exhaustive; that is, it is the meaning of "including but not limited to".
[0089] In the description of the present application, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. In addition, in the description of the present application, unless otherwise stated, the meaning of "a plurality of" is two or more.
[0090] In the prior art, for medium and small-sized applications, due to their small scale and few developers, an AI-assisted solution is required to develop the above-mentioned medium and small-sized applications to meet user requirements. The above AI-assisted solution refers to the process of using artificial intelligence technology to assist software development. For example, a large model is used to generate code that meets user requirements. However, there are limitations when using a large model to generate code. For example, in the current AI requirement implementation solutions such as RAG code refactoring and RAG front-end page code generation, only simple requirements and small-scale code libraries can be processed during processing. When dealing with complex modifications and large-scale historical code libraries, the quality of the generated code is poor. Therefore, how to improve the quality of the code generated by the large model is a problem that needs to be solved currently.
[0091] In the embodiments of the present invention, the large model can also be referred to as an Artificial Intelligence (AI) model or a Large Language Model (LLM). Among them, the large language model is a deep learning model based on a transformer architecture, capable of processing and generating natural language text. It is usually trained on a large amount of text data and has the ability to understand and generate language, and is widely used in dialogue systems, text generation, and other natural language processing tasks.
[0092] In the embodiments of the present invention, to solve the above problems, a method for code generation is proposed, specifically as Figure 1 shown, the method includes:
[0093] Step S101, obtain user requirement information.
[0094] Specifically, the user can input the requirement information into the front-end interface of a pre-designed code generation system. For example, the user's requirement information is "find all the projects submitted by a user within one year". After the user inputs the above requirement information, the code generation system obtains the user's requirement information. The requirement information here is only for illustrative purposes and is determined according to the actual situation.
[0095] Step S102, recall multiple first data in parallel in at least two databases according to the user requirement information to generate a target language file set.
[0096] Among them, the target language file set includes multiple first data.
[0097] In a possible implementation, the first data is the first code snippet, change request data, or implementation class data. The recall of the first data according to the user requirement information to generate a target language file set specifically includes: recalling a plurality of first data from a code library, a change request history data set, and an interface library according to the user requirement information, and forming the target language file set with the plurality of first data.
[0098] In a possible implementation, the target language may be a Java language, a TypeScript language, a Kotlin language, etc., which is specifically determined according to the actual situation.
[0099] Specifically, as Figure 2 shown, after the code generation system receives the requirement information sent by the user, it performs the first data recall in parallel, recalls the first code snippet in the code library, recalls the change request data in the change request (CR) history data set, and recalls the implementation class data in the interface library respectively.
[0100] Specifically, when recalling the first code snippet in the code library, first determine the encoding of the code library, and then retrieve the code snippet related to the requirement information in the code library corresponding to the encoding as the above-mentioned first code snippet; when recalling the change request data in the CR history data set, first determine the encoding of the code library, then determine the CR history data set in the code library corresponding to the encoding, and then perform a similarity scoring service on the historical data in the CR history data set, and use the top 5 code review historical data as the change request data. It is also possible to select the top 3, top 8, or top 10 code review historical data as the change request data, which is specifically determined according to the actual situation; when recalling the implementation class data in the interface library, first determine the interface library in the code platform, and then retrieve the implementation class in the interface library; the above-mentioned code generation system integrates the retrieved first code snippet, change request data, or implementation (impl) class data to generate a Java file set, where the impl class refers to the specific class for implementing a certain interface or abstract class.
[0101] In the embodiment of the present invention, both the code library and the interface library are stored in the code platform. Through the above multi-dimensional parallel recall, the project structure and history can be comprehensively understood, the recall speed is improved, and relatively rich information is provided for the subsequent steps.
[0102] Step S103: Obtain the files corresponding to each of the first data in the target language file set, and generate a summary result set.
[0103] Among them, the abstract result set includes the abstract corresponding to each of the said documents.
[0104] Specifically, obtain the document corresponding to each of the first data in the target language document set; compress each document in parallel to generate multiple abstracts; and form the multiple abstracts into the abstract result set.
[0105] For example, as Figure 3 described, after obtaining the multiple documents corresponding to the target language document set, the multiple documents are processed in parallel by multiple threads in a thread pool, and each thread extracts the abstract of one document to obtain the abstract, that is, obtain the key structure information of each document; assume that the target language is the Java language and the target language document set is a set of Java files. Thread 1 extracts the abstract of Java file 1 to generate abstract 1; Thread 2 extracts the abstract of Java file 2 to generate abstract 2; Thread 3 extracts the abstract of Java file 3 to generate abstract 3; and so on. Thread n extracts the abstract of Java file n to generate abstract n; merge the above abstract 1, abstract 2, abstract 3, and abstract n to generate the abstract result set; the key structure information includes information such as class signatures, member variables, methods, and inner static classes. Through the above method, the processing time can be reduced, more structured and refined information can be provided, which helps the subsequent large model to more accurately understand the file structure and function.
[0106] Step S104: Group the abstract result set, perform a preset Chain of Thought (CoT) analysis on each group of abstracts, and screen out the selected files in each group of abstracts to generate a selected file set.
[0107] In a possible implementation manner, the preset Chain of Thought (CoT) analysis on each group of abstracts to screen out the selected files in each group of abstracts specifically includes: inputting each group of abstracts into a second large language model, and screening out the selected files in each group of abstracts according to the prompt words corresponding to the preset Chain of Thought CoT analysis in the second large language model. The preset Chain of Thought CoT analysis specifically includes one or more of requirement analysis, historical change analysis, call analysis, and plan analysis. The CoT analysis is a prompt word set by the user to guide the large prediction model to think about the existing code, and the specific setting is determined according to the actual situation.
[0108] Specifically, such as Figure 4As shown, the files in the abstract result set are divided into multiple groups. Suppose there are 30 abstract files in the abstract result set, and every 10 abstract files are divided into one group, then a total of three groups are formed. The above three groups of abstract files are processed in parallel, and each group of abstract files (which can also be called each group of abstracts) undergoes chain-of-thought analysis. The grouping here is only for illustrative purposes, and the grouping is specifically determined according to the actual situation; the CoT analysis includes requirement analysis, historical change analysis, call analysis, and plan analysis. After the above CoT analysis, the available selected files are selected from each group of abstracts, and the selected files in each group are combined to generate a set of selected files; that is, group 1 undergoes COT analysis to obtain the selected files in group 1; group 2 undergoes COT analysis to obtain the selected files in group 2; group 3 undergoes COT analysis to obtain the selected files in group 3; the selected files in group 1, the selected files in group 2, and the selected files in group 3 are combined to generate a set of selected files.
[0109] Through the above parallel CoT analysis, while improving the processing speed, it also significantly improves the accuracy of abstract file selection through fine-grained analysis, effectively solving the selection bias problem in large-scale file processing and ensuring that no potentially useful abstract files are missed.
[0110] Step S105: Input the set of selected files into the first large language model to generate target code.
[0111] In the embodiment of the present invention, through the above method, target code can be quickly generated, and the generation speed and quality of the target code are improved through parallel recall. On this basis, the quality of the target code can be further improved, specifically as Figure 5 shown, after step S104, it specifically includes:
[0112] Step S106: Perform a second data recall according to each selected file in the set of selected files to generate a set of recall results.
[0113] Among them, the set of recall results includes multiple pieces of second data.
[0114] Specifically, as Figure 6As shown in the figure, for multiple selected files in the selected file set, the multiple selected files are processed in parallel by multiple threads in the thread pool. Each thread recalls multiple second data for one selected file in the code library. Suppose thread 1 recalls code in the code library for selected file 1 and generates at least one code snippet 1; thread 2 recalls code in the code library for selected file 2 and generates at least one code snippet 2; thread 3 recalls code in the code library for selected file 3 and generates at least one code snippet 3; and so on. Thread n recalls code in the code library for selected file n and generates at least one code snippet n. The above code snippets 1, code snippet 2, code snippet 3, and code snippet n are merged to generate a recall result set. Among them, the second data is a second code snippet. The multiple second data are combined into a recall result set, and the recall results of each selected file are aggregated to generate a recall result set.
[0115] Since there is a limit on the number of recalled segments when recalling the first data, the number of recalled segments for each file is small. After file screening, only the selected files are left for recalling the second data. Although the number limit is the same, the number of segments allocated to each selected file is greater than the number of segments when recalling the first data, which not only significantly reduces the recall time but also provides richer and more relevant context information for each selected file.
[0116] Step S107: Generate complete context information according to the recall result set.
[0117] Specifically, complete context information is generated according to the recall result set and a pre-set prompt template.
[0118] In a possible implementation, generating complete context information according to the recall result set and a pre-set prompt template specifically includes: inputting the recall result set and the pre-set prompt template into a third large language model; performing requirement analysis, historical analysis, call analysis, incremental analysis, and execution plan formulation on the recall result set and the pre-set prompt template by the third large language model. Specifically, requirement analysis results, historical analysis results, call analysis results, and incremental analysis results are respectively obtained, and finally complete context information is generated through execution plan formulation. Among them, the above historical analysis and call analysis can be executed in parallel to improve efficiency.
[0119] In the embodiment of the present invention, through the above construction process of context file information, a comprehensive and coherent context information can be created, enabling the AI model to better understand requirements, project history, and code structure, thereby generating higher-quality code that better conforms to the project style.
[0120] Step S108: Input the complete context information into the first large language model to generate target code.
[0121] In the embodiments of the present invention, each of the above steps can be regarded as steps that different modules in the code generation system need to execute. Modularizing the above code generation system can make the above code generation system easy to adjust and expand for different projects or requirements.
[0122] Through the above embodiments, a multi-dimensional parallel recall mechanism is respectively implemented, improving the comprehensiveness and speed of recall; parallel summary extraction of Java files, using a thread pool to process multiple Java files in parallel, extracting key structure information to generate a summary, which not only retains complete information but also reduces context occupancy; grouped parallel file selection: grouping files to perform CoT analysis in parallel, including requirement analysis, historical change analysis, etc., improving the accuracy of large-scale file processing; parallel secondary recall, performing separate code recall on each selected file in parallel, enriching context information; structured context construction, designing a structured context construction process including steps such as requirement analysis and historical analysis, with some steps using parallel processing, improving processing efficiency; specialized optimization for small and medium-sized applications, focusing on high-frequency requirement scenarios of small and medium-sized applications, and using their characteristics for targeted optimization; effective utilization of historical CR data, integrating historical code review data into the solution, improving the relevance and accuracy of the generated code; In summary, through the above process and corresponding advantages, an efficient and accurate AI-assisted development system is jointly constructed, improving the quality and efficiency of the generated code, applicable to high-frequency requirement implementation scenarios of small and medium-sized applications, and also applicable to large-scale and complex projects.
[0123] In a possible implementation manner, when recalling the first data in sequence serially, although the recall speed will be reduced, it will not affect the quality of the generated target code. The specific process is as Figure 7 shown, including the following steps:
[0124] Step S701: Obtain user requirement information.
[0125] Step S702: According to the user requirement information, recall multiple first data in sequence serially in at least two databases to generate a target language file set.
[0126] Among them, the target language file set includes the multiple first data.
[0127] Specifically, recall the first code snippet, change request data, or implementation class data in the code library, change request history data set, and interface library in sequence according to the user requirement information. Among them, the order of the code library, change request history data set, and interface library can be changed arbitrarily, and is specifically determined according to the actual situation.
[0128] Step S703: Obtain the files corresponding to each of the first data in the target language file set, and generate a summary result set.
[0129] Wherein, the summary result set includes the summaries corresponding to each of the files.
[0130] Step S704: Group the summary result set, perform a preset Chain of Thought (CoT) analysis on each group of summaries, and screen out the selected files in each group of summaries to generate a selected file set.
[0131] Step S705: Perform a second data recall based on each selected file in the selected file set to generate a recall result set, where the recall result set includes a plurality of second data.
[0132] Step S706: Generate complete context information based on the recall result set.
[0133] Step S707: Input the complete context information into a first large language model to generate target code.
[0134] In an embodiment of the present invention, a code generation system is designed, and the system includes: a terminal side and a cloud server;
[0135] Wherein, the terminal side is used to receive user requirement information and send the user requirement information to the cloud server; the cloud server is used to execute the method described in any one of the above embodiments and send the generated target code to the terminal side.
[0136] In an embodiment of the present invention, a code generation device is provided, as Figure 8 shown, and specifically includes: an acquisition unit 801, a recall unit 802, and a generation unit 803;
[0137] Among them, the obtaining unit 801 is used to obtain user requirement information; the recall unit 802 is used to recall multiple first data in at least two databases in parallel according to the user requirement information, and generate a target language file set, where the target language file set includes the multiple first data; the generating unit 803 is used to obtain the file corresponding to each first data in the target language file set, and generate a summary result set, where the summary result set includes the summary corresponding to each file; the generating unit 803 is further used to group the summary result set, perform a preset Chain of Thought (CoT) analysis on each group of summaries, screen out the selected files in each group of summaries, and generate a selected file set; the recall unit 802 is further used to recall second data according to each selected file in the selected file set, and generate a recall result set, where the recall result set includes multiple second data; the generating unit 803 is further used to generate complete context information according to the recall result set; the generating unit 803 is further used to input the complete context information into a first large language model to generate target code.
[0138] Further, the recall unit is specifically used for:
[0139] Recall multiple first data in parallel from a code library, a change request history data set, and an interface library according to the user requirement information, where the first data is a first code snippet, change request data, or implementation class data;
[0140] Form the multiple first data into a target language file set.
[0141] Further, the generating unit is specifically used for:
[0142] Obtain the file corresponding to each first data in the target language file set;
[0143] Compress each file in parallel to generate multiple summaries;
[0144] Form the multiple summaries into the summary result set.
[0145] Further, the generating unit is specifically used for:
[0146] Input each group of summaries into a second large language model, and screen out the selected files in each group of summaries according to the prompt words corresponding to the preset Chain of Thought (CoT) analysis in the second large language model.
[0147] Further, the preset Chain of Thought (CoT) analysis specifically includes:
[0148] One or more of requirements analysis, historical change analysis, call analysis, and plan analysis.
[0149] Further, the recall unit is specifically configured to:
[0150] Parallelly recall multiple second data in the code library for each selected file, where the second data is a second code snippet;
[0151] Form the multiple second data into a recall result set.
[0152] Further, the generation unit is specifically configured to:
[0153] Generate complete context information according to the recall result set and a preset prompt template.
[0154] Further, the generation unit is specifically configured to:
[0155] Input the recall result set and a preset prompt template into a third large language model;
[0156] Perform requirements analysis, historical analysis, call analysis, incremental analysis, and execution plan formulation on the recall result set and the preset prompt template according to the third large language model, and generate complete context information.
[0157] Figure 9 It is a schematic structural diagram of the electronic device in the embodiment of the present invention. As Figure 9 shown, it includes a general computer hardware structure, which at least includes a processor 901 and a memory 902. The processor 901 and the memory 902 are connected through a bus 903. The memory 902 is adapted to store instructions or programs executable by the processor 901. The processor 901 can be an independent microprocessor or a set of one or more microprocessors. Thus, the processor 901 executes the instructions stored in the memory 902, thereby executing the method flow of the embodiment of the present invention as described above to implement data processing and control of other devices. The bus 903 connects the above-mentioned multiple components together, and at the same time connects the above-mentioned components to a display controller 904, a display device, and an input / output (I / O) device 905. The input / output (I / O) device 905 can be a mouse, a keyboard, a modem, a network interface, a touch input device, a somatosensory input device, a printer, and other devices well known in the art. Typically, the input / output (I / O) device 905 is connected to the system through an input / output (I / O) controller 906.
[0158] Among them, the instructions stored in the memory 902 are executed by at least one processor 901 to achieve: obtaining user requirement information; parallelly retrieving multiple first data from at least two databases according to the user requirement information to generate a target language file set, where the target language file set includes the multiple first data; obtaining the files corresponding to each of the first data in the target language file set to generate a summary result set, where the summary result set includes the summary corresponding to each file; grouping the summary result set, performing a preset chain of thought CoT analysis on each group of summaries, screening out the selected files in each group of summaries to generate a selected file set; retrieving second data according to each selected file in the selected file set to generate a retrieval result set, where the retrieval result set includes multiple second data; generating complete context information according to the retrieval result set; inputting the complete context information into a first large language model to generate target code.
[0159] Specifically, the electronic device includes: one or more processors 901 and a memory 902. Figure 9 Taking one processor 901 as an example. The processor 901 and the memory 902 can be connected through a bus or other means. Figure 9 Taking the connection through the bus as an example. The memory 902, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The processor 901 executes various functional applications and data processing of the device by running the non-volatile software programs, instructions, and modules stored in the memory 902, that is, implementing the method for determining code generation described above.
[0160] The memory 902 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store option lists, etc. In addition, the memory 902 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 902 can optionally include a memory remotely set relative to the processor 901, and these remote memories can be connected to external devices through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0161] One or more modules are stored in the memory 902 and, when executed by one or more processors 901, execute the method for code generation in any of the above method embodiments.
[0162] As those skilled in the art will realize, various aspects of the embodiments of the present invention can be implemented as a system, a method, or a computer program product. Accordingly, various aspects of the embodiments of the present invention may take the form of: a full hardware implementation, a full software implementation (including firmware, resident software, microcode, etc.), or an implementation combining software aspects and hardware aspects that are generally referred to herein as "circuits", "modules", or "systems". In addition, various aspects of the embodiments of the present invention may take the form of a computer program product implemented in one or more computer-readable media having computer-readable program code embodied thereon.
[0163] Any combination of one or more computer-readable media may be utilized. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium would include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of the embodiments of the present invention, a computer-readable storage medium may be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0164] A computer-readable signal medium may include a propagated digital signal having computer-readable program code embodied therein, such as in baseband or as part of a carrier wave. Such a propagated signal may take any of a variety of forms, including but not limited to: electromagnetic, optical, or any suitable combination thereof. A computer-readable signal medium may be any computer-readable medium that is not a computer-readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0165] Any suitable medium may be used to transmit the program code embodied on the computer-readable medium, including but not limited to wireless, wireline, fiber optic cable, RF, etc., or any suitable combination of the foregoing.
[0166] Computer program code for performing operations for aspects of the embodiments of the present invention may be written in any combination of one or more programming languages, including: object-oriented programming languages such as Java, Smalltalk, C++; and conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code may execute entirely on the user's computer as a stand-alone software package, partly on the user's computer, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0167] The flowcharts and / or block diagrams of the methods, apparatus (systems) and computer program products according to the embodiments of the present invention described above depict various aspects of the embodiments of the present invention. It will be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general purpose computer, a special purpose computer, or other programmable data processing device to produce a machine, such that the instructions, executed by the processor of the computer or other programmable data processing device, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0168] These computer program instructions may also be stored in a computer-readable medium that can direct a computer, other programmable data processing device, or other apparatus to operate in a particular manner, such that the instructions stored in the computer-readable medium produce an article of manufacture including instructions for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0169] The computer program instructions may also be loaded onto a computer, other programmable data processing device, or other apparatus to cause a series of operational steps to be performed on the computer, other programmable device, or other apparatus to produce a computer-implemented process, such that the instructions executed on the computer or other programmable device provide a process for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0170] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
[0171] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse. If a user refuses to process personal information other than the necessary information required for basic functions, it will not affect the user's use of basic functions.
Claims
1. A method for code generation, characterized in that: The method comprises: Obtain information about user needs; Recalling a plurality of first data in parallel in at least two databases according to the user demand information to generate a target language file set, wherein the target language file set includes the plurality of first data; Acquire a file corresponding to each of the first data in the target language file set, and generate a summary result set, wherein the summary result set includes a summary corresponding to each of the files; Grouping the summary result set, performing a preset thought chain CoT analysis on each group of summaries, screening out selected files in each group of summaries, and generating a selected file set; Recalling the second data according to each selected file in the selected file set to generate a recall result set, wherein the recall result set includes a plurality of second data; generating complete context information according to the recall result set; The complete context information is input into a first large language model to generate a target code.
2. The method according to claim 1, characterized in that: The step of recalling a plurality of first data in parallel in at least two databases according to the user demand information to generate a target language file set specifically includes: Recalling a plurality of first data in parallel from a code library, a change request history data set, and an interface library according to the user demand information, wherein the first data is a first code snippet, change request data, or implementation class data; The plurality of first data are combined into a target language file set.
3. The method according to claim 1, characterized in that The acquiring of the file corresponding to each of the first data in the target language file set and generating a summary result set specifically includes: Obtaining a file corresponding to each first data in the target language file set; Compress each file in parallel to generate multiple summaries; The multiple summaries are combined into the summary result set.
4. The method according to claim 1, characterized in that: The performing of a preset CoT analysis on each set of abstracts to filter out the selected files in each set of abstracts specifically includes: Input each group of abstracts into the second large language model, and filter out the selected files in each group of abstracts according to the prompt words corresponding to the thought chain CoT analysis preset in the second large language model.
5. The method according to claim 1, characterized in that The preset thinking chain CoT analysis specifically includes: One or more of requirements analysis, historical change analysis, call analysis, and plan analysis.
6. The method according to claim 1, characterized in that The performing second data recall according to each selected file in the selected file set to generate a recall result set specifically includes: Recalling a plurality of second data in the code base for each selected file in parallel, wherein the second data is a second code fragment; The multiple second data are combined into a recall result set.
7. The method according to claim 1, characterized in that Generating complete context information according to the recall result set specifically includes: Complete context information is generated according to the recall result set and a preset prompt word template.
8. The method according to claim 7, characterized in that Generate complete context information based on the recall result set and the preset prompt word template, including: Inputting the recall result set and the preset prompt word template into a third large language model; According to the third large language model, demand analysis, history analysis, call analysis, incremental analysis and execution plan formulation are performed on the recall result set and the preset prompt word template to generate complete context information.
9. A method for code generation, characterized in that: The method comprises: Obtain information about user needs; Recalling a plurality of first data in parallel in at least two databases according to the user demand information to generate a target language file set, wherein the target language file set includes the plurality of first data; Acquire a file corresponding to each of the first data in the target language file set, and generate a summary result set, wherein the summary result set includes a summary corresponding to each of the files; Grouping the summary result set, performing a preset thought chain CoT analysis on each group of summaries, screening out selected files in each group of summaries, and generating a selected file set; The selected file set is input into a first large language model to generate a target code.
10. A method for code generation, characterized in that: The method comprises: Obtain information about user needs; Recalling a plurality of first data in sequence and serially in at least two databases according to the user demand information to generate a target language file set, wherein the target language file set includes the plurality of first data; Acquire a file corresponding to each of the first data in the target language file set, and generate a summary result set, wherein the summary result set includes a summary corresponding to each of the files; Grouping the summary result set, performing a preset thought chain CoT analysis on each group of summaries, screening out selected files in each group of summaries, and generating a selected file set; Recalling the second data according to each selected file in the selected file set to generate a recall result set, wherein the recall result set includes a plurality of second data; generating complete context information according to the recall result set; The complete context information is input into a first large language model to generate a target code.
11. A device for code generation, characterized in that: The device comprises: An acquisition unit, used for acquiring user demand information; A recall unit, configured to recall a plurality of first data in parallel from at least two databases according to the user demand information to generate a target language file set, wherein the target language file set includes the plurality of first data; A generating unit, configured to obtain a file corresponding to each of the first data in the target language file set, and generate a summary result set, wherein the summary result set includes a summary corresponding to each of the files; The generating unit is further used to group the summary result set, perform a preset thought chain CoT analysis on each group of summaries, filter out selected files in each group of summaries, and generate a selected file set; The recall unit is further used to recall the second data according to each selected file in the selected file set to generate a recall result set, wherein the recall result set includes a plurality of second data; The generating unit is further used to generate complete context information according to the recall result set; The generating unit is further configured to input the complete context information into a first large language model to generate a target code.
12. A device for code generation, characterized in that: The device comprises: An acquisition unit, used for acquiring user demand information; A recall unit, configured to recall a plurality of first data in parallel from at least two databases according to the user demand information to generate a target language file set, wherein the target language file set includes the plurality of first data; A generating unit, configured to obtain a file corresponding to each of the first data in the target language file set, and generate a summary result set, wherein the summary result set includes a summary corresponding to each of the files; The generating unit is further used to group the summary result set, perform a preset thought chain CoT analysis on each group of summaries, filter out selected files in each group of summaries, and generate a selected file set; The generating unit is further configured to input the selected file set into a first large language model to generate a target code.
13. A code generation device, characterized in that: The device comprises: An acquisition unit, used for acquiring user demand information; A recall unit, configured to recall a plurality of first data in sequence and serially from at least two databases according to the user demand information to generate a target language file set, wherein the target language file set includes the plurality of first data; A generating unit, configured to obtain a file corresponding to each of the first data in the target language file set, and generate a summary result set, wherein the summary result set includes a summary corresponding to each of the files; The generating unit is further used to group the summary result set, perform a preset thought chain CoT analysis on each group of summaries, filter out selected files in each group of summaries, and generate a selected file set; The recall unit is further used to recall the second data according to each selected file in the selected file set to generate a recall result set, wherein the recall result set includes a plurality of second data; The generating unit is further used to generate complete context information according to the recall result set; The generating unit is further configured to input the complete context information into a first large language model to generate a target code.
14. A code generation system, characterized in that: The system comprises: Terminal side and cloud server; Wherein, the terminal side is used to receive user demand information and send the user demand information to the cloud server; The cloud server is used to execute the method described in any one of claims 1 to 10 above, and send the generated target code to the terminal side.
Citation Information
Patent Citations
Method and system for automatically generating reusable API based on code snippets
CN117892031A
Code generation method and device based on cloud service
CN118092923A
Method and system for supporting graphical creative programming in classroom environment
CN118860362A
Language model loosely coupled cascade unit test case generation method and device
CN119065988A
Code generation method and device, equipment and storage medium
CN119271211A