Software code generation method based on retrieval enhancement generation and large model fine tuning
By constructing a Q&A dataset and low-rank adaptation technology, and combining Lora technology, the problem of high error rate and low development efficiency in generating Abaqus Python code in large language models is solved, and the accuracy and reliability of code generation is improved.
Patent Information
- Application Number
- CN202510398207.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-11
AI Technical Summary
When generating Abaqus Python code in the prior art, large language models have problems such as high code error rate and still cannot run after multiple iterations. They lack in-depth understanding of the Abaqus API, resulting in low development efficiency.
Using a method based on search enhancement generation and large-model fine-tuning, a large language model is fine-tuned by building a question-and-answer dataset and low-rank adaptation technology, and combining Lora technology, a search enhancement generation process is designed to generate comprehensive prompt words to improve the accuracy of code generation.
It significantly improves the accuracy of code generation, reduces the cost of iterative modifications for users, enhances the model's understanding of the Abaqus API, and improves the reliability of code generation.
Smart Images

Figure CN120295609A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of code generation, and more precisely, it relates to a software code generation method based on retrieval-augmented generation and large model fine-tuning. Background Art
[0002] With the rapid development of computer simulation technology, Abaqus, as a simulation software widely used in the engineering field, its Python API provides powerful customization functions for developers. However, when directly generating Abaqus Python code using large language models (such as GPT, etc.), there are often problems such as a high code error rate and the inability to run after multiple iterations. This is mainly because large language models lack sufficient domain knowledge and context information when dealing with code generation tasks in specific domains. Summary of the Invention
[0003] The object of the present invention is to address the deficiencies of the prior art and propose a software code generation method based on retrieval-augmented generation and large model fine-tuning.
[0004] In a first aspect, there is provided a software code generation method based on retrieval-augmented generation and large model fine-tuning, including:
[0005] S1. Construct a question-answer pair data set;
[0006] S2. Based on the question-answer pair data set, fine-tune a pre-trained large language model using the low-rank adaptation technique;
[0007] S3. Design a retrieval-augmented generation process;
[0008] S4. Generate a comprehensive prompt based on the retrieval-augmented generation result.
[0009] Preferably, in S1, the construction of the question-answer pair data set includes: parsing natural language question descriptions and their corresponding script codes from the script use case documents of the target simulation software to generate structured question-script pairs.
[0010] Preferably, in S2, the fine-tuning of the pre-trained large language model using the low-rank adaptation technique includes: using the question-answer pair data set to adjust the parameters of the large language model through the low-rank adaptation technique so that it adapts to the API and simulation tasks of the target simulation software.
[0011] Preferably, in S3, the design of the retrieval-augmented generation process includes:
[0012] According to the natural language query input by the user, retrieve script use cases in the question-answer pair data set that are semantically matched with the query, and extract the associated API documents in the script use cases.
[0013] Preferably, in S4, generating the comprehensive prompt includes:
[0014] Integrate the user query, retrieved script cases, and associated API documents into a prompt and input it into the fine-tuned large language model to generate the executable code of the target simulation software.
[0015] In a second aspect, a software code generation system based on retrieval-augmented generation and large model fine-tuning is provided for performing any of the methods described in the first aspect, including:
[0016] A construction module for constructing a Q&A pair dataset;
[0017] A fine-tuning module for fine-tuning the pre-trained large language model based on the Q&A pair dataset using low-rank adaptation technology;
[0018] A design module for designing a retrieval-augmented generation process;
[0019] A generation module for generating a comprehensive prompt based on the retrieval-augmented generation result.
[0020] In a third aspect, a computer storage medium is provided, in which a computer program is stored; when the computer program runs on a computer, the computer is enabled to execute any of the methods described in the first aspect.
[0021] In a fourth aspect, an electronic device is provided, including:
[0022] A memory for storing the computer program;
[0023] A processor for executing the computer program to implement any of the methods described in the first aspect.
[0024] The beneficial effects of the present invention are:
[0025] 1. The present invention can improve the accuracy of code generation: By RAG and Few-Shot learning, combined with the Lora fine-tuning technology, the accuracy of the large language model in generating Abaqus code is significantly improved.
[0026] 2. The present invention can reduce the iteration cost: The improved accuracy of the generated code reduces the cost for users to iteratively modify the code multiple times.
[0027] 3. The present invention can enhance domain knowledge: By retrieving relevant Abaqus script cases and API documents, the understanding of the Abaqus API by the model is enhanced, and the reliability of code generation is improved. Description of the Drawings
[0028] Figure 1Flowchart of the software code generation method based on retrieval-augmented generation and large model fine-tuning provided by this application. Detailed implementation manners
[0029] The present invention will be further described below in conjunction with embodiments. The description of the following embodiments is only used to help understand the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several modifications can still be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
[0030] Embodiment 1:
[0031] The existing technical solutions mainly rely on large language models to directly generate Abaqus Python code. The user inputs a simulation task described in natural language, and the large language model generates the corresponding Python code according to the input. The advantage of this method is that it can generate code quickly, but since the large language model lacks an in-depth understanding of the Abaqus API, the generated code often has syntax errors or logical errors, resulting in the code being unable to run.
[0032] Specifically, the code generated by the large language model has a high error rate. Especially when dealing with complex simulation tasks, the accuracy of the code is difficult to guarantee. Moreover, due to code generation errors, the user needs to iterate and modify the code multiple times, resulting in low development efficiency. In addition, when generating code, the large language model lacks an in-depth understanding of the Abaqus API and cannot make full use of the API documentation and script use cases.
[0033] Another existing technology is to generate Abaqus Python code by fine-tuning the large language model. Specifically, developers will use a large number of Abaqus scripts and API documents to fine-tune the large language model so that it can better understand the Abaqus API and simulation tasks.
[0034] However, fine-tuning the large language model requires a large amount of training data and computing resources, and the training cost is relatively high. And although the fine-tuned model performs well on specific tasks, its generalization ability is poor when dealing with unseen tasks.
[0035] To solve the problems of the existing technology, Embodiment 1 of this application provides a method based on Retrieval-Augmented Generation (RAG) and Few-Shot learning, combined with the Low-Rank Adaptation (Lora) fine-tuning technology, to improve the accuracy and reliability of the large language model when generating Abaqus Python code.
[0036] Specifically, as Figure 1As shown in the figure, a software code generation method based on retrieval-augmented generation and large model fine-tuning includes:
[0037] S1. Construct a question-answer pair dataset.
[0038] In S1, the construction of the question-answer pair dataset includes: parsing natural language question descriptions and their corresponding script codes from the script use case documents of the target simulation software to generate structured question-script pairs.
[0039] Exemplarily, question-answer pairs of questions-scripts are extracted from the script use case documents of Abaqus for subsequent Few-Shot learning and RAG retrieval.
[0040] For example, the user's question may be "How to create a simple beam model?", and then obtain the corresponding answer script.
[0041] In this way, a large number of question-answer pairs are generated for subsequent Few-Shot learning and RAG retrieval.
[0042] S2. Fine-tune the pre-trained large language model based on the question-answer pair dataset using low-rank adaptation technology.
[0043] In S2, the fine-tuning of the pre-trained large language model based on low-rank adaptation technology includes: using the question-answer pair dataset to adjust the parameters of the large language model through low-rank adaptation technology so that it can better understand the API and simulation tasks of Abaqus.
[0044] The specific steps are as follows:
[0045] Prepare training data: Take the extracted question-answer pairs as training data and input them into the large language model.
[0046] Fine-tune the model: Use the Lora technology to fine-tune the large language model so that it can better understand the API and simulation tasks of Abaqus.
[0047] S3. Design a retrieval-augmented generation process.
[0048] S4. Generate a comprehensive prompt according to the retrieval-augmented generation result.
[0049] Example 2:
[0050] On the basis of Example 1, Example 2 of this application provides a more specific software code generation method based on retrieval-augmented generation and large model fine-tuning, including:
[0051] S1. Construct a question-answer pair dataset.
[0052] S2. Based on the Q&A pair dataset, fine-tune the pre-trained large language model using low-rank adaptation technology.
[0053] S3. Design a retrieval-augmented generation process.
[0054] In S3, the design of the retrieval-augmented generation process includes:
[0055] Retrieve script cases semantically matching the query from the Q&A pair dataset according to the natural language query input by the user, and extract the associated API documents in the script cases.
[0056] Exemplarily, when the user inputs a query, retrieve the top K most similar Abaqus script cases through RAG, and extract the relevant Abaqus Python API documents from these script cases.
[0057] For example, when the user inputs a query, retrieve the top K most similar Abaqus script cases through RAG. The specific steps are as follows:
[0058] User inputs a query: The user inputs a natural language description of the simulation task, such as "How to create a simple beam model?".
[0059] Retrieve top K script cases: Retrieve the top K Abaqus script cases most similar to the user query through vector retrieval technology (such as FAISS).
[0060] Extract relevant API documents: Extract the relevant Abaqus Python API documents from the retrieved script cases.
[0061] S4. Generate a comprehensive prompt based on the retrieval-augmented generation result.
[0062] In S4, the generation of the comprehensive prompt includes:
[0063] Integrate the user query, the retrieved script cases, and the associated API documents into a prompt and input it into the fine-tuned large language model to generate the executable code of the target simulation software. Exemplarily, put the user question, the retrieved top K script cases, and the relevant API documents together into the prompt to guide the fine-tuned large language model to generate the corresponding Abaqus code.
[0064] In addition, the following alternative solutions can also be considered:
[0065] Rule-based code generation: Generate Abaqus code through predefined rules and templates. The advantage of this method is that the generated code has high accuracy, but the disadvantage is that it has poor flexibility and is difficult to handle complex simulation tasks.
[0066] Code Generation Based on Reinforcement Learning: Training a large language model through reinforcement learning so that it can adjust according to feedback when generating code. The advantage of this method is that it can dynamically adjust the code generation strategy, but the disadvantage is that the training cost is relatively high.
[0067] It should be noted that the same or similar parts in this embodiment and Embodiment 1 can be referred to each other and will not be elaborated in this application.
[0068] Embodiment 3:
[0069] Based on Embodiment 2, Embodiment 3 of this application provides a software code generation system based on retrieval-augmented generation and large model fine-tuning, including:
[0070] A construction module for constructing a Q&A pair dataset;
[0071] A fine-tuning module for fine-tuning a pre-trained large language model based on the low-rank adaptation technique according to the Q&A pair dataset;
[0072] A design module for designing a retrieval-augmented generation process;
[0073] A generation module for generating a comprehensive prompt according to the retrieval-augmented generation result.
[0074] It should be noted that the system provided in this embodiment is the system corresponding to the method provided in Embodiment 2. Therefore, the same or similar parts in this embodiment and Embodiment 2 can be referred to each other and will not be elaborated in this application.
Claims
1. A software code generation method based on retrieval-augmented generation and large model fine-tuning, characterized in that Including: S1. Construct a question-answer pair dataset; S2. Based on the question-answer pair dataset, fine-tune a pre-trained large language model using the low-rank adaptation technique; S3. Design a retrieval-augmented generation process; S4. Generate a comprehensive prompt according to the retrieval-augmented generation result.
2. The software code generation method based on retrieval-augmented generation and large model fine-tuning according to claim 1, wherein In S1, the construction of the question-answer pair dataset includes: parsing natural language question descriptions and their corresponding script codes from the script use case document of the target simulation software to generate structured question-script pairs.
3. The software code generation method based on retrieval-augmented generation and large model fine-tuning according to claim 2, wherein In S2, the fine-tuning of the pre-trained large language model using the low-rank adaptation technique includes: using the question-answer pair dataset to adjust the parameters of the large language model through the low-rank adaptation technique to make it adapt to the API and simulation tasks of the target simulation software.
4. The software code generation method based on retrieval-augmented generation and large model fine-tuning according to claim 3, wherein In S3, the design of the retrieval-augmented generation process includes: According to the natural language query input by the user, retrieve script use cases semantically matching the query from the question-answer pair dataset, and extract the associated API documents in the script use cases.
5. The software code generation method based on retrieval-augmented generation and large model fine-tuning according to claim 4, wherein, In S4, the generation of the comprehensive prompt includes: Integrate the user query, the retrieved script use cases and the associated API documents into a prompt and input it into the fine-tuned large language model to generate executable code for the target simulation software.
6. A software code generation system based on retrieval-augmented generation and large model fine-tuning, characterized in that, For executing the method according to any one of claims 1 to 5, including: A construction module for constructing a question-answer pair dataset; A fine-tuning module for fine-tuning a pre-trained large language model based on the question-answer pair dataset using the low-rank adaptation technique; A design module for designing a retrieval-augmented generation process; A generation module for generating a comprehensive prompt according to the retrieval-augmented generation result.
7. A computer storage medium, characterized in that, The computer storage medium stores a computer program; when the computer program runs on a computer, the computer is caused to execute the method according to any one of claims 1 to 5.
8. An electronic device, characterized in that, Including: A memory for storing a computer program; A processor for executing the computer program to implement the method according to any one of claims 1 to 5.
Citation Information
Cited By
Multi-source heterogeneous code conversion method and device for cross-language development
CN121143797A
A multi-source heterogeneous code conversion method and device for cross-language development
CN121143797B
Method and device for processing computer algorithm
CN121478390A
Techniques for data-driven computation to enable the automatic generation of data-driven insights in response to natural language queries
US12669984B1