Code generation and quality optimization method based on static analysis and large language model
By combining static analysis and code generation methods of large language models, the problems of uneven code quality and insufficient context understanding capabilities in the prior art are solved, and high-quality and maintainable code generation is achieved.
Patent Information
- Application Number
- CN202510150718.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-11
AI Technical Summary
Existing code automatic generation technology has problems such as uneven code quality, low development efficiency, difficult to maintain the generated code, and low correlation ability of class-level code context understanding.
Using code generation and quality optimization methods based on static analysis and large language models, we ensure the generated code quality and context understanding through UML model definition and mapping, refined prompt word engineering, static code analysis and quality feedback iteration.
Improve the quality and context understanding of generated code, ensure the syntax accuracy and maintainability of the code, and improve development efficiency.
Smart Images

Figure CN120010858A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to code generation technology for computer software engineering, and in particular to a code generation and quality optimization method based on static analysis and a large language model. Background Art
[0002] Automatic code generation is an important research direction in software engineering and is widely used in software development, database management, Internet of Things, finance and other fields. Among many methods, large language models have attracted widespread attention from researchers because of their powerful natural language understanding capabilities, which can better understand the natural language descriptions provided by developers or non-technical personnel, and require less manpower and time costs than other methods. However, using large language models to generate accurate code is still a challenging task. The main difficulty lies in the fact that the accuracy and quality of the code generated by large language models cannot be guaranteed. To overcome this shortcoming, existing code automatic generation methods based on large language models usually improve the code generation quality of large language models through prompt word engineering and dataset fine-tuning, but this also increases the difficulty of model construction.
[0003] Currently, for model-driven code generation tasks, the traditional UML-to-code mapping method mainly generates skeleton code for classes, which is insufficient in terms of complexity and number of methods. The method of using large language models for code generation is more inclined to generate method / function level code, which has not yet effectively solved the challenge of context understanding and association ability when directly generating class-level code, and cannot guarantee the grammatical correctness of the generated code, making it unreliable for use. Summary of the invention
[0004] The present invention aims to solve the technical problems existing in the existing automatic code generation technology, such as uneven code quality, low development efficiency, difficult to maintain the generated code, and low ability to understand the context association of the generated class-level code. A code generation and quality optimization method based on static analysis and a large language model is proposed.
[0005] The technical solution adopted by the present invention is:
[0006] A code generation and quality optimization method based on static analysis and a large language model comprises the following steps:
[0007] S1. UML model definition, mapping and verification:
[0008] According to the requirements and functional description, the UML model is mapped into a preliminary code framework based on UML forward engineering and verified. This process involves the definition of class structure, method responsibilities and relationships between classes, and mapping class definitions, package structures, attributes, interfaces, enumerations and basic methods through class diagrams.
[0009] S2. Set up a detailed prompt word project:
[0010] Set a refined prompt word strategy based on a large language model, interact with the model through natural language, and generate preliminary completed class-level code. The prompt word design covers input semantics, output customization, error recognition, prompt improvement, and context control to optimize the accuracy, security, and generalization ability of the generated code.
[0011] S3, static code analysis:
[0012] Perform static analysis on the class-level code generated in step S2, build an abstract syntax tree (AST), analyze the entities and dependencies in the code, scan the incorrect calls in the code, and identify potential incorrect call problems by checking whether the called entity exists in the class to which it belongs.
[0013] S4. Iterative improvement of code based on quality feedback:
[0014] According to the results of static analysis, multiple rounds of interaction with the large language model are carried out based on quality feedback through refined prompt words to continuously optimize and correct the code, and static analysis is repeated until incorrect calls and other problems in the code are completely eliminated.
[0015] S5. Non-functional quality assessment:
[0016] After the error check and repair, the code is evaluated for non-functional quality, including maintainability, cyclomatic complexity and performance indicators. If the evaluation result does not meet the predetermined standards, it will continue to interact with the large language model for optimization, and finally generate high-quality code that meets the quality requirements.
[0017] Furthermore, the step S1 comprises:
[0018] Through UML forward engineering (UML->code), the UML model is mapped to the target implementation language to generate class skeleton code. This process defines the class structure, method responsibilities and the relationship between classes, and maps class definitions, package structures, attributes, interfaces, enumerations and basic methods through class diagrams.
[0019] Furthermore, the step S2 comprises:
[0020] Use multiple prompt word modes, including input semantics, output customization, error recognition, prompt improvement, interaction optimization, and context control, to set refined prompt word strategies. Optimize prompt words based on the three dimensions of code generation accuracy, security, and generalization ability to ensure that the large language model preliminarily improves the skeleton code generated in step S1.
[0021] Furthermore, the step S3 comprises:
[0022] The preliminary class-level code generated in step S2 is analyzed through static analysis, its abstract syntax tree is constructed, the entities and dependencies in the code are parsed, and it is determined whether the called entity exists in the class to which it belongs, so as to identify and fix the incorrect calling problem.
[0023] Furthermore, the step S4 comprises:
[0024] The error call results in step S3 are combined with the large language model, and the interaction quality with the large language model is improved by refining the prompt word method. The generated modification results are applied to the original code, and then static analysis and iterative interaction are performed again until the number of error calls is reduced to zero.
[0025] Furthermore, the step S5 comprises:
[0026] Perform non-functional quality assessment on the error-checked code, including maintainability, cyclomatic complexity, etc. If the assessment score does not meet the standard, further optimization is performed through interaction with the large language model to improve code quality.
[0027] Compared with the prior art, the present invention has the following advantages and technical effects:
[0028] The present invention realizes the combination of data model and large language model, maps UML diagram to code of class skeleton, and then automatically processes UML function description and code at class skeleton level by large language model, so that it automatically generates relatively complete code at class level. At the same time, the present invention combines advanced prompt word engineering and quality assessment feedback to ensure the correctness of code method function, and improves the quality of generated code. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the embodiments of the present invention or the drawings of related technical solutions in the prior art are introduced below. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0030] Figure 1 It is a flowchart of the steps of a code generation and quality optimization method based on static analysis and a large language model according to an embodiment of the present invention;
[0031] Figure 2 It is a flow chart of an embodiment of the present invention for interacting with a large language model to construct a complete skeleton code prompt word;
[0032] Figure 3 is a flow chart of an example of an erroneous method call in quality assessment feedback in an embodiment of the present invention;
[0033] Figure 4 is a flow chart of an embodiment of the present invention for interactively constructing and repairing incorrect call prompt words with a large language model;
[0034] Figure 5 is a graph of experimental measurement results of three code optimization levels in the code generation process of an embodiment of the present invention;
[0035] Figure 6 It is a flowchart of a code generation and quality optimization method based on static analysis and a large language model according to an embodiment of the present invention. DETAILED DESCRIPTION
[0036] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0037] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the protection scope of the present invention.
[0038] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application may be combined with each other.
[0039] In view of the existing technical problems, the present invention focuses on combining the data model with the large language model, mapping the UML diagram to the code of the class skeleton, and then automatically processing the UML functional description and the code at the class skeleton level by the large language model to automatically generate relatively complete code at the class level. At the same time, the present invention combines advanced prompt word engineering and quality assessment feedback to improve the accuracy of code generation and quality optimization based on static analysis and the large language model.
[0040] The method is explained in detail below with reference to the accompanying drawings and specific embodiments.
[0041] like Figure 6 As shown, this embodiment is a code generation and quality optimization method based on static analysis and a large language model, comprising the following steps:
[0042] The first step is to construct the class diagram and generate the code mapping according to the UML description. Figure 1 The S1 section:
[0043] S1. Define, map and verify the UML model. Complete the preliminary code framework mapping and verification of the model according to the requirements and functional description, and complete the generation of the class skeleton code.
[0044] In this embodiment, the structure of the class, the responsibilities of the methods, and the relationship between the classes are defined, and the mapping of the class definition, package, attribute, interface, enumeration, and basic method to the code is completed through the UML class diagram. According to the requirements and functional description of the task to be completed, combined with the UML model definition, through UML forward engineering, that is, UML->CODE, the mapping of UML to language is realized and the model is converted into the code at the class skeleton level.
[0045] The second step is to judge the class diagram code and complete the preliminary completion code of the large language model. Figure 1 The S2 section:
[0046] S2. For the skeleton code obtained in step S1, a refined prompt word project is set to interact with the large language model to complete the class-level code after preliminary completion.
[0047] like Figure 2 As shown, for the initial setting of the large language model, "role playing" is first performed to set an identity for the large language model. Then a specific interaction template is selected for application, and then the interaction is carried out, and a customized template is selected to complete the interaction output. Templated language is used throughout the process, and the constructed prompt words are universal and can cope with a variety of code generation situations. The preliminary code constructed in step S1 is iteratively modified with the large model to interact with the class skeleton level code, and finally the class level code after the preliminary method body is completed is obtained.
[0048] The third step is to perform static analysis of the code to check for incorrect calls. Figure 1 The S3 part:
[0049] S3. Perform static code analysis on the class-level code obtained in S2, complete the construction of the abstract syntax tree, entity and dependency analysis, and scan for incorrect call problems.
[0050] Specifically, it is necessary to verify the correctness of the class-level code obtained in step S2, use the entity dependency detection tool ENRE to scan the generated code, and perform static analysis on the class-level code in step S2 to build the corresponding abstract syntax tree to obtain variable declarations, method declarations, method calls, etc., and further obtain the entities and dependencies in the code to determine whether there are erroneous calling relationships in the code.
[0051] like Figure 3As shown in the figure, by using the entity dependency call tool ENRE to complete the AST scanning of the code, build the entity dependency graph, and make specific judgments on the wrong call. For checking the wrong call, first perform a scope check to obtain the entity after static analysis. The entity can be a method, function or other code block in the code. Then check whether the entity is of scope type, that is, whether it belongs to a specific scope. If the entity is of scope type, further perform an existence check to check whether the entity contains a call call. This requires identifying whether there are calls to other methods or functions inside the entity. Then get all the call calls contained in the entity. Then perform a parameter and return value compatibility check to determine whether the entity called by the entity exists, ensure that the called method or function is valid in the code base, and record the wrong call. If the called entity does not exist, record the wrong call. Finally, record all the wrong calls. It is convenient for the subsequent construction of the prompt word project to interact with the large language model to fix the wrong syntax and improve the correctness of the code.
[0052] The fourth step is to optimize the code based on the feedback from the large language model. Figure 1 The S4 section:
[0053] S4. Based on the feedback from static checking, the code is improved through continuous interaction and multiple iterations in combination with a large language model.
[0054] Specifically, according to the error call analysis data obtained in step S3, a prompt word for interacting with the large language model is further constructed. Figure 4 As shown, for the prompt word project at this time, the same "role playing" is first performed, an identity is set for the large language model, the large model is allowed to learn the pre-information content, and the prompt word template is applied to the large model for interaction, and then the prompt word is improved by inputting specific error call code related fragments. Finally, an output template is set for the large language model to facilitate the subsequent automatic processing of data storage code. This process continues, and continuous entity dependency analysis is performed to detect errors in the scanned code, and the code is optimized and improved by interacting with the large language model for automatic verification, and finally the interaction is stopped when the number of error calls in step S3 is zero.
[0055] The fifth step is to evaluate the code quality and generate high-quality code. Figure 1 The S5 part:
[0056] S5. Perform non-functional quality assessment on the improved code to ensure that the final high-quality code is generated.
[0057] Specifically, the final code is evaluated based on non-functional indicators, and indicators such as the maintainability and complexity of the code are analyzed to improve the final code.
[0058] Specifically, Figure 5 As shown, the code generation effect is evaluated and analyzed at three levels in this implementation process: class skeleton code, large language model optimized class skeleton code, and code after static analysis to detect error calls and optimization. Among them, CLOC, CM-LOC, WMC, RFC, and CBO refer to the number of lines of code, the number of lines of code without comments, the number of weighted methods per class, the number of responses of a class, and the degree of coupling between objects. According to the results, the code skeleton has less code, fewer explanatory comments on the generated code, fewer methods, and a moderate degree of coupling. The code that is finally optimized has a significantly increased amount of code and comments, a significantly increased number of methods and responses, and a further reduced degree of coupling.
[0059] From the example results, it can be seen that the target code corresponding to the requirements and functional description can be successfully generated through the application method.
[0060] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
Claims
1. A code generation and quality optimization method based on static analysis and large language model, characterized in that: The following steps are involved: S1. UML model definition, mapping and verification, complete the preliminary code framework mapping and verification of the model according to the requirements and functional description; S2. Set up a refined prompt word project. Based on the large language model, set up a refined prompt word strategy, use natural language to interact with the model, and generate preliminary completed class-level code; S3, static code analysis, completes the construction of abstract syntax tree, entity and dependency analysis, and scans for incorrect call problems in the code; S4. Iterative improvement of code based on quality feedback; S5. Non-functional quality assessment to ensure that the final high-quality code is generated.
2. The code generation and quality optimization method based on static analysis and large language model according to claim 1, characterized in that: In step S1, the UML model definition, mapping and verification include: Based on UML forward engineering, the UML model is mapped to the target implementation language to define the class structure, method responsibilities, relationships between classes, and the basic code framework; Use class diagrams to map class definitions, package structures, attributes, interfaces, enumerations, and base methods, and perform preliminary validation of the generated code to ensure consistency.
3. The code generation and quality optimization method based on static analysis and large language model according to claim 1, characterized in that: In step S2, the setting of the refined prompt word project includes: Provide fine-grained guidance for the code generation process through a variety of prompt word strategies, including input semantics, output customization, error identification, prompt improvement, interaction optimization, and context control; In the prompt word design, optimization is performed based on the three dimensions of code generation accuracy, security, and generalization ability to improve the class-level code generated in step S1 and achieve preliminary completion of the code structure.
4. The code generation and quality optimization method based on static analysis and large language model according to claim 1, characterized in that: In step S3, the static code analysis completes the construction of the abstract syntax tree, entity and dependency analysis, and scans for incorrect call problems in the code, including: Perform static analysis on the preliminary class-level code generated in step S2, construct the corresponding abstract syntax tree (AST), and parse the entities and dependencies in the code; Through static analysis, we can determine whether the entity calls in the code are legal, that is, we can identify incorrect calls by detecting whether the called entity exists in the class to which it belongs, and ensure the correctness of the code structure and dependencies.
5. The code generation and quality optimization method based on static analysis and large language model according to claim 1, characterized in that: In step S4, the code is iteratively improved based on the quality feedback, including: Based on the error calls identified in step S3 and other static analysis feedback results, multiple rounds of interaction are performed with the large language model through a refined prompt word method to repair and optimize code errors; The generated modification results are applied to the initial code, and then the static analysis and error checking of step S3 are repeated until the erroneous calls and other static checking problems are eliminated; through continuous interaction and context control, the code is iteratively improved and high-quality code is finally generated.
6. The code generation and quality optimization method based on static analysis and large language model according to claim 1, characterized in that: In step S5, non-functional quality assessment is performed on the improved code to determine whether the final high-quality code is generated, including: After completing the error call check, conduct non-functional quality assessment on the code, including evaluation of indicators such as code maintainability, cyclomatic complexity, performance, and security; If the evaluation result does not meet the predetermined standards, further interaction is performed through the large language model to optimize and improve the code, ultimately generating high-quality code that meets the predetermined non-functional quality requirements.
7. The code generation and quality optimization method based on static analysis and large language model according to claim 1, characterized in that: The large language model supports context management functions, which can adjust the generation results based on interaction records and problem feedback to improve the consistency and continuity of the generated code.
8. The code generation and quality optimization method based on static analysis and large language model according to claim 1, characterized in that: Use the code entity dependency analysis tool ENRE to customize analysis rules and error detection standards according to the characteristics of different languages.
9. The code generation and quality optimization method based on static analysis and large language model according to claim 1, characterized in that: The non-functional quality assessment supports user customized configuration, including adjustment of the weights of performance and security indicators to meet the needs of specific scenarios.
Citation Information
Patent Citations
Code generation method and system based on large language model
CN118210489A
Bidirectional intelligent question and answer code supplementing method and system based on large language model
CN118426825A
Automatic code defect repairing method based on static code analysis tool and artificial intelligence
CN118860864A
Cited By
Code generation task adaptive reasoning method, device and equipment
CN120560664A
Large model testing system, method, equipment and medium
CN121501637A