Code Quality Optimization Method, Device and Storage Medium Based on Causal Reasoning and LLM
By applying causal reasoning and large language models in code optimization, a causal relationship model is constructed to identify modifications that affect code quality and automatically generate optimization solutions, the problems of low efficiency and dependency of code optimization in the existing technology are solved, and efficient and accurate code quality optimization is achieved.
Patent Information
- Application Number
- CN202510436553.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-09
AI Technical Summary
In the code optimization, it is difficult to identify code modifications that affect code quality in code optimization, and rely on the correlation of historical data, and it is impossible to automatically generate optimization solutions that conform to the code style, resulting in engineers spending a lot of time to optimize and troubleshoot.
The code quality optimization method based on causal reasoning and large language model (LLM) is adopted to construct a causal relationship model, identify code modifications that affect the quality of the code, and automatically generate optimization solutions that conform to the corresponding code style.
It realizes the identification of code modifications that affect the quality of the code and automatically generates optimization solutions, saving engineers' time and energy and improving the efficiency and accuracy of code optimization.
Smart Images

Figure CN119938061B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a code quality optimization method, device, and storage medium based on causal reasoning and LLM. Background Art
[0002] With long-term technological accumulation, artificial intelligence technology has developed rapidly in the past two years. One of the main achievements is the large language model (LLM). Large language model products such as GPT, DeepSeek, Grok, etc. have been put into operation. The large language model is essentially an analysis tool. At present, enterprises and institutions have invested huge costs in the training work of the large language model. The further problem lies in how to apply it to specific practical applications, so as to solve the problems in the actual production process and improve the operation efficiency in relevant application scenarios.
[0003] Currently, one of the main possible application scenarios is the analysis and optimization of code. In the past, the analysis and optimization of program code usually required a large number of personnel to participate. Although there are many automated script programs as code analysis tools, their flexibility and adaptability are often insufficient, and ultimately they can only play a part of the auxiliary role, and it is difficult to further save labor costs. For example: The current rule-based static analysis scheme (such as SonarQube) is highly dependent on the fixed rules set by senior engineers and cannot adapt to different Java development scenarios. Another example: Code optimization based on machine learning. This scheme is mainly based on statistical correlation, and actually provides only statistical results, which cannot be used to explain the causal mechanism of optimization decisions. Engineers still need to spend a lot of time to determine the specific optimization direction based on the statistical results. The most advanced current method is to optimize code through LLM, but although it can provide optimization suggestions at present, it is often based on pattern matching, and it is difficult to determine whether the optimized code will cause new problems. Engineers still need to spend a lot of time to find and fix bugs.
[0004] Therefore, it is necessary to re-develop the current code optimization scheme based on LLM, so that it not only depends on the correlation in historical data, but can identify code modifications that affect code quality, and can also automatically generate optimization schemes that conform to the corresponding code style, thereby saving the time and energy of engineers. Summary of the Invention
[0005] Embodiments of the present invention provide a code quality optimization method, device, and storage medium based on causal reasoning and LLM, which can identify code modifications that affect code quality, and can also automatically generate optimization schemes that conform to the corresponding code style, thereby saving the time and energy of engineers.
[0006] To achieve the above object, the embodiments of the present invention adopt the following technical solutions:
[0007] In a first aspect, an embodiment of the present invention provides a code quality optimization method based on causal reasoning and LLM, which is characterized by including:
[0008] S1. Parse the target code data and establish a corresponding causal relationship model;
[0009] S2. Use the causal relationship model to simulate the causal intervention effects of different optimization strategies on code quality;
[0010] S3. Use a large language model to generate an optimization plan according to the obtained simulation results;
[0011] S4. Record the modifications to the target code data, and perform regression testing and verify the optimization effect.
[0012] In a second aspect, an embodiment of the present invention provides a code quality optimization device based on causal reasoning and LLM, which is characterized by including:
[0013] An analysis engine module for parsing the target code data and establishing a corresponding causal relationship model;
[0014] A simulation module for using the causal relationship model to simulate the causal intervention effects of different optimization strategies on code quality;
[0015] A plan generation module for using a large language model to generate an optimization plan according to the obtained simulation results;
[0016] An update module for recording the modifications to the target code data, and performing regression testing and verifying the optimization effect.
[0017] In a third aspect, an embodiment of the present invention provides a storage medium storing a computer program or instruction, which when the computer program or instruction is run, implements the method in this embodiment.
[0018] In the embodiments of the present invention, various influencing factors in the code are analyzed through causal reasoning, such as code complexity, duplication rate, defect density, etc., and a causal relationship model (or causal relationship network) is constructed. Through the causal relationship model, it is identified which Java code modifications truly affect code quality, rather than simply relying on the correlations in historical data. Moreover, in Java code optimization, for different types of code modules, especially the core business code part (such as Java network communication, database operations, service layer logic, etc.), customized optimization strategies can be adopted. For example, for modules with high complexity, redundant code can be reduced, the logic can be simplified, and maintainability can be improved; while for core business logic code, the optimization should focus on performance improvement and stability guarantee, avoiding large-scale refactoring and unnecessary modifications to reduce the possibility of introducing potential risks. In addition, during the optimization process, optimization suggestions that conform to the Java code style are generated by combining the LLM. Thus, it realizes the identification of code modifications that affect code quality and is also able to automatically generate optimization solutions that conform to the corresponding code style, thereby saving the time and effort of engineers. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary engineers in the field, other accompanying drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 It is a schematic flowchart of the method provided by the embodiments of the present invention;
[0021] Figure 2 It is a schematic structural diagram of the device provided by the embodiments of the present invention;
[0022] Figure 3 It is a network diagram of the training process provided by the embodiments of the present invention;
[0023] Figure 4 It is a topological diagram of the algorithm operation provided by the embodiments of the present invention;
[0024] Figure 5 It is a schematic diagram of a specific example provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] To enable engineers in this field to better understand the technical solution of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention. Engineers in this technical field can understand that unless specifically stated, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or coupling. The phrase "and / or" used herein includes any and all combinations of one or more of the associated listed items. Engineers in this technical field can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of an ordinary engineer in the field to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless defined as herein.
[0026] This embodiment is mainly applied to the optimization of Java code quality. The general design purpose is as follows: a Java code quality optimization method combining causal reasoning and LLM. By constructing a causal graph, quantify the impact of Java code optimization on quality metrics, and generate optimization suggestions with the help of LLM to achieve efficient, accurate, and interpretable Java code quality optimization. Among them, analyze the causal relationship between Java variables through causal reasoning technology, identify which code modifications really affect code quality, rather than relying only on the correlations in historical data. Combining LLM and causal reasoning can not only provide Java code optimization suggestions, but also explain the rationality of these suggestions, helping developers make more informed decisions.
[0027] In practical applications, the implementation process of this embodiment can be roughly divided into:
[0028] 1. Construct a causal relationship network: First, analyze Java code quality data and extract core metrics such as code complexity, maintainability, defect density, etc. This data can come from static analysis tools (such as SonarQube), code review records, and historical project data. Causal relationship modeling: Use Bayesian networks, structural equation models (SEM), or causal diagrams to construct the causal relationships between Java code quality factors. For example, the influence chain between factors such as code complexity, duplication rate, and function length. Calculate causal impacts: Calculate the impacts of various factors on target variables (such as maintainability, code reliability, etc.) through causal reasoning, and screen out key influencing factors.
[0029] 2. Calculate optimization strategies based on causal reasoning: Causal intervention simulation: Use do - calculus to simulate the causal intervention effects of different optimization strategies on code quality. For example, simulate the impacts on maintainability and defect density after reducing the code duplication rate or lowering the function complexity. Counterfactual analysis: Through counterfactual analysis, evaluate the potential effects of different optimization schemes and rule out schemes that may lead to negative optimization (such as certain optimizations may cause a decline in code performance). Customized optimization strategies: According to different types of code (such as core business code, algorithm core modules, UI components, etc.), formulate more targeted optimization strategies. For modules with higher complexity, focus on code refactoring and modularization optimization; for core functional modules, avoid large - scale modifications and focus on performance improvement and stability assurance.
[0030] 3. Use LLM to generate optimization suggestions and generate optimization plans: Based on the results of causal reasoning, use LLM (such as GPT, CodeT5, etc.) to generate optimization plans that conform to Java code style and best practices. The LLM will propose more feasible optimization suggestions based on context analysis. Explanation of optimization plans: When the LLM generates optimization suggestions, it will provide detailed explanations of the expected effects before and after optimization, possible side effects, and how to balance different optimization goals (such as the trade - off between maintainability and performance). Generation of optimization reports: Generate a detailed report containing optimization strategies, expected effects, potential risks, and implementation suggestions to help developers understand the optimization process and results.
[0031] 4. Code Optimization Execution and Feedback: Execution Optimization: Apply the optimization solutions generated by the LLM to the Java code and conduct regression testing to ensure that the optimization does not affect existing functions or introduce new problems. Quality Assessment: Verify the optimization effect through static analysis tools (such as SonarQube) or dynamic testing to ensure that the optimized Java code meets the expected quality standards (such as maintainability, performance, etc.). Feedback and Learning: The LLM records the optimization results, analyzes the developers' feedback, and dynamically adjusts the optimization strategy to improve the accuracy and feasibility of subsequent optimization suggestions.
[0032] 5. Continuous Improvement: Incremental Learning: The optimization strategy continuously makes feedback and adjustments based on historical optimization data, combines the actual usage scenarios and feedback of developers, and continuously improves the intelligence level of the optimization model and optimization suggestions. Evolution of the Optimization Strategy: With the accumulation of more projects and optimization cases, the optimization strategy will gradually evolve to form more accurate and personalized optimization solutions.
[0033] Combined with the actual application scenario, the preferred solution for the design of the embodiment is, for example Figure 1 , 3 , as shown in Figure 4, including:
[0034] S1. Parse the target code data and establish a corresponding causal relationship model.
[0035] S2. Use the causal relationship model to simulate the causal intervention effects of different optimization strategies on code quality.
[0036] In this embodiment, more targeted optimization strategies are formulated according to different types of code (such as core business code, algorithm core modules, UI components, etc.); for modules with higher complexity, key optimization can be carried out on code refactoring and modularization; for core function modules, large-scale modifications should be avoided, and emphasis should be placed on performance improvement and stability assurance. For example: Core business code refers to the code part that carries the core business logic. Usually, this code has a high degree of complexity and is directly related to the function and performance of the system. When optimizing core business code, the focus is on code refactoring, improving maintainability, and optimizing performance, but frequent and large-scale modifications need to be avoided to ensure the stability of the business logic.
[0037] Specific optimization strategies include: Code refactoring: By analyzing the functions and classes in the core business modules, identify complex and lengthy parts and refactor them. For example: Split long methods: If a function is very complex and contains multiple responsibilities, it should be split into multiple small functions with single responsibilities. Simplify conditional statements: Avoid excessive nested conditions and use design patterns such as the strategy pattern or state pattern to simplify complex business logic. Improve modularity: Split complex business logic into multiple small modules or services to make the code easier to understand and modify. Performance optimization: Core business modules usually need to process a large amount of data or perform high-frequency calculations, so performance improvement should be a key focus during optimization. Optimization methods include: Algorithm optimization: Check the efficiency of existing algorithms and use more efficient algorithms (e.g., optimize sorting algorithms, improve data queries, etc.). Cache optimization: Cache repeated calculations to avoid unnecessary repeated calculations.
[0038] The algorithm core module is the core code part in the system for data processing, calculation, or transformation. When optimizing the algorithm module, the most important thing is to improve the efficiency of the algorithm, especially in scenarios where large-scale data is processed or high real-time performance is required. The optimization strategies include: Optimize algorithm complexity: By analyzing the time complexity and space complexity of existing algorithms, find more efficient algorithms. For example, replace brute-force search algorithms with hash table lookups, or use dynamic programming to reduce redundant calculations. Reduce resource consumption: Optimize memory usage, avoid memory leaks and resource waste, especially when dealing with big data. Parallel computing and distributed processing: Perform computationally intensive tasks in parallel or through distributed computing to improve processing power. Algorithm stability and robustness: Ensure that the algorithm runs stably under various boundary conditions and avoid algorithm crashes caused by non-standard inputs or extreme cases.
[0039] UI components are responsible for interacting with users. When optimizing the code of UI components, the main focuses are response speed, interface smoothness, user experience, and maintainability. UI optimization is not just about reducing code complexity but also about improving the user experience. The optimization strategies include: Improve rendering efficiency: For complex UI interfaces, optimize the rendering process to avoid unnecessary re-rendering. You can use virtualized lists, lazy loading, or incremental rendering to reduce performance bottlenecks during rendering. Reduce DOM operations: For front-end web applications, reduce frequent operations on the DOM and avoid causing a large number of reflows and repaints of DOM elements for each interaction. UI componentization: Split complex UI interfaces into multiple reusable UI components to improve the maintainability and extensibility of the code. Responsive design: Ensure that the UI can run smoothly on different devices and adapt to different screen sizes and resolutions.
[0040] S3. Utilize large language models to generate optimization solutions based on the obtained simulation results.
[0041] Among them, the code optimization dataset is used to train large language models. The code optimization dataset is designed to generate, improve, and optimize the code quality specifically for training large language models. This dataset contains multiple dimensions related to code optimization, aiming to help large language models learn how to generate high-quality code optimization strategies according to different project requirements and goals. The dataset will include a large number of comparisons between original code and optimized code. These codes can come from different application scenarios, such as core business logic, algorithm core modules, UI components, etc. These data are sourced from open-source code libraries, technical documents, and programming tutorials to ensure high accuracy and practicality when the model generates and understands code. The code optimization dataset will also include various optimization strategies and implementation methods. These strategies are targeted at different types of code and modules. After each code comparison record, metrics related to code quality will be attached, such as complexity, defect density, and maintainability score. The code optimization dataset also includes examples of error fixing and prevention, covering the repair of common programming errors (such as memory leaks, null pointer exceptions, deadlocks, etc.) and enhancing the robustness of the code, especially error fixing in high-concurrency and distributed environments. Large language models (LLMs) are commonly used in various application scenarios, such as text generation, translation, and dialogue systems. When training these models, the training datasets used are mainly natural language texts, while the unique feature of the code optimization dataset is that it focuses on the optimization of the structure, performance, and quality at the code level.
[0042] S4. Record the modifications to the target code data, and perform regression testing and verify the optimization effect.
[0043] Among them, modify the target code data according to the generated optimization plan, and perform regression testing and verify the optimization effect. For example: The large language model (LLM) will generate an optimization plan based on the analyzed code optimization dataset and project goals. Developers or automated tools or the large language model automatically optimize and modify the target code according to the optimization plan.
[0044] Specifically, in S1 of this embodiment, it includes: obtaining the core metric information of the target code data; using the core metric information to establish a causal relationship model, and the causal relationship model is used to calculate the impact of each core metric on the target variable.
[0045] Specifically, after calculating the impact of each core metric on the target variable (such as code quality), some key influencing factors are obtained. The subsequent effects of the key influencing factors include the design of optimization strategies, resource allocation, and priority setting. For example: If code complexity (C) is a key influencing factor and its increase has a negative impact on code quality, then the optimization strategy might be to refactor complex functions or split long methods to reduce code complexity, thereby improving code quality. If defect density (d) is determined to be a key factor and it has a negative impact on code quality, then the strategy might be to strengthen code reviews and unit tests to reduce the number of defects, thereby improving code quality. If maintainability (M) is a key factor affecting code quality and, compared to other factors, its improvement can maximize the enhancement of code quality, then the R & D team should prioritize optimizing the maintainability of the code. After generating an optimization plan through S3, the technology R & D team can focus resources on key influencing factors such as code complexity and defect density, rather than wasting excessive time and resources on other variables with less impact. At the same time, the computer records the modifications made by the technical staff to the target code data. During subsequent development, the key influencing factors can also serve as monitoring and evaluation metrics. The R & D team can continuously monitor changes in these factors and adjust development and optimization strategies based on the changes. Through continuous evaluation of the key influencing factors, a feedback loop can be formed to ensure continuous improvement and the continuous adjustment of optimization strategies. After implementing certain optimization strategies, the R & D team can use the model to predict the intervention effect. By predicting how the key influencing factors in the model change, the R & D team can anticipate the future code quality at the initial stage of the project and make appropriate interventions.
[0046] It should be noted that in this embodiment, the optimization plan generated through S3 can be directly provided to the technical staff of the R & D team in text form. After the technical staff make code modifications, the computer records the modifications made by the technical staff to the target code data and executes S4 to perform regression testing and verify the optimization effect.
[0047] Specifically, the core metric information for obtaining the target code data includes: scanning the code library of the target code data through a static analysis tool (such as SonarQube) to obtain the code complexity of the target code data, and determining the maintainability and defect density; among them, the metrics corresponding to the code complexity include: obtaining the cyclomatic complexity, function complexity, and nesting depth of the code of the corresponding module, class, and method from the target code, where the cyclomatic complexity is expressed as V(G)=E-N+2P, the function complexity is expressed as Fc=L+B, the nesting depth N' represents the number of layers of the most nested conditional statement or loop statement, V(G) represents the cyclomatic complexity of the graph (i.e., the cyclomatic complexity), E is the number of edges in the control flow graph, N is the number of nodes in the control flow graph, P is the number of independent connected regions in the program (usually 1, representing the main region of the program), Fc is the function complexity, L is the number of lines of code in the function (for example, the number of lines of code does not include blank lines and comment lines), and B is the number of branches in the function (for example, the number of branch statements such as if, switch, etc.).
[0048] Among them, Cyclomatic Complexity: Measures the number of independent paths in the code, usually using the McCabe Cyclomatic Complexity metric. A higher complexity value may mean that the code is difficult to understand and test. Function Complexity: Measures the number of lines of code, branch structure, and loop structure inside the function, etc. Larger functions usually mean higher complexity, and it is recommended to split them into multiple smaller functions. Nesting Depth: Measures the nesting depth of conditional statements and loop statements. Deep nesting increases the difficulty of understanding and maintaining the code, especially during debugging. Code Duplication: Duplicate code increases the complexity of the code because it requires changing multiple pieces of code when modifying, which is prone to maintenance errors.
[0049] Cyclomatic complexity is used to measure the number of independent paths in the code. By analyzing the control flow graph, the number of paths in the program is calculated to obtain the complexity of the code. For each function or method, analyze its control flow graph (ControlFlowGraph, CFG), calculate the number of edges and nodes in it, and substitute them into the formula to calculate the complexity. If the complexity value is too high (usually greater than 10), the code needs to be split or refactored. For example: In the process of calculating the cyclomatic complexity V(G)=E-N+2P = 2, if in the target code, there are two main nodes, such as: the while(true) loop condition judgment, and the operations in the loop body (memory allocation, printing, pausing). Therefore, the number of nodes in the control flow graph N = 2 can be obtained; if in the target code, the loop condition judgment is while(true), which means there is one edge in E, and the operations in the loop body (memory allocation, printing, pausing, etc.) are also one edge, so the number of edges in the control flow graph E = 2; if the structure of this program is a single infinite loop without other independent regions, then the number of independent connected regions P = 1. Another example: Suppose the control flow graph of a certain function contains 7 nodes, 8 edges, and the function is a single connected region (P = 1), then its cyclomatic complexity is: V(G)=8−7+2(1)=3.
[0050] Function complexity mainly considers the size of the function (the number of lines of code) and its internal control structure. It can be evaluated by calculating the number of lines of the function and some other measurement criteria. Perform static analysis on the function, count the number of lines of its code (L) and the number of branches (B). If the number of lines of the function exceeds the set threshold (for example, 20 lines), or there are too many branches, it should be considered to split or simplify it. For example: The main structure in the target code is an infinite loop, then the main number of lines of the code after removing blank lines and comments is 6 lines, that is, L = 6; if there is only one while(true) loop in the code and no other conditional branches (such as if statements, switch statements, etc.), then B = 1, and the function complexity Fc = L + B = 7. Another example: A function contains 25 lines of code and has 4 conditional branches (if, else), then its complexity is: Fc = 25 + 4 = 29.
[0051] The nesting depth measures the levels of conditional statements or loop statements in the code. An excessive nesting depth usually means that the logic of the code is difficult to follow, increasing the complexity of maintenance: N' = max(Nesting Depth of all branches), where N' is the maximum nesting depth. The nesting depth refers to the number of levels of the most nested conditional statement or loop statement. Specifically, analyze each control flow, calculate the nesting depth of each branch and loop, and find the maximum value. If the nesting depth exceeds the set threshold (e.g., 3 levels), the code structure should be optimized. For example: if there is only an infinite loop while(true) without any other nested loops or conditional statements, then the nesting depth N' = 1. Another example: Suppose there is the following nesting in a function: the first layer: if(condition1), the second layer: if(condition2), the third layer: for(i = 0; i < 10; i++), then the nesting depth is 3, indicating a relatively high complexity of this function.
[0052] The code duplication rate refers to the existence of multiple duplicate code blocks in the program. This not only increases the code complexity but also may lead to difficulties in maintenance and errors: D = Number of duplicate lines / Total lines of code * 100%, where D is the code duplication rate. The "number of duplicate lines" refers to the total number of code lines that appear multiple times. The "total lines of code" is the number of code lines in the entire project or module. Specifically, perform code analysis, find duplicate code blocks, and calculate the duplication rate. If the duplication rate is too high (usually exceeding 10%), it is necessary to eliminate the duplicate code by extracting functions or modularization. For example: Suppose the total number of code lines in a project is 1000 lines, and 120 of them are duplicate code lines, then the duplication rate is: D = 120 / 1000 * 100% = 12%. If the duplication rate is too high, it is recommended to refactor and abstract the duplicate code.
[0053] In this embodiment, in order to comprehensively evaluate the complexity and quality of the code, multiple complexity metrics can be combined to give a comprehensive score. The code complexity metric is expressed as C = w1 * V(G) + w2 * Fc + w3 * N' + w4 * D, where w1, w2, w3, and w4 are the weight coefficients corresponding to the four code complexity metrics, and the preferred solutions are shown in Table 1.
[0054] Table 1
[0055] Project scenario type <![CDATA[Cyclomatic complexity (w1)]]> <![CDATA[Function complexity (w2)]]> <![CDATA[Nesting depth (w3)]]> <![CDATA[Code duplication degree (w4)]]> Explanation Rapid development project (prototype, MVP) 0.2 0.3 0.2 0.3 Priority is given to development speed, with lower requirements for complexity and maintainability, and acceptable duplicate code. Enterprise-level application (long-term maintenance system) 0.3 0.4 0.2 0.1 Emphasis is on maintainability and code quality, with lower complexity and code duplication, and stability is the most crucial. Performance optimization project (high-performance computing) 0.4 0.2 0.3 0.1 Performance optimization is the core, complexity and nesting depth affect performance, and function complexity and code duplication are less important. Rapid iteration product (SaaS, Internet application) 0.25 0.25 0.25 0.25 Pay attention to the balance between development speed and code quality, and the weights of various indicators are relatively balanced.
[0056] Maintainability metrics include: cyclomatic complexity, duplicate code. In addition, there are also: function length (i.e., the number of lines in a function), module coupling (Coupling), where coupling refers to the degree of dependence between modules. Code with too high a coupling is difficult to modify because changing one module may affect other modules. Coupling = Number of dependencies between modules / Number of modules. Comments are crucial for helping developers understand the code. Good comments can significantly improve the maintainability of the code. CR = Lines of Comments / Total Lines of Code * 100%, where CR is the comment rate, the number of comment lines is the number of comments in the code, and the total number of code lines is the total number of code lines in the entire codebase.
[0057] For the comprehensive score of maintainability, the above metrics can be combined and a comprehensive score can be given through a weighted method to quantify the maintainability of the code. The maintainability metric is expressed as M = w’1 * V(G) + w’2 * L + w’3 * D + w’4 * C’ + w’5 * CR, where M is the maintainability score of the code, D is the code duplication rate, C' is the module coupling degree, CR is the comment rate, and w’1, w’2, w’3, w’4, w’5 represent the weight coefficients corresponding to the five maintainability metrics. The preferred solutions are shown in Table 2.
[0058] Table 2
[0059] Project scenario type <![CDATA[Cyclomatic complexity (w’1)]]> <![CDATA[Function length (w’2)]]> <![CDATA[Code duplication degree (w’3)]]> <![CDATA[Module coupling degree (w’4)]]> <![CDATA[Annotation rate (w’5)]]> Explanation Rapid development project (prototype, MVP) 0.3 0.2 0.2 0.15 0.15 The main goal is development speed, allowing a certain degree of code complexity and duplication, and focusing on function implementation and rapid iteration. Enterprise-level application (long-term maintenance system) 0.35 0.3 0.15 0.1 0.1 Requirements for long-term maintenance, code quality and maintainability are crucial, focus on loop complexity and function length, and reduce code duplication. Performance optimization project (high-performance computing) 0.4 0.2 0.15 0.15 0.1 High performance requirements, code complexity and coupling directly affect performance optimization, so the weights of loop complexity and module coupling are relatively high. Rapid iteration product (SaaS, Internet application) 0.25 0.25 0.2 0.2 0.1 Balance development speed and code quality, moderately pay attention to code complexity and module coupling, with a lower comment rate and moderate code duplication.
[0060] Defect Density is one of the important metrics for measuring code quality. It represents the number of defects per thousand lines of code (KLOC). It is used to measure the defects and problems in the code and is usually used to reflect the quality level in software development. A higher defect density generally means poorer code quality, and conversely, a lower defect density means better code quality. Design of defect density: Defect Density = Number of Defects / Number of Lines of Code * 1000. Number of Defects: The number of defects found in the code, usually sourced from test reports, code reviews, or bug tracking systems (such as JIRA, Bugzilla). Number of Lines of Code: The total number of lines in the code file (usually excluding blank lines and comment lines). 1000: To make the unit of defect density "per thousand lines of code" (KLOC). Causal Inference is used to identify the causal relationships between code quality factors, rather than just statistical correlations. For example: Code Complexity → Increased Maintenance Cost, Duplicate Code → Increased Code Redundancy → Reduced Readability, Defect Density → Reduced Code Reliability. In this invention, a Causal Graph is used to model Java code quality, and Do-Calculus and Counterfactual Analysis are used to optimize code quality. These causal inference techniques can help in deeply understanding the causal mechanisms of code optimization and ensuring that every optimization decision has a reasonable explanation. In the process of establishing a causal relationship model using the core metric information in this embodiment, a structure learning algorithm (such as the PC algorithm, GIES algorithm) is used to automatically construct the causal relationship network of code quality factors. The variables included in the causal graph can be: X1: Number of Lines of Code, X2: Code Duplication Rate, X3: Cyclomatic Complexity, X4: Defect Density, X5: Maintainability. Example causal relationships: X1→X2→X3→X4→X5, X2→X5, X3→X5, or, X1→X2→X3→X4→X5, that is, the number of lines of code affects the code duplication rate, which in turn affects the code complexity, and ultimately affects the maintainability of the code. This causal graph can be modeled through a Bayesian Network or a Structural Equation Model (SEM). For each code optimization suggestion, it is necessary to evaluate whether the code modification (intervention) can actually improve the quality, rather than blindly optimizing based only on statistical correlations as in the prior art. When performing code optimization, Do-Calculus is used to simulate the causal impact of code modification (intervention) on Java code quality. In this way, the potential effects of different optimization schemes can be evaluated. For example, simulate the impact of reducing duplicate code on maintainability.Example: Suppose we want to evaluate the impact of reducing duplicate code on the maintainability of Java code: P(X5∣do(X2 = 0.1)), which represents the change in maintainability X5 after reducing the code duplication rate (intervention X2). Through counterfactual analysis, we can assess questions like "If a certain optimization is not carried out, will the code quality deteriorate?" For example, if the duplicate code is not reduced, will the maintenance cost of the code be higher? This kind of analysis helps identify which optimizations are effective and which may lead to negative optimizations. This method can be used to screen the most effective optimization strategies and avoid negative optimizations (such as over-optimization resulting in decreased code readability).
[0061] Furthermore, we can further optimize the strategy screening. Through causal reasoning calculations, we can screen out the most effective optimization strategies and avoid negative optimizations (such as decreased code readability after optimization). For core modules (such as database access code, service layer logic, etc.), over-optimization should be avoided, and instead, focus should be placed on ensuring performance and stability. For code modules with high complexity or high duplication rate, their structures can be optimized, redundant code can be reduced, and the logic can be simplified; for core code that has been fully verified, more attention should be paid to performance improvement and maintainability, and large-scale refactoring should be avoided.
[0062] Based on the above ideas, in this embodiment, we deeply analyze the causal relationship, simulate the impact of different optimization strategies on code quality, and use Do-Calculus to represent the causal model. In the actual causal graph:
[0063] C→Q (Code complexity affects code quality);
[0064] d→Q (Defect density affects code quality);
[0065] M→Q (Maintainability affects code quality);
[0066] Causal reasoning can be performed through Do-Calculus: P(Q∣do(C = c o ), do(d = d o ), do(M = m o ))), which represents the probability distribution of Q when C, d, and M are specific values c o , d o , m o respectively. Among them, c o , d o , m oThe specific value can be set or changed dynamically. Using this causal reasoning method, the intervention effects of different variables can be simulated, that is, "if a certain indicator is changed, how will the code quality change". In causal reasoning, variable intervention is usually carried out to observe its impact on the target variable. Suppose we want to analyze the impact of code complexity (C) on code quality (Q). The intervention model can be expressed as follows: Q = β0 + β1*C + β2*d + β3*M + ϵ2, and we can simulate the change of Q when C changes, that is, through Do-Calculus simulation: P(Q∣do(C = c)). Through this method, we can evaluate how the values of different code complexities affect code quality and help optimize the code structure. Considering the multiple causal relationships and feedback mechanisms between variables, in the finally established causal reasoning model (Do-Calculus), code quality is expressed as Q = β0 + β1*C + β2*d + β3*M + ϵ1, C = γ1*d + γ2*M + ϵ2, d = λ1*M + ϵ3, where Q is the code quality (target variable), β0 is the constant term (intercept), β1, β2, β3 are three regression coefficients to be estimated, used to represent the causal impact of core indicators on code quality, γ1, γ2 represent the influence degree coefficients of defect density and maintainability on code complexity, λ1 represents the influence degree coefficient of maintainability on defect density, and ϵ1~ϵ3 represent the first to third error terms.
[0067] Specifically, in S2 of this embodiment, it includes: simulating the causal intervention effects of different optimization strategies on code quality through the Do-Calculus method and counterfactual analysis; when the large language model provides the modified code, causal analysis can simulate the impact of the modified code on the target variable through Do-Calculus or counterfactual analysis. That is, deduce the causal intervention effects.
[0068] Simulate the impact of the modified code on the target variable: Use the Do-Calculus method to simulate the change of the target variable (such as code quality) after modifying the code. For example: P(Q∣do(C = c′), do(d = d′), do(M = m′)), where c′, d′, m′ represent the modified code complexity, defect density, and maintainability. Through this simulation, the model can predict the change of code quality after modifying the code.
[0069] The counterfactual analysis is used to rule out strategies that would lead to negative optimization, expressed as: P(Q∣C,d,M) vs P(Q∣C′,d′,M′), where C′, d′, and M′ represent the modified code complexity, defect density, and maintainability in the counterfactual analysis. The counterfactual analysis is used to compare the results before and after the modification. For example, the counterfactual analysis can be used to compare the differences between the code before the modification (such as C, d, M) and after the modification (such as C′, d′, M′) to evaluate whether the modification is effective. The counterfactual model can analyze the effect of the modification by comparing the actual result with the hypothetical result: P(Q∣C,d,M) vs P(Q∣C′,d′,M′). By comparing the code quality before and after the modification, the counterfactual analysis helps to evaluate whether the modification strategy has brought the expected positive impact or whether it may introduce negative effects.
[0070] To illustrate with a clearer example, in combination with the above example, in the target code, there are two main nodes: the while(true) loop condition judgment and the operations within the loop (memory allocation, printing, pausing). Then, the number of nodes N of the control flow graph is 2; for the loop condition judgment while(true), the operations within the loop also form one edge, so the number of edges E of the control flow graph is 2; if the structure of the program is a single infinite loop without other independent regions, then the number of independent connected regions P is 1; the main structure in the target code is an infinite loop, so L = 6; if there is only one while(true) loop in the code and no other conditional branches (such as if statements, switch statements, etc.), then B = 1, and the function complexity Fc = L + B = 7; there is only one infinite loop while(true) without any other nested loops or conditional statements, so the nesting depth N' = 1. In the scenario of this example, if the project scenario type is selected as "rapid development project", then substituting into the formula: the code complexity C = w1*V(G) + w2*Fc + w3*N' + w4*D = 2.7; the maintainability index M = w’1*V(G) + w’2*L + w’3*D + w’4*C’ + w’5*CR = 2.42; the defect density d = the number of defects (1) / the number of code lines (6) * 1000 = 166.67.
[0071] If the application scenario of the "Rapid Development Project" project scenario type is selected, a general causal model is constructed to represent the causal relationship between the input, processing process, and output of the program. Among them, the static analysis program analyzes the code to define the causal relationship between various variables in the code and establish a causal model. Then, the main variables can be found as follows: Memory allocation function (memory allocation): Allocate 1MB of memory multiple times and add it to the memoryLeakList; Memory usage (memoryLeakList.size()); Memory-related problems include out-of-memory errors (OutOfMemoryError). From this, a causal diagram as shown in Figure 5 can be generated.
[0072] Continuing with the above Figure 5 example, in the intervention analysis, 1. Limit the size of the memoryLeakList to 1000, and stop adding new elements when 1000 elements are reached. And: The code complexity C = 3.2; The maintainability index M = 2.925; The defect density d = 111.11. 2. Regularly clean the elements in the memoryLeakList. And: The code complexity C = 3.4; The maintainability index M = 3.185; The defect density d = 100. It can be represented by the program as: If L(t)≥MemoryLimit, then stop adding new elements or clear the list.
[0073] In counterfactual reasoning, first calculate the quality of the current code, which can be represented by the program as: Qcurrent = -0.5 * 2.7 + 0.3 * 166.67 + 0.2 * 2.42 = -1.35 + 50.00 + 0.484 = 49.134.
[0074] Moreover, the above counterfactual analysis process can be divided into multiple stages. For example: The first stage is vulnerability repair, which focuses on ensuring the repair of vulnerabilities in the code. Repairing vulnerabilities may not directly improve performance or code quality, but it can prevent errors from occurring and reduce problems such as unexpected crashes and memory overflows. There are two optimization strategies given in the first stage: Optimization Strategy 1-1, Qoptimized1 = -0.5 * 3.2 + 0.3 * 111.11 + 0.2 * 2.925 = -1.6 + 33.33 + 0.585 = 32.315; Optimization Strategy 1-2, Qoptimized2 = -0.5 * 3.4 + 0.3 * 100 + 0.2 * 3.185 = -1.7 + 30.00 + 0.637 = 28.937. According to the results of the counterfactual reasoning of the two optimization strategies: Optimization Strategy 1 (limiting the size of memoryLeakList to 1000) and Optimization Strategy 2 (periodically cleaning the elements in memoryLeakList) did not improve the code quality and actually led to a decline in code quality. Since this problem involves memory leakage, the optimization strategy is necessary. According to the calculation, Optimization Strategy 1 is significantly better than Optimization Strategy 2, so Optimization Strategy 1 is adopted for subsequent solution optimization.
[0075] The second stage is to optimize the performance bottleneck in the code, that is, after the vulnerability problem is solved, the performance can be further improved, such as reducing the code complexity and improving the computing efficiency. There is one optimization strategy given in the first stage, Optimization Strategy 2-1, Qoptimized1’ = -0.5 * 3.2 + 0.3 * 125 + 0.2 * 2.925 = -1.6 + 37.5 + 0.585 = 36.485.
[0076] In this embodiment, the modification of the target code data can be operated by a technician on a computer. The computer records the modification traces and verifies the optimization effect after modification. For example: The effect is evaluated in each stage. For example, in Stage 1 (after repairing the vulnerability), the code stability can be evaluated first to check whether the repair goal is achieved, and then enter the next stage. If too much complexity is not introduced in Stage 1 and the problem is solved, the optimization process will continue. Stage 1: After repairing the vulnerability, run the program to ensure that the vulnerability is solved. If the code quality has slightly declined but the program has run stably, enter Stage 2. Stage 2: After introducing performance optimization, conduct a re-evaluation to detect whether the performance has improved and no additional complexity has been brought.
[0077] In actual Do-Calculus simulations, for technicians who optimize operating computers, the model can automatically generate optimization strategies. Based on the model's analysis of code complexity, the model may suggest splitting a complex function into multiple small functions to reduce loop complexity. According to the defect density, the model may suggest conducting more unit tests on certain high-risk areas or performing more rigorous code reviews. The model may suggest improving maintainability by enhancing the quality of code comments or simplifying the code structure. The model can automatically analyze which past optimization strategies have had a significant impact on code quality improvement and generate similar optimization plans.
[0078] In this embodiment, a code quality optimization device based on causal reasoning and LLM is also provided, as Figure 2 shown, including:
[0079] An analysis engine module for parsing the target code data and establishing a corresponding causal relationship model;
[0080] A simulation module for using the causal relationship model to simulate the causal intervention effects of different optimization strategies on code quality;
[0081] A solution generation module for using a large language model to generate an optimization plan based on the obtained simulation results;
[0082] An update module for recording the modifications to the target code data, performing regression tests, and verifying the optimization effects.
[0083] This embodiment also provides a storage medium, including: storing a computer program or instruction, which, when the computer program or instruction is run, implements the method described in this embodiment. For example: The code for implementing the algorithm for causal reasoning can be stored in the storage medium:
[0084] The following is an example code for performing causal reasoning calculations using the dowhy library in Python, demonstrating how to use a causal reasoning model to evaluate the impact of code duplication rate on the maintainability of Java code.
[0085] from dowhy import CausalModel
[0086] import pandas as pd
[0087] # Example code quality data
[0088] data = pd.DataFrame({
[0089] "code_lines": [100, 200, 150, 300],
[0090] "duplication": [0.2, 0.3, 0.15, 0.4],
[0091] "complexity": [10, 15, 12, 20],
[0092] "defect_density": [0.05, 0.1, 0.07, 0.15],
[0093] "maintainability": [0.8, 0.6, 0.75, 0.5]
[0094] [[ID=13}}
[0095] # Construct the causal graph
[0096] model = CausalModel(
[0097] data = data,
[0098] treatment = "duplication", # Intervention variable
[0099] outcome = "maintainability", # Target variable
[0100] graph = "digraph{duplication->complexity;complexity->maintainability}" )
[0102] # Estimate the causal impact
[0103] identified_estimand = model.identify_effect()
[0104] estimate = model.estimate_effect(identified_estimand, method_name = "backdoor.linear_regression")
[0105] print(estimate.value).
[0106] This embodiment is mainly applied to the optimization of Java code quality. The optimization process targets the key factors in Java code. First, it analyzes various influencing factors in the code through causal reasoning, such as code complexity, duplication rate, defect density, etc., and constructs a causal relationship model (or called a causal relationship network). Through this causal relationship model, it can identify which Java code modifications truly affect code quality, rather than simply relying on the correlations in historical data. Moreover, in Java code optimization, for different types of code modules, especially the core business code part (such as Java network communication, database operations, service layer logic, etc.), customized optimization strategies can be adopted. For example, for modules with high complexity, redundant code can be reduced, the logic can be simplified, and maintainability can be improved; while for core business logic code, the optimization should focus on performance improvement and stability guarantee, avoiding large-scale refactoring and unnecessary modifications to reduce the possibility of introducing potential risks. In addition, during the optimization process, by combining the LLM to generate optimization suggestions that conform to the Java code style, and ensuring the rationality and feasibility of the optimization suggestions through the results of causal reasoning. The LLM will generate a detailed optimization explanation report to help developers understand the reasons, expected effects, and possible side effects of the optimization plan, and provide optimization suggestions on how to balance performance and maintainability. In summary, the present invention provides a precise, interpretable, and efficient Java code quality optimization method, which can help developers improve code quality while reducing unnecessary optimization risks, and further enhance the optimization effect through real-time feedback and adjustment. Through the combination of causal reasoning and the LLM, targeted optimization strategies can be achieved to ensure that the improvement of Java code quality meets the actual needs of the project.
[0107] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiment. As described above, only the specific implementation manners of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by an engineer familiar with the technical field of the present invention within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A code quality optimization method based on causal reasoning and LLM, characterized in that: include: S1. Analyze the target code data and establish the corresponding causal relationship model; S2. Using the causal relationship model, simulate the causal intervention effects of different optimization strategies on code quality; S3, using the large language model to generate an optimization plan based on the obtained simulation results; S4, recording the modification to the target code data and verifying the optimization effect; In S1, it includes: obtaining the core indicator information of the target code data; establishing a causal relationship model using the core indicator information, wherein the causal relationship model is used to calculate the influence of each core indicator on the target variable; The obtaining of the core indicator information of the target code data includes: scanning the code base of the target code data by a static analysis tool to obtain the code complexity of the target code data, and determining maintainability and defect density; Among them, the corresponding code complexity indicators include: the cyclomatic complexity, function complexity and nesting depth of the code of the corresponding modules, classes and methods obtained from the target code, wherein the cyclomatic complexity is expressed as V(G)=E-N+2P, the function complexity is expressed as Fc=L+B, V(G) represents the loop complexity of the graph, E is the number of edges in the control flow graph, N is the number of nodes in the control flow graph, P is the number of independent connected areas in the program, Fc is the function complexity, L is the number of lines of code in the function, and B is the number of branches in the function; the code complexity indicator is expressed as C=w1*V(G)+w2*Fc+w 3*N'+w4*D, w1, w2, w3, w4 are weight coefficients corresponding to the four code complexity indicators, D is the code duplication, and the value of the nesting depth N' is equal to the number of layers of the maximum nested conditional statements or loop statements; the maintainability index is expressed as M=w'1*V(G)+w'2*L+w'3*D+w'4*C'+w'5*CR, M is the maintainability score of the code, C' is the module coupling, CR is the comment rate, w'1, w'2, w'3, w'4, w'5 are the weight coefficients corresponding to the five maintainability indicators; defect density d=number of defects / number of lines of code*1000; The use of the core indicator information to establish a causal relationship model includes: establishing a causal reasoning model Do-Calculus, wherein the causal graph shows: C → Q, used to indicate that code complexity affects code quality, d → Q is used to indicate that defect density affects code quality, and M → Q is used to indicate that maintainability affects code quality; code quality is expressed as Q=β0+β1*C+β2*d+β3*M+ϵ1, C=γ1*d+γ2*M+ϵ2, d=λ1*M+ϵ3, Q is code quality, β0 is a constant term, β1, β2, β3 are three regression coefficients to be estimated, which are used to indicate the causal influence of the core indicator on code quality, γ1, γ2 indicate the influence degree coefficients of defect density and maintainability on code complexity, λ1 indicates the influence degree coefficient of maintainability on defect density, and ϵ1~ ϵ3 indicate the first to third error terms.
2. The code quality optimization method based on causal reasoning and LLM according to claim 1 is characterized in that: The corresponding relationship between the weight coefficients of the four code complexity indicators and the project scenario types includes: For rapid development project scenarios: w1=0.2,w2=0.3,w3=0.2,w4=0.3; For enterprise-level application scenarios: w1=0.3, w2=0.4, w3=0.2, w4=0.1; For performance optimization project scenarios: w1=0.4, w2=0.2, w3=0.3, w4=0.1; For fast iteration product scenarios: w1=0.25,w2=0.25,w3=0.25,w4=0.
25.
3. The code quality optimization method based on causal reasoning and LLM according to claim 1, characterized in that: The corresponding relationship between the weight coefficients of the five maintainability indicators and the project scenario types includes: For rapid development project scenarios: w'1=0.3, w'2=0.2, w'3=0.2, w'4=0.15, w'5=0.15; For enterprise-level application scenarios: w'1=0.35, w'2=0.3, w'3=0.15, w'4=0.1, w'5=0.1; For performance optimization project scenarios: w'1=0.4, w'2=0.2, w'3=0.15, w'4=0.15, w'5=0.1; For fast iteration product scenarios: w'1=0.25,w'2=0.25,w'3=0.2,w'4=0.2,w'5=0.
1.
4. The code quality optimization method based on causal reasoning and LLM according to claim 1, characterized in that: In S2, it includes: Through the Do-Calculus method and Counterfactual Analysis, the causal intervention effects of different optimization strategies on code quality are simulated; The Do-Calculus method is used to deduce the causal intervention effect, which is expressed as: P(Q|do(C=c′),do(d=d′),do(M=m′)), where c′, d′, m′ represent the modified code complexity, defect density and maintainability in the Do-Calculus method; The counterfactual analysis is used to exclude strategies that will lead to negative optimization, expressed as: P(Q|C,d,M)vsP(Q|C′,d′,M′), where C′,d′,M′ represent the modified code complexity, defect density and maintainability in the counterfactual analysis.
5. A code quality optimization device based on causal reasoning and LLM, characterized in that: include: An analysis engine module, used to parse the target code data and establish a corresponding causal relationship model; A simulation module, used to simulate the causal intervention effects of different optimization strategies on code quality by using the causal relationship model; A solution generation module is used to generate an optimization solution based on the obtained simulation results using a large language model; An update module, used to record modifications to the target code data, perform regression testing, and verify optimization effects; The analysis engine module is specifically used to obtain the core indicator information of the target code data; use the core indicator information to establish a causal relationship model, and the causal relationship model is used to calculate the impact of each core indicator on the target variable; The obtaining of the core indicator information of the target code data includes: scanning the code base of the target code data by a static analysis tool to obtain the code complexity of the target code data, and determining maintainability and defect density; Among them, the corresponding code complexity indicators include: the cyclomatic complexity, function complexity and nesting depth of the code of the corresponding modules, classes and methods obtained from the target code, wherein the cyclomatic complexity is expressed as V(G)=E-N+2P, the function complexity is expressed as Fc=L+B, V(G) represents the loop complexity of the graph, E is the number of edges in the control flow graph, N is the number of nodes in the control flow graph, P is the number of independent connected areas in the program, Fc is the function complexity, L is the number of lines of code in the function, and B is the number of branches in the function; the code complexity indicator is expressed as C=w1*V(G)+w2*Fc+w 3*N'+w4*D, w1, w2, w3, w4 are weight coefficients corresponding to the four code complexity indicators, D is the code duplication, and the value of the nesting depth N' is equal to the number of layers of the maximum nested conditional statements or loop statements; the maintainability index is expressed as M=w'1*V(G)+w'2*L+w'3*D+w'4*C'+w'5*CR, M is the maintainability score of the code, C' is the module coupling, CR is the comment rate, w'1, w'2, w'3, w'4, w'5 are the weight coefficients corresponding to the five maintainability indicators; defect density d=number of defects / number of lines of code*1000; The use of the core indicator information to establish a causal relationship model includes: establishing a causal reasoning model Do-Calculus, wherein the causal graph shows: C → Q, used to indicate that code complexity affects code quality, d → Q is used to indicate that defect density affects code quality, and M → Q is used to indicate that maintainability affects code quality; code quality is expressed as Q=β0+β1*C+β2*d+β3*M+ϵ1, C=γ1*d+γ2*M+ϵ2, d=λ1*M+ϵ3, Q is code quality, β0 is a constant term, β1, β2, β3 are three regression coefficients to be estimated, which are used to indicate the causal influence of the core indicator on code quality, γ1, γ2 indicate the influence degree coefficients of defect density and maintainability on code complexity, λ1 indicates the influence degree coefficient of maintainability on defect density, and ϵ1~ ϵ3 indicate the first to third error terms.
6. A storage medium, characterized in that: include: A computer program or instruction is stored, and when the computer program or instruction is executed, the method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Model validation as a service
US12086053B1
Systems and methods for generating natural language using language models trained on computer code
US20240020116A1